All guides

pandas says "Error tokenizing data. C error: Expected 5 fields, saw 7" — find the row and fix the file

The message means one line of the file split into more pieces than the heading row did: a comma inside an address that was never quoted, a line break in the middle of a note, or a heading row shorter than the data. pandas stops at the first such line and reports its number. Skipping bad lines makes the error go away and takes those rows with it; finding the lines and fixing the file keeps them.

What the message is counting

read_csv takes the number of fields in the first line as the width of the table, then reads every other line expecting the same count. "Expected 5 fields in line 8, saw 7" means line 8 had two extra separators — most often a value with a comma in it, such as "Smith, John" or "12, High Street", written without the quotes that would mark it as one field. A note with a line break inside it does the opposite, and the two halves become two short lines.

A heading row with fewer names than the data has columns produces the error on the very first data line, which is the clue: when every line is "wrong", the first line is.

Find the rows here

The health check reads the whole file and lists the rows whose field count differs from the heading, with their line numbers and what it saw — the first five, and how many more there are, so a file with a hundred bad lines is fixed five at a time rather than one at a time. It also reports the separator it detected, which catches the case where the file is semicolon-separated and the commas are inside the values.

With the line numbers in hand, the editor opens the file with every value as written, so the offending cell can be quoted or the stray break removed and the file downloaded with only that change.

Fix the file, not the code

on_bad_lines="skip" makes the error disappear by dropping every row it cannot read, silently. For a one-off look that is fine; for a file that feeds anything else, the rows are gone and nothing downstream knows. Quoting the field that contains the separator fixes the row for every program that will ever read the file.

When the file uses a separator other than the one you told read_csv, the delimiter fixer here rewrites it as a standard comma-separated file with proper quoting, so the code needs no special arguments at all.

Questions

Why does it say saw 7 when my rows have 5 columns?
Because that line contains two extra separators inside values — a comma in a name or an address, unquoted — so it was split into seven pieces. The health check here shows the line and the pieces it saw.
Is on_bad_lines="skip" a fix?
It makes the error stop, by dropping the rows it cannot read. If those rows matter, they are silently gone; quote the field or fix the line and every row survives.
How do I find every bad line, not only the first?
Run the health check here. It reads the whole file, counts every row whose field count differs from the heading, and shows the first five with their line numbers; fix those, run it again, and the next five appear. pandas stops at the first.
Could the separator be wrong rather than the rows?
Yes, and the health check reports which separator it detected. A semicolon-separated file read with commas produces this error on any line whose values contain commas.
Can I fix the rows without a script?
Yes. Open the file in the editor here, go to the line numbers the health check gave, quote the field or join the broken line, and download the file with only those edits applied.
Is my file uploaded anywhere?
No. The health check counts the fields on every row inside this page, so the rows pandas stopped on are found and numbered without the file going anywhere.