All guides

Which separator and encoding does my CSV use?

“CSV” names a family, not a format. The separator can be a comma, a semicolon, a tab or a pipe, and the bytes can be UTF-8, an old Windows code page, or UTF-16 straight out of Excel. Both are visible in the first line if you know what to look for, and both can be rewritten without touching a single value.

Telling the separator from the first line

Open the file in a text editor rather than a spreadsheet — Notepad, TextEdit, VS Code — and look at the header. The character between the column names is the separator: a comma, a semicolon, a wide gap that is a tab, or a vertical bar. If a spreadsheet showed you everything in one column, the separator is whatever character is sitting inside those long cells.

When a line does not make it obvious, count: the separator is the character that appears the same number of times on every line. Quoted values can contain the other characters, which is why a comma inside "Smith, Jane" does not make the file semicolon-separated. The delimiter fixer does this count for you and shows what it found, with an override for the one case that fools counting — a file of decimal commas with semicolons between fields.

Telling the encoding from the damage

The encoding shows itself when a program reads the bytes as the wrong one. é where é should be, or ’ for an apostrophe, is UTF-8 read as Windows-1252. A � in place of a letter is the reverse: Windows-1252 bytes read as UTF-8. A space between every letter, or a file that opens as gibberish, is UTF-16 — Excel writes it when you choose "Unicode Text". Three invisible characters at the very start, shown as  by some programs, are the UTF-8 byte order mark that Excel’s "CSV UTF-8" adds.

None of this means the data is damaged. The bytes are what they always were; the reading is wrong. The health check reports garbled text and a byte order mark when it sees them, so you know which case you have before deciding what to do.

Fixing both at once

The delimiter fixer decodes the file by what it actually is — UTF-8 with or without the byte order mark, or Windows-1252, the usual case for an old export — says which it found and which separator, and writes a clean UTF-8 copy with the separator you choose, quoting values that need it. Every value is carried as text: a customer number keeps its zeros and a long ID stays a long ID. Decimal commas can be rewritten to points at the same time, counted before you download, and switched off if you would rather they stayed.

UTF-16 is the one encoding it recognises but does not rewrite: it stops and says so. Open the file in Excel or a text editor and save it as CSV UTF-8, then the fixer takes it from there. Excel only writes UTF-16 from "Unicode Text", so the re-save is the whole fix.

If the file only needs its separator changed and is already UTF-8, the tab-to-comma converter is the direct route for a tab-separated file; the fixer covers every other combination.

The sample file

Auftrag;Kunde;Betrag;Status
S-1001;Müller;1.234,56;bezahlt
S-1002;Bär;89,50;bezahlt
S-1003;Schulz;32,10;erstattet
S-1004;König;410,00;bezahlt
S-1005;Vogel;55,25;offen

Obviously made-up data, small enough to read. Download it and drop it on the tool to see the fix before trying your own file.

Questions

Why did Excel save my CSV with semicolons?
Because your system uses a comma as the decimal mark, so Excel uses the list separator — a semicolon in most of Europe — between values. The file is correct for your locale and wrong for tools that expect commas.
How do I know if a file is UTF-8?
If every accented letter and symbol looks right when opened as UTF-8 — most editors say the encoding in the status bar — it is. é and ’ mean it is UTF-8 being read as something else; � means it is something else being read as UTF-8.
What is the  at the start of my file?
The UTF-8 byte order mark, three bytes Excel writes at the front of a "CSV UTF-8" file. Most programs skip it; some read it as part of the first column name. The fixer and the editor here both leave it out of the download.
Excel shows é instead of é — is my data corrupted?
No. The file is UTF-8 and Excel opened it as Windows-1252. Open it through Data › From Text with UTF-8 chosen, or run it through the fixer here, which writes a copy Excel reads correctly.
My file is UTF-16 — can the fixer read it?
It recognises UTF-16 and stops, naming the way out rather than guessing: save the file as CSV UTF-8 from Excel or a text editor, and the fixer reads that copy and sorts out the separator.
Can a CSV use tabs or pipes?
Yes, and many exports do — tab-separated files are common from databases and pipes from systems whose data contains commas. The fixer reads all of them and writes whichever you choose.
Is my file uploaded anywhere?
No. The separator and the encoding are read from the file’s own bytes inside this page. Open the Network tab in your developer tools before you drop the file, and click through what appears: not one request carries a byte of it.