Common causes
1. A UTF-8 CSV without a byte-order mark opened in Excel
Excel on Windows assumes the system code page unless the file starts with a UTF-8 BOM. Export with a BOM (Python: encoding="utf-8-sig"; Excel: “CSV UTF-8”), or import through Data > From Text/CSV and pick 65001: Unicode (UTF-8).
df.to_csv("customers.csv", index=False)df.to_csv("customers.csv", index=False, encoding="utf-8-sig")2. Text that was already double-encoded
If a pipeline read UTF-8 as Latin-1 and saved the result as UTF-8, the mojibake is now stored in the file. Reverse the mistake once: encode as Latin-1 (or cp1252) and decode as UTF-8.
name,city
José,Zürichname,city
José,Zürich3. Curly quotes and dashes from Word
Typographic characters are three bytes in UTF-8, so a misread ’ becomes ’ and “ becomes “. The fix is the same as for accented letters: read the file as UTF-8.
id,note
1,Customer’s orderid,note
1,Customer’s order4. A Windows-1252 file read as UTF-8
Excel’s plain “CSV (Comma delimited)” format writes Windows-1252 on Windows. Reading it as UTF-8 fails on the first accented character. Tell the reader the real encoding.
df = pd.read_csv("export.csv")df = pd.read_csv("export.csv", encoding="cp1252")Frequently asked questions
Why does the file look fine in a text editor but not in Excel?
Modern editors detect UTF-8 automatically. Excel on Windows does not unless the file starts with a byte-order mark, so the same bytes are shown differently.
How do I fix mojibake that is already in my data?
In Python, text.encode(“cp1252”).decode(“utf-8”) reverses one round of the mistake. The ftfy library detects and repairs mixed or repeated mojibake automatically.
Does adding a BOM break other tools?
Some strict parsers treat the BOM as part of the first column name. Most CSV readers strip it, and PasteKit does too. See the CSV and Excel encoding guide for a full comparison.