Skip to content

Windows-1252 text is decoded with the wrong charset ("Café crème" becomes "Cafﻠ crﻟme") #2678

Description

Bug

Text files in Windows-1252 (or Latin-1) with French, Spanish, Portuguese or Nordic text come out garbled. There is no error, so the wrong text goes straight into the Markdown. The same happens to .csv files and to text files inside a .zip.

import io

from markitdown import MarkItDown

md = MarkItDown()
for text in [
    "Café crème",
    "Le café crème coûte 3 € à la boulangerie près de l'église.",
    "Señor Muñoz, año próximo",
    "Blåbærgrød og rødgrød med fløde på ø-hoppet i Århus.",
]:
    data = text.encode("cp1252")
    print(md.convert_stream(io.BytesIO(data), file_extension=".txt").markdown)

Output on main (4cc9fa1), charset_normalizer 3.5.2:

Cafﻠ crﻟme
Le café crčme coűte 3 € ŕ la boulangerie prčs de l'église.
Seńor Muńoz, ańo próximo
Blĺbćrgrřd og rřdgrřd med flřde pĺ ř-hoppet i Ĺrhus.

German text is not affected, because its letters are the same in cp1250 and cp1252.

Cause

_get_stream_info_guesses (and the plain text and CSV fallbacks) use charset_normalizer.from_bytes(data).best(). For these samples, cp1252 is in the match list with exactly the same chaos and coherence as the first match, and best() returns the first one, which is cp1250 (or cp1006 for "Café crème"). For example, for the French sentence:

cp1250   chaos=0.000 coherence=0.524   decodes correctly: no
cp1252   chaos=0.000 coherence=0.524   decodes correctly: yes

The charset_normalizer maintainer has said such close calls can't be fixed in the detector, and advised callers to choose among the top matches themselves (jawah/charset_normalizer#524, jawah/charset_normalizer#571).

Proposed fix

When cp1252 ties with the best match in both chaos and coherence, use cp1252, unless the data contains S or Z with caron (bytes 0x8A 0x8E 0x9A 0x9E). Those bytes are the same in cp1250 and cp1252, are common in Croatian and Slovene, and are rare in Western European text.

I compared the charset picked with and without this rule on 26 samples in 20 languages:

Result Samples
Fixed French (3), Spanish (2), Portuguese, Danish, Icelandic
Correct before and after German, Polish, Czech, Hungarian, Croatian, Slovene, Turkish, Greek, Russian
Wrong before and after Italian, Dutch, Romanian, Finnish, Lithuanian, Croatian without š/ž, short Turkish and Russian samples
Correct before, wrong after Bosnian sentence without š/ž ("Ćevapi u Sarajevu su čuveni, a đaci ih vole.")

The last row is the trade-off: Croatian, Slovene and Bosnian text without š or ž is byte-for-byte as ambiguous as the French sample, and today it only decodes correctly because cp1250 sorts first. I think preferring cp1252 on a tie is the better default, but I'm happy to narrow or change the rule if you see it differently.

Related

I have a fix with tests and will open a PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions