Bug
Text files in Windows-1252 (or Latin-1) with French, Spanish, Portuguese or Nordic text come out garbled. There is no error, so the wrong text goes straight into the Markdown. The same happens to .csv files and to text files inside a .zip.
import io
from markitdown import MarkItDown
md = MarkItDown()
for text in [
"Café crème",
"Le café crème coûte 3 € à la boulangerie près de l'église.",
"Señor Muñoz, año próximo",
"Blåbærgrød og rødgrød med fløde på ø-hoppet i Århus.",
]:
data = text.encode("cp1252")
print(md.convert_stream(io.BytesIO(data), file_extension=".txt").markdown)
Output on main (4cc9fa1), charset_normalizer 3.5.2:
Cafﻠ crﻟme
Le café crčme coűte 3 € ŕ la boulangerie prčs de l'église.
Seńor Muńoz, ańo próximo
Blĺbćrgrřd og rřdgrřd med flřde pĺ ř-hoppet i Ĺrhus.
German text is not affected, because its letters are the same in cp1250 and cp1252.
Cause
_get_stream_info_guesses (and the plain text and CSV fallbacks) use charset_normalizer.from_bytes(data).best(). For these samples, cp1252 is in the match list with exactly the same chaos and coherence as the first match, and best() returns the first one, which is cp1250 (or cp1006 for "Café crème"). For example, for the French sentence:
cp1250 chaos=0.000 coherence=0.524 decodes correctly: no
cp1252 chaos=0.000 coherence=0.524 decodes correctly: yes
The charset_normalizer maintainer has said such close calls can't be fixed in the detector, and advised callers to choose among the top matches themselves (jawah/charset_normalizer#524, jawah/charset_normalizer#571).
Proposed fix
When cp1252 ties with the best match in both chaos and coherence, use cp1252, unless the data contains S or Z with caron (bytes 0x8A 0x8E 0x9A 0x9E). Those bytes are the same in cp1250 and cp1252, are common in Croatian and Slovene, and are rare in Western European text.
I compared the charset picked with and without this rule on 26 samples in 20 languages:
| Result |
Samples |
| Fixed |
French (3), Spanish (2), Portuguese, Danish, Icelandic |
| Correct before and after |
German, Polish, Czech, Hungarian, Croatian, Slovene, Turkish, Greek, Russian |
| Wrong before and after |
Italian, Dutch, Romanian, Finnish, Lithuanian, Croatian without š/ž, short Turkish and Russian samples |
| Correct before, wrong after |
Bosnian sentence without š/ž ("Ćevapi u Sarajevu su čuveni, a đaci ih vole.") |
The last row is the trade-off: Croatian, Slovene and Bosnian text without š or ž is byte-for-byte as ambiguous as the French sample, and today it only decodes correctly because cp1250 sorts first. I think preferring cp1252 on a tie is the better default, but I'm happy to narrow or change the rule if you see it differently.
Related
I have a fix with tests and will open a PR.
Bug
Text files in Windows-1252 (or Latin-1) with French, Spanish, Portuguese or Nordic text come out garbled. There is no error, so the wrong text goes straight into the Markdown. The same happens to
.csvfiles and to text files inside a.zip.Output on
main(4cc9fa1), charset_normalizer 3.5.2:German text is not affected, because its letters are the same in cp1250 and cp1252.
Cause
_get_stream_info_guesses(and the plain text and CSV fallbacks) usecharset_normalizer.from_bytes(data).best(). For these samples,cp1252is in the match list with exactly the same chaos and coherence as the first match, andbest()returns the first one, which iscp1250(orcp1006for "Café crème"). For example, for the French sentence:The charset_normalizer maintainer has said such close calls can't be fixed in the detector, and advised callers to choose among the top matches themselves (jawah/charset_normalizer#524, jawah/charset_normalizer#571).
Proposed fix
When
cp1252ties with the best match in both chaos and coherence, usecp1252, unless the data contains S or Z with caron (bytes0x8A 0x8E 0x9A 0x9E). Those bytes are the same in cp1250 and cp1252, are common in Croatian and Slovene, and are rare in Western European text.I compared the charset picked with and without this rule on 26 samples in 20 languages:
The last row is the trade-off: Croatian, Slovene and Bosnian text without š or ž is byte-for-byte as ambiguous as the French sample, and today it only decodes correctly because
cp1250sorts first. I think preferring cp1252 on a tie is the better default, but I'm happy to narrow or change the rule if you see it differently.Related
I have a fix with tests and will open a PR.