Repository navigation
[BUG] Claude Code does not respect file encoding, corrupts Windows-1252 files #7134
Description
Activity
- addedduplicateThis issue or pull request already existsThis issue or pull request already existsplatform:windowsIssue specifically occurs on WindowsIssue specifically occurs on Windows
on Sep 4, 2025 Found 3 possible duplicate issues:
- Character Encoding Handling Fails for WINDOWS-1252 ADVPL Files #3416
- [BUG] Edit tool corrupts Windows-1252 encoding to UTF-8 #5518
- File Encoding Limitation: Edit Tool Overwrites Non-UTF-8 Files #6485
This issue will be automatically closed as a duplicate in 3 days.
- If your issue is a duplicate, please close it and 👍 the existing issue instead
- To prevent auto-closure, add a comment or 👎 this comment
🤖 Generated with Claude Code
Reacted by Edson Lyra, Clemens Gruber, Vladimir Yudin and pdfbrewerPlease give your issue a descriptive title
- changed the title
[-][BUG][/-][+][BUG] Claude Code does not respect file encoding, corrupts Windows-1252 files[/+]on Sep 4, 2025 fixing the errors generated by this issue consumes a lot of tokens
2 workarounds are on this older ticket : #3416
The best is probably to use this : https://github.com/devslimbr/cc-tools
Reacted by Edson Lyra and brice-tThis issue has been inactive for 30 days. If the issue is still occurring, please comment to let us know. Otherwise, this issue will be automatically closed in 30 days for housekeeping purposes.
- addedautocloseIssue will be closed automaticallyIssue will be closed automatically
on Dec 7, 2025 Still an issue. Same with Win-1251 encoding
Reacted by Edson Lyra- removedautocloseIssue will be closed automaticallyIssue will be closed automatically
on Dec 15, 2025 Still actively affecting users. This issue is part of a larger UTF-8/encoding bug pattern in Claude Code's Edit/Write tools.
I've posted a consolidated impact analysis at #13939 documenting 6+ related issues spanning 6+ months. This is a systemic problem affecting all international users.
Cross-linking for visibility: #13939, #13080, #7332, #7335
🤖 Generated with Claude Code
Reacted by Edson Lyra, VoronkovIvan, Enrique Blanco, Clemens Gruber, Evandro Lucas Figueiredo Teixeira and sideshowbarkerIf anyone needs a workaround today I wrote an MCP server that handles encoding-aware file reads and writes CP1251, CP1252, ISO-8859, KOI8 with auto-detection. Single Go binary, no runtime dependencies.
https://github.com/dimitar-grigorov/mcp-file-toolsReacted by 陪我灬去流浪13 remaining items
Facing the same problem, would be nice to have native solution.
AI nowadays could do anything, but they can't solve this single issue.
Facing the same problem, would be nice to have native solution.
Adding another data point:
Environment: Claude Code via AWS Bedrock
Codebase: Legacy PL/I mainframe sources stored in Latin-1 / ISO-8859-1 (not UTF-8). The encoding is mandated by the toolchain and can't be changed.
Symptoms (same mechanism described above):
- Read decodes the file as UTF-8, so umlauts (ä ö ü ß) come back as the U+FFFD replacement character �.
- Edit writes the file back as UTF-8, re-encoding every pre-existing umlaut. Lines I never touched show up as changes, flooding
git diffwith false encoding noise, and the result is invalid for a toolchain that requires Latin-1.
Current workaround — bypassing Read/Edit with a Python shim
Confirming this on a legacy Delphi (Windows-1252/ANSI) codebase, with a minimal, deterministic reproduction that isolates the tool from any editor.
Key finding: an ASCII-only edit — a change that touches no non-ASCII characters — still corrupts every non-ASCII byte in the whole file. So this isn't about the edited span; the tool round-trips the entire file through a lossy UTF-8 decode/encode.
Minimal reproduction
-
Create a 3-line file encoded as Windows-1252 containing 7 non-ASCII bytes on line 2:
$enc = [System.Text.Encoding]::GetEncoding(1252) $ae=[char]0xE4;$oe=[char]0xF6;$ue=[char]0xFC;$ss=[char]0xDF;$Ae=[char]0xC4;$Oe=[char]0xD6;$Ue=[char]0xDC $txt = "line 1: ASCII marker MARKER_A`r`n// umlauts: $ae $oe $ue $ss $Ae $Oe $Ue`r`nline 3: ASCII end`r`n" [System.IO.File]::WriteAllText("repro.txt", $txt, $enc) # -> 7 valid cp1252 bytes (E4 F6 FC DF C4 D6 DC), 0x EF BF BD sequences, 80 bytes
-
Use the Read tool on
repro.txt. It displays the 7 umlauts as the replacement character�(U+FFFD) — i.e. the original bytes are already lost at read time. -
Use the Edit tool to change only ASCII on line 1:
MARKER_A→MARKER_B(line 2 is untouched). -
Inspect the bytes:
$b=[System.IO.File]::ReadAllBytes("repro.txt") ($b | Where-Object { $_ -in 0xE4,0xF6,0xFC,0xC4,0xD6,0xDC,0xDF }).Count # valid cp1252 umlauts $c=0; for($i=0;$i -lt $b.Length-2;$i++){ if($b[$i]-eq0xEF -and $b[$i+1]-eq0xBF -and $b[$i+2]-eq0xBD){$c++} }; $c # EF BF BD sequences
Result
valid cp1252 umlaut bytes EF BF BDsequencessize before edit 7 0 80 after ASCII-only edit 0 7 94 Every umlaut (
E4,F6, …) becameEF BF BD(the 3-byte UTF-8 encoding of U+FFFD), and the file grew by 14 bytes (7 × 2). In a Windows-1252 editor these now show as�.Why this is bad
- The corruption is silent and irreversible: because Read already decodes to U+FFFD, the identity of each character (ä vs ö vs ü vs ß …) is gone — it can't be recovered without re-typing from context.
- It corrupts the entire file on any edit, so a single unrelated change to a large source file destroys hundreds of characters at once.
- It breaks compiler-relevant strings and comments in Delphi/Pascal, C/C++, etc. that must stay in Windows-1252/ANSI.
Note re: editors
The Zed editor itself recently added legacy-encoding open/save with auto-detection (zed-industries/zed PR #44819, merged 2025-12), and it preserves Windows-1252 on save. But the agent Edit/Write tools still corrupt the file (reproduced above), which suggests they write directly rather than through the encoding-aware path.
Suggested fix
Detect the file's encoding on read and preserve it on write; or, at minimum, don't re-encode bytes that weren't part of the edit. A per-file/target-encoding override would also work.
Environment
- OS: Windows 10 (
Pro 22H2) - Agent:
Claude Code (CLI) 2.1.193, accessed viaZed 1.10.0 - File encoding: Windows-1252 (Delphi/Pascal codebase)
-
Can confirm the same root cause with a different legacy codepage: Windows-1250 (CP1250), Polish Java source files.
One detail worth adding to the report: the corruption isn't limited to the characters/lines actually touched by the edit. Edit/Write appear to decode the entire file as UTF-8 on read (lossy — any CP1250/CP1252 high byte that isn't valid UTF-8 becomes U+FFFD) and then re-serialize the entire file as UTF-8 on write. So a single-line edit corrupts every accented character anywhere in the file, including lines nowhere near the diff. Repeated edits compound this, since each subsequent Edit call re-reads the now-partially-UTF-8/partially-corrupted file and repeats the process.
Recovery required manually reconstructing the original text from conversation history/context and re-saving with an external tool (PowerShell [System.Text.Encoding]::GetEncoding(1250)) — Claude Code itself has no way to write anything but UTF-8, so there's no in-tool fix once a non-UTF-8 file has been touched.
+1 for detecting/preserving the original encoding (or at minimum failing loudly on a lossy decode instead of silently substituting U+FFFD and persisting the loss).
Please try Kilo Code. It's the only plugin/AI assistant I could find that supports legacy encoding.
Confirming this bug with a large-scale case, plus a reliable detection method and a recovery strategy that may help others who have already been hit.
Environment
Field Value Claude Code version OS Windows 10 File types .pas,.dfm,.dpr(Delphi source)Original encoding ANSI / Windows-1252, no BOM Encoding after edit UTF-8 without BOM, accented characters replaced by U+FFFD What happened
We are migrating a ~1.4M-line ERP from Delphi 5 to Delphi 11. The sources are ANSI/CP1252 with Portuguese text (
ç ã é í ó â ê õ à ü) in string literals, comments and form captions.Over one working session, Claude Code applied a number of small, targeted edits — adding a unit to a
usesclause, qualifying a function call, fixing a missing comma. Each edit touched one or two lines.The result:
Files Corrupted characters .pas348 17,947 .dfm81 3,968 The damage is not proportional to the edit. A one-character fix destroys every accented character in the file, because the whole file is rewritten on save.
Why this is worse than it first appears
-
It is completely silent. The corrupted literals are still valid literals. The project compiled with 0 errors. The problem only became visible when the application was launched and the menus read
Gest�o,Servi�os,Pa�ses. -
It reaches production data. In our case the corrupted text includes error messages, report captions and strings written to the database.
-
The information is unrecoverable from the file alone.
U+FFFDdoes not record which character it replaced. Without a clean copy,configura��ocannot be resolved toconfiguraçãoautomatically. -
Consecutive accented characters can collapse. When two accented bytes are adjacent, the UTF-8 decoder sometimes emits a single U+FFFD for both, so the corrupted text is not even the same length as the original. This matters for any automated repair.
Detection: the intuitive check does not work
This tripped us up for a while and is worth documenting.
You cannot detect the corruption by opening the file and looking for
?/`` characters, and you cannot detect it by reading the file as UTF-8 in a script — a clean ANSI file read as UTF-8 also shows U+FFFD, because its accented bytes are invalid UTF-8. Both the healthy file and the corrupted file look identical that way.The only reliable test is to look for the byte sequence
EF BF BD(the UTF-8 encoding of U+FFFD) in the raw bytes:$b = [System.IO.File]::ReadAllBytes($path) for ($i = 0; $i -lt $b.Length - 2; $i++) { if ($b[$i] -eq 0xEF -and $b[$i+1] -eq 0xBF -and $b[$i+2] -eq 0xBD) { "CORRUPTED: $path"; break } }
A clean CP1252 file will never contain that sequence. A corrupted one will.
Recovery, for anyone already affected
Recovery is possible without losing the edits, provided you have a clean copy from before the damage, and provided the edits themselves are ASCII-only.
The corrupted file is the clean file with two independent layers on top:
- (a) the intended edits, which are ASCII;
- (b) every non-ASCII byte replaced by U+FFFD.
If you normalize the clean copy by replacing every non-ASCII character with U+FFFD, the lines you did not edit become byte-identical to the corrupted ones. Those can be restored verbatim from the clean copy. The lines that differ are your real edits and are left untouched.
Applying this to our codebase recovered 21,613 of 21,777 corrupted lines (99.2%). The remaining 164 were lines that were both edited and contained accents, plus lines with consecutive accented characters (see point 4 above) — all reported in a CSV for manual review rather than guessed at.
Two details that matter if you implement this:
- Match on repeated lines as well.
.dfmfiles are full of duplicated captions; if all candidate lines in the clean copy are identical, there is no ambiguity even though the line occurs many times. Requiring a unique occurrence leaves ~25% of recoverable lines behind. - Use ordinal, case-sensitive string comparison. PowerShell's
Sort-Object -Uniqueis case-insensitive and will silently merge lines that differ only in case.
Suggested fix
Detect the encoding on read and preserve it on write, at minimum for the no-BOM CP1252 case. Failing that, a per-project setting (e.g.
"fileEncoding": "windows-1252") or refusing to write a file whose original encoding cannot be preserved would both be far better than silent data loss.A warning at edit time would already have prevented this entire incident.
-
Problem persisting with windows-1252 files
Still reproducing on 2.1.263 (latest as of today, 2026-09-08), Windows 11.
Hit this twice today on the same C# file in two different solutions. The file is Windows-1252, no BOM, ~177 KB, with accented characters in comments and string literals (Italian codebase). A single Edit to
an unrelated line rewrote the whole file as UTF-8, turning every pre-existing è/à/é into è/à /é. Had to revert and reapply manually both times.Three things that make this worse than a plain encoding mismatch:
- It's silent. The tool reports success. Nothing in the result indicates the file was re-encoded.
- The damage extends beyond the edited region. The entire file is rewritten, so untouched lines elsewhere get corrupted too.
- It defeats the obvious verification. Reading the file back after the edit shows correct content, because it's read in the same encoding it was just written in. The corruption only appears when the file is
read in its actual codepage. I initially declared the file clean after checking it this way — it wasn't.
The only workaround is to bypass Edit/Write entirely and go through a shell command with an explicit codepage:
$enc = [System.Text.Encoding]::GetEncoding(1252)
$lines = [System.IO.File]::ReadAllLines($path, $enc)...modify...
[System.IO.File]::WriteAllLines($path, $lines, $enc)
That works, but it's error-prone and depends on remembering to do it on every single subsequent edit to that file.
Given this has been open for a year: even without full encoding detection, refusing the write with a clear error when the target file's encoding can't be preserved would be a big improvement over silently
corrupting it. Right now the failure mode is "looks fine, ships broken".When that issue will be fixed ? Now we have to use external tools like github.com/dimitar-grigorov/mcp-file-tools, but a real harness should have native support for that.
There are lots of work-arounds, -one is to instruct Claude Code in always converting files to UTF8 before making changes, but still in some cases it doesn't do that, and it introduce garbage characters in the text files.
We have loads of code files still in ansi format.- added a commit that references this issue
on Oct 10, 2026 Update for anyone following this: 2.1.296 changes Edit from silently corrupting to refusing, per its changelog. Checked on Windows 11 with a Windows-1252 C# file (Italian accents in a comment and a string literal, CRLF, 97 bytes), via
claude -p(Sonnet).Edit now refuses and writes nothing. The file stays byte-for-byte unchanged, and the tool returns:
File is not valid UTF-8. It may use a legacy encoding such as Windows-1252, Shift-JIS or GBK, or be binary. This tool saves the whole file as UTF-8, which would replace every byte it cannot decode with U+FFFD. Nothing was written. Make the change with a shell command that reads and writes the file in its own encoding, or ask the user whether to convert the file to UTF-8 first.
Two gaps remain:
- Write is not covered. Write on the same existing Windows-1252 file (content given verbatim, with the real accented letters) returns "updated successfully" with no warning and saves it as UTF-8 with LF:
èE8 → C3 A8, CR 5 → 0. A model that reaches for Write instead of Edit still re-encodes the file silently. - Read still shows every accented letter as U+FFFD (5 of 5 here:
// Commento: � gi� cos�, perch� funziona). With shell tools available, the model saw that, skipped Edit, and went through Bash; in 2 of 2 runssed -istripped every CR on the way (97 → 92 bytes), the model noticed and repaired it byte by byte in PowerShell, and the end result was the correct one-byte change in Windows-1252. It got there, but by a detour.
So whole-file corruption through Edit looks fixed in 2.1.296. Native support for legacy encodings, which most comments here ask for, is not part of the fix.
- Write is not covered. Write on the same existing Windows-1252 file (content given verbatim, with the real accented letters) returns "updated successfully" with no warning and saves it as UTF-8 with LF:
Environment
Bug Description
Claude Code does not respect the original file encoding when editing files. It assumes UTF-8 encoding for all operations, which causes character corruption when working with files encoded in other formats like Windows-1252.
Steps to Reproduce
Expected Behavior
Claude Code should:
Current Workaround
Users must either:
Impact
This affects any codebase that uses non-UTF-8 encoding, particularly:
Technical Details
Environment
This is a critical issue for maintaining legacy codebases that require specific encodings for compiler compatibility.