Skip to content

[BUG] Claude Code does not respect file encoding, corrupts Windows-1252 files #7134

Description

@edlyra

Environment

  • Platform (select one):
    • [ X] Anthropic API
    • AWS Bedrock
    • Google Vertex AI
    • Other:
  • Claude CLI version: 1.0.103 (Claude Code)
  • Operating System: Windows 11
  • Terminal: VSCode

Bug Description

Claude Code does not respect the original file encoding when editing files. It assumes UTF-8 encoding for all operations, which causes character corruption when working with files encoded in other formats like Windows-1252.

Steps to Reproduce

  1. Open a file that is encoded in Windows-1252 (common for Delphi/Pascal projects)
  2. Use Claude Code's Edit tool to modify text containing accented characters (e.g., Portuguese: configuração, não)
  3. The file becomes corrupted with characters like configuração instead of configuração

Expected Behavior

Claude Code should:

  • Detect the original file encoding
  • Preserve the same encoding when making edits
  • Or provide a way to configure the target encoding for edits

Current Workaround

Users must either:

  1. Accept corrupted characters and manually fix them
  2. Maintain separate UTF-8 clones of their repositories specifically for Claude Code usage
  3. Use pre-commit scripts to convert between encodings

Impact

This affects any codebase that uses non-UTF-8 encoding, particularly:

  • Legacy Delphi/Pascal projects (Windows-1252)
  • Projects with specific encoding requirements
  • International projects with accented characters

Technical Details

  • Claude Code's Read tool correctly displays Windows-1252 files
  • The Edit/Write tools force UTF-8 output regardless of input encoding
  • VSCode correctly shows Windows-1252 files when configured properly
  • The issue only occurs when Claude Code modifies the files

Environment

  • Claude Code version: [current version]
  • File encoding: Windows-1252 (ANSI - Latin 1)
  • IDE: Delphi XE
  • Language: Portuguese (Brazilian) with accented characters

This is a critical issue for maintaining legacy codebases that require specific encodings for compiler compatibility.

Activity

  1. github-actions commented on Sep 4, 2025

    @github-actions

    Found 3 possible duplicate issues:

    1. Character Encoding Handling Fails for WINDOWS-1252 ADVPL Files #3416
    2. [BUG] Edit tool corrupts Windows-1252 encoding to UTF-8 #5518
    3. File Encoding Limitation: Edit Tool Overwrites Non-UTF-8 Files #6485

    This issue will be automatically closed as a duplicate in 3 days.

    • If your issue is a duplicate, please close it and 👍 the existing issue instead
    • To prevent auto-closure, add a comment or 👎 this comment

    🤖 Generated with Claude Code

  2. corneliusroemer commented on Sep 4, 2025

    @corneliusroemer

    Please give your issue a descriptive title

  3. changed the title [-][BUG][/-] [+][BUG] Claude Code does not respect file encoding, corrupts Windows-1252 files[/+] on Sep 4, 2025
  4. edlyra commented on Oct 7, 2025

    @edlyra
    Author

    fixing the errors generated by this issue consumes a lot of tokens

  5. ohuet commented on Oct 13, 2025

    @ohuet

    2 workarounds are on this older ticket : #3416

    The best is probably to use this : https://github.com/devslimbr/cc-tools

  6. github-actions commented on Dec 7, 2025

    @github-actions

    This issue has been inactive for 30 days. If the issue is still occurring, please comment to let us know. Otherwise, this issue will be automatically closed in 30 days for housekeeping purposes.

  7. VoronkovIvan commented on Dec 15, 2025

    @VoronkovIvan

    Still an issue. Same with Win-1251 encoding

  8. xrchz commented on Jan 12, 2026

    @xrchz

    Still actively affecting users. This issue is part of a larger UTF-8/encoding bug pattern in Claude Code's Edit/Write tools.

    I've posted a consolidated impact analysis at #13939 documenting 6+ related issues spanning 6+ months. This is a systemic problem affecting all international users.

    Cross-linking for visibility: #13939, #13080, #7332, #7335


    🤖 Generated with Claude Code

  9. dimitar-grigorov commented on Feb 10, 2026

    @dimitar-grigorov

    If anyone needs a workaround today I wrote an MCP server that handles encoding-aware file reads and writes CP1251, CP1252, ISO-8859, KOI8 with auto-detection. Single Go binary, no runtime dependencies.
    https://github.com/dimitar-grigorov/mcp-file-tools

  10. 13 remaining items

  11. mloeffler123 commented on Jun 7, 2026

    @mloeffler123

    Facing the same problem, would be nice to have native solution.

  12. myonlylonely commented on Jun 25, 2026

    @myonlylonely

    AI nowadays could do anything, but they can't solve this single issue.

  13. mloeffler123 commented on Jun 25, 2026

    @mloeffler123

    Facing the same problem, would be nice to have native solution.

  14. havmar commented on Jul 9, 2026

    @havmar

    Adding another data point:

    Environment: Claude Code via AWS Bedrock

    Codebase: Legacy PL/I mainframe sources stored in Latin-1 / ISO-8859-1 (not UTF-8). The encoding is mandated by the toolchain and can't be changed.

    Symptoms (same mechanism described above):

    • Read decodes the file as UTF-8, so umlauts (ä ö ü ß) come back as the U+FFFD replacement character �.
    • Edit writes the file back as UTF-8, re-encoding every pre-existing umlaut. Lines I never touched show up as changes, flooding git diff with false encoding noise, and the result is invalid for a toolchain that requires Latin-1.

    Current workaround — bypassing Read/Edit with a Python shim

  15. viaregio commented on Jul 9, 2026

    @viaregio

    Confirming this on a legacy Delphi (Windows-1252/ANSI) codebase, with a minimal, deterministic reproduction that isolates the tool from any editor.

    Key finding: an ASCII-only edit — a change that touches no non-ASCII characters — still corrupts every non-ASCII byte in the whole file. So this isn't about the edited span; the tool round-trips the entire file through a lossy UTF-8 decode/encode.

    Minimal reproduction

    1. Create a 3-line file encoded as Windows-1252 containing 7 non-ASCII bytes on line 2:

      $enc = [System.Text.Encoding]::GetEncoding(1252)
      $ae=[char]0xE4;$oe=[char]0xF6;$ue=[char]0xFC;$ss=[char]0xDF;$Ae=[char]0xC4;$Oe=[char]0xD6;$Ue=[char]0xDC
      $txt = "line 1: ASCII marker MARKER_A`r`n// umlauts: $ae $oe $ue $ss $Ae $Oe $Ue`r`nline 3: ASCII end`r`n"
      [System.IO.File]::WriteAllText("repro.txt", $txt, $enc)
      # -> 7 valid cp1252 bytes (E4 F6 FC DF C4 D6 DC), 0x EF BF BD sequences, 80 bytes
    2. Use the Read tool on repro.txt. It displays the 7 umlauts as the replacement character � (U+FFFD) — i.e. the original bytes are already lost at read time.

    3. Use the Edit tool to change only ASCII on line 1: MARKER_A → MARKER_B (line 2 is untouched).

    4. Inspect the bytes:

      $b=[System.IO.File]::ReadAllBytes("repro.txt")
      ($b | Where-Object { $_ -in 0xE4,0xF6,0xFC,0xC4,0xD6,0xDC,0xDF }).Count   # valid cp1252 umlauts
      $c=0; for($i=0;$i -lt $b.Length-2;$i++){ if($b[$i]-eq0xEF -and $b[$i+1]-eq0xBF -and $b[$i+2]-eq0xBD){$c++} }; $c   # EF BF BD sequences

    Result

    valid cp1252 umlaut bytes EF BF BD sequences size
    before edit 7 0 80
    after ASCII-only edit 0 7 94

    Every umlaut (E4, F6, …) became EF BF BD (the 3-byte UTF-8 encoding of U+FFFD), and the file grew by 14 bytes (7 × 2). In a Windows-1252 editor these now show as �.

    Why this is bad

    • The corruption is silent and irreversible: because Read already decodes to U+FFFD, the identity of each character (ä vs ö vs ü vs ß …) is gone — it can't be recovered without re-typing from context.
    • It corrupts the entire file on any edit, so a single unrelated change to a large source file destroys hundreds of characters at once.
    • It breaks compiler-relevant strings and comments in Delphi/Pascal, C/C++, etc. that must stay in Windows-1252/ANSI.

    Note re: editors

    The Zed editor itself recently added legacy-encoding open/save with auto-detection (zed-industries/zed PR #44819, merged 2025-12), and it preserves Windows-1252 on save. But the agent Edit/Write tools still corrupt the file (reproduced above), which suggests they write directly rather than through the encoding-aware path.

    Suggested fix

    Detect the file's encoding on read and preserve it on write; or, at minimum, don't re-encode bytes that weren't part of the edit. A per-file/target-encoding override would also work.

    Environment

    • OS: Windows 10 (Pro 22H2)
    • Agent: Claude Code (CLI) 2.1.193, accessed via Zed 1.10.0
    • File encoding: Windows-1252 (Delphi/Pascal codebase)
  16. amajkowski commented on Aug 4, 2026

    @amajkowski

    Can confirm the same root cause with a different legacy codepage: Windows-1250 (CP1250), Polish Java source files.

    One detail worth adding to the report: the corruption isn't limited to the characters/lines actually touched by the edit. Edit/Write appear to decode the entire file as UTF-8 on read (lossy — any CP1250/CP1252 high byte that isn't valid UTF-8 becomes U+FFFD) and then re-serialize the entire file as UTF-8 on write. So a single-line edit corrupts every accented character anywhere in the file, including lines nowhere near the diff. Repeated edits compound this, since each subsequent Edit call re-reads the now-partially-UTF-8/partially-corrupted file and repeats the process.

    Recovery required manually reconstructing the original text from conversation history/context and re-saving with an external tool (PowerShell [System.Text.Encoding]::GetEncoding(1250)) — Claude Code itself has no way to write anything but UTF-8, so there's no in-tool fix once a non-UTF-8 file has been touched.

    +1 for detecting/preserving the original encoding (or at minimum failing loudly on a lossy decode instead of silently substituting U+FFFD and persisting the loss).

  17. manueljmgomes commented on Aug 21, 2026

    @manueljmgomes

    Problem still persists - same problem when using either claude code cli or claude code inside, as soon as it touches one file it corrupts all the international characters immediately,

    Before:
    Image

    After:
    Image

    All international chars aare replaced with �

  18. myonlylonely commented on Aug 21, 2026

    @myonlylonely

    Please try Kilo Code. It's the only plugin/AI assistant I could find that supports legacy encoding.

  19. odacirx commented on Aug 23, 2026

    @odacirx

    Confirming this bug with a large-scale case, plus a reliable detection method and a recovery strategy that may help others who have already been hit.

    Environment

    Field Value
    Claude Code version
    OS Windows 10
    File types .pas, .dfm, .dpr (Delphi source)
    Original encoding ANSI / Windows-1252, no BOM
    Encoding after edit UTF-8 without BOM, accented characters replaced by U+FFFD

    What happened

    We are migrating a ~1.4M-line ERP from Delphi 5 to Delphi 11. The sources are ANSI/CP1252 with Portuguese text (ç ã é í ó â ê õ à ü) in string literals, comments and form captions.

    Over one working session, Claude Code applied a number of small, targeted edits — adding a unit to a uses clause, qualifying a function call, fixing a missing comma. Each edit touched one or two lines.

    The result:

    Files Corrupted characters
    .pas 348 17,947
    .dfm 81 3,968

    The damage is not proportional to the edit. A one-character fix destroys every accented character in the file, because the whole file is rewritten on save.

    Why this is worse than it first appears

    1. It is completely silent. The corrupted literals are still valid literals. The project compiled with 0 errors. The problem only became visible when the application was launched and the menus read Gest�o, Servi�os, Pa�ses.

    2. It reaches production data. In our case the corrupted text includes error messages, report captions and strings written to the database.

    3. The information is unrecoverable from the file alone. U+FFFD does not record which character it replaced. Without a clean copy, configura��o cannot be resolved to configuração automatically.

    4. Consecutive accented characters can collapse. When two accented bytes are adjacent, the UTF-8 decoder sometimes emits a single U+FFFD for both, so the corrupted text is not even the same length as the original. This matters for any automated repair.

    Detection: the intuitive check does not work

    This tripped us up for a while and is worth documenting.

    You cannot detect the corruption by opening the file and looking for ?/`` characters, and you cannot detect it by reading the file as UTF-8 in a script — a clean ANSI file read as UTF-8 also shows U+FFFD, because its accented bytes are invalid UTF-8. Both the healthy file and the corrupted file look identical that way.

    The only reliable test is to look for the byte sequence EF BF BD (the UTF-8 encoding of U+FFFD) in the raw bytes:

    $b = [System.IO.File]::ReadAllBytes($path)
    for ($i = 0; $i -lt $b.Length - 2; $i++) {
        if ($b[$i] -eq 0xEF -and $b[$i+1] -eq 0xBF -and $b[$i+2] -eq 0xBD) {
            "CORRUPTED: $path"; break
        }
    }

    A clean CP1252 file will never contain that sequence. A corrupted one will.

    Recovery, for anyone already affected

    Recovery is possible without losing the edits, provided you have a clean copy from before the damage, and provided the edits themselves are ASCII-only.

    The corrupted file is the clean file with two independent layers on top:

    • (a) the intended edits, which are ASCII;
    • (b) every non-ASCII byte replaced by U+FFFD.

    If you normalize the clean copy by replacing every non-ASCII character with U+FFFD, the lines you did not edit become byte-identical to the corrupted ones. Those can be restored verbatim from the clean copy. The lines that differ are your real edits and are left untouched.

    Applying this to our codebase recovered 21,613 of 21,777 corrupted lines (99.2%). The remaining 164 were lines that were both edited and contained accents, plus lines with consecutive accented characters (see point 4 above) — all reported in a CSV for manual review rather than guessed at.

    Two details that matter if you implement this:

    • Match on repeated lines as well. .dfm files are full of duplicated captions; if all candidate lines in the clean copy are identical, there is no ambiguity even though the line occurs many times. Requiring a unique occurrence leaves ~25% of recoverable lines behind.
    • Use ordinal, case-sensitive string comparison. PowerShell's Sort-Object -Unique is case-insensitive and will silently merge lines that differ only in case.

    Suggested fix

    Detect the encoding on read and preserve it on write, at minimum for the no-BOM CP1252 case. Failing that, a per-project setting (e.g. "fileEncoding": "windows-1252") or refusing to write a file whose original encoding cannot be preserved would both be far better than silent data loss.

    A warning at edit time would already have prevented this entire incident.

  20. Andrekarma commented on Sep 2, 2026

    @Andrekarma

    Problem persisting with windows-1252 files

  21. dariocigna commented on Sep 8, 2026

    @dariocigna

    Still reproducing on 2.1.263 (latest as of today, 2026-09-08), Windows 11.

    Hit this twice today on the same C# file in two different solutions. The file is Windows-1252, no BOM, ~177 KB, with accented characters in comments and string literals (Italian codebase). A single Edit to
    an unrelated line rewrote the whole file as UTF-8, turning every pre-existing è/à/é into è/à /é. Had to revert and reapply manually both times.

    Three things that make this worse than a plain encoding mismatch:

    1. It's silent. The tool reports success. Nothing in the result indicates the file was re-encoded.
    2. The damage extends beyond the edited region. The entire file is rewritten, so untouched lines elsewhere get corrupted too.
    3. It defeats the obvious verification. Reading the file back after the edit shows correct content, because it's read in the same encoding it was just written in. The corruption only appears when the file is
      read in its actual codepage. I initially declared the file clean after checking it this way — it wasn't.

    The only workaround is to bypass Edit/Write entirely and go through a shell command with an explicit codepage:

    $enc = [System.Text.Encoding]::GetEncoding(1252)
    $lines = [System.IO.File]::ReadAllLines($path, $enc)

    ...modify...

    [System.IO.File]::WriteAllLines($path, $lines, $enc)

    That works, but it's error-prone and depends on remembering to do it on every single subsequent edit to that file.

    Given this has been open for a year: even without full encoding detection, refusing the write with a clear error when the target file's encoding can't be preserved would be a big improvement over silently
    corrupting it. Right now the failure mode is "looks fine, ships broken".

  22. simeonbodurov commented on Sep 12, 2026

    @simeonbodurov

    When that issue will be fixed ? Now we have to use external tools like github.com/dimitar-grigorov/mcp-file-tools, but a real harness should have native support for that.

  23. LasseHolm commented on Sep 25, 2026

    @LasseHolm

    There are lots of work-arounds, -one is to instruct Claude Code in always converting files to UTF8 before making changes, but still in some cases it doesn't do that, and it introduce garbage characters in the text files.
    We have loads of code files still in ansi format.

  24. tsurutanmen commented on Oct 11, 2026

    @tsurutanmen

    Update for anyone following this: 2.1.296 changes Edit from silently corrupting to refusing, per its changelog. Checked on Windows 11 with a Windows-1252 C# file (Italian accents in a comment and a string literal, CRLF, 97 bytes), via claude -p (Sonnet).

    Edit now refuses and writes nothing. The file stays byte-for-byte unchanged, and the tool returns:

    File is not valid UTF-8. It may use a legacy encoding such as Windows-1252, Shift-JIS or GBK, or be binary. This tool saves the whole file as UTF-8, which would replace every byte it cannot decode with U+FFFD. Nothing was written. Make the change with a shell command that reads and writes the file in its own encoding, or ask the user whether to convert the file to UTF-8 first.

    Two gaps remain:

    1. Write is not covered. Write on the same existing Windows-1252 file (content given verbatim, with the real accented letters) returns "updated successfully" with no warning and saves it as UTF-8 with LF: è E8 → C3 A8, CR 5 → 0. A model that reaches for Write instead of Edit still re-encodes the file silently.
    2. Read still shows every accented letter as U+FFFD (5 of 5 here: // Commento: � gi� cos�, perch� funziona). With shell tools available, the model saw that, skipped Edit, and went through Bash; in 2 of 2 runs sed -i stripped every CR on the way (97 → 92 bytes), the model noticed and repaired it byte by byte in PowerShell, and the end result was the correct one-byte change in Windows-1252. It got there, but by a detour.

    So whole-file corruption through Edit looks fixed in 2.1.296. Native support for legacy encodings, which most comments here ask for, is not part of the fix.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:toolsbugSomething isn't workingduplicateThis issue or pull request already existsplatform:windowsIssue specifically occurs on Windows

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions