Skip to content

Scrub long transient values without a regex and find short base64 - #1684

Merged
danielgwilson merged 9 commits into
mainfrom
fix/scrub-remaining-1646
Oct 8, 2026
Merged

danielgwilson merged 9 commits into
mainfrom
fix/scrub-remaining-1646

Conversation

@danielgwilson

@danielgwilson danielgwilson commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Closes #1646. #1674 merged the main hardening; this PR does the three items left after it, then the fixes from the Codex security review of this PR. idiom.md in the 10-08 audit ("fast-check for #1674") calls these items property-shaped; each now has a property test, examples, or both.

What changed

Literal scrub (src/evidence/secret-scrub.ts, src/run/transient-comms-secrets.ts, item 1 and review finding 4)

  • New interface in the known-value scrub module: scrubValuesAsWritten(values)(text), the same shape as scrubSecretValues.
    • It builds the alternation regex scrubTransientCommsText used to build, longest value first, and runs it.
    • V8 compiles a regex at its first use. When the regex alternates and a value has 32,768 characters or more (Node 24.12.0; 32,767 compile), V8 refuses it there. The scrub catches that and switches to replaceLongestFirst, with the same output.
    • replaceLongestFirst uses eachOccurrence (Deepen text readings and test the scrub against an independent model #1674's linear search) to find every occurrence of each value. It keeps the longest value at each start position and resumes after each replacement, so it never searches a marker it wrote.
  • Before this PR, a received link of that length, registered next to the code it carries, failed the scope with TRANSIENT_NARRATION_SECRET_LIMIT. The analysis then failed with analysis_validation_failed_unexpected.
  • The first revision of this PR ran the linear search on every call. The review measured it at the registration limits on 64 KiB of text: 729 ms and 2,065 ms, against 3.1 and 1.7 ms for the regex. The transient scrub runs the literal pass twice per field. The regex is back on the common path.
  • The transient scope caches the scrub for its current values (literal), as it cached the regex. The try/catch that turned a regex failure into the limit error is deleted.
  • The registration limits stay: 8192 values, 1 MiB in total, 65,536 bytes per value. They bound memory and the known-value scrub's cost per text. The regex limit was not their reason:
    • the per-value limit is twice what the regex accepts;
    • 8192 values of 128 bytes (1 MiB) compile.

Short base64 (src/evidence/secret-scrub.ts, item 2): implemented.

  • MIN_ENCODED_FORM is 6 (was 8).
  • encodedForms lists unpadded standard base64 itself. base64url supplied it only for a value whose base64 has no + or /.
  • Found now: abcd written YWJjZA, ???? written Pz8/Pw, and a 5-byte value at every byte offset.
  • Still not found: a 4-byte value inside a longer base64 run. Only 4 or 5 characters of its encoding are independent of the bytes around it. The contract states this and a test records it.
  • The cost, from review finding 3: an identifier that holds a short value's base64 loses that part. aMTIzNAz becomes a[REDACTED_SECRET]z while the code 1234 (MTIzNA) is registered. docs/contracts/study-analysis.md states this next to the minimum, and an example test records it. The minimum stays at 6 on the corpus evidence below.
  • RunSecrets (src/run/secrets.ts) reads encodedForms, so run evidence gets the same forms.

Evidence, read in place from the 0.114.0 and 0.115.0 smoke runs:

  • Run text: 9 runs, 128 .json/.ndjson/.md files, 2.48 M characters.
  • Separately, 9 observer/index.html files: 11.40 M characters of the Observer bundle, whose embedded base64 no scrub reads.
Values Count 6- and 7-character forms Values hit in run text Values hit in index.html
4-digit codes 10,000 10,000 0 0
5-digit codes 100,000 399,999 0 29
6-digit codes 1,000,000 2,000,000 0 28
words of 4 to 6 letters from the run text, as written, lower case and capitalized 1,718 4,026 0 0
random alphanumeric tokens of 4 to 6 characters 100,000 231,446 0 0
  • Run text has 14,992 base64-alphabet runs of exactly 6 characters and 21,821 of exactly 7. None mixes upper case, lower case and digits.
  • None of the 1,452 distinct ones is the unpadded base64 of printable ASCII, so no registered text value of 4 or 5 bytes would redact one.
  • In index.html, 36 of the 10,000 4-digit codes have a 4-character form present and 11 have a 5-character form. Their 6-character forms have no hits. That is why the floor is 6.

RunSecrets final check (src/run/secrets.ts, review finding 1)

  • With the 6-character forms, values ["abcd","SECRET]x"] on SECRET]x YWJjZAx came out as [REDACTED_SECRET] [REDACTED_SECRET]x. The marker written for YWJjZA spells SECRET]x again.
  • RunSecrets.scrub now ends with the holdsSecretValue check the shared and transient scrubs have. If the result still holds a held value in any readingsOf, the whole text becomes the marker.
  • The check is rebuilt after add holds a new value, because the scrub reads the values on each call.
  • Routing RunSecrets through scrubSecretValues instead would remove no copy. That scrub returns decoded text, and RunSecrets keeps every other character as written.

Terminal escape sequences in one pass (src/run/escape-sequences.ts, review finding 5, which predates this PR)

  • RunSecrets found the stretches a value's view drops or decodes with one regex. Its operating-system command branch, \x1b\][^\x07]*(?:\x07|\x1b\\), searched to the end of the text and back from every start with no bell character after it.
  • New module escapeSequences(text) returns the same stretches in one pass.
    • The six shapes start with different characters, so at most one can match at a position.
    • A command ends at the first bell character after it, and each stretch of text is searched for one only once.
    • With no bell character after it, a command ends at the last string terminator in the text. That is what the regex's greedy match gives, including the text between two commands.
  • viewSpans reads it. The regex is deleted from src and kept in the test model as the reference.
  • The same regex shape is in src/observer/data.ts:356, outside this PR's files.

Values that are part of a marker or hold one (item 3): no change.

  • Refusing such a value at registration leaves it unscrubbed.
    • Registered values include email subjects. CODE, LOBBY, TEXT and PATH are each part of a marker humanish writes.
    • A refused subject would appear as written in every analysis field.
    • Failing the scope instead would turn an ordinary subject into a failed analysis.
  • A warning changes nothing: the scope scrubs the value the same way whether or not it warns.
  • No leak remains for a refusal to close within the one-extra-decoding check.
  • Outputs of transientCommsKnownValueScrub on this branch:
    • ["CODE"], The subject CODE arrived; the lobby showed [REDACTED_LOBBY_CODE]. gives The subject [REDACTED_SECRET] arrived; the lobby showed [REDACTED_LOBBY_[REDACTED_SECRET]].
    • ["[REDACTED_SECRET] for you"], Subject: [REDACTED_SECRET] for you gives Subject: [REDACTED_SECRET].
    • ["Re: [REDACTED_SECRET]"], Opened Re: [REDACTED_SECRET] and Re%3A%20%5BREDACTED_SECRET%5D gives Opened [REDACTED_SECRET] and [REDACTED_SECRET].
  • The properties draw values from marker text (markerDrawn) and passed at 10,000 cases on two seeds.

Tests

  • tests/helpers/scrub-model.ts:
    • modelLiteralScrub(values)(text), the regex alternation on the platform's RegExp;
    • modelEscapeSequences, the regex RunSecrets used;
    • MIN_BINARY_FORM 6.
  • tests/helpers/scrub-arbitraries.ts:
    • literalInputs(): values over few characters, regex syntax, a lone surrogate half, marker pieces, and prefixes and suffixes of one string.
    • escapeTexts(): raw and JSON-escaped sequence pieces, whole and cut, and percent pieces. Command starts and ends are drawn more often.
  • Properties:
    • scrubTransientCommsText equals modelLiteralScrub, through V8's regex and through the fallback. The fallback is forced by registering a 32,768-character value that no generated text holds.
    • escapeSequences equals modelEscapeSequences.
  • Timing:
    • The literal scrub at the registration limits on 64 KiB is timed in turn with a regex built in the test, fastest of five, ratio under 10 with a 5 ms floor. The two registries come from the review.
    • RunSecrets.scrub on three escape shapes at 8 KiB and 64 KiB: ratio under 16 with a 20 ms floor.
  • Examples:
    • long links next to a code;
    • YWJjZA, Pz8/Pw, NzQzOTI, and a 5-byte value at offsets 0 to 2;
    • a 4-byte value inside a longer run;
    • the identifier aMTIzNAz;
    • RunSecrets with ["abcd","SECRET]x"];
    • commands ending at a bell character or the last terminator, control sequences and percent runs, with literal expected spans.

Checked

  • Red first:

    • 32,768- and 65,536-character links threw TRANSIENT_NARRATION_SECRET_LIMIT.
    • Short base64 examples survived.
    • The timing test at the limits failed at ratios 185 and 530 on the first revision's linear search.
    • RunSecrets returned [REDACTED_SECRET] [REDACTED_SECRET]x.
    • The command-scaling test failed at ratio 61.
  • Properties at 10,000 cases on seeds 271828 and 99991: 0 failures. This covers all scrub properties, both literal paths, readingsOf and escapeSequences. escapeSequences also passed 50,000 cases on seed 5.

  • The equivalence of the linear search with the regex was first shown against the regex still in production. The review then ran 329,500 differential comparisons with no difference.

  • Mutations:

    • Keeping the first value found instead of the longest fails the fallback-path property: {"values":["aaaaaa","aaaa","aaaaa"],"text":"aaaaa"}.
    • Merging overlaps fails it: {"values":["|.|."],"text":"|.|.|.|."}.
    • A minimum of 8, or no unpadded standard form, fails the scrub properties.
    • Ending a command at the first string terminator fails the escapeSequences property ("\u001b]\u001b\\\u001b\\") once command pieces are drawn more often. It passed 400 cases before that.
    • Dropping > from the two-byte escapes fails it too.
  • Timings, Node 24.12.0, fastest of several, on this machine:

    Case Before After
    Literal pass, 8192 values that nearly match, 64 KiB 644 ms (first revision's linear search) 1.6 ms
    Literal pass, 8192 values that overlap everywhere, 64 KiB 1,019 ms (same) 0.5 ms
    Escape sequences, 16 KiB of unterminated commands 114 ms (old regex) 0.28 ms
    Escape sequences, 64 KiB of unterminated commands 1,846 ms (old regex) 0.54 ms
    RunSecrets.scrub, 64 KiB of unterminated commands 3.4 s (review, main) 0.92 ms
    • The regex at the limits costs about 72 ms to compile at its first use for a given set of values.
    • The new final check costs 1.5 ms of RunSecrets.scrub's 4.3 ms on 63 KiB of colored terminal text with 4 values.
  • pnpm release:check on 601b1f5 (merged with origin/main 3adf086): exit 0. 7687 root tests passed (12 skipped) and 158 TUI tests. Lint at 439, its cap. Public API proof and public-surface scan passed.

  • run, routes, actors, study, comms, evidence, analysis and verify tests passed after the RunSecrets change: 5165.

Out of scope

  • Review finding 2 (P3): two further decodings at a marker edge, such as ["abcd","T]]x"] with T]]x [REDACTED_YWJjZA]%252578. The output shows T]]x only after two more percent decodings, beyond the documented one-extra-decoding check. The marker manufactures the occurrence. The same class as Codex round 2 finding 3 on Deepen text readings and test the scrub against an independent model #1674.

Not verified

  • No live run, and no received email with a long link. None of the 9 corpus runs received email, so corpus words and random tokens stood in for subjects and link tokens.
  • The regex limit and the timings were measured on Node 24.12.0 only. CI runs the tests on Node 22 (node-floor).
  • The fallback search's cost grows with the number of values times the text length. Only values of 32,768 characters or more reach it.
  • RunSecrets.spans, which the terminal route uses to replace values across chunks, has no final check. The transcript built from the chunks goes through scrub afterwards.
  • The final check can now replace whole texts of run evidence that hold a value in a reading RunSecrets does not decode, such as HTML references. No count was taken of how often real evidence does that.
  • Nested markers when a value is part of a marker are unchanged: value SECRET makes [REDACTED_SECRET] read [REDACTED_[REDACTED_[REDACTED_SECRET]]].

🤖 Generated with Claude Code

danielgwilson and others added 2 commits October 8, 2026 16:58
scrubTransientCommsText built one global regex alternation of the
scope's values. V8 refuses a literal of 32,768 characters once the
regex alternates, so a received link that long, registered with the code
it carries, made the scope fail with TRANSIENT_NARRATION_SECRET_LIMIT and
the analysis fail.

scrubValuesAsWritten in secret-scrub.ts is the new interface: it finds
every occurrence of each value with eachOccurrence (from #1674), keeps
the longest value per start position, and sweeps once from the left,
resuming after each replaced value. That is the regex's output without
its size limit. The transient scope loses its pattern cache and the
try/catch that turned a regex failure into the limit error.

The registration limits stay. They bound memory and the known-value
scrub's cost per text; the regex limit was not their reason: the per-value
limit is 65,536 bytes, twice what the regex accepts, and 8192 values of
128 bytes compile.

Checked: the old regex is kept as modelLiteralScrub in the test model;
the property that scrubTransientCommsText equals it passed 10,000 cases
on seeds 271828 and 99991. Mutations (first value found instead of the
longest; merged overlaps) fail it. A 32,768- and a 65,536-character
link next to a code are scrubbed (red first: TRANSIENT_NARRATION_SECRET_LIMIT).
Not checked: a live run.

Part of #1646.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A value of 4 or 5 bytes, such as a 4- or 5-digit code, has an unpadded
base64 form under the 8-character minimum, so `abcd` written `YWJjZA`
survived both known-value scrubs. MIN_ENCODED_FORM is now 6, and the
unpadded standard base64 form is listed explicitly (base64url covered it
only for values without `+` or `/`). A 5-byte value is now found at every
byte offset; a 4-byte value inside a longer base64 run keeps only 4 or 5
stable characters and stays unfound. The contract says so and a test
records it.

Measured on the text files of the 0.114.0 and 0.115.0 smoke runs (9
runs, 2.48 M characters, read in place): no 6- or 7-character form of
any 4-, 5- or 6-digit code, 1,718 word values or 100,000 random tokens
occurs. The only hits are inside observer/index.html's embedded base64,
which no scrub reads. None of the 1,452 distinct 6- and 7-character
tokens in that text is the unpadded base64 of printable ASCII.

RunSecrets uses encodedForms too, so run evidence gets the same forms.

Checked: red first on `YWJjZA`, `Pz8/Pw` and `NzQzOTI`; the reference
model's minimum moved to 6 and all properties passed 10,000 cases on
seeds 271828 and 99991. Restoring 8, or dropping the unpadded standard
form, fails the properties. evidence, run, comms, analysis, verify and
routes tests: 2989 passed.

Part of #1646.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
humanish Ready Ready Preview Oct 8, 2026 6:36pm UTC

Request Review

danielgwilson and others added 7 commits October 8, 2026 17:37
A Codex review of #1684 measured the linear search at the registration
limits (8192 values, 1 MiB) on 64 KiB of text: 729 ms for values that
nearly match and 2,065 ms for values that overlap at every position,
where the regex took 3.1 and 1.7 ms. The transient scrub runs the
literal pass twice per field.

scrubValuesAsWritten now builds the alternation regex as before and runs
it. V8 compiles a regex at its first use and refuses it there once the
regex alternates and a value has 32,768 characters; the scrub catches
that and finds each value itself (replaceLongestFirst), with the same
output. The transient scope caches the scrub per set of values again,
as it cached the regex.

Checked: a timing test at the limits times the scrub and a regex built
in the test in turn, fastest of five, and bounds the ratio at 10 with a
5 ms floor; red first at ratios 185 and 530. The literal property runs
through both paths, the second forced by registering a 32,768-character
value no text holds; 10,000 cases on seeds 271828 and 99991 each.
Keeping the first value found instead of the longest fails the fallback
path's property.

Part of #1646.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
RunSecrets.scrub had no final check. With the 6-character base64 forms,
values ["abcd", "SECRET]x"] on "SECRET]x YWJjZAx" came out as
"[REDACTED_SECRET] [REDACTED_SECRET]x": the marker written for `YWJjZA`
spells `SECRET]x` again. The shared and transient scrubs catch this
with holdsSecretValue; RunSecrets now does the same and returns its
marker in place of the whole text. The check is rebuilt after `add`
holds a new value, since the scrub reads the values on each call.

Routing RunSecrets through scrubSecretValues would remove no copy: that
scrub returns decoded text, and RunSecrets keeps every other character
as written.

Checked: example test with the review's input, red first; a value added
later is checked too. run, routes, actors, study, comms, evidence,
analysis and verify tests: 5165 passed.

Part of #1646.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
RunSecrets found the stretches a value's view drops or decodes with one
regex. Its operating-system command branch, `\x1b\][^\x07]*(?:\x07|\x1b\\)`,
searched to the end of the text and back from every start that had no
BEL after it, so "\x1b]".repeat(n / 2) took 361 ms at 16 KiB and 3.4 s
at 64 KiB (Codex review of #1684).

escapeSequences (src/run/escape-sequences.ts) returns the same matches
in time linear in the text: the six shapes start with different
characters, a command ends at the first BEL after it, found once per
stretch, or else at the last string terminator in the text, which is
what the regex's greedy match gives. viewSpans reads it; the regex is
deleted from src and kept in the test model.

Checked: escapeSequences equals the old regex on generated text of
escape pieces, 10,000 cases on seeds 271828 and 99991 and 50,000 on seed
5; ending a command at the first terminator fails it once the generator
draws command pieces more often. Scaling tests through RunSecrets.scrub
at 8 KiB and 64 KiB: the command shape failed at ratio 61 before, and two
other shapes guard the rest. run, terminal route and evidence tests pass.

Part of #1646.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A Codex review of #1684 built identifiers that hold a short value's
base64 middle, such as `aMTIzNAz` with the code `1234` registered; the
scrub redacts that part. The contract now states this cost next to the
6-character minimum, with the corpus count that set it, and an example
test records the case. CHANGELOG lines cover the three follow-up fixes.

Part of #1646.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
prose:check counts BEL as all-caps emphasis and "used to" as history.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@danielgwilson
danielgwilson merged commit 59b5fc2 into main Oct 8, 2026
10 checks passed
@danielgwilson
danielgwilson deleted the fix/scrub-remaining-1646 branch October 8, 2026 18:45
This was referenced Oct 9, 2026

This branch was successfully deployed

1 active deployment
Preview — 996a7bde Deployed Oct 8, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Harden the shared secret scrubber against encoded and marker-shaped values

1 participant