Repository navigation
Redact study warnings before the run records them (#1747) - #1748
Merged
Merged
Conversation
The warning that a participant instruction reads like a script quotes up to three of its lines. runPublisher wrote study warnings into events.ndjson, run.json and observer-data.json without redactText, so a line naming a sandbox URL other than the subject's (an inbox page, a second app) reached the bundle, verify graded it blocked, and a hosted run failed with "Run bundle failed verification". Study warnings now go through redactText, as run failures already do. Checked: tests/run/study-warning-redaction.test.ts (a dry run whose mission names an e2b.app inbox URL records the warning redacted and verifies share_ready) failed before the change and passes after; full vitest suite 8034 passed. Not checked: a live hosted run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Merged
danielgwilson
added a commit
that referenced
this pull request
Oct 9, 2026
#1748 (fixes #1747) merged to main after the 0.119.0 cut: study warnings recorded in the run's bundle go through redactText, so a hosted study whose participant instruction names a sandbox URL other than the subject's verifies share_ready again. It ships in 0.119.0, and the ruleset needs the branch up to date with main. CHANGELOG.md conflict: Unreleased stays empty. #1748's Fixed line moves into the 0.119.0 release notes (Fixes), where the rest of the cut Unreleased body went; the CHANGELOG entry keeps only the title, the opening paragraph and the release link, as docs/release/publish.md step 2 says. The title and the opening paragraph gain the warning redaction. Checked: no conflict markers; the merged tree differs from b382916 only in package.json, cli.mdx, CHANGELOG.md and the benchmark files. release:check runs on the branch head next. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
danielgwilson
added a commit
that referenced
this pull request
Oct 9, 2026
The benchmark ran on the release commit 53dbf23. The branch then merged #1748, which redacts study warnings before the run records them. That changes what the bundle records and nothing a participant receives or does, so the result stands for 0.119.0; the README paragraph says so. Checked: node scripts/check-code-prose.mjs (no count rose). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
danielgwilson
added a commit
that referenced
this pull request
Oct 9, 2026
* Release 0.119.0 Bump the version to 0.119.0, move the Unreleased notes into the release notes, and keep the CHANGELOG entry's title, opening paragraph and release link. pnpm docs:generate writes 0.119.0 into cli.mdx. Contents: #1745 and #1742, cut from main 6405a7c. Checked: a live smoke on a tarball packed from main 6405a7c runs alongside this commit (try-live with analysis, 24 participants arriving 10 s apart with their cohort analysis, a late participant past the study budget, analyze --run on a 24-participant bundle). pnpm release:check runs on this commit next. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add the 0.119.0 Taskly benchmark result pnpm bench with its defaults (3 runs per arm, openai-computer-use, neutral mission) on the release commit 53dbf23, as docs/release/publish.md asks for a minor release: report recall 10/15, analysis recall 10/15, 0 invented on either arm, $6.12 estimated. 0.118.0 scored 6/15 and 6/15 at $5.38. D1 and D5 account for the rise: two planted reports and two planted analyses name each, where none did on 0.118.0. D3 and D4 stayed 3/3. Two unresolved planted lines describe D5 (planted 1's report, planted 3's analysis), so recall read by hand is 11/15 and 11/15. No planted participant typed more than 27 characters in one action, so none met D2. The analysis cap was the dry run's admittedCostUsd, $1.25, against bills of $0.60 to $0.86; none was refused or skipped. Every participant ran threaded, ended on its own and gave 5 or 6 impressions. E2B listed 0 tagged sandboxes after it. Issue #1749 reads the decline from 0.114.0 to 0.118.0 as noise at three runs per arm (same-hour A/B of npm 0.114.0 and main d183d1a: 52/80 each). The README paragraph states that finding without the issue number, since the evidence prose check caps issue references at 102. Checked: the summary's claims to check by hand, against the results file and the two unresolved findings' text in the analyses; typed lengths from the traces' action titles; node scripts/check-code-prose.mjs. Not checked: which controls planted participants clicked, read from trace coordinates as #1749 does; three runs per arm cannot separate the rise from run-to-run variation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Note that the 0.119.0 benchmark ran before the warning fix The benchmark ran on the release commit 53dbf23. The branch then merged #1748, which redacts study warnings before the run records them. That changes what the bundle records and nothing a participant receives or does, so the result stands for 0.119.0; the README paragraph says so. Checked: node scripts/check-code-prose.mjs (no count rose). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #1747. The warning that a participant instruction reads like a script quotes up to three of its lines.
runPublisher(src/run/run.ts) wrote study warnings intoevents.ndjson,run.jsonandobserver-data.jsonwithoutredactText, so a line naming a sandbox URL other than the subject's (a run inbox page, a second app) reached the bundle.verifythen graded the bundleblocked, and a hosted run failed with "Run bundle failed verification". Study warnings now go throughredactText, as run failures already do. The terminal still prints the warning as written.Checked
tests/run/study-warning-redaction.test.ts: a dry run of a computer-use study whose mission names ane2b.appinbox URL records the warning with a redaction marker,events.ndjsonholds noe2b.app, andverifygrades the runshare_ready. It failed before the change (the warning held the raw URL) and passes after.Not checked
🤖 Generated with Claude Code