Skip to content

Release 0.119.0 - #1751

Merged
danielgwilson merged 4 commits into
mainfrom
release/0.119.0
Oct 9, 2026
Merged

danielgwilson merged 4 commits into
mainfrom
release/0.119.0

Conversation

@danielgwilson

@danielgwilson danielgwilson commented Oct 9, 2026 •

Copy link
Copy Markdown
Owner

Cuts 0.119.0 from main, a minor release with #1745, #1742 and #1748 merged since 0.118.0. Commit 53dbf23 sets the version on main 6405a7c, moves the Unreleased body out of CHANGELOG.md and regenerates cli.mdx. Commit cd8bf63 adds the Taskly benchmark result that docs/release/publish.md asks for on a minor release, run on 53dbf23. Merge d390396 brings in main b382916 (#1748, fixes #1747); it moves #1748's Fixed line into these notes and adds it to the CHANGELOG title and opening paragraph. Commit 558a295 notes in the benchmark README that the run predates #1748, which changes only how study warnings are recorded. Nothing else merged to main after the cut. The draft release notes follow; the GitHub release body is the same text from the opening paragraph down.

humanish 0.119.0 lets a computer-use or shared-world study start its participants over time, and
analyses a run of 17 to 100 participants in one report. participants[].startAfterMs starts a
participant up to 24 hours after the run starts its participants, and a count group's
startEveryMs spreads its members. Each participant's desktop is created at its start, the plan and
study check print the first and last start and the most participants at once, and run.json records
each participant's arrival. A participant whose start comes after the study crossed
caps.maxTotalUsd is skipped before its desktop is created. The analysis of more than 16
participants splits them into cohorts of at most 16, analyses each cohort with its own request, and
merges the cohort reports into the run's one report with one more request; 0.118.0 skipped the
automatic analysis of such a run. Admission, the plan's analysis line, study check and the
analyze --dry-run worst case count every request, and analyze --json gives their number as
admission.requests. At the default $3 cap the analysis of 17 or more participants is refused even
with no evidence, so it needs a higher review.analysis.maxCostUsd. A provisioned shared world gets
an app sandbox that lives until the last start or wave can end. A study warning recorded in the
run's bundle is redacted as a run failure is, so a hosted study whose participant instruction names
a sandbox URL other than the subject's verifies share_ready again, as on 0.111.2.

This is a minor release because it adds two study fields and a JSON field, changes how a run of
more than 16 participants is analysed, and changes outputs for the same input. A study that sets
startAfterMs or startEveryMs runs on 0.119.0, and npm 0.118.0 refuses it when read. A run of 17
to 100 participants, whose automatic analysis 0.118.0 skipped with
AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT, is analysed in cohorts when its cap admits it; at the
default $3 cap its default analysis is skipped with AUTOMATIC_ANALYSIS_ADMISSION_REFUSED.
AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT is no longer recorded, and study check --json no longer
carries analysis.skip; 0.118.0 added both. analyze --run on such a run sends three or more
requests, and its expected cost is about three times that of 0.118.0's partial analysis ($6.20
against $2.20 on the smoke's 24-participant copy). analyze --json adds
admission.requests. A participant whose start comes after the study budget is spent is recorded
blocked with no desktop, where it made one model request in 0.118.0. A provisioned shared world
whose waves outlast the sandbox ceiling is refused at plan time. A hosted study whose participant
instruction names a sandbox URL other than the subject's verifies share_ready again, where
0.113.0 to 0.118.0 graded it blocked and failed the run. The package exports the same 13
values and 31 types as 0.118.0; the exported RunBundle's participant records gain an optional
arrival. It removes or renames no command, flag, study field or bundle field. These outputs
change for the same input, which a script could notice:

# 0.118.0 0.119.0 What to do
1 A study that sets participants[].startAfterMs or startEveryMs was refused when read: "Unknown study field in participants[0]: startEveryMs. Known fields: ...", HUMANISH_STUDY_INVALID, exit 2. It plans and runs. study check adds a row, "- ok schedule: the first participant starts at +0s and the last at +3m 50s; at most 18 run at once when every session uses its 3m budget." (a checks[] entry named schedule in --json); the plan prints the same after schedule:; run.json records each participant's arrival (startAfterMs, and on a live run scheduledAt and startedAt). A start past 24 hours or not a whole number, startEveryMs without count, a start on a study of one participant, a group whose last member passes 24 hours and a later start for the external-public host are refused with HUMANISH_STUDY_INVALID, each in its own sentence. A study that declares no start prints and records the same as before. (#1745) A study that declares starts needs 0.119.0. npm 0.118.0 reads such a run's bundle with the same output, but not its study file, so its review of a run whose study sets review.analysis: false says "No analysis has run for this run."
2 A participant whose start came after the study crossed caps.maxTotalUsd (a later wave) had its desktop created, made one model request and stopped. It is skipped before its desktop is created: blocked in run.json and the result, with the summary "Late was skipped: study budget reached before this participant started: the run's estimated model spend $0.0021 crossed caps.maxTotalUsd=$0.001; no desktop was created.", and counted in the fan-out summary's skipped participants. In the smoke, two participants, the second due at +1 min: exit 2, "Fan-out run failed: 0/2 participants passed (1 skipped, 0 harness errors, 0 without engagement).", one sandbox. (#1745) Count a blocked participant whose summary says it was skipped as not run.
3 A live run of more than 16 participants recorded its automatic analysis as skipped with AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT ("analysis: not run, because the study has more than 16 participants ..."): exit 0 without a declared review.analysis, 2 with one. run, study check and the TUI printed "After live runs: automatic analysis will not run. This study has 24 participants, and automatic analysis reads at most 16. ..." One analysis covers every participant, in cohorts of at most 16 and one merge request. At the default $3 cap the line reads "refused before it starts for 17 participants in 2 cohort requests of at most 16 participants and one merge request even with no evidence (expected $3.15), since both its worst case and its expected cost plus a 10% margin are over $3", and a default analysis is skipped with AUTOMATIC_ANALYSIS_ADMISSION_REFUSED; a declared one is refused and the run exits 2. At a cap that admits it the line gives a range: "expected $5.14 to $9.25 for 24 participants in 2 cohort requests of at most 16 participants and one merge request ..." at $10. The smoke's 24-participant run at $10 was admitted at an expected $6.20 and billed $2.65. (#1742) Set review.analysis.maxCostUsd from the plan's range, or review.analysis: false. Match AUTOMATIC_ANALYSIS_ADMISSION_REFUSED; AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT is no longer recorded.
4 study check --json of more than 16 participants carried analysis.skip: "AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT". doctor --study printed "- note post-run analysis: Will not run: this study has 17 participants, more than automatic analysis reads. The participants still run." analysis.refusedFromUsd (3.150188 for 17 at $3) or analysis.expectedCostUsd, as for any study, and no skip. doctor prints the post-run analysis row a study of 16 gets. (#1742) Read refusedFromUsd or expectedCostUsd.
5 analyze --run on a run of more than 16 participants sent one request with the first 16 participants' evidence and saved a partial analysis listing the rest as omitted. On the smoke's 24-participant copy: "Worst case: $2.44, if the analyst writes its whole 32768-token output allowance, reasoning included.", expected $2.20, billed $1.36, 16 participant reviews. It sends one request per cohort and a merge request, and saves a complete analysis. On the same copy: "Expected cost: $6.20. Worst case: $7.07, if every request (2 cohort requests of at most 16 participants and one merge request) writes its whole 32768-token output allowance, reasoning included.", billed $2.79, 24 participant reviews. It prints an Analysis requesting line for each request; the merge's reads "0 evidence items, 0 captures." (#1742) Pass the dry run's worst case as --max-cost.
6 analyze --json's admission had no requests. admission.requests: 1 up to 16 participants, and the cohorts plus one past 16 (3 for 17 to 32). The other keys are unchanged. (#1742) Nothing; read it for the number of requests.
7 A stored analysis held at most 2,000 evidence items, 64 captures and 4 MiB. At most 6,400, 320 and 16 MiB, eight cohorts' packets. npm 0.118.0 reads the smoke's cohort analyses (351 evidence items, 24 captures) with the same output; one past its bounds was not tried. (#1742) Open a cohort analysis of more than 64 captures with 0.119.0.
8 A provisioned shared world that ran in waves planned with an app sandbox deadline of one session plus provisioning, which a later wave could outlast: 40 members at the default limit planned as "participants: 40, up to 19 at once". The app sandbox lives until the last start or wave can end, and a study that puts that past HUMANISH_E2B_MAX_SANDBOX_MINUTES is refused at plan time, dry runs too: "The participants use the app for up to 30m when every session uses its 10m budget, because 21 participants wait for a free slot. The subject sandbox, which serves the app until every participant ends, then needs a 70m deadline, ... set HUMANISH_E2B_MAX_SANDBOX_MINUTES to 70 or more.", HUMANISH_SHARED_WORLD_INVALID. (#1745) Set the setting to your plan's limit (Pro allows 1440), run more members at once, or lower execution.timeoutMs.
9 A study whose mission or participant instruction names a sandbox URL other than the subject's, in a line the script-like warning quotes, recorded the warning with the URL in events.ndjson, run.json and observer-data.json. verify graded the run blocked (PUBLIC_SAFETY_FINDINGS) and the run failed, dry runs too: "humanish run failed: Run bundle failed verification.", HUMANISH_COMPUTER_USE_FAILED, exit 2. This started in 0.113.0. The recorded warning reads [REDACTED_SECRET] in place of the URL, as the recorded mission already did, and the run verifies share_ready and exits 0. In the smoke, one live participant whose mission names https://3000-...e2b.app/inbox, with blurred screenshots: share_ready, 16 of 16, the URL in no run file; npm 0.118.0: exit 2, blocked. The terminal still prints the warning with the URL. (#1748, fixes #1747) Nothing; a study 0.113.0 to 0.118.0 failed this way runs again.

Added

  • A computer-use or shared-world study can start its participants over time:
    participants[].startAfterMs starts a participant that many milliseconds after the run starts
    its participants (up to 24 hours), and a group's startEveryMs spreads its members, member k at
    startAfterMs + (k - 1) * startEveryMs. Each participant's desktop is created at its start. One
    whose start comes while every slot is taken waits for the next free one. study check and the
    run's plan print the first and last start and the most participants at once, and run.json
    records each participant's arrival with its scheduled and actual start. A shared-world plan
    with a schedule prints its worst case in sandbox-minutes. A provisioned shared world's app
    sandbox serves until the last participant ends, and a study whose last start puts that past
    HUMANISH_E2B_MAX_SANDBOX_MINUTES is refused before anything is created. A rerun
    (--rerun-failed-from) starts its selected participants together (row 1). (Start participants over time from a declared schedule #1745, from A study can't simulate a full day: computer-use studies stop at 16 participants #1737)

Changed

  • A computer-use or shared-world participant whose start comes after the study crossed
    caps.maxTotalUsd, in a later wave or later in its schedule, is skipped before its desktop is
    created. It is recorded as blocked with a reason naming the budget and counted among the fan-out
    summary's skipped participants. It used to create its desktop, make one model request and stop
    (row 2). (Start participants over time from a declared schedule #1745, from A study can't simulate a full day: computer-use studies stop at 16 participants #1737)
  • One analysis of a run with more than 16 participants covers all of them, up to 128. They are
    split into cohorts of at most 16, each analysed by its own request under the full packet limits,
    and one more request merges the cohort reports into the run's one report. In 0.118.0 a live run
    of more than 16 participants recorded its automatic analysis as skipped with
    AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT, and analyze read the first 16 participants and listed
    the rest as omitted; that reason is no longer recorded. Admission, the expected cost, the worst
    case and the study check range cover every request. The plan's analysis line and the analyze --dry-run worst case name the cohort and merge requests, and admission.requests in --json
    counts them. The default $3 cap refuses the analysis of 17 participants even for a run that
    keeps no evidence, so a default analysis of more than 16 participants is skipped with
    AUTOMATIC_ANALYSIS_ADMISSION_REFUSED unless review.analysis.maxCostUsd admits it. analyze
    prints an Analysis requesting line for each request. A cohort or merge request that fails
    fails the attempt with no findings and a warning, and its usage counts every request sent. The
    humanish.study-analysis.v1 schema is unchanged; its stored limits grew to 6,400 evidence
    items, 320 captures and 16 MiB (rows 3 to 7). (Analyse a run of more than 16 participants in cohorts merged into one report #1742, from A study can't simulate a full day: computer-use studies stop at 16 participants #1737)

Fixes

Known gaps

Package

humanish 0.119.0 packs to 2.5 MB, 7.9 MB unpacked, 972 files (npm pack on release/0.119.0 at
558a295, shasum 9c884092b9d8dca9d1396e2699772752f617a691; the same file list as at 53dbf23).
Against 0.118.0 it adds dist/analysis/cohorts.js, dist/analysis/responses.js and
dist/study/arrivals.js, each with its .d.ts, and drops dist/study/automatic-analysis-plan.js
and .d.ts. The package exports only ., with the same 13 values and 31 types as 0.118.0 (pnpm api:proof); the exported RunBundle's participant records gain an optional arrival. Its
dependencies, optional peer and engines are the same as 0.118.0's: Node 22.19 or newer, and
@e2b/desktop 2.3.2 or newer as an optional peer.

Checked before release

A live smoke on a tarball packed from main 6405a7c, and its #1748 row on the release tarball from
558a295, in fresh projects under /tmp with a fresh HOME and npm cache, DO_NOT_TRACK=1, and
@e2b/desktop 2.4.0 beside humanish. Keys came from the user key store through key discovery, with
XDG_CONFIG_HOME pointing at it; no --dotenv was passed. The store's OpenAI organization is
non-ZDR, so every OpenAI participant ran threaded; its E2B key is on a Pro plan. Controls ran the
same study files or commands on npm 0.118.0, read copies of the main runs, or analysed a second copy
of the 0.118.0 smoke's 24-participant run. Priced spend: $8.54 for the smoke ($7.57 of it in four
analyses, $1.36 of those on npm 0.118.0; $0.02 in the npm 0.118.0 run of the #1748 row), then $0.22
for the dogfood run and $6.12 for the benchmark. The dogfood run, the benchmark and the consumer
check ran on the release commit 53dbf23, before main b382916 (#1748) was merged into the branch;
#1748 changes how study warnings are recorded in the bundle and nothing a participant receives or
does.

  • init --yes writes the same files as npm 0.118.0, and doctor (the concurrent-sandboxes row
    unset, at 100 and at lots), doctor --study try-live, keys, study check try-live and
    analyze --help print the same apart from the cwd line.
  • try-live, as init writes it: reached the goal in 11 participant turns, input growing every
    turn from 2,888 to 19,839 tokens, then 20,170 for the impressions request; cached input on each
    turn from the second was the previous turn's input minus 3; threaded on all 12 requests. $0.28
    run. Its automatic analysis, one request, was admitted at an expected $1.04 (admittedCostUsd
    $1.15, worst case $2.03, cap $3) and billed $0.78, with 4 findings and 4 design findings whose 30
    cited evidence ids all resolve. It gave 6 impressions (unclear, liked, unclear, unfinished,
    liked, liked).
  • 24 participants arriving 10 s apart (count: 24, startEveryMs: 10000) on a one-paragraph
    static page, HUMANISH_E2B_MAX_CONCURRENT_SANDBOXES=100, 3-minute sessions,
    review.analysis.maxCostUsd: 10 (rows 1 and 3). The plan read "the first participant starts at
    +0s and the last at +3m 50s; at most 18 run at once when every session uses its 3m budget."
    run.json recorded each arrival: participant 1 and participants 6 to 24 started within 55 ms of
    scheduledAt (19 of the 20 within 1 ms); participants 2 to 5, due at +10 s to +40 s, started
    at +44 s, when the first participant cleared the local-tree pipeline gate. Sandbox receipts are
    timed from the first start to 0.4 s after the last; E2B listed at most 7 of the run's sandboxes
    at once. 24 of 24 reached the goal; $0.62 run. The analysis line read "expected $5.14 to $9.25
    for 24 participants in 2 cohort requests of at most 16 participants and one merge request"; the
    automatic analysis was admitted at an expected $6.20 (admittedCostUsd $6.82, worst case $7.07,
    requests 3) and billed $2.65, 5 min 33 s. The report covers 24 of 24 participants with 24
    participant reviews, 2 findings, 2 design findings and 4 concern reviews; its 172 distinct cited
    evidence ids all resolve in the stored packet, and its 24 cited captures exist with their
    recorded sha256.
  • The Redact study warnings before the run records them (#1747) #1748 row (row 9): one participant on the bakery page with blurred screenshots, a mission
    that names https://3000-isbxsmokewarning0119.e2b.app/inbox, and no analysis. On the release
    tarball from 558a295 the run exited 0 and verified share_ready, 16 of 16; its recorded warning
    and mission read [REDACTED_SECRET] in place of the URL, no run file holds it, and the terminal
    printed the warning with the URL. $0.04. On npm 0.118.0 the same study exited 2, "humanish run
    failed: Run bundle failed verification.", HUMANISH_COMPUTER_USE_FAILED, blocked with
    PUBLIC_SAFETY_FINDINGS, and the URL in run.json, events.ndjson and observer-data.json. $0.02.
    Dry runs of the same study, and of one whose participant instruction names such a URL, gave the
    same grades.
  • The budget skip (row 2): two participants, the second due at +60 s, caps.maxTotalUsd 0.001.
    The first crossed the budget on its first request ($0.0021). The second was recorded blocked
    with "no desktop was created", its arrival with scheduledAt and no startedAt; the run had
    one sandbox receipt, E2B listed one sandbox for it, and the run ended at the second's scheduled
    time. Exit 2, "0/2 participants passed (1 skipped, ...)". $0.006.
  • analyze --run on two copies of the 0.118.0 smoke's 24-participant run, with each version's
    dry-run worst case rounded up as --max-cost (row 5): npm 0.118.0 sent one request of 233
    evidence items and 16 captures and saved a partial analysis of 16 participants, 8 omitted
    (expected $2.20, billed $1.36). The candidate sent two cohort requests of 175 and 176 items with
    12 captures each and a merge request, and saved a complete analysis of 24 (expected $6.20, billed
    $2.79, 6 min 17 s); its 195 distinct cited ids all resolve and its 24 cited captures exist.
  • verify, review and runs (text and --json) on copies of the three live runs and the
    candidate-analysed 24-participant copy: byte-identical on main and npm 0.118.0, except review
    of the budget-skip run, whose study file npm 0.118.0 cannot parse (row 1). verify and review
    of the arrivals run with the release tarball from 558a295 print the same as with main's.
  • Dry runs and study check of the arrivals and budget-skip studies plan on main and are refused
    as unknown fields on npm 0.118.0. A provisioned shared world of 40 members and one whose second
    member starts 3 hours in are refused on main (row 8); npm 0.118.0 plans the 40.
  • After every live run, reclaim --check exited 0 with clean, and verify exited 0 with 16 of 16
    checks, local_only. Raw sandbox ids appear only in sandbox-receipts.ndjson. E2B listed 0
    tagged sandboxes after the smoke and after the Redact study warnings before the run records them (#1747) #1748 pair (2 before, from other sessions' runs)
    and 0 after the benchmark. No key value from the store appears in any run or kit file, as written,
    base64 or hex (16,446 files).

On the release branch:

  • pnpm release:check on the release head 558a295, after the merge of main b382916 (Redact study warnings before the run records them (#1747) #1748) and
    the benchmark note: exit 0 in 3.9 min, 513 test files passed and 10 skipped, 8,034 tests passed
    and 12 skipped (Redact study warnings before the run records them (#1747) #1748 adds one); TUI 192 tests; lint 426 warnings at its cap of 426; api:proof "13
    values and 31 types match"; public-surface scan of 3,063 text files and 527 binary assets; npm pack --dry-run 972 files with shasum 9c884092b9d8dca9d1396e2699772752f617a691. On the benchmark
    commit cd8bf63 before the merge: exit 0 in 3.7 min, 3,062 text files, shasum
    b10a60b8c14494d4035e7bf353211a1001b21d93. On the release commit 53dbf23: exit 0 in 3.9 min, 512
    test files passed and 10 skipped, 8,033 tests passed and 12 skipped; TUI 192 tests; lint 426
    warnings at its cap of 426; api:proof "13 values and 31 types match"; public-surface scan of 3,060
    text files and 527 binary assets; npm pack --dry-run 972 files with shasum
    b10a60b8c14494d4035e7bf353211a1001b21d93. pnpm docs:generate there wrote 0.119.0 into cli.mdx;
    the merge leaves cli.mdx as it is. The first run on the benchmark commit, before an amend
    (f0ac0892), exited 1 in 11 s on prose.evidence.issue-refs: 103 (cap 102, over by 1): the new
    benchmark README paragraph named Benchmark recall decline 0.114.0 to 0.118.0 is n=3 reach noise; same-hour A/B ties at 52/80 #1749. The amend states the finding without the issue number.
  • pnpm release:dogfood on 53dbf23, before the merge of Redact study warnings before the run records them (#1747) #1748 (the script run with Node, keys read
    from the user store by Node's --env-file into its environment, and a temporary HOME):
    "release:dogfood ok", verdict blocked (blocked_approval), no-spend satisfied, $0.22. The
    participant ran init, the free four-participant dry run, verify ("all 16 evidence checks",
    share_ready) and review on the installed 0.119.0, and stopped at the keys and the local VM
    devices the gate withholds; the script reads that verdict as the expected stop. verify on a copy
    of the run: share_ready, 16 of 16; reclaim --check clean.
  • The post-publish check on the release tarball from 558a295 (TMPDIR=/tmp, HUMANISH_SPEC the
    tarball): 123 of 123 passed. Its 6 new rows cover the schedule from startAfterMs and
    startEveryMs in study check, the plan and run.json (Start participants over time from a declared schedule #1745); four start refusals (Start participants over time from a declared schedule #1745); a
    3-hour start planned on computer use and refused on a provisioned shared world (Start participants over time from a declared schedule #1745); the
    17-participant analysis line at a $10 cap and analysis.refusedFromUsd at $3 (Analyse a run of more than 16 participants in cohorts merged into one report #1742);
    admission.requests 1 and 3 (Analyse a run of more than 16 participants in cohorts merged into one report #1742); and a dry run whose participant instruction names a
    non-subject e2b URL, which verifies share_ready with the URL in no run file and the warning on
    the terminal (Redact study warnings before the run records them (#1747) #1748). Row 20bz now expects the 17-participant refusal line in place of 0.118.0's
    skip. The same script on npm 0.118.0 passed 116 and failed the 6 new rows and 20bz. An earlier
    pass of the 122 rows before Redact study warnings before the run records them (#1747) #1748 passed 122 on the 53dbf23 tarball and 116 on npm 0.118.0.
  • The consumer check: the 0.118.0 check's library and CLI consumer, plus 14 CLI lines for this
    release, on the release tarball from 53dbf23, before the merge of Redact study warnings before the run records them (#1747) #1748. On that tarball the
    consumer typechecks with 0 errors, as on npm 0.118.0, and
    needs no migration. The library steps are the same on both versions, and of the 124 CLI lines
    from the 0.118.0 check, 121 are the same; the three that differ are admission.requests in
    analyze --json (row 6), and the 17-participant study check line and doctor --study note
    (rows 3 and 4). The new lines differ as rows 1 and 3 to 8 say.
  • pnpm bench on 53dbf23, before the merge of Redact study warnings before the run records them (#1747) #1748, with its defaults (3 runs per arm,
    openai-computer-use, neutral mission), committed in cd8bf63: report recall 10/15, analysis
    recall 10/15, 0 invented on either arm, $6.12 estimated. 0.118.0 scored 6/15 and 6/15 at $5.38. D1
    and D5 account for the rise: two planted reports and two planted analyses name each, where none
    did on 0.118.0; D3 and D4 stayed 3/3. Two unresolved planted lines describe D5 (planted 1's
    report, planted 3's analysis), so recall read by hand is 11/15 and 11/15. Benchmark recall decline 0.114.0 to 0.118.0 is n=3 reach noise; same-hour A/B ties at 52/80 #1749 found no harness
    regression behind the decline from 0.114.0 to 0.118.0: a same-hour A/B of npm 0.114.0 and main
    d183d1a scored 52/80 planted-report recall each over 16 planted participants (25/40 and 27/40 in
    its first round), and three runs per arm cannot separate this rise from the same noise. The
    analysis cap was the dry run's admittedCostUsd, $1.25, against bills of $0.60 to $0.86; none was
    refused or skipped. No planted participant typed more than 27 characters in one action, so none
    met D2. Every participant ran threaded, ended on its own and gave 5 or 6 impressions.

Not verified

  • macOS and Windows. Every check ran on Linux x64 with Node 24.12.
  • A live run on the published package: the live smoke ran on a main tarball, and the post-publish
    rows are offline dry runs, refusals and local runs.
  • A participant waiting for a free slot on a live run: at 100 sandboxes nobody waited, and the
    smoke's late starts came from the pipeline gate. A live shared world with a declared schedule,
    external-public followers on a schedule, and a rerun of a scheduled study.
  • A start hours after the run starts: the smoke's latest start was +3m 50s.
  • The default-cap refusal of a live run of more than 16 participants
    (AUTOMATIC_ANALYSIS_ADMISSION_REFUSED) and its exit codes; tests cover them.
  • How many requests the automatic analysis sent: the bundle does not record it, and its usage is
    consistent with 3. analyze --run printed its 3.
  • A run of 40 or 100 participants analysed live (3 or more cohorts), a failed cohort or merge
    request live, and a Codex account analysis of more than 16 participants.
  • npm 0.118.0 reading a cohort analysis past its stored bounds (2,000 evidence items, 64
    captures, 4 MiB).
  • A live run on an E2B Hobby key.
  • The dogfood run, the consumer check and the benchmark on the final head 558a295: they ran on
    53dbf23, before the merge of Redact study warnings before the run records them (#1747) #1748. The live smoke apart from the Redact study warnings before the run records them (#1747) #1748 row ran on main
    6405a7c's tarball.
  • Redact study warnings before the run records them (#1747) #1748 on a live run whose participant instruction, as opposed to its mission, names the URL: the
    smoke ran that case as a dry run.

🤖 Generated with Claude Code

danielgwilson and others added 2 commits October 9, 2026 05:46
Bump the version to 0.119.0, move the Unreleased notes into the release
notes, and keep the CHANGELOG entry's title, opening paragraph and
release link. pnpm docs:generate writes 0.119.0 into cli.mdx.

Contents: #1745 and #1742, cut from main 6405a7c.

Checked: a live smoke on a tarball packed from main 6405a7c runs
alongside this commit (try-live with analysis, 24 participants arriving
10 s apart with their cohort analysis, a late participant past the study
budget, analyze --run on a 24-participant bundle). pnpm release:check
runs on this commit next.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pnpm bench with its defaults (3 runs per arm, openai-computer-use,
neutral mission) on the release commit 53dbf23, as
docs/release/publish.md asks for a minor release: report recall 10/15,
analysis recall 10/15, 0 invented on either arm, $6.12 estimated.
0.118.0 scored 6/15 and 6/15 at $5.38.

D1 and D5 account for the rise: two planted reports and two planted
analyses name each, where none did on 0.118.0. D3 and D4 stayed 3/3.
Two unresolved planted lines describe D5 (planted 1's report, planted
3's analysis), so recall read by hand is 11/15 and 11/15. No planted
participant typed more than 27 characters in one action, so none met
D2. The analysis cap was the dry run's admittedCostUsd, $1.25, against
bills of $0.60 to $0.86; none was refused or skipped. Every participant
ran threaded, ended on its own and gave 5 or 6 impressions. E2B listed
0 tagged sandboxes after it.

Issue #1749 reads the decline from 0.114.0 to 0.118.0 as noise at three
runs per arm (same-hour A/B of npm 0.114.0 and main d183d1a: 52/80
each). The README paragraph states that finding without the issue
number, since the evidence prose check caps issue references at 102.

Checked: the summary's claims to check by hand, against the results
file and the two unresolved findings' text in the analyses; typed
lengths from the traces' action titles; node
scripts/check-code-prose.mjs.
Not checked: which controls planted participants clicked, read from
trace coordinates as #1749 does; three runs per arm cannot separate the
rise from run-to-run variation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
humanish Ignored Ignored Preview Oct 9, 2026 6:45am UTC

Request Review

danielgwilson and others added 2 commits October 9, 2026 06:30
#1748 (fixes #1747) merged to main after the 0.119.0 cut: study warnings
recorded in the run's bundle go through redactText, so a hosted study
whose participant instruction names a sandbox URL other than the
subject's verifies share_ready again. It ships in 0.119.0, and the
ruleset needs the branch up to date with main.

CHANGELOG.md conflict: Unreleased stays empty. #1748's Fixed line moves
into the 0.119.0 release notes (Fixes), where the rest of the cut
Unreleased body went; the CHANGELOG entry keeps only the title, the
opening paragraph and the release link, as docs/release/publish.md
step 2 says. The title and the opening paragraph gain the warning
redaction.

Checked: no conflict markers; the merged tree differs from b382916
only in package.json, cli.mdx, CHANGELOG.md and the benchmark files.
release:check runs on the branch head next.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The benchmark ran on the release commit 53dbf23. The branch then
merged #1748, which redacts study warnings before the run records them.
That changes what the bundle records and nothing a participant receives
or does, so the result stands for 0.119.0; the README paragraph says so.

Checked: node scripts/check-code-prose.mjs (no count rose).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@danielgwilson
danielgwilson merged commit 1c2fc87 into main Oct 9, 2026
12 of 13 checks passed
@danielgwilson
danielgwilson deleted the release/0.119.0 branch October 9, 2026 07:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A study warning that quotes a sandbox URL makes the run fail its own verification

1 participant