Repository navigation
Release 0.119.0 - #1751
Merged
Merged
Release 0.119.0#1751
Conversation
Bump the version to 0.119.0, move the Unreleased notes into the release notes, and keep the CHANGELOG entry's title, opening paragraph and release link. pnpm docs:generate writes 0.119.0 into cli.mdx. Contents: #1745 and #1742, cut from main 6405a7c. Checked: a live smoke on a tarball packed from main 6405a7c runs alongside this commit (try-live with analysis, 24 participants arriving 10 s apart with their cohort analysis, a late participant past the study budget, analyze --run on a 24-participant bundle). pnpm release:check runs on this commit next. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pnpm bench with its defaults (3 runs per arm, openai-computer-use, neutral mission) on the release commit 53dbf23, as docs/release/publish.md asks for a minor release: report recall 10/15, analysis recall 10/15, 0 invented on either arm, $6.12 estimated. 0.118.0 scored 6/15 and 6/15 at $5.38. D1 and D5 account for the rise: two planted reports and two planted analyses name each, where none did on 0.118.0. D3 and D4 stayed 3/3. Two unresolved planted lines describe D5 (planted 1's report, planted 3's analysis), so recall read by hand is 11/15 and 11/15. No planted participant typed more than 27 characters in one action, so none met D2. The analysis cap was the dry run's admittedCostUsd, $1.25, against bills of $0.60 to $0.86; none was refused or skipped. Every participant ran threaded, ended on its own and gave 5 or 6 impressions. E2B listed 0 tagged sandboxes after it. Issue #1749 reads the decline from 0.114.0 to 0.118.0 as noise at three runs per arm (same-hour A/B of npm 0.114.0 and main d183d1a: 52/80 each). The README paragraph states that finding without the issue number, since the evidence prose check caps issue references at 102. Checked: the summary's claims to check by hand, against the results file and the two unresolved findings' text in the analyses; typed lengths from the traces' action titles; node scripts/check-code-prose.mjs. Not checked: which controls planted participants clicked, read from trace coordinates as #1749 does; three runs per arm cannot separate the rise from run-to-run variation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
#1748 (fixes #1747) merged to main after the 0.119.0 cut: study warnings recorded in the run's bundle go through redactText, so a hosted study whose participant instruction names a sandbox URL other than the subject's verifies share_ready again. It ships in 0.119.0, and the ruleset needs the branch up to date with main. CHANGELOG.md conflict: Unreleased stays empty. #1748's Fixed line moves into the 0.119.0 release notes (Fixes), where the rest of the cut Unreleased body went; the CHANGELOG entry keeps only the title, the opening paragraph and the release link, as docs/release/publish.md step 2 says. The title and the opening paragraph gain the warning redaction. Checked: no conflict markers; the merged tree differs from b382916 only in package.json, cli.mdx, CHANGELOG.md and the benchmark files. release:check runs on the branch head next. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The benchmark ran on the release commit 53dbf23. The branch then merged #1748, which redacts study warnings before the run records them. That changes what the bundle records and nothing a participant receives or does, so the result stands for 0.119.0; the README paragraph says so. Checked: node scripts/check-code-prose.mjs (no count rose). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Cuts 0.119.0 from main, a minor release with #1745, #1742 and #1748 merged since 0.118.0. Commit 53dbf23 sets the version on main 6405a7c, moves the Unreleased body out of CHANGELOG.md and regenerates cli.mdx. Commit cd8bf63 adds the Taskly benchmark result that docs/release/publish.md asks for on a minor release, run on 53dbf23. Merge d390396 brings in main b382916 (#1748, fixes #1747); it moves #1748's Fixed line into these notes and adds it to the CHANGELOG title and opening paragraph. Commit 558a295 notes in the benchmark README that the run predates #1748, which changes only how study warnings are recorded. Nothing else merged to main after the cut. The draft release notes follow; the GitHub release body is the same text from the opening paragraph down.
humanish 0.119.0 lets a computer-use or shared-world study start its participants over time, and
analyses a run of 17 to 100 participants in one report.
participants[].startAfterMsstarts aparticipant up to 24 hours after the run starts its participants, and a
countgroup'sstartEveryMsspreads its members. Each participant's desktop is created at its start, the plan andstudy checkprint the first and last start and the most participants at once, and run.json recordseach participant's
arrival. A participant whose start comes after the study crossedcaps.maxTotalUsdis skipped before its desktop is created. The analysis of more than 16participants splits them into cohorts of at most 16, analyses each cohort with its own request, and
merges the cohort reports into the run's one report with one more request; 0.118.0 skipped the
automatic analysis of such a run. Admission, the plan's analysis line,
study checkand theanalyze --dry-runworst case count every request, andanalyze --jsongives their number asadmission.requests. At the default $3 cap the analysis of 17 or more participants is refused evenwith no evidence, so it needs a higher
review.analysis.maxCostUsd. A provisioned shared world getsan app sandbox that lives until the last start or wave can end. A study warning recorded in the
run's bundle is redacted as a run failure is, so a hosted study whose participant instruction names
a sandbox URL other than the subject's verifies
share_readyagain, as on 0.111.2.This is a minor release because it adds two study fields and a JSON field, changes how a run of
more than 16 participants is analysed, and changes outputs for the same input. A study that sets
startAfterMsorstartEveryMsruns on 0.119.0, and npm 0.118.0 refuses it when read. A run of 17to 100 participants, whose automatic analysis 0.118.0 skipped with
AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT, is analysed in cohorts when its cap admits it; at thedefault $3 cap its default analysis is skipped with
AUTOMATIC_ANALYSIS_ADMISSION_REFUSED.AUTOMATIC_ANALYSIS_PARTICIPANT_LIMITis no longer recorded, andstudy check --jsonno longercarries
analysis.skip; 0.118.0 added both.analyze --runon such a run sends three or morerequests, and its expected cost is about three times that of 0.118.0's partial analysis ($6.20
against $2.20 on the smoke's 24-participant copy).
analyze --jsonaddsadmission.requests. A participant whose start comes after the study budget is spent is recordedblockedwith no desktop, where it made one model request in 0.118.0. A provisioned shared worldwhose waves outlast the sandbox ceiling is refused at plan time. A hosted study whose participant
instruction names a sandbox URL other than the subject's verifies
share_readyagain, where0.113.0 to 0.118.0 graded it
blockedand failed the run. The package exports the same 13values and 31 types as 0.118.0; the exported
RunBundle's participant records gain an optionalarrival. It removes or renames no command, flag, study field or bundle field. These outputschange for the same input, which a script could notice:
participants[].startAfterMsorstartEveryMswas refused when read: "Unknown study field inparticipants[0]: startEveryMs. Known fields: ...",HUMANISH_STUDY_INVALID, exit 2.study checkadds a row, "- ok schedule: the first participant starts at +0s and the last at +3m 50s; at most 18 run at once when every session uses its 3m budget." (achecks[]entry namedschedulein--json); the plan prints the same afterschedule:; run.json records each participant'sarrival(startAfterMs, and on a live runscheduledAtandstartedAt). A start past 24 hours or not a whole number,startEveryMswithoutcount, a start on a study of one participant, a group whose last member passes 24 hours and a later start for the external-publichostare refused withHUMANISH_STUDY_INVALID, each in its own sentence. A study that declares no start prints and records the same as before. (#1745)reviewof a run whose study setsreview.analysis: falsesays "No analysis has run for this run."caps.maxTotalUsd(a later wave) had its desktop created, made one model request and stopped.blockedin run.json and the result, with the summary "Late was skipped: study budget reached before this participant started: the run's estimated model spend $0.0021 crossed caps.maxTotalUsd=$0.001; no desktop was created.", and counted in the fan-out summary's skipped participants. In the smoke, two participants, the second due at +1 min: exit 2, "Fan-out run failed: 0/2 participants passed (1 skipped, 0 harness errors, 0 without engagement).", one sandbox. (#1745)blockedparticipant whose summary says it was skipped as not run.AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT("analysis: not run, because the study has more than 16 participants ..."): exit 0 without a declaredreview.analysis, 2 with one.run,study checkand the TUI printed "After live runs: automatic analysis will not run. This study has 24 participants, and automatic analysis reads at most 16. ..."AUTOMATIC_ANALYSIS_ADMISSION_REFUSED; a declared one is refused and the run exits 2. At a cap that admits it the line gives a range: "expected $5.14 to $9.25 for 24 participants in 2 cohort requests of at most 16 participants and one merge request ..." at $10. The smoke's 24-participant run at $10 was admitted at an expected $6.20 and billed $2.65. (#1742)review.analysis.maxCostUsdfrom the plan's range, orreview.analysis: false. MatchAUTOMATIC_ANALYSIS_ADMISSION_REFUSED;AUTOMATIC_ANALYSIS_PARTICIPANT_LIMITis no longer recorded.study check --jsonof more than 16 participants carriedanalysis.skip: "AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT".doctor --studyprinted "- note post-run analysis: Will not run: this study has 17 participants, more than automatic analysis reads. The participants still run."analysis.refusedFromUsd(3.150188 for 17 at $3) oranalysis.expectedCostUsd, as for any study, and noskip.doctorprints the post-run analysis row a study of 16 gets. (#1742)refusedFromUsdorexpectedCostUsd.analyze --runon a run of more than 16 participants sent one request with the first 16 participants' evidence and saved a partial analysis listing the rest as omitted. On the smoke's 24-participant copy: "Worst case: $2.44, if the analyst writes its whole 32768-token output allowance, reasoning included.", expected $2.20, billed $1.36, 16 participant reviews.Analysis requestingline for each request; the merge's reads "0 evidence items, 0 captures." (#1742)--max-cost.analyze --json'sadmissionhad norequests.admission.requests: 1 up to 16 participants, and the cohorts plus one past 16 (3 for 17 to 32). The other keys are unchanged. (#1742)HUMANISH_E2B_MAX_SANDBOX_MINUTESis refused at plan time, dry runs too: "The participants use the app for up to 30m when every session uses its 10m budget, because 21 participants wait for a free slot. The subject sandbox, which serves the app until every participant ends, then needs a 70m deadline, ... set HUMANISH_E2B_MAX_SANDBOX_MINUTES to 70 or more.",HUMANISH_SHARED_WORLD_INVALID. (#1745)execution.timeoutMs.verifygraded the runblocked(PUBLIC_SAFETY_FINDINGS) and the run failed, dry runs too: "humanish run failed: Run bundle failed verification.",HUMANISH_COMPUTER_USE_FAILED, exit 2. This started in 0.113.0.[REDACTED_SECRET]in place of the URL, as the recorded mission already did, and the run verifiesshare_readyand exits 0. In the smoke, one live participant whose mission nameshttps://3000-...e2b.app/inbox, with blurred screenshots:share_ready, 16 of 16, the URL in no run file; npm 0.118.0: exit 2,blocked. The terminal still prints the warning with the URL. (#1748, fixes #1747)Added
participants[].startAfterMsstarts a participant that many milliseconds after the run startsits participants (up to 24 hours), and a group's
startEveryMsspreads its members, member k atstartAfterMs + (k - 1) * startEveryMs. Each participant's desktop is created at its start. Onewhose start comes while every slot is taken waits for the next free one.
study checkand therun's plan print the first and last start and the most participants at once, and run.json
records each participant's
arrivalwith its scheduled and actual start. A shared-world planwith a schedule prints its worst case in sandbox-minutes. A provisioned shared world's app
sandbox serves until the last participant ends, and a study whose last start puts that past
HUMANISH_E2B_MAX_SANDBOX_MINUTESis refused before anything is created. A rerun(
--rerun-failed-from) starts its selected participants together (row 1). (Start participants over time from a declared schedule #1745, from A study can't simulate a full day: computer-use studies stop at 16 participants #1737)Changed
caps.maxTotalUsd, in a later wave or later in its schedule, is skipped before its desktop iscreated. It is recorded as blocked with a reason naming the budget and counted among the fan-out
summary's skipped participants. It used to create its desktop, make one model request and stop
(row 2). (Start participants over time from a declared schedule #1745, from A study can't simulate a full day: computer-use studies stop at 16 participants #1737)
split into cohorts of at most 16, each analysed by its own request under the full packet limits,
and one more request merges the cohort reports into the run's one report. In 0.118.0 a live run
of more than 16 participants recorded its automatic analysis as skipped with
AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT, andanalyzeread the first 16 participants and listedthe rest as omitted; that reason is no longer recorded. Admission, the expected cost, the worst
case and the
study checkrange cover every request. The plan's analysis line and theanalyze --dry-runworst case name the cohort and merge requests, andadmission.requestsin--jsoncounts them. The default $3 cap refuses the analysis of 17 participants even for a run that
keeps no evidence, so a default analysis of more than 16 participants is skipped with
AUTOMATIC_ANALYSIS_ADMISSION_REFUSEDunlessreview.analysis.maxCostUsdadmits it.analyzeprints an
Analysis requestingline for each request. A cohort or merge request that failsfails the attempt with no findings and a warning, and its usage counts every request sent. The
humanish.study-analysis.v1schema is unchanged; its stored limits grew to 6,400 evidenceitems, 320 captures and 16 MiB (rows 3 to 7). (Analyse a run of more than 16 participants in cohorts merged into one report #1742, from A study can't simulate a full day: computer-use studies stop at 16 participants #1737)
Fixes
The warning that a participant instruction reads like a script quotes up to three of its lines,
and a line that named a sandbox URL other than the subject's (a run inbox page, a second app)
put that URL in
events.ndjson,run.jsonandobserver-data.json.verifythen graded thebundle
blockedand a hosted run failed with "Run bundle failed verification", where 0.111.2verified it
share_ready. The terminal still prints the warning as written (row 9). (Redact study warnings before the run records them (#1747) #1748,fixes A study warning that quotes a sandbox URL makes the run fail its own verification #1747)
the last wave can end, and is refused when that passes
HUMANISH_E2B_MAX_SANDBOX_MINUTES. Itsapp sandbox lived one session plus provisioning, which a second wave could outlast (row 8).
(Start participants over time from a declared schedule #1745, from A study can't simulate a full day: computer-use studies stop at 16 participants #1737; the waves came in 0.118.0 with Let a study have 100 participants and run them within the E2B plan #1739)
Known gaps
humanish watchat 40 and 100 participants are the last slice of A study can't simulate a full day: computer-use studies stop at 16 participants #1737, inCount, filter and page the Observer grid and bound watch output at 100 participants #1744, which is not in this release. (A study can't simulate a full day: computer-use studies stop at 16 participants #1737, Count, filter and page the Observer grid and bound watch output at 100 participants #1744)
spend.requestscounts an analysis attempt as one request. The smoke's24-participant run priced 3 requests, and its
study-analysis.jsonreads"requests": 1. Thebundle does not record how many requests an attempt sent; its receipt holds their summed usage.
(Observer spend.requests counts a cohort analysis as one request #1743)
local-treeorclonesubject, the first participant holds the others at the pipeline gateuntil its desktop is ready. Participants due before the gate opens start when it opens, late,
and the plan's schedule does not count the gate: in the smoke, participants due at +10 s to
+40 s all started at +44 s. A failed first desktop skips every other participant. (One participant's failed desktop skips every other participant in the study #1741)
budget is crossed early idles until its last start, with no sandbox.
runprints such aparticipant as "blocked · unknown ending" and leaves out the reason, which run.json holds.
(Start participants over time from a declared schedule #1745)
plan's peak and the provisioned shared world's app deadline count each session at its full
budget and leave out each desktop's start-up time; with several waves on a provisioned shared
world, that start-up comes out of the 10-minute teardown buffer. (Start participants over time from a declared schedule #1745)
doctor --studyof a study of more than 16 participants prints "ok post-run analysis" and doesnot say that the default $3 cap refuses its analysis;
runandstudy checksay so. (Analyse a run of more than 16 participants in cohorts merged into one report #1742)cohorts". The merge request's
Analysis requestingline reads "0 evidence items, 0 captures."(Analyse a run of more than 16 participants in cohorts merged into one report #1742)
rundoes not print the warning that namesHUMANISH_E2B_MAX_CONCURRENT_SANDBOXESwhen thelimit holds participants to waves; the run records it as a
study.warningevent, and theterminal shows only the plan line. (Let a study have 100 participants and run them within the E2B plan #1739)
run100warning lines, among them 24 persona-background warnings for a persona with no background.
(Let a study have 100 participants and run them within the E2B plan #1739)
runprints its analysis line from the declared participant count, so a--countoverride doesnot change it, and a
--countabove 64 with real email receiving still fails at run start.(Let a study have 100 participants and run them within the E2B plan #1739)
humanish tuican still confirm an armed action on its first key repeat. The400 ms floor ignores repeats that come sooner, and a system's initial repeat delay is often
longer (500 ms by default in GNOME). (TUI: the first key repeat of a held Enter can still confirm an armed action #1730)
serve.startthat bash cannot parse still waits outreadyTimeoutMs. The scripted andshared-world routes give a failed build their generic failure code; the message carries bash's
error. (Fail a detached step bash cannot parse at once, with bash's error #1734)
guest holds 8 utterances between two screenshots, so a long wait in a talking call can overflow
it; a wait there stays 30 seconds unless the study raises
actor.maxWaitMs. (A long wait no longer ends a participant's session #1724)(A host that sleeps mid-run is reported as dozens of product failures #1740)
same user, to swap a directory in the run between two filesystem calls: temporary-file cleanup
still unlinks by path (item 1), and the reviewer-notes listing still lists
notes/by path andcan follow a swapped directory (item 3). Storage validation and the scan still list each folder
whole before the 10,000-entry cap applies, verify has no total read budget, and a legitimate
sandbox-receipts.ndjsonover 32 MiB is refused. (Bind contained-file writes, reads and listings to directory descriptors #1669)humanish tui. (Coding agents never tell their person about humanish tui, and bare humanish does not name it #1701)such as a line break and the next header after
cwd:. Nothing leaks; text a reviewer needs islost. (Path redaction deletes the text after a path in terminal transcripts #1702)
é,€) askeys; it has to type them as text. (Translate E2B desktop key names through the shared key table (#1697) #1713)
E2B_DOMAINfromprocess.envon create and ignore one inRunStudyOptions.env. (Kill and list E2B sandboxes with the run's API key #1717)decodings at a marker edge are not read,
RunSecrets.spans(the terminal route's per-chunkreplacement) has no final check, and a value that is part of a marker nests markers. (Harden the shared secret scrubber against encoded and marker-shaped values #1646)
humanish tuireads "1 run, none live".visual critique whatever its persona. (Ask for impressions in the persona's own terms #1654)
runtime setup --memory --cpus.Package
humanish 0.119.0 packs to 2.5 MB, 7.9 MB unpacked, 972 files (
npm packon release/0.119.0 at558a295, shasum 9c884092b9d8dca9d1396e2699772752f617a691; the same file list as at 53dbf23).
Against 0.118.0 it adds
dist/analysis/cohorts.js,dist/analysis/responses.jsanddist/study/arrivals.js, each with its.d.ts, and dropsdist/study/automatic-analysis-plan.jsand
.d.ts. The package exports only., with the same 13 values and 31 types as 0.118.0 (pnpm api:proof); the exportedRunBundle's participant records gain an optionalarrival. Itsdependencies, optional peer and engines are the same as 0.118.0's: Node 22.19 or newer, and
@e2b/desktop2.3.2 or newer as an optional peer.Checked before release
A live smoke on a tarball packed from main 6405a7c, and its #1748 row on the release tarball from
558a295, in fresh projects under /tmp with a fresh HOME and npm cache,
DO_NOT_TRACK=1, and@e2b/desktop2.4.0 beside humanish. Keys came from the user key store through key discovery, withXDG_CONFIG_HOMEpointing at it; no--dotenvwas passed. The store's OpenAI organization isnon-ZDR, so every OpenAI participant ran
threaded; its E2B key is on a Pro plan. Controls ran thesame study files or commands on npm 0.118.0, read copies of the main runs, or analysed a second copy
of the 0.118.0 smoke's 24-participant run. Priced spend: $8.54 for the smoke ($7.57 of it in four
analyses, $1.36 of those on npm 0.118.0; $0.02 in the npm 0.118.0 run of the #1748 row), then $0.22
for the dogfood run and $6.12 for the benchmark. The dogfood run, the benchmark and the consumer
check ran on the release commit 53dbf23, before main b382916 (#1748) was merged into the branch;
#1748 changes how study warnings are recorded in the bundle and nothing a participant receives or
does.
init --yeswrites the same files as npm 0.118.0, anddoctor(the concurrent-sandboxes rowunset, at 100 and at
lots),doctor --study try-live,keys,study check try-liveandanalyze --helpprint the same apart from the cwd line.initwrites it: reached the goal in 11 participant turns, input growing everyturn from 2,888 to 19,839 tokens, then 20,170 for the impressions request; cached input on each
turn from the second was the previous turn's input minus 3;
threadedon all 12 requests. $0.28run. Its automatic analysis, one request, was admitted at an expected $1.04 (
admittedCostUsd$1.15, worst case $2.03, cap $3) and billed $0.78, with 4 findings and 4 design findings whose 30
cited evidence ids all resolve. It gave 6 impressions (unclear, liked, unclear, unfinished,
liked, liked).
count: 24,startEveryMs: 10000) on a one-paragraphstatic page,
HUMANISH_E2B_MAX_CONCURRENT_SANDBOXES=100, 3-minute sessions,review.analysis.maxCostUsd: 10(rows 1 and 3). The plan read "the first participant starts at+0s and the last at +3m 50s; at most 18 run at once when every session uses its 3m budget."
run.json recorded each
arrival: participant 1 and participants 6 to 24 started within 55 ms ofscheduledAt(19 of the 20 within 1 ms); participants 2 to 5, due at +10 s to +40 s, startedat +44 s, when the first participant cleared the local-tree pipeline gate. Sandbox receipts are
timed from the first start to 0.4 s after the last; E2B listed at most 7 of the run's sandboxes
at once. 24 of 24 reached the goal; $0.62 run. The analysis line read "expected $5.14 to $9.25
for 24 participants in 2 cohort requests of at most 16 participants and one merge request"; the
automatic analysis was admitted at an expected $6.20 (
admittedCostUsd$6.82, worst case $7.07,requests3) and billed $2.65, 5 min 33 s. The report covers 24 of 24 participants with 24participant reviews, 2 findings, 2 design findings and 4 concern reviews; its 172 distinct cited
evidence ids all resolve in the stored packet, and its 24 cited captures exist with their
recorded sha256.
that names
https://3000-isbxsmokewarning0119.e2b.app/inbox, and no analysis. On the releasetarball from 558a295 the run exited 0 and verified
share_ready, 16 of 16; its recorded warningand mission read
[REDACTED_SECRET]in place of the URL, no run file holds it, and the terminalprinted the warning with the URL. $0.04. On npm 0.118.0 the same study exited 2, "humanish run
failed: Run bundle failed verification.",
HUMANISH_COMPUTER_USE_FAILED,blockedwithPUBLIC_SAFETY_FINDINGS, and the URL in run.json, events.ndjson and observer-data.json. $0.02.Dry runs of the same study, and of one whose participant instruction names such a URL, gave the
same grades.
caps.maxTotalUsd0.001.The first crossed the budget on its first request ($0.0021). The second was recorded
blockedwith "no desktop was created", its
arrivalwithscheduledAtand nostartedAt; the run hadone sandbox receipt, E2B listed one sandbox for it, and the run ended at the second's scheduled
time. Exit 2, "0/2 participants passed (1 skipped, ...)". $0.006.
analyze --runon two copies of the 0.118.0 smoke's 24-participant run, with each version'sdry-run worst case rounded up as
--max-cost(row 5): npm 0.118.0 sent one request of 233evidence items and 16 captures and saved a partial analysis of 16 participants, 8 omitted
(expected $2.20, billed $1.36). The candidate sent two cohort requests of 175 and 176 items with
12 captures each and a merge request, and saved a complete analysis of 24 (expected $6.20, billed
$2.79, 6 min 17 s); its 195 distinct cited ids all resolve and its 24 cited captures exist.
verify,reviewandruns(text and--json) on copies of the three live runs and thecandidate-analysed 24-participant copy: byte-identical on main and npm 0.118.0, except
reviewof the budget-skip run, whose study file npm 0.118.0 cannot parse (row 1).
verifyandreviewof the arrivals run with the release tarball from 558a295 print the same as with main's.
study checkof the arrivals and budget-skip studies plan on main and are refusedas unknown fields on npm 0.118.0. A provisioned shared world of 40 members and one whose second
member starts 3 hours in are refused on main (row 8); npm 0.118.0 plans the 40.
reclaim --checkexited 0 withclean, andverifyexited 0 with 16 of 16checks,
local_only. Raw sandbox ids appear only insandbox-receipts.ndjson. E2B listed 0tagged sandboxes after the smoke and after the Redact study warnings before the run records them (#1747) #1748 pair (2 before, from other sessions' runs)
and 0 after the benchmark. No key value from the store appears in any run or kit file, as written,
base64 or hex (16,446 files).
On the release branch:
pnpm release:checkon the release head 558a295, after the merge of main b382916 (Redact study warnings before the run records them (#1747) #1748) andthe benchmark note: exit 0 in 3.9 min, 513 test files passed and 10 skipped, 8,034 tests passed
and 12 skipped (Redact study warnings before the run records them (#1747) #1748 adds one); TUI 192 tests; lint 426 warnings at its cap of 426; api:proof "13
values and 31 types match"; public-surface scan of 3,063 text files and 527 binary assets;
npm pack --dry-run972 files with shasum 9c884092b9d8dca9d1396e2699772752f617a691. On the benchmarkcommit cd8bf63 before the merge: exit 0 in 3.7 min, 3,062 text files, shasum
b10a60b8c14494d4035e7bf353211a1001b21d93. On the release commit 53dbf23: exit 0 in 3.9 min, 512
test files passed and 10 skipped, 8,033 tests passed and 12 skipped; TUI 192 tests; lint 426
warnings at its cap of 426; api:proof "13 values and 31 types match"; public-surface scan of 3,060
text files and 527 binary assets;
npm pack --dry-run972 files with shasumb10a60b8c14494d4035e7bf353211a1001b21d93.
pnpm docs:generatethere wrote 0.119.0 into cli.mdx;the merge leaves cli.mdx as it is. The first run on the benchmark commit, before an amend
(f0ac0892), exited 1 in 11 s on
prose.evidence.issue-refs: 103 (cap 102, over by 1): the newbenchmark README paragraph named Benchmark recall decline 0.114.0 to 0.118.0 is n=3 reach noise; same-hour A/B ties at 52/80 #1749. The amend states the finding without the issue number.
pnpm release:dogfoodon 53dbf23, before the merge of Redact study warnings before the run records them (#1747) #1748 (the script run with Node, keys readfrom the user store by Node's
--env-fileinto its environment, and a temporary HOME):"release:dogfood ok", verdict
blocked(blocked_approval), no-spend satisfied, $0.22. Theparticipant ran
init, the free four-participant dry run,verify("all 16 evidence checks",share_ready) andreviewon the installed 0.119.0, and stopped at the keys and the local VMdevices the gate withholds; the script reads that verdict as the expected stop.
verifyon a copyof the run:
share_ready, 16 of 16;reclaim --checkclean.TMPDIR=/tmp,HUMANISH_SPECthetarball): 123 of 123 passed. Its 6 new rows cover the schedule from
startAfterMsandstartEveryMsinstudy check, the plan and run.json (Start participants over time from a declared schedule #1745); four start refusals (Start participants over time from a declared schedule #1745); a3-hour start planned on computer use and refused on a provisioned shared world (Start participants over time from a declared schedule #1745); the
17-participant analysis line at a $10 cap and
analysis.refusedFromUsdat $3 (Analyse a run of more than 16 participants in cohorts merged into one report #1742);admission.requests1 and 3 (Analyse a run of more than 16 participants in cohorts merged into one report #1742); and a dry run whose participant instruction names anon-subject e2b URL, which verifies
share_readywith the URL in no run file and the warning onthe terminal (Redact study warnings before the run records them (#1747) #1748). Row 20bz now expects the 17-participant refusal line in place of 0.118.0's
skip. The same script on npm 0.118.0 passed 116 and failed the 6 new rows and 20bz. An earlier
pass of the 122 rows before Redact study warnings before the run records them (#1747) #1748 passed 122 on the 53dbf23 tarball and 116 on npm 0.118.0.
release, on the release tarball from 53dbf23, before the merge of Redact study warnings before the run records them (#1747) #1748. On that tarball the
consumer typechecks with 0 errors, as on npm 0.118.0, and
needs no migration. The library steps are the same on both versions, and of the 124 CLI lines
from the 0.118.0 check, 121 are the same; the three that differ are
admission.requestsinanalyze --json(row 6), and the 17-participantstudy checkline anddoctor --studynote(rows 3 and 4). The new lines differ as rows 1 and 3 to 8 say.
pnpm benchon 53dbf23, before the merge of Redact study warnings before the run records them (#1747) #1748, with its defaults (3 runs per arm,openai-computer-use,neutralmission), committed in cd8bf63: report recall 10/15, analysisrecall 10/15, 0 invented on either arm, $6.12 estimated. 0.118.0 scored 6/15 and 6/15 at $5.38. D1
and D5 account for the rise: two planted reports and two planted analyses name each, where none
did on 0.118.0; D3 and D4 stayed 3/3. Two unresolved planted lines describe D5 (planted 1's
report, planted 3's analysis), so recall read by hand is 11/15 and 11/15. Benchmark recall decline 0.114.0 to 0.118.0 is n=3 reach noise; same-hour A/B ties at 52/80 #1749 found no harness
regression behind the decline from 0.114.0 to 0.118.0: a same-hour A/B of npm 0.114.0 and main
d183d1a scored 52/80 planted-report recall each over 16 planted participants (25/40 and 27/40 in
its first round), and three runs per arm cannot separate this rise from the same noise. The
analysis cap was the dry run's
admittedCostUsd, $1.25, against bills of $0.60 to $0.86; none wasrefused or skipped. No planted participant typed more than 27 characters in one action, so none
met D2. Every participant ran
threaded, ended on its own and gave 5 or 6 impressions.Not verified
rows are offline dry runs, refusals and local runs.
smoke's late starts came from the pipeline gate. A live shared world with a declared schedule,
external-public followers on a schedule, and a rerun of a scheduled study.
(
AUTOMATIC_ANALYSIS_ADMISSION_REFUSED) and its exit codes; tests cover them.consistent with 3.
analyze --runprinted its 3.request live, and a Codex account analysis of more than 16 participants.
captures, 4 MiB).
53dbf23, before the merge of Redact study warnings before the run records them (#1747) #1748. The live smoke apart from the Redact study warnings before the run records them (#1747) #1748 row ran on main
6405a7c's tarball.
smoke ran that case as a dry run.
🤖 Generated with Claude Code