Repository navigation
Analyse a run of more than 16 participants in cohorts merged into one report - #1742
Merged
Merged
Conversation
One analysis now covers up to 128 participants. captureEvidence selects each cohort of at most EVIDENCE_LIMITS.cohortParticipants (16, now in analysis-limits.ts with the other packet limits) under the full packet limits, and runAnalysis sends one request per cohort, four at a time, then one merge request over the cohort reports. The merged report keeps the cohorts' participant reviews and is checked by the same validator against the evidence the cohort reports cite. Admission, the expected cost, the worst case and the study check range sum every request. A failed request fails the attempt with no findings and a warning. Response checking and usage moved to responses.ts so execute.ts keeps the prompts, admission and the request loop. Stored limits grew to hold eight cohorts. EVIDENCE_LIMITS.participants is now 128, the stream limit the source parser already had. Checked: 11 tests in tests/analysis/cohorts.test.ts (17- and 40-participant bundles, admission worked example, merge through an injected fetch); full vitest suite; one live analysis of a 24-participant run (3 requests, expected $6.08, billed $2.65, 447 citations resolved). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The merge request saw no evidence, so its findings, design findings and concern reviews may now cite only evidence the cohort reports cite in theirs; a capture cited only in a participant review was accepted before. checkMergedResponse in cohorts.ts owns that check and replaces citedInput and mergedResponse. cohortRequest and mergeRequest build what each request sends, and admission prices the same objects, so the two cannot drift. The merge instructions now say how to stay under 30 observations per finding. The plan line counts cohorts of at most 128 participants. A cohort input no longer recomputes a digest nothing reads. Tests: 17 in tests/analysis/cohorts.test.ts. The two merge-citation cases fail on the previous commit. New coverage for a fifth cohort that never starts, a failed merge request, a cancellation before the merge, a summed bill over the summed worst case, and Codex requests one at a time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ticipant With EVIDENCE_LIMITS.participants at 128, a study of 17 to 100 participants gets its automatic analysis, so the four tests that pinned the interim AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT skip now check that the analysis is planned, named in the study check line and requested after a live run of 17. The skip itself, in src/study/automatic-analysis-plan.ts, is left for its owner: no study reaches 128 participants, so it no longer fires. Its Unreleased CHANGELOG entry is replaced by the cohort entry, which now says a default analysis of 17 participants is over the $3 cap. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
With EVIDENCE_LIMITS.participants at 128 and a study capped at 100, the AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT skip from #1739 could not fire. src/study/automatic-analysis-plan.ts is deleted, and its callers go back to the analysis owner: the route shell calls completeAutomaticAnalysis, the CLI exit rule and the run envelope use automaticAnalysisSucceeded, run, study check, the study summary the TUI reads and doctor use automaticAnalysisBudget and formatAutomaticAnalysisBudget, and the plan no longer carries a skip or the participant count it was read from. The doctor note, the CLI reason text and the study-files sentence go too. The analyze dry run's worst case now says every request writes its whole allowance and names the cohort and merge requests, in words from analysisRequestsText, which also writes the plan line. The count is analysisRequestCount in analysis-limits.ts; admission reports it as admission.requests (additive in --json; the analyze golden gains it). Checked: no code, test or doc mentions the skip (git grep). New tests: the dry run of a 17-participant preview run names 2 cohort requests and one merge request (red on the previous commit), and admission counts 1, 3 and 4 requests for 16, 17 and 40 participants, null when refused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Conflicts with #1735, resolved to keep both: - automatic-config.ts: admissionRule words the rule; the participants phrase ("24 participants in 2 cohort requests of at most 16 participants and one merge request") feeds both the range form and the refused-even-with-no-evidence form. - execute.ts: the request builders stay; analysisCostRange keeps #1735's bisection and largerOutput, and each probe prices every cohort request and the merge request at that share of the evidence limits. - doctor.ts: imports admissionRule and automaticAnalysisBudget; the deleted skip module stays deleted. - automatic-analysis.md, analyze-cost.test.ts: both sides' text and tests kept. - CHANGELOG: 0.118.0 from main; the cohort entry under a fresh Unreleased, saying 0.118.0 skipped analysis past 16 participants. Tests: the plan-line tests now expect the refused form at the default $3 cap and the range form at $20; a new test checks the 40-participant range ends at a $10 cap. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # CHANGELOG.md # src/routes/shared-world/plan.ts
Merged
danielgwilson
added a commit
that referenced
this pull request
Oct 9, 2026
* Release 0.119.0 Bump the version to 0.119.0, move the Unreleased notes into the release notes, and keep the CHANGELOG entry's title, opening paragraph and release link. pnpm docs:generate writes 0.119.0 into cli.mdx. Contents: #1745 and #1742, cut from main 6405a7c. Checked: a live smoke on a tarball packed from main 6405a7c runs alongside this commit (try-live with analysis, 24 participants arriving 10 s apart with their cohort analysis, a late participant past the study budget, analyze --run on a 24-participant bundle). pnpm release:check runs on this commit next. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add the 0.119.0 Taskly benchmark result pnpm bench with its defaults (3 runs per arm, openai-computer-use, neutral mission) on the release commit 53dbf23, as docs/release/publish.md asks for a minor release: report recall 10/15, analysis recall 10/15, 0 invented on either arm, $6.12 estimated. 0.118.0 scored 6/15 and 6/15 at $5.38. D1 and D5 account for the rise: two planted reports and two planted analyses name each, where none did on 0.118.0. D3 and D4 stayed 3/3. Two unresolved planted lines describe D5 (planted 1's report, planted 3's analysis), so recall read by hand is 11/15 and 11/15. No planted participant typed more than 27 characters in one action, so none met D2. The analysis cap was the dry run's admittedCostUsd, $1.25, against bills of $0.60 to $0.86; none was refused or skipped. Every participant ran threaded, ended on its own and gave 5 or 6 impressions. E2B listed 0 tagged sandboxes after it. Issue #1749 reads the decline from 0.114.0 to 0.118.0 as noise at three runs per arm (same-hour A/B of npm 0.114.0 and main d183d1a: 52/80 each). The README paragraph states that finding without the issue number, since the evidence prose check caps issue references at 102. Checked: the summary's claims to check by hand, against the results file and the two unresolved findings' text in the analyses; typed lengths from the traces' action titles; node scripts/check-code-prose.mjs. Not checked: which controls planted participants clicked, read from trace coordinates as #1749 does; three runs per arm cannot separate the rise from run-to-run variation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Note that the 0.119.0 benchmark ran before the warning fix The benchmark ran on the release commit 53dbf23. The branch then merged #1748, which redacts study warnings before the run records them. That changes what the bundle records and nothing a participant receives or does, so the result stands for 0.119.0; the README paragraph says so. Checked: node scripts/check-code-prose.mjs (no count rose). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
#1737 slice 3 of 4: one
humanish.study-analysis.v1report for a run of 17 to 128 participants.Design: cohorts merged by one request
A run with more than 16 participants is split into the fewest cohorts of at most 16 (round-robin over the roster: 17 make 9 and 8, 24 make 12 and 12, 40 make 14, 13 and 13, 100 make seven). Each cohort gets its own request with today's instructions, schema, validator and packet limits (800 entries, 40 captures, 160 KiB, 20 MiB). One more request reads the cohort reports, never the evidence, and writes the run's summary, findings, design findings, concern reviews and limitations. The participant reviews are the cohorts' own.
Why (a) over (b), one packet sampled to fit:
What changed
captureEvidence(src/analysis/evidence.ts), same interface: selection runs per cohort under the full packet limits; evidence ids stay global and in bundle order.EVIDENCE_LIMITSmoved to src/analysis/analysis-limits.ts (re-exported from evidence.ts for src/run/costs.ts and src/feedback/export.ts).participantsis now 128, the most streams the source parser already read, so it means "the most participants one run's analysis covers" (agreed with the participants-cap session, which reads it for its skip).cohortParticipants: 16is the one owner of the 16, andanalysisCohortsis the one owner of the split, used by packet building, the request loop, the cost range and the plan line.runAnalysis(src/analysis/execute.ts), same interface: onebeforeDispatch(one start receipt), cohort requests four at a time for OpenAI (four requests the size of a billed one at the packet limits ask for about 590,000 tokens, under the Build tier's 1,000,000 TPM; Codex one at a time), then the merge request. Usage is the sum, priced with each response as its own request so the long-context tier applies per request.checkMergedResponse, which adds the cohorts' participant reviews, runs the standard check and refuses a merged citation that no cohort finding, design finding or concern review made (analysis_validation_failed_merge_reference_invalid). Deletion test: without it the projection, the packet and the citation rule reappear in the request loop and admission.cohortRequestandmergeRequest(execute.ts) build what each request sends; admission prices the same objects, so what is priced and what is sent cannot drift.recordProviderUsage(now over several responses) and the Codex response warnings, so execute.ts holds the prompts, admission and the request loop.origin/main).ANALYSIS_PROMPT_VERSIONstaysstudy-evidence-8: no earlier study-evidence-8 artifact has cohorts, and a 16-participant run's request is unchanged, so its reuse key still holds.Admission, plan time, failures, schema
analysisCostRange(thestudy checkrange) sums the same requests. The plan line names them past 16: at a $20 cap,expected $5.14 to $11.24 for 24 participants in 2 cohort requests of at most 16 participants and one merge request, depending on how much evidence the run keeps. At the default $3 cap the analysis of 17 or more participants is refused even with no evidence, so a default analysis of them is skipped withAUTOMATIC_ANALYSIS_ADMISSION_REFUSED.humanish.study-analysis.v1is unchanged in shape. Its stored bounds grew to hold eight cohorts: 6,400 evidence items, 320 captures, 16 MiBanalysis.json, andvalidateAnalysisEvidenceallows 20 MiB of images per cohort.Checked
origin/main(17 participants packed to 16, one request, the old cost signature). Later additions, each red on the commit before it: the 20 MiB image bound on validation, the plan line, and the two merge-citation cases (a capture cited only in a participant review, and one cited nowhere). Coverage tests for existing behavior: a fifth cohort that never starts after a failure, a failed merge request, a cancellation before the merge, a summed bill over the summed worst case, and Codex requests one at a time (mutation-checked: the test fails with Codex at four at a time). tests/analysis/cohorts.test.ts has 17 tests; values come from worked examples (cost: two cohorts and a merge, $4.6585 expected, $6.0737 worst; captures per participant: 4 to 5 of 9 for 17, at least 2 of 3 for 40) and from the captured wire envelope's usage.AUTOMATIC_ANALYSIS_PARTICIPANT_LIMITskip now check that 17 participants are planned, named in thestudy checkline as 2 cohorts, and analysed after a live fake-desktop run; the explicit-request skip test is removed, since no study reaches 128 participants.pnpm vitest run tests/analysis/: 24 files, 504 tests.cua-2026-10-09T03-17-16-059Z-0654b949, example.com):analyze --dry-run --max-cost 30: admitted at $6.689375 (expected $6.08, worst $6.96).--max-cost 7is that worst case rounded up, the cap humanish's refusal suggests.completein 5 min 6 s, billed estimate $2.649317 (91,342 input, 31,523 output tokens, usage complete). Expected was 2.3 times the bill.--rerun, final code and merge instructions): 3 requests,completein 5 min 48 s, billed $1.919481 (91,838 input of which 74,654 cached, 32,601 output).review --jsongiveshumanish.analysis-findings.v1readywith 143 cited evidence items and 0 missing capture files.verifystill grades the runlocal_only, as before analysis; the Observer refreshed without a warning.pnpm release:check: exit 0 on 4e6dcab (508 files, 8,000 tests passed, 12 skipped; TUI 192 passed).Lead review changes
AUTOMATIC_ANALYSIS_PARTICIPANT_LIMITskip from Let a study have 100 participants and run them within the E2B plan #1739:src/study/automatic-analysis-plan.ts, the plan'sanalysis.skipand the participant countplanBaseread it from (computer-use and shared-world planners), the route shell's skip branch, the CLI exit rule's skip pass, the CLI reason text, doctor's "Will not run" note and the study-files sentence. Callers go back to the analysis owner:completeAutomaticAnalysis,automaticAnalysisSucceeded,automaticAnalysisBudgetandformatAutomaticAnalysisBudget(run, study check, the study summary the TUI reads, doctor). Each restored hunk matches its pre-Let a study have 100 participants and run them within the E2B plan #1739 form.git grepfinds no mention of the skip.analyze --dry-runworst case now readsif every request (2 cohort requests of at most 16 participants and one merge request) writes its whole 32768-token output allowancefor a cohort analysis. The words come fromanalysisRequestsText(automatic-config.ts), which also writes the plan line, and the count fromanalysisRequestCount(analysis-limits.ts), reported as the additiveadmission.requestsin--json. One request keeps the old sentence. The analyze golden gains"requests": 1in each admission and nothing else.Merge with 0.118.0 (#1735)
automatic-config.ts: Fix the four TUI gaps from the 0.116.0 release notes #1735'sadmissionRuleis the one text of the rule. The participants phrase carries the requests (24 participants in 2 cohort requests of at most 16 participants and one merge request) and feeds both of Fix the four TUI gaps from the 0.116.0 release notes #1735's forms: the range, and "refused before it starts for ... even with no evidence".execute.ts: Fix the four TUI gaps from the 0.116.0 release notes #1735'sanalysisCostRangebisection andlargerOutputstay; each probe prices every cohort request at that share of its evidence limits plus the merge request, so the range ends at what admission admits for the whole analysis. With the default $3 cap, 17 or more participants are refused even with no evidence; at $10, 40 participants range from $7.05 to $10.00.doctor.ts: importsadmissionRuleandautomaticAnalysisBudget; the deleted skip module stays deleted.docs/product/automatic-analysis.md,tests/cli/analyze-cost.test.ts: both sides kept.## 0.118.0from main. The cohort entry is under a fresh Unreleased and says that 0.118.0 recordedAUTOMATIC_ANALYSIS_PARTICIPANT_LIMITpast 16 participants, which shipped in that release.Not verified
humanish watchat 40 and 100 participants (slice 4).spend.requestsandstatscount one analysis attempt as one request: filed as Observer spend.requests counts a cohort analysis as one request #1743 (observer/ belongs to another session).🤖 Generated with Claude Code