Skip to content

Analyse a run of more than 16 participants in cohorts merged into one report - #1742

Merged
danielgwilson merged 8 commits into
mainfrom
feat/analysis-past-16
Oct 9, 2026
Merged

danielgwilson merged 8 commits into
mainfrom
feat/analysis-past-16

Conversation

@danielgwilson

@danielgwilson danielgwilson commented Oct 9, 2026 •

Copy link
Copy Markdown
Owner

#1737 slice 3 of 4: one humanish.study-analysis.v1 report for a run of 17 to 128 participants.

Design: cohorts merged by one request

A run with more than 16 participants is split into the fewest cohorts of at most 16 (round-robin over the roster: 17 make 9 and 8, 24 make 12 and 12, 40 make 14, 13 and 13, 100 make seven). Each cohort gets its own request with today's instructions, schema, validator and packet limits (800 entries, 40 captures, 160 KiB, 20 MiB). One more request reads the cohort reports, never the evidence, and writes the run's summary, findings, design findings, concern reviews and limitations. The participant reviews are the cohorts' own.

Why (a) over (b), one packet sampled to fit:

  • (b) at today's limits gives each of 100 participants 8 entries and 0.4 captures; a 16-participant run gives 50 and 2.5.
  • (b) at today's per-participant quality is about 712K input tokens for 100 participants. gpt-6-astra takes it (1,050,000-token context, 128,000 output tokens, developers.openai.com/api/docs/models/gpt-6-astra, read 2026-10-09), but every prompt over 272K is billed at 2x input and 1.5x output, and 100 participant reviews do not fit the 32,768-token output allowance that also holds the reasoning. A 12-participant request already used 17,289 output tokens (docs/evidence/computer-use/analysis-reliability-2026-09-19.md).
  • (a) reuses the request, prompt and validator unchanged per cohort, and balanced cohorts give each participant at least the share of a 16-participant run (tests below).

What changed

  • captureEvidence (src/analysis/evidence.ts), same interface: selection runs per cohort under the full packet limits; evidence ids stay global and in bundle order.
  • EVIDENCE_LIMITS moved to src/analysis/analysis-limits.ts (re-exported from evidence.ts for src/run/costs.ts and src/feedback/export.ts). participants is now 128, the most streams the source parser already read, so it means "the most participants one run's analysis covers" (agreed with the participants-cap session, which reads it for its skip). cohortParticipants: 16 is the one owner of the 16, and analysisCohorts is the one owner of the split, used by packet building, the request loop, the cost range and the plan line.
  • runAnalysis (src/analysis/execute.ts), same interface: one beforeDispatch (one start receipt), cohort requests four at a time for OpenAI (four requests the size of a billed one at the packet limits ask for about 590,000 tokens, under the Build tier's 1,000,000 TPM; Codex one at a time), then the merge request. Usage is the sum, priced with each response as its own request so the long-context tier applies per request.
  • src/analysis/cohorts.ts (new): a cohort's input, the merge packet, the merge response schema, and checkMergedResponse, which adds the cohorts' participant reviews, runs the standard check and refuses a merged citation that no cohort finding, design finding or concern review made (analysis_validation_failed_merge_reference_invalid). Deletion test: without it the projection, the packet and the citation rule reappear in the request loop and admission.
  • cohortRequest and mergeRequest (execute.ts) build what each request sends; admission prices the same objects, so what is priced and what is sent cannot drift.
  • src/analysis/responses.ts: the response check (moved unchanged from execute.ts), recordProviderUsage (now over several responses) and the Codex response warnings, so execute.ts holds the prompts, admission and the request loop.
  • The instructions are split into named paragraphs; the merge instructions reuse eight of them. The single-request instructions are byte-identical (checked by rejoining them against origin/main). ANALYSIS_PROMPT_VERSION stays study-evidence-8: no earlier study-evidence-8 artifact has cohorts, and a 16-participant run's request is unchanged, so its reuse key still holds.

Admission, plan time, failures, schema

  • Admission prices every request before the first is sent: each cohort request as before, and the merge request with its instructions, schema and packet plus each cohort report at that request's whole output allowance; it is expected to write what one request covering every participant would. The post-hoc check compares the summed bill with the summed worst case and the cap, and any response's output with the allowance.
  • analysisCostRange (the study check range) sums the same requests. The plan line names them past 16: at a $20 cap, expected $5.14 to $11.24 for 24 participants in 2 cohort requests of at most 16 participants and one merge request, depending on how much evidence the run keeps. At the default $3 cap the analysis of 17 or more participants is refused even with no evidence, so a default analysis of them is skipped with AUTOMATIC_ANALYSIS_ADMISSION_REFUSED.
  • A cohort or merge request that fails or is rejected fails the attempt with that request's code: no further request starts (requests already sent finish and are counted), no merge request after a cohort failure, no findings, and a warning names the cohort's size. The usage counts every request sent. A report on part of the panel is not kept: the findings view and the Observer have no state for a participant who was not analysed.
  • humanish.study-analysis.v1 is unchanged in shape. Its stored bounds grew to hold eight cohorts: 6,400 evidence items, 320 captures, 16 MiB analysis.json, and validateAnalysisEvidence allows 20 MiB of images per cohort.

Checked

  • Red first: 9 tests in tests/analysis/cohorts.test.ts failed on origin/main (17 participants packed to 16, one request, the old cost signature). Later additions, each red on the commit before it: the 20 MiB image bound on validation, the plan line, and the two merge-citation cases (a capture cited only in a participant review, and one cited nowhere). Coverage tests for existing behavior: a fifth cohort that never starts after a failure, a failed merge request, a cancellation before the merge, a summed bill over the summed worst case, and Codex requests one at a time (mutation-checked: the test fails with Codex at four at a time). tests/analysis/cohorts.test.ts has 17 tests; values come from worked examples (cost: two cohorts and a merge, $4.6585 expected, $6.0737 worst; captures per participant: 4 to 5 of 9 for 17, at least 2 of 3 for 40) and from the captured wire envelope's usage.
  • The participants-cap tests that pinned the interim AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT skip now check that 17 participants are planned, named in the study check line as 2 cohorts, and analysed after a live fake-desktop run; the explicit-request skip test is removed, since no study reaches 128 participants.
  • pnpm vitest run tests/analysis/: 24 files, 504 tests.
  • Live with OpenAI, on a copy of the participants-cap session's 24-participant computer-use run (cua-2026-10-09T03-17-16-059Z-0654b949, example.com):
    • analyze --dry-run --max-cost 30: admitted at $6.689375 (expected $6.08, worst $6.96). --max-cost 7 is that worst case rounded up, the cap humanish's refusal suggests.
    • Run 1 (first commit): 3 requests (cohorts of 134 and 133 evidence items with 12 captures each, then the merge), complete in 5 min 6 s, billed estimate $2.649317 (91,342 input, 31,523 output tokens, usage complete). Expected was 2.3 times the bill.
    • Run 2 (--rerun, final code and merge instructions): 3 requests, complete in 5 min 48 s, billed $1.919481 (91,838 input of which 74,654 cached, 32,601 output).
    • Run 2's report: 24 participant reviews, 2 findings (one affecting 24 of 24, one 3 of 24), 2 design findings, 4 concern reviews. All 466 citations resolve in the stored packet and all 24 capture files exist (run 1: 447, all resolve). review --json gives humanish.analysis-findings.v1 ready with 143 cited evidence items and 0 missing capture files. verify still grades the run local_only, as before analysis; the Observer refreshed without a warning.
    • Total live spend: $4.57 of the $15 budget. The analysed copies are in the private handoff folder, not in this repo.
  • pnpm release:check: exit 0 on 4e6dcab (508 files, 8,000 tests passed, 12 skipped; TUI 192 passed).

Lead review changes

  • Deleted the AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT skip from Let a study have 100 participants and run them within the E2B plan #1739: src/study/automatic-analysis-plan.ts, the plan's analysis.skip and the participant count planBase read it from (computer-use and shared-world planners), the route shell's skip branch, the CLI exit rule's skip pass, the CLI reason text, doctor's "Will not run" note and the study-files sentence. Callers go back to the analysis owner: completeAutomaticAnalysis, automaticAnalysisSucceeded, automaticAnalysisBudget and formatAutomaticAnalysisBudget (run, study check, the study summary the TUI reads, doctor). Each restored hunk matches its pre-Let a study have 100 participants and run them within the E2B plan #1739 form. git grep finds no mention of the skip.
  • The analyze --dry-run worst case now reads if every request (2 cohort requests of at most 16 participants and one merge request) writes its whole 32768-token output allowance for a cohort analysis. The words come from analysisRequestsText (automatic-config.ts), which also writes the plan line, and the count from analysisRequestCount (analysis-limits.ts), reported as the additive admission.requests in --json. One request keeps the old sentence. The analyze golden gains "requests": 1 in each admission and nothing else.
  • Tests: a CLI dry run of a 17-participant preview run (red on the previous head), admission request counts for 16, 17, 40 and a refused input.

Merge with 0.118.0 (#1735)

  • automatic-config.ts: Fix the four TUI gaps from the 0.116.0 release notes #1735's admissionRule is the one text of the rule. The participants phrase carries the requests (24 participants in 2 cohort requests of at most 16 participants and one merge request) and feeds both of Fix the four TUI gaps from the 0.116.0 release notes #1735's forms: the range, and "refused before it starts for ... even with no evidence".
  • execute.ts: Fix the four TUI gaps from the 0.116.0 release notes #1735's analysisCostRange bisection and largerOutput stay; each probe prices every cohort request at that share of its evidence limits plus the merge request, so the range ends at what admission admits for the whole analysis. With the default $3 cap, 17 or more participants are refused even with no evidence; at $10, 40 participants range from $7.05 to $10.00.
  • doctor.ts: imports admissionRule and automaticAnalysisBudget; the deleted skip module stays deleted.
  • docs/product/automatic-analysis.md, tests/cli/analyze-cost.test.ts: both sides kept.
  • CHANGELOG: ## 0.118.0 from main. The cohort entry is under a fresh Unreleased and says that 0.118.0 recorded AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT past 16 participants, which shipped in that release.

Not verified

  • Codex account analysis of more than 16 participants live; the sequential path is covered by a test with an injected Codex provider only.
  • A cohort or merge failure live, and a run of 40 or 100 participants live. The merge request's estimate rests on two live runs of one 24-participant bundle (2.3 and 3.2 times the bill).
  • The Observer and humanish watch at 40 and 100 participants (slice 4).
  • The Observer's spend.requests and stats count one analysis attempt as one request: filed as Observer spend.requests counts a cohort analysis as one request #1743 (observer/ belongs to another session).

🤖 Generated with Claude Code

danielgwilson and others added 4 commits October 9, 2026 03:56
One analysis now covers up to 128 participants. captureEvidence selects
each cohort of at most EVIDENCE_LIMITS.cohortParticipants (16, now in
analysis-limits.ts with the other packet limits) under the full packet
limits, and runAnalysis sends one request per cohort, four at a time,
then one merge request over the cohort reports. The merged report keeps
the cohorts' participant reviews and is checked by the same validator
against the evidence the cohort reports cite. Admission, the expected
cost, the worst case and the study check range sum every request. A
failed request fails the attempt with no findings and a warning.

Response checking and usage moved to responses.ts so execute.ts keeps
the prompts, admission and the request loop. Stored limits grew to hold
eight cohorts. EVIDENCE_LIMITS.participants is now 128, the stream limit
the source parser already had.

Checked: 11 tests in tests/analysis/cohorts.test.ts (17- and 40-participant
bundles, admission worked example, merge through an injected fetch);
full vitest suite; one live analysis of a 24-participant run (3 requests,
expected $6.08, billed $2.65, 447 citations resolved).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The merge request saw no evidence, so its findings, design findings and
concern reviews may now cite only evidence the cohort reports cite in
theirs; a capture cited only in a participant review was accepted
before. checkMergedResponse in cohorts.ts owns that check and replaces
citedInput and mergedResponse.

cohortRequest and mergeRequest build what each request sends, and
admission prices the same objects, so the two cannot drift. The merge
instructions now say how to stay under 30 observations per finding. The
plan line counts cohorts of at most 128 participants. A cohort input no
longer recomputes a digest nothing reads.

Tests: 17 in tests/analysis/cohorts.test.ts. The two merge-citation cases
fail on the previous commit. New coverage for a fifth cohort that never
starts, a failed merge request, a cancellation before the merge, a summed
bill over the summed worst case, and Codex requests one at a time.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ticipant

With EVIDENCE_LIMITS.participants at 128, a study of 17 to 100
participants gets its automatic analysis, so the four tests that pinned
the interim AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT skip now check that the
analysis is planned, named in the study check line and requested after a
live run of 17. The skip itself, in src/study/automatic-analysis-plan.ts,
is left for its owner: no study reaches 128 participants, so it no longer
fires. Its Unreleased CHANGELOG entry is replaced by the cohort entry,
which now says a default analysis of 17 participants is over the $3 cap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
humanish Ignored Ignored Preview Oct 9, 2026 5:26am UTC

Request Review

With EVIDENCE_LIMITS.participants at 128 and a study capped at 100, the
AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT skip from #1739 could not fire.
src/study/automatic-analysis-plan.ts is deleted, and its callers go back
to the analysis owner: the route shell calls completeAutomaticAnalysis,
the CLI exit rule and the run envelope use automaticAnalysisSucceeded,
run, study check, the study summary the TUI reads and doctor use
automaticAnalysisBudget and formatAutomaticAnalysisBudget, and the plan
no longer carries a skip or the participant count it was read from.
The doctor note, the CLI reason text and the study-files sentence go too.

The analyze dry run's worst case now says every request writes its whole
allowance and names the cohort and merge requests, in words from
analysisRequestsText, which also writes the plan line. The count is
analysisRequestCount in analysis-limits.ts; admission reports it as
admission.requests (additive in --json; the analyze golden gains it).

Checked: no code, test or doc mentions the skip (git grep). New tests:
the dry run of a 17-participant preview run names 2 cohort requests and
one merge request (red on the previous commit), and admission counts 1,
3 and 4 requests for 16, 17 and 40 participants, null when refused.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Conflicts with #1735, resolved to keep both:
- automatic-config.ts: admissionRule words the rule; the participants
  phrase ("24 participants in 2 cohort requests of at most 16
  participants and one merge request") feeds both the range form and
  the refused-even-with-no-evidence form.
- execute.ts: the request builders stay; analysisCostRange keeps #1735's
  bisection and largerOutput, and each probe prices every cohort request
  and the merge request at that share of the evidence limits.
- doctor.ts: imports admissionRule and automaticAnalysisBudget; the
  deleted skip module stays deleted.
- automatic-analysis.md, analyze-cost.test.ts: both sides' text and
  tests kept.
- CHANGELOG: 0.118.0 from main; the cohort entry under a fresh
  Unreleased, saying 0.118.0 skipped analysis past 16 participants.

Tests: the plan-line tests now expect the refused form at the default $3
cap and the range form at $20; a new test checks the 40-participant
range ends at a $10 cap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
danielgwilson and others added 2 commits October 9, 2026 05:22
# Conflicts:
#	CHANGELOG.md
#	src/routes/shared-world/plan.ts
#1745 deleted src/run/concurrency.ts when its last caller moved to the
arrivals schedule; #1742's cohort requests call it. Restored unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@danielgwilson
danielgwilson merged commit 6405a7c into main Oct 9, 2026
10 checks passed
@danielgwilson
danielgwilson deleted the feat/analysis-past-16 branch October 9, 2026 05:35
@danielgwilson danielgwilson mentioned this pull request Oct 9, 2026
danielgwilson added a commit that referenced this pull request Oct 9, 2026
* Release 0.119.0

Bump the version to 0.119.0, move the Unreleased notes into the release
notes, and keep the CHANGELOG entry's title, opening paragraph and
release link. pnpm docs:generate writes 0.119.0 into cli.mdx.

Contents: #1745 and #1742, cut from main 6405a7c.

Checked: a live smoke on a tarball packed from main 6405a7c runs
alongside this commit (try-live with analysis, 24 participants arriving
10 s apart with their cohort analysis, a late participant past the study
budget, analyze --run on a 24-participant bundle). pnpm release:check
runs on this commit next.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add the 0.119.0 Taskly benchmark result

pnpm bench with its defaults (3 runs per arm, openai-computer-use,
neutral mission) on the release commit 53dbf23, as
docs/release/publish.md asks for a minor release: report recall 10/15,
analysis recall 10/15, 0 invented on either arm, $6.12 estimated.
0.118.0 scored 6/15 and 6/15 at $5.38.

D1 and D5 account for the rise: two planted reports and two planted
analyses name each, where none did on 0.118.0. D3 and D4 stayed 3/3.
Two unresolved planted lines describe D5 (planted 1's report, planted
3's analysis), so recall read by hand is 11/15 and 11/15. No planted
participant typed more than 27 characters in one action, so none met
D2. The analysis cap was the dry run's admittedCostUsd, $1.25, against
bills of $0.60 to $0.86; none was refused or skipped. Every participant
ran threaded, ended on its own and gave 5 or 6 impressions. E2B listed
0 tagged sandboxes after it.

Issue #1749 reads the decline from 0.114.0 to 0.118.0 as noise at three
runs per arm (same-hour A/B of npm 0.114.0 and main d183d1a: 52/80
each). The README paragraph states that finding without the issue
number, since the evidence prose check caps issue references at 102.

Checked: the summary's claims to check by hand, against the results
file and the two unresolved findings' text in the analyses; typed
lengths from the traces' action titles; node
scripts/check-code-prose.mjs.
Not checked: which controls planted participants clicked, read from
trace coordinates as #1749 does; three runs per arm cannot separate the
rise from run-to-run variation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Note that the 0.119.0 benchmark ran before the warning fix

The benchmark ran on the release commit 53dbf23. The branch then
merged #1748, which redacts study warnings before the run records them.
That changes what the bundle records and nothing a participant receives
or does, so the result stands for 0.119.0; the README paragraph says so.

Checked: node scripts/check-code-prose.mjs (no count rose).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@danielgwilson danielgwilson mentioned this pull request Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant