Repository navigation
Count, filter and page the Observer grid and bound watch output at 100 participants - #1744
Merged
Merged
Conversation
Measured first with a 40- and a 100-participant run built by the producer from tests/golden/labs/live.json (observer/tests/scale-fixture.ts), in Chromium at 1440 and 390 px: the grid, paging (36 a page), status filter, pins, comparison and report all worked, and page 1 painted in 316-486 ms. Three things were hard to use: - Next page from the bottom of a page kept the scroll position, so page 2 opened at its end, 9877 px past its first card at 390 px. - Reaching the blocked participants took the options popover and a select. - A persona could only be found by typing its id into search. Changes: - observer/lib/grid-filters.ts is the one module for the grid's filter fields: GridFilters (persona added, optional so a preference saved before it still reads), NO_FILTERS, isGridFilters, filterParticipants, activeFilterCount and statusCounts. The default, guard and predicate leave app.tsx; the badge count and clear literal leave grid-options.tsx. - GridStatusSummary: one toggle per status label with its count, shown when participants ended more than one way. It counts statusLabel, the key the status filter matches, so a count equals the cards it shows. - A Persona select in the view options when there is more than one persona, reading stream.sim.personaId. No observer-data change. - The pager scrolls the grid section to the top of its scroller, and a filter change returns to page 1. Checked: observer/tests/scale.test.tsx (12 tests through the rendered App at 40 and 100) was red for the status row, persona filter and page reset before the change. Mutations caught: page reset removed, persona clause removed, persona required by the guard. The new grid-pages-phone browser case fails without the scroll and passes with it. Observer suite 491 passed; the four observer proofs passed (browser 74 cases). Artifact 1,185,473 -> 1,187,030 bytes, budget 1,500,000. Not checked: a real 100-participant live bundle; polling cost is unchanged (about 75 ms of long tasks per 5 s poll at 100). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`humanish watch` (and `run`) printed one line per participant, each with the participant's whole closing message, and one copy of the screenshots warning per participant. Measured with a 100-participant dry run (cap patched locally, then reverted): 220 stdout lines. formatCuaStudyHuman now builds its participant section with participantLines: up to 16 participants, one line each as before; above 16, a count line by status, the first 16 that did not pass, and a `not listed:` line. Every participant line keeps its closing message on one line of at most 160 characters. warningLines prints a repeated warning once, in all four route formatters. --json output is unchanged. The same dry run now prints 122 stdout lines, on main with the cap at 100 (#1739). 100 of them are the persona-background warning per participant from src/study/warnings.ts, and stderr keeps the plan's line per participant from participant-runs.ts; both files belong to the participants-cap session, which was told. A live-shaped 100-participant result prints 31 lines, the same as 40. Checked: tests/cli/computer-use-participants-output.test.ts (6 tests: five on results built from tests/golden/routes/computer-use-fanout-live.json, one through `humanish watch <study> --dry-run` at 40 participants) was red before the change; mutations caught for the 16 bound and the dedupe. vocabulary.lane fell 91 -> 90, cap lowered. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This was referenced Oct 9, 2026
Merged
…ticipants # Conflicts: # CHANGELOG.md
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Slice 4 of #1737: the Observer grid and
humanish watchat 40 and 100 participants. Nohumanish.observer-data.v1change.Measurements
Fixture:
observer/tests/scale-fixture.tsbuilds a finished computer-use run of N participants with the producer (buildObserverData) from the recorded one-participant run intests/golden/labs/live.json(50 turns each, own screenshot folder each). Personas rotate over 3 ids; every ten participants end as 6 reported complete and 1 each blocked, gave up, interrupted, failed. Headless Chromium, built artifact served over loopback. The measurement script is scratch and not committed.The timing rows come from one alternating before/after run on a shared, loaded box. An earlier unloaded run measured 316-486 ms before the change. Before and after are within noise. Served polling at 100 participants costs about 300 ms of long tasks per 20 s (four 5 s polls of the 5.35 MB JSON) before and after.
humanish watch <study> --dry-runwith 100 participants. The before column was measured with the cap patched locally to 100 and then reverted. The after column was measured on main with the cap at 100 (#1739), unpatched:not listed:)A live-shaped result (
tests/cli/computer-use-participants-output.test.ts) prints 31 stdout lines at both 40 and 100 participants. Before this PR the same 100-participant result printed one participant line per participant with its whole closing message, plus 100 copies of the screenshots warning.What changed
observer/lib/grid-filters.tsis a new module that owns the grid's filter fields. Its interface isGridFilters(withpersonaoptional, so a preference saved before persona existed still reads),NO_FILTERS,isGridFilters,filterParticipants,activeFilterCountandstatusCounts. The default, the storage guard and the predicate move out ofapp.tsx, and the badge count and clear literal move out ofgrid-options.tsx. Before this PR a new field touched those five places.GridStatusSummary(incomponents/grid-options.tsx) adds one toggle button per status label with its count, shown when participants ended more than one way. Pressing a count filters the grid, and pressing it again clears the filter. It countsstatusLabel, the key the status filter matches, so a count equals the cards it shows. StudyGrid renders it through a newstatusSummaryslot.stream.sim.personaId, which the contract already carries.formatCuaStudyHuman(src/cli/commands/study-format.ts), the output ofwatchandrun:not listed:line.warningLinesprints a repeated warning once, in all four route formatters.--jsonoutput is unchanged.Checks
Red first, through the rendered
Appin jsdom:observer/tests/scale.test.tsx, 12 tests at 40 and 100 participants. Before the change, the status row, persona and page-reset tests failed (6 of 11 at that point). Mutations caught: page reset removed, persona clause removed, persona required by the guard.tests/cli/computer-use-participants-output.test.ts, 6 tests, red before the change:tests/golden/routes/computer-use-fanout-live.json;humanish watch <study> --dry-runat 40 participants.Mutations caught: the bound raised to 17 or to 100, the dedupe removed.
New browser case
grid-pages-phone(80 participants at 390 px): it fails without the scroll and passes with it.pnpm --filter humanish-observer typecheck, then the observer suite after a build: 491 passed.pnpm build, then the four observer proofs passed: iframe; browser, 74 cases; chrome, 4; reliability, 2. The observer suite and the proofs last ran on this branch rebased onto 41a72af. The two later rebases (Let a study have 100 participants and run them within the E2B plan #1739, Fix the four TUI gaps from the 0.116.0 release notes #1735) changed nothing underobserver/.Artifact: 1,185,473 to 1,187,030 bytes before data, against a 1,500,000-byte budget.
pnpm release:checkpassed on the pushed head: vitest 7987 passed and 12 skipped, TUI 192 passed.vocabulary.lanefell from 91 to 90, so this PR lowers its cap.Needs attention
src/study/warnings.ts) and the plan line per participant on stderr (src/routes/computer-use/participant-runs.ts) still grow with the roster. Both files belong to the participants-cap session (A study can't simulate a full day: computer-use studies stop at 16 participants #1737 slice 1), which was told.AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT, Let a study have 100 participants and run them within the E2B plan #1739) only in the CLI result.withParticipantLimitSkipwrites no automatic job record, so the Observer of that run has no Findings view and no sentence saying why. I found this by readingsrc/run/route-shell.ts; no run was made. The Observer has no detail text for that code either. The cohort analysis of slice 3 removes the skip.Not verified
watchmeasurement at 100 used a locally patched cap, because the cap was 16 on main at the time.observer-data.jsonevery 5 s was measured, not changed.🤖 Generated with Claude Code