Skip to content

Count, filter and page the Observer grid and bound watch output at 100 participants - #1744

Merged
danielgwilson merged 3 commits into
mainfrom
feat/observer-100-participants
Oct 9, 2026
Merged

danielgwilson merged 3 commits into
mainfrom
feat/observer-100-participants

Conversation

@danielgwilson

Copy link
Copy Markdown
Owner

Slice 4 of #1737: the Observer grid and humanish watch at 40 and 100 participants. No humanish.observer-data.v1 change.

Measurements

Fixture: observer/tests/scale-fixture.ts builds a finished computer-use run of N participants with the producer (buildObserverData) from the recorded one-participant run in tests/golden/labs/live.json (50 turns each, own screenshot folder each). Personas rotate over 3 ids; every ten participants end as 6 reported complete and 1 each blocked, gave up, interrupted, failed. Headless Chromium, built artifact served over loopback. The measurement script is scratch and not committed.

40 @1440 40 @390 100 @1440 100 @390
artifact with data, before / after 3,329,516 / 3,331,001 B same 6,538,638 / 6,540,123 B same
observer-data.json 2.14 MB 2.14 MB 5.35 MB 5.35 MB
page 1 painted, median of 5 alternating loads, before / after 756 / 731 ms 533 / 493 ms 763 / 726 ms 701 / 691 ms
grid page 1 36 cards 36 36 36
paging 36 + 4 36 + 4 36 + 36 + 28 36 + 36 + 28
status counts above the grid, before / after none / 24 · 4 · 4 · 4 · 4 same none / 60 · 10 · 10 · 10 · 10 same
persona, before / after search text only, 13 / select, 13 same search only, 33 / select, 33 same
status filter Blocked 4 cards 4 10 10
pin on last page moves to page 1 same same same
comparison of 3 opens opens opens opens
report (analysis citing 360 / 900 entries) opens opens opens opens
first card after Next page from the bottom, before / after in view (page 2 is 4 cards) 9877 px above / in view 5389 px above / in view 9877 px above / in view

The timing rows come from one alternating before/after run on a shared, loaded box. An earlier unloaded run measured 316-486 ms before the change. Before and after are within noise. Served polling at 100 participants costs about 300 ms of long tasks per 20 s (four 5 s polls of the 5.35 MB JSON) before and after.

humanish watch <study> --dry-run with 100 participants. The before column was measured with the cap patched locally to 100 and then reverted. The after column was measured on main with the cap at 100 (#1739), unpatched:

before after
stdout lines 220 122
of which participant lines 100 2 (count line, not listed:)
of which persona-background warnings 100 100 (src/study, not in this PR)
stderr lines (plan, one per participant) 103 103 plus 2 lines of the Observer artifact build (participant-runs.ts, not in this PR)

A live-shaped result (tests/cli/computer-use-participants-output.test.ts) prints 31 stdout lines at both 40 and 100 participants. Before this PR the same 100-participant result printed one participant line per participant with its whole closing message, plus 100 copies of the screenshots warning.

What changed

  • observer/lib/grid-filters.ts is a new module that owns the grid's filter fields. Its interface is GridFilters (with persona optional, so a preference saved before persona existed still reads), NO_FILTERS, isGridFilters, filterParticipants, activeFilterCount and statusCounts. The default, the storage guard and the predicate move out of app.tsx, and the badge count and clear literal move out of grid-options.tsx. Before this PR a new field touched those five places.
  • GridStatusSummary (in components/grid-options.tsx) adds one toggle button per status label with its count, shown when participants ended more than one way. Pressing a count filters the grid, and pressing it again clears the filter. It counts statusLabel, the key the status filter matches, so a count equals the cards it shows. StudyGrid renders it through a new statusSummary slot.
  • A Persona select in the view options, shown when the run has more than one persona. It reads stream.sim.personaId, which the contract already carries.
  • Paging: the pager scrolls the grid section to the top of its scroller, and a filter change returns to page 1. Windowing was not needed: 36 cards per page render in under 0.5 s, and paging already worked.
  • formatCuaStudyHuman (src/cli/commands/study-format.ts), the output of watch and run:
    • Up to 16 participants print one line each, as before.
    • Above 16, it prints one count line by status, the first 16 participants that did not pass, and a not listed: line.
    • Every participant line keeps its closing message on one line of at most 160 characters.
    • warningLines prints a repeated warning once, in all four route formatters.
    • --json output is unchanged.

Checks

  • Red first, through the rendered App in jsdom: observer/tests/scale.test.tsx, 12 tests at 40 and 100 participants. Before the change, the status row, persona and page-reset tests failed (6 of 11 at that point). Mutations caught: page reset removed, persona clause removed, persona required by the guard.

  • tests/cli/computer-use-participants-output.test.ts, 6 tests, red before the change:

    • five on results built from tests/golden/routes/computer-use-fanout-live.json;
    • one through humanish watch <study> --dry-run at 40 participants.

    Mutations caught: the bound raised to 17 or to 100, the dedupe removed.

  • New browser case grid-pages-phone (80 participants at 390 px): it fails without the scroll and passes with it.

  • pnpm --filter humanish-observer typecheck, then the observer suite after a build: 491 passed.

  • pnpm build, then the four observer proofs passed: iframe; browser, 74 cases; chrome, 4; reliability, 2. The observer suite and the proofs last ran on this branch rebased onto 41a72af. The two later rebases (Let a study have 100 participants and run them within the E2B plan #1739, Fix the four TUI gaps from the 0.116.0 release notes #1735) changed nothing under observer/.

  • Artifact: 1,185,473 to 1,187,030 bytes before data, against a 1,500,000-byte budget.

  • pnpm release:check passed on the pushed head: vitest 7987 passed and 12 skipped, TUI 192 passed.

  • vocabulary.lane fell from 91 to 90, so this PR lowers its cap.

Needs attention

  • The persona-background warning per participant (src/study/warnings.ts) and the plan line per participant on stderr (src/routes/computer-use/participant-runs.ts) still grow with the roster. Both files belong to the participants-cap session (A study can't simulate a full day: computer-use studies stop at 16 participants #1737 slice 1), which was told.
  • The shared-world formatter still lists every participant on its own line.
  • A live run over 16 participants records its analysis skip (AUTOMATIC_ANALYSIS_PARTICIPANT_LIMIT, Let a study have 100 participants and run them within the E2B plan #1739) only in the CLI result. withParticipantLimitSkip writes no automatic job record, so the Observer of that run has no Findings view and no sentence saying why. I found this by reading src/run/route-shell.ts; no run was made. The Observer has no detail text for that code either. The cohort analysis of slice 3 removes the skip.

Not verified

  • No live run and no real 40- or 100-participant bundle. The fixtures replicate one recorded participant.
  • The before watch measurement at 100 used a locally patched cap, because the cap was 16 on main at the time.
  • The browser timing comes from one shared box under load from other sessions.
  • Served polling of a 5 MB observer-data.json every 5 s was measured, not changed.
  • The report was checked with a synthetic analysis citing up to 900 entries. The cohort analysis of slice 3 does not exist yet.

🤖 Generated with Claude Code

danielgwilson and others added 2 commits October 9, 2026 04:12
Measured first with a 40- and a 100-participant run built by the producer
from tests/golden/labs/live.json (observer/tests/scale-fixture.ts), in
Chromium at 1440 and 390 px: the grid, paging (36 a page), status filter,
pins, comparison and report all worked, and page 1 painted in 316-486 ms.
Three things were hard to use:

- Next page from the bottom of a page kept the scroll position, so page 2
  opened at its end, 9877 px past its first card at 390 px.
- Reaching the blocked participants took the options popover and a select.
- A persona could only be found by typing its id into search.

Changes:

- observer/lib/grid-filters.ts is the one module for the grid's filter
  fields: GridFilters (persona added, optional so a preference saved
  before it still reads), NO_FILTERS, isGridFilters, filterParticipants,
  activeFilterCount and statusCounts. The default, guard and predicate
  leave app.tsx; the badge count and clear literal leave grid-options.tsx.
- GridStatusSummary: one toggle per status label with its count, shown
  when participants ended more than one way. It counts statusLabel, the
  key the status filter matches, so a count equals the cards it shows.
- A Persona select in the view options when there is more than one
  persona, reading stream.sim.personaId. No observer-data change.
- The pager scrolls the grid section to the top of its scroller, and a
  filter change returns to page 1.

Checked: observer/tests/scale.test.tsx (12 tests through the rendered App
at 40 and 100) was red for the status row, persona filter and page reset
before the change. Mutations caught: page reset removed, persona clause
removed, persona required by the guard. The new grid-pages-phone browser
case fails without the scroll and passes with it. Observer suite 491
passed; the four observer proofs passed (browser 74 cases). Artifact
1,185,473 -> 1,187,030 bytes, budget 1,500,000.

Not checked: a real 100-participant live bundle; polling cost is
unchanged (about 75 ms of long tasks per 5 s poll at 100).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`humanish watch` (and `run`) printed one line per participant, each with
the participant's whole closing message, and one copy of the screenshots
warning per participant. Measured with a 100-participant dry run (cap
patched locally, then reverted): 220 stdout lines.

formatCuaStudyHuman now builds its participant section with
participantLines: up to 16 participants, one line each as before; above
16, a count line by status, the first 16 that did not pass, and a
`not listed:` line. Every participant line keeps its closing message on
one line of at most 160 characters. warningLines prints a repeated
warning once, in all four route formatters. --json output is unchanged.

The same dry run now prints 122 stdout lines, on main with the cap at 100
(#1739). 100 of them are the persona-background warning per participant
from src/study/warnings.ts, and stderr keeps the plan's line per
participant from participant-runs.ts; both files belong to the
participants-cap session, which was told. A live-shaped 100-participant
result prints 31 lines, the same as 40.

Checked: tests/cli/computer-use-participants-output.test.ts (6 tests:
five on results built from tests/golden/routes/computer-use-fanout-live.json,
one through `humanish watch <study> --dry-run` at 40 participants) was
red before the change; mutations caught for the 16 bound and the dedupe.
vocabulary.lane fell 91 -> 90, cap lowered.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
humanish Ready Ready Preview Oct 9, 2026 7:04am UTC

Request Review

@danielgwilson
danielgwilson merged commit 474563a into main Oct 9, 2026
10 checks passed
@danielgwilson
danielgwilson deleted the feat/observer-100-participants branch October 9, 2026 07:13

This branch was successfully deployed

1 active deployment
Preview — 7d22ee78 Deployed Oct 9, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant