Skip to content

A long wait no longer ends a participant's session - #1724

Merged
danielgwilson merged 6 commits into
mainfrom
claude/codex-wait-setting
Oct 9, 2026
Merged

danielgwilson merged 6 commits into
mainfrom
claude/codex-wait-setting

Conversation

@danielgwilson

@danielgwilson danielgwilson commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Fixes #920. A Codex participant's whole session ended when one action asked for a wait over 30 seconds. Waiting is normal participant behavior: sitting in a visit lobby, waiting for an email code, waiting for a reply. After this, a long wait no longer ends the session.

What changes

  • A wait longer than one desktop call is sent as several shorter steps, up to actor.maxWaitMs. If a step stalls, the rest of the wait is skipped and the session goes on.
  • actor.maxWaitMs is a documented study setting (1 s to 10 min). On a desktop with speech, the default stays 30 s unless the study raises it, since heard speech is only collected between steps.
  • Terminal, scripted and preview studies read no waits; a study for them that sets the field is refused, like other fields a route does not read.
  • Library callers get a RangeError for a value outside the range.
Proof
  • New tests: long waits split into steps, a stalled step skips the rest (fake timers; checks exactly one desktop call), the speech default, the Codex tool description, the range check, and the inert-field row. Each fails when its fix is reverted.
  • format, lint, knip, prose, vocabulary, docs, site-css, typecheck and build pass. pnpm test and pnpm tui:test show the same failures as the parent commit on this Mac (temp paths, local Codex and Lima).
  • Live on a hosted E2B desktop with one Codex participant (actor.maxWaitMs: 90000): the page showed a code only after 60 s. The participant asked for one 70 s wait, the desktop ran it as 30 s, 30 s and 10 s, and the participant then entered the code and finished all 3 tasks (goal_satisfied). No wait shortened, stalled or skipped notice. humanish verify: local_only (raw screenshots kept). On the parent commit the same wait is cut to 30 s with a wait shortened notice.
  • Not checked live: a desktop with speech, and the shared-world route.

🤖 Generated with Claude Code

danielgwilson and others added 2 commits October 8, 2026 14:38
#920: a Codex participant waiting in a call for another person asked
for a 60 s wait. The 30 s limit is BROWSER_CONTROL_LIMITS.waitMs: one
browser-control EXECUTE must be answered within its 35 s request
deadline, so the wire refuses a longer wait. Before #1590 the Codex
tool parser threw on it, the restricted transport failed the native run
with codex_tool_call, and the session ended with protocol_error. #1590
kept the session by shortening the wait to 30 s, so the participant
spent another model turn to keep waiting. A hosted desktop runs a wait
as `sleep` through E2B commands.run, whose default timeout is 60 s, so a
longer wait from any provider could fail there too (read from the SDK,
not run live).

The loop now owns the limit for every provider. runActionBatch shortens
a wait above the session's maxWaitMs, records a `wait shortened` notice
and tells the participant on its next request. dispatchAction sends a
wait as consecutive desktop calls of at most CUA_WAIT_LIMITS.stepMs
(30 s, the wire's limit); a stalled step skips the rest. The study sets
the limit with actor.maxWaitMs, a whole number of ms from 1000 to
600000, default 120000. The parser and the computer-use and shared-world
planners check it; terminal, scripted and preview studies refuse it.
The Codex humanish_ui schema admits any non-negative wait and its
description states the study's limit. CuaTurn.shortenedWaits and the
policy-level shortening are removed. The v2 migrate key table leaves the
new field out, so v2 refusals read as before.

Checked: tests/actors/computer-use/long-wait.test.ts (loop steps,
shortening, default, stalled step), tests/actors/codex/
participant-long-wait.test.ts (the real participant through the loop:
a 60 s wait runs as two 30 s calls with no notice; 300 s with a 90 s
limit runs three calls, notice and hint), tests/study/wait-limit.test.ts
(parse, plan, route refusals), route wiring in
tests/routes/desktop-participant-deps.test.ts, and admission-order
goldens. 30 of these fail on origin/main. format:check, lint (437, at
cap), knip, prose:check, vocabulary:check, docs:check, typecheck and
build pass. The other failures in the affected folders also fail on
origin/main on this Mac (tmp path realpath, local Codex release, Lima).

Not checked: a live Codex or hosted run. The guest's heard-speech queue
holds 8 utterances between screenshots, and a longer wait in a talking
call can overflow it; participant-media.mdx says so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review fixes for the actor.maxWaitMs change.

Speech: heard speech reaches the participant only with a screenshot, and
the guest holds 8 utterances between two screenshots; a 9th ends the
desktop session. A wait takes no screenshot, so the 120 s default made a
session-ending overflow likelier in the call scenario this change was
for. defaultMaxWaitMs(speechEnabled) now gives 30 s (one desktop call)
when the executor is speechEnabled and the study sets no
actor.maxWaitMs. LoopSession and the Codex humanish_ui description use
the same default. A study that sets maxWaitMs keeps it.

LoopSession throws a RangeError for a maxWaitMs outside 1000 to 600000
(NaN, negative, fractional), before the session starts. NaN used to turn
shortening off and a negative value ended the session on the wire.

actor.maxWaitMs on a terminal, scripted or preview study is now an
inert-field row in warnings.ts: the parser refuses it as "route: X does
not read actor.maxWaitMs" and a library caller gets the inert warning.
waitLimitValidationReason keeps only the range check and drops its
readsWaits parameter.

The stalled-step test stalled the last step, so nothing was left to
skip. It now stalls the first of three steps under fake timers and
checks that one desktop call was made.

Checked: tests/actors/computer-use/long-wait.test.ts,
tests/actors/codex/participant-long-wait.test.ts and
tests/study/wait-limit.test.ts. Each fix was reverted in turn and its
test failed. format:check, lint, knip, prose:check, vocabulary:check,
docs:check, site-css:check, typecheck and build pass. The full vitest
suite (98 failures) and tui:test (55) fail on the same tests at the
parent commit on this Mac (tmp path realpath, local Codex release,
Lima).

Not checked: a live local Codex run of a long wait.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
humanish Ignored Ignored Preview Oct 9, 2026 1:31am UTC

Request Review

danielgwilson and others added 2 commits October 8, 2026 16:01
The guest desktop proof copies a fixed set of compiled files, and the
protocol's new import of the wait limits was not among them, so the guest
could not start. The protocol keeps its own 30 s wait bound and a test
pins the loop's step to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@danielgwilson
danielgwilson merged commit 53d63f5 into main Oct 9, 2026
12 checks passed
danielgwilson added a commit that referenced this pull request Oct 9, 2026
* Release 0.117.0

Bump the version to 0.117.0, move the Unreleased notes into the release
notes, and keep the CHANGELOG entry's title, opening paragraph and
release link. pnpm docs:generate writes 0.117.0 into cli.mdx.

Contents: #1724, #1726, #1727, #1728, #1734 and #1736, cut from main
882d47f. #1733 and #1735 are held until this release merges.

Checked: a live smoke on a tarball packed from main 882d47f runs
alongside this commit (try-live with analysis, the lobby wait on npm
0.116.0 and on main, a Codex 70 s wait, a broken and a fixed
serve.build, the TUI under a pseudo-terminal). pnpm release:check runs
on this commit next.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add the 0.117.0 Taskly benchmark result

pnpm bench with its defaults (3 runs per arm, openai-computer-use,
neutral mission) on the release commit af0b19f, as
docs/release/publish.md asks for a minor release: report recall 8/15,
analysis recall 7/15, 0 invented on either arm, $6.04 estimated.

Each analysis's cap was the dry run's admittedCostUsd: caps $1.14 to
$1.23 against bills of $0.66 to $0.76; none was refused or skipped. One
unresolved planted analysis line ("Repeated clearing attempts left
completed tasks visible") is D1, which puts analysis recall at 8/15
read by hand; report recall stays 8/15. The third clean participant gave
up after a reload emptied the list, so its run reads as failed; it was
scored as the others were. Every wait the participants sent followed
another action in the same turn and lasted 500 ms, as on 0.116.0, so
the idle-turn wait length (#1736) changed no wait here: 10 waits in
all, against 14 on 0.116.0. No planted participant typed more than 30
characters in one action, so none met D2.

Checked: the summary's claims to check by hand, format:check,
prose:check, vocabulary:check and docs:check.
Not checked: whether any number moved because of a code change; three
runs per arm cannot separate it from run-to-run variation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A Codex participant's session ends when one action asks for a wait over 30 seconds

1 participant