Repository navigation
Conversation
🦋 Changeset detectedLatest commit: b01ea98 The changes in this PR will be included in the next version bump. This PR includes changesets to release 2 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
# Conflicts: # packages/agents/src/tasks/tasks.ts # packages/agents/src/tests/capabilities/tasks.ts
agents
@cloudflare/ai-chat
@cloudflare/codemode
hono-agents
@cloudflare/shell
@cloudflare/think
@cloudflare/voice
@cloudflare/worker-bundler
commit: |
🔴 agents import sizesMeasured 346 runtime imports as minified bundles. The primary size is gzip; raw minified size is included for diagnosis. An existing import growing by more than 10% is marked red. This report is informational.
Compared Changed imports (60)
All 346 current runtime imports
Reported by agent-think[bot]. |
# Conflicts: # packages/think/src/think.ts
There was a problem hiding this comment.
🔴 Slow routed failures strand pending runs
When routed onJob dispatch exceeds five seconds and later fails, it completes the wake before observing the failure. trackAlarmWork does not retry ordinary RPC failures, leaving an unclaimed facet run without a root wake.
(Refers to this code)
Learn more
A routed run lives on a facet, but its alarm job lives on the root. The root races the dispatch RPC against a five-second timer. When the timer wins, the root returns an empty job outcome; applyOutcome deletes the job unless the facet has already pushed a newer wake. If the RPC then fails before claiming the run, trackAlarmWork observes the rejection but only applies recovery for memory-limit resets. The pending run remains durable without a wake.
Example: Root dispatch to a cold facet takes six seconds during startup, then the facet's startup hook rejects. The five-second budget deletes its root wake. No claim or replacement wake was written, so the pending run never starts.
Recommended fix: Preserve or defer the root job until the facet confirms it has claimed the run or pushed its own successor wake. For a detached RPC failure, explicitly restore a wake from the authoritative facet state; do not treat a timed-out dispatch as completed solely because its RPC remains outstanding.
Was this helpful? React with 👍 or 👎 to provide feedback.
| Agent restores persisted facet routing identity synchronously during | ||
| construction, before these capability hooks run. Its host startup phase then | ||
| hydrates facet connection state without changing the route identity capabilities | ||
| already observed. |
There was a problem hiding this comment.
| ## Public API | ||
|
|
||
| ```ts | ||
| await tasks.sendEvent( | ||
| runId, | ||
| "approval", | ||
| { approved: true }, | ||
| { idempotencyKey: "approval:1" } | ||
| ); | ||
| ``` |
Summary
Add durable external events to
agents/tasks.Tasks.sendEvent()sends run-scoped, optionally idempotent events.step.waitForEvent()waits for one event, indefinitely or with a timeout.step.takeEvents()drains currently buffered events in bounded FIFO batches.recover root alarm ownership across eviction.
Input timing
The run mailbox supports input regardless of handler timing.
If a run is already waiting for an event, delivery wakes it immediately. If
the run is pending or actively working, the event remains buffered until a
later
waitForEvent()ortakeEvents()step consumes it.New input must still target a known, non-terminal run. This is not a
general-purpose inbox. An identical idempotent retry may recover a retained
receipt after the run settles.
Example
Architecture
Consumption and journaling either commit together or do not commit. After
process loss, replay returns the journaled event instead of consuming another
mailbox entry.
Facet startup ordering
Task runs may live on routed sub-agent facets, but only the root Lifecycle owns
the physical alarm. On a cold start, Lifecycle starts capabilities before the
Agent's asynchronous host-startup hook. Previously that hook also restored
persisted facet identity, so Tasks could temporarily see no route source,
mistake the facet for an alarm-owning root, and reconcile its wake on the wrong
object. Because a facet cannot drive a physical alarm, that ordering could
leave work stranded after eviction.
The fix is deliberately split across two integration points:
dynamic-agents.tsrestores the persisted facet marker, name, and parentpath synchronously.
index.tscalls that restoration from the Agent constructor before routetransport and Lifecycle capability startup.
Virtual WebSocket connection hydration remains asynchronous because Tasks only
needs routing identity at startup. The final implementation reads the marker
once for roots and new agents, and reads all three identity keys once for cold
facets. This changes no dynamic-agent route format or persisted key name.
Behavior
nullon expiry.takeEvents()defaults to 100 events and allows at most 1,000.TaskEventIdempotencyConflictError.retrying the same retained event and idempotency key returns the original
receipt with
accepted: false.null.task:${runId}. Routed wakes use a separate internalnamespace and preserve stable identity while refreshing transport data.
Correctness
sendEvent()resolves, but may also havecommitted when later wake synchronization rejects.
JavaScript memory.
facet identity before Tasks reconciles wakes, then coalesces alarm rearming.
Storage
Adds
cf_agents_task_eventsfor durable mailbox entries.Write amplification and intended use
The mailbox is optimized for durable approval and message delivery, not as a
high-throughput message queue. Cloudflare bills writes to indexed columns as a
table-row write plus an additional row write for each affected index. An
individually consumed event therefore costs approximately:
That is roughly seven row writes per individually consumed event before
AUTOINCREMENTmetadata and run/wake bookkeeping. Batching amortizes thejournal write, but does not remove event and index writes.
Under the current Durable Objects SQLite pricing,
the Workers Free plan allows 100,000 rows written per day. Seven writes per
event gives a theoretical ceiling of roughly 14,000 individually consumed
events per day if nothing else writes; real Tasks workloads have less headroom.
The Workers Paid plan includes 50 million rows written per month and then
charges $1 per million rows, or roughly $7 per million individually consumed
events beyond the included allocation.
The indexes avoid full mailbox scans and keep exact-type FIFO lookup bounded.
That tradeoff is reasonable for the intended approval/message workload, but
callers needing event-stream throughput should use a queue-oriented primitive.
The existing step journal gains
wait_eventandtake_eventsstep kinds plusevent_typefor replay validation. Existing v1 step rows migrate withoutmodification. Rolling back to older SDK code after writing event steps is
unsupported.
Deep dive
See Tasks event mailbox
for the complete storage model, delivery and wait flows, replay behavior, and
module responsibilities.
Testing
failures are
agent-think/tsconfig.json,agent-think/tests-e2e/tsconfig.json, andexamples/context-overflow-recovery/tsconfig.json.The SIGKILL coverage verifies both an indefinite wait surviving restart and a
consumed event replaying from its journal after another process loss.