Skip to content

feat(agents): peer-to-peer messaging and idle notice (#16948, #16949) - #16996

Merged
mrveiss merged 4 commits into
mainfrom
issue-16948-16949-peer-messaging
Sep 19, 2026
Merged

mrveiss merged 4 commits into
mainfrom
issue-16948-16949-peer-messaging

Conversation

@mrveiss

@mrveiss mrveiss commented Sep 18, 2026 •

Copy link
Copy Markdown
Owner

Thinking Path

#16947 gave every agent kind a live, named presence entry with a busy/idle
signal but no way to reach another entry directly -- a peer could only be
observed, not addressed. #16948/#16949 add that address book and the two
things it needs to be safe: delivery that respects tenancy and turn
boundaries, and a push notice for the busy->idle transition so a waiter
doesn't have to poll presence.

The two hard calls, both made explicit in the assignment and carried through
here:

  • A peer message is never human approval. protocols/peer_inbox.py
    delivers into a per-(kind, tenant_id, name) inbox with no side channel
    into any approval path. A sensitive tool call a peer message triggers goes
    through the exact same gate a human-triggered one does --
    AI_STACK's enforce_work_item_approval (already unconditional on every
    tool call, so it needed no new gating code) and SESSION's own
    approval_handler/assess_command_risk, unchanged.
  • Delivery happens at the real per-turn boundary, not mid-task. The drain
    call sits at the one seam each kind's production callers all share:
    AI_STACK's chat_workflow/tool_handler.py::_dispatch_tool_call and
    SESSION's services/agent_terminal/service.py::execute_command -- the same
    seams mapped and confirmed for feat(agents): live presence — one registry of named agents with busy/idle state, across all three kinds #16947's busy/idle wiring, reused rather than
    re-derived. A message that arrives after one drain sits queued until the
    next.

protocols/idle_notice.py (#16949) is built on the existing events/bus.py
transport rather than a new one: AgentPresenceRegistry.report() now
publishes EVT_AGENT_IDLE on exactly a busy=True -> busy=False transition,
and wait_for_idle() is the one-shot subscriber -- resolves immediately for
an already-idle target, otherwise waits for one matching event or a bounded
timeout (AUTOBOT_IDLE_NOTICE_TIMEOUT_SECONDS), always unsubscribing.

What Changed

  • protocols/peer_inbox.py (new): PeerMessageEntry, PeerInbox (deliver/
    drain), PeerInboxDirectory.send() -- addresses by the presence registry's
    own (kind, name), resolves the recipient's real tenant from the matched
    PresenceEntry (never from the caller), refuses when the name isn't live
    (RecipientNotAddressableError); EXTERNAL and UNKNOWN_TENANT are never
    reachable because list_live() never surfaces them.
  • chat_workflow/tool_dispatch_guards.py: new enforce_peer_messages(ctx) --
    drains the chat AI_STACK inbox into ctx.context["peer_messages"], a
    no-op with no ctx.
  • chat_workflow/tool_handler.py: _dispatch_tool_call calls the new guard
    first, alongside the existing six unconditional guards.
  • services/agent_terminal/service.py: new _drain_peer_messages(session),
    called at the top of execute_command before command assessment; keys the
    inbox lookup by session.tenant_id or UNKNOWN_TENANT, mirroring
    sync_session_presence's own derivation exactly (a raw None key would
    silently miss anything send() actually authorized).
  • protocols/idle_notice.py (new): notify_agent_idle() (fire-and-forget
    publish, safe no-op with no running loop), wait_for_idle(),
    IdleWaitExpiredError.
  • protocols/agent_presence.py: report() now calls notify_agent_idle()
    on a busy: True -> False transition only.
  • autobot_shared/env_registry_agent_runtime.py: registers
    AUTOBOT_IDLE_NOTICE_TIMEOUT_SECONDS (default 300s); docs/developer/ENV_VARS.md
    regenerated from the registry.
  • New tests: protocols/peer_inbox_test.py (9), protocols/idle_notice_test.py
    (9), chat_workflow/peer_messages_dispatch_16948_test.py (4, including the
    AC4 mid-task-vs-next-boundary case), services/agent_terminal/peer_messages_dispatch_16948_test.py
    (5, including an execute_command-level integration test and the
    UNKNOWN_TENANT key-matching regression).
  • scripts/python_file_size_known_large.py / repo_tests/python_file_size_ratchet_baseline.py:
    services/agent_terminal/service.py ceiling lowered 956 -> 951 (comment
    trimming offset the new drain call and helper).

Verification

PYTHONPATH="..:.:$PYTHONPATH" python3 -m pytest autobot-backend/protocols/ \
  autobot-backend/chat_workflow/peer_messages_dispatch_16948_test.py \
  autobot-backend/services/agent_terminal/peer_messages_dispatch_16948_test.py \
  autobot-backend/services/agent_terminal/running_command_task_busy_signal_16947_test.py \
  autobot-backend/chat_workflow/manager_is_processing_16947_test.py \
  autobot-backend/services/agent_terminal/session_manager_tenant_16975_test.py -q
# 63 passed
  • python3 scripts/check_python_file_size.py --audit-ceilings -- 5563 files
    scanned, 488 grandfathered, all live and at size.
  • python3 pipeline-scripts/check_env_var_registry.py -- exit 0.
  • python3 -m black --check, isort --check, flake8, bandit -q on every
    changed/new file -- all clean.
  • bash pipeline-scripts/detect-hardcoded-values.sh on every changed/new file --
    SSOT Coverage: pass (new=0, ...).
  • detect-secrets scan (no baseline write) on every changed/new file -- no
    findings; .secrets.baseline unmodified.
  • Local jscpd@5.0.6 reproduction of the duplication-guard's exact invocation:
    11718 duplicated lines, under the 11723 pin, unchanged from the pre-PR
    baseline -- no new duplication.
  • Verified wait_for_idle()/notify_agent_idle() end to end against the real
    EventBus (not mocked): caught and fixed a real bug where the in-process
    listener payload is nested as {"type", "payload"} by
    EventManager.publish, which the first draft's listener didn't unwrap and
    would have silently never matched in production.
  • Pre-push hook re-ran the touched-test set and confirmed green before this
    push.

Risks

Model Used

Claude Sonnet 5 (claude-sonnet-5)

Issue Link

Refs #16948

Changelog fragment

  • Added changelog/unreleased/16948-16949-peer-messaging.md

Checklist

  • Tests added/updated
  • Docs regenerated where applicable (ENV_VARS.md)
  • No hardcoded values, no secrets, no new duplication (see Verification)
  • File-size ratchet respected (lowered, never raised)

🤖 Generated with Claude Code

Single-issue rationale

This PR closes no issue yet. #16949 (the idle-notice half) has been split
out into #17015, which lands independently. #16948's own send side is
blocked by #16986 (protocols/agent_communication's consolidation, per
the review's "consolidate, never fork" ruling on peer_inbox.py) --
tracked as a blocked_by edge on issue #16948. This PR stays on Refs
until #16986 lands and an end-to-end test through the real transport
passes.

A peer agent can now reach another by its stable presence name
(protocols/peer_inbox.py) instead of only through a human: delivery is
scoped by tenant (never crosses tenants, and UNKNOWN_TENANT / EXTERNAL
identities are never addressable) and drained at each kind's real
per-turn boundary, never mid-task -- AI_STACK's
chat_workflow/tool_handler.py::_dispatch_tool_call and SESSION's
services/agent_terminal/service.py::execute_command, the same seams
mapped for #16947's busy/idle signal. A peer message carries no
approval of its own: a sensitive tool call it triggers still goes
through the existing human-approval gate unchanged.

protocols/idle_notice.py adds the other half of #16949: a one-shot
notice the moment a busy peer's presence transitions to idle, built on
the existing EventBus (events/bus.py) rather than a new transport, with
a bounded wait (AUTOBOT_IDLE_NOTICE_TIMEOUT_SECONDS) so a waiter is
told explicitly on expiry instead of blocking forever.
@coderabbitai

coderabbitai Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • main
  • release

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 31f26a5f-1dc7-4cff-bd11-9a47e4f1e662

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Code review on PR #16996 found that sync_ai_stack_presence (#16947)
reports every AgentHealthRegistry role -- "chat", "rag" and
"system_commands" -- as its own addressable AI_STACK presence entry,
but the only drain wired for #16948 is chat_workflow's "chat" seam.
Addressing "rag" or "system_commands" succeeded silently and then lost
the message forever, since nothing ever drains those inboxes: rag_agent.py
and system_command_agent.py run their own StandardizedAgent.process_request
flow, with no relationship to chat_workflow's dispatch seam.

PeerInboxDirectory.send() now refuses any AI_STACK name outside the
addressable set with the same RecipientNotAddressableError already used
for tenant/EXTERNAL refusals, so misaddressing fails loud instead of
silently. Filed #16997 (sub-issue of #16946) to map rag/system_commands'
own real turn boundary and lift the restriction once wired.
@mrveiss

mrveiss commented Sep 18, 2026

Copy link
Copy Markdown
Owner Author

Code-reviewer pass complete. One HIGH finding, fixed:

AI_STACK peer-message drain was hardcoded to "chat"; "rag"/"system_commands" silently lost messages. sync_ai_stack_presence (#16947) reports every AgentHealthRegistry role as its own addressable AI_STACK presence entry, but only "chat" has a drain wired at chat_workflow/tool_dispatch_guards.py::enforce_peer_messages -- rag_agent.py/system_command_agent.py run their own StandardizedAgent.process_request flow, unrelated to that seam. send() succeeded for "rag"/"system_commands" (they're live in presence) but nothing ever drained their inbox -- silent message loss, not even a refusal.

Fixed in 4d831776b: PeerInboxDirectory.send() now refuses any AI_STACK name outside {"chat"} with RecipientNotAddressableError, so misaddressing fails loud instead of losing the message. Filed #16997 (sub-issue of #16946) to map rag_agent.py/system_command_agent.py's own real turn boundary and lift the restriction once wired -- that's a separate seam-mapping exercise, the same shape as #16992's for COMPANY_OS, not a same-PR fix.

Everything else the review checked out sound: the recipient-tenant-key resolution in send(), the report() lock/notify race, and the EventManager payload-unwrap in idle_notice.py's listener (all three were things caught and fixed during development, and the review independently confirmed each is correct). One LOW/observational note (pre-existing PersistStrategy.NONE channel fan-out behavior, not introduced by this PR) needs no action here.

Re-verified after the fix: 64 tests passing, file-size ratchet clean (--audit-ceilings), duplication unchanged at 11718/11723, black/isort/flake8/bandit clean, no secrets, no new hardcoded values.

… issue-16948-16949-peer-messaging-merge

# Conflicts:
#	autobot_shared/env_registry_agent_runtime.py
@github-actions

Copy link
Copy Markdown
Contributor

Notice: 29 open PRs — past the runaway threshold (25)

There is no PR queue limit, and this is not a request to defer this PR. Work proceeds one issue at a time without a cap on open PRs; review capacity is the constraint.

This notice only means the count is high enough to be worth a glance for a runaway — something opening PRs in a loop, or a merge pipeline that has stalled so nothing is draining.

Currently open:

If the queue is draining normally, ignore this. Otherwise:

  1. Check whether CI is dispatching at all — see the ci-dispatch-watchdog status on these PRs
  2. Merge the ones whose CI has finished and review has passed: gh pr merge <number> --squash --delete-branch
  3. Look for a loop opening near-identical PRs

Warn-only runaway detector — .github/workflows/pr-queue-gate.yml. It never blocks a merge.

This was referenced Sep 18, 2026
@mrveiss mrveiss added this to the v0.9.0 milestone Sep 19, 2026
Base automatically changed from issue-16975-session-tenant to main September 19, 2026 22:10
@mrveiss
mrveiss merged commit 92f3874 into main Sep 19, 2026
19 of 31 checks passed
@mrveiss
mrveiss deleted the issue-16948-16949-peer-messaging branch September 19, 2026 22:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant