Skip to content

Umbrella: agent-to-agent coordination alongside A2A — discovery, messaging, idle notice, shared resources #16946

Description

@mrveiss

Goal

Agents inside AutoBot can find each other, message each other by name, learn when a peer is free, and coordinate access to resources they compete for — without any agent being able to borrow another's permissions.

It complements A2A, it does not replace it. A2A (autobot-backend/a2a/) is an external, task-shaped federation protocol: an outside system submits one task and polls for one result. ACP (autobot-backend/acp/) is one editor driving one agent over stdio. Neither is a way for AutoBot's own agents to talk to each other.

Scope — decided by the owner, 2026-09-18

Which agents: all three kinds, which today have three separate, unconnected representations:

  • Company OS agents — the org-chart agents with roles and heartbeat runs
  • AI-stack role agents — the chat, RAG and system-command agents that protocols/agent_communication.py addresses by static ID
  • Live sessions — concurrent interactive or terminal sessions working in parallel

Plus resource contention: when several agents compete for the same limited resource — API rate limits, CPU time, work queues — they coordinate rather than collide. This is what turns messaging into coordination.

What exists today (audited against main)

Capability State Where
Reply addressing Exists protocols/agent_communication.py:116-125 — MessageHeader carries sender, reply_to, correlation_id
Discovery Partial llc/api/agents.py:41 lists Company OS agents with their last run status — a poll, not live, and Company OS only. Four separate registries exist, none of them live named presence
Direct messaging by name Broken → fixing BaseAgent.send_message_to_agent did not work: both channels delivered to the sender, which answered itself (#16986, fixed by #17001). Recipients are static IDs, and delivery runs the handler immediately, interrupting the recipient mid-task (#16948)
Delivery at the next turn Pattern exists, unused here agent_loop/loop.py:203-206 drains a steering inbox only at the top of each iteration — built for human steering, never for peers
Idle notification Absent nothing subscribes to "tell me when agent X is next free"
Trust boundary Partial A2A gates an external peer by trust level, but then runs every admitted peer as one identity (task_executor.py:114, agent_id="a2a-executor"), and the internal channel has no trust gate at all

Recommended home: extend protocols/agent_communication.py — it already has the envelope, the addressed request/response path and pluggable transports. Not A2A (an external, spec-governed protocol) and not ACP (single client, single agent).

Design rule that must hold throughout

A message from a peer is that peer's request, never the human's approval. A peer cannot grant another peer a permission it does not hold, and an agent refused an action must not be able to get it done by asking another agent to do it. Today that second property does not hold — see the trust-boundary child.

Children

Tracked as native sub-issues below.

Activity

  1. added this to the v0.10.0 milestone on Sep 18, 2026
  2. removed this from the v0.10.0 milestone on Sep 18, 2026
  3. added this to the v0.9.0 milestone on Sep 18, 2026
  4. mrveiss commented on Sep 18, 2026

    @mrveiss
    OwnerAuthor

    Owner direction, 2026-09-18: agent coordination is top priority. Moved with all five children into v0.9.0 and marked priority: critical.

    Build order, because three children share one foundation. Presence (#16947), next-turn messaging (#16948) and the trust fix (#16950) all need the same model of what an agent's identity is, across Company OS agents, AI-stack role agents and live sessions. Built independently they would produce three incompatible identity schemes — which is precisely how this codebase came to have four separate agent registries. So:

    1. First — the shared identity and addressing model, published on this issue. Every child consumes it; none invents its own.
    2. In parallel, the parts that do not depend on it: the inventory of existing resource primitives for feat(agents): coordinate agents competing for shared resources — API limits, CPU time, queues #16951 (leases, budget tracker, the unwired rate limiters), and a failing test for security(agents): an agent refused an action can get another agent to do it — no per-peer identity #16950 that demonstrates the permission laundering on main today.
    3. Then feat(agents): live presence — one registry of named agents with busy/idle state, across all three kinds #16947 builds presence on the model, and feat(agents): deliver peer messages at the recipient's next turn, not mid-task #16948 and feat(agents): subscribe once to be told when a peer is next idle #16949 follow it (both are blocked_by feat(agents): live presence — one registry of named agents with busy/idle state, across all three kinds #16947).
  5. mrveiss commented on Sep 18, 2026

    @mrveiss
    OwnerAuthor

    Shared agent identity and addressing model

    The foundation every child of this umbrella builds on. Presence (#16947), next-turn messaging (#16948), idle notices (#16949), the trust fix (#16950) and resource coordination (#16951) all consume it. None invents its own identity scheme — this codebase already has six overlapping places holding agent identity, and the point is to reduce that, not add a seventh.

    Code claims below were verified against main. Sections marked Review correction changed the first draft.

    1. One identity type — extend AgentIdentity, don't add a new one

    protocols/agent_communication.py:102-113 AgentIdentity is already the type MessageHeader.sender uses, and send_message already stamps it from the sender's bound identity rather than caller data (:413). Extend it additively:

    Field Meaning
    kind COMPANY_OS, AI_STACK or SESSION — new; lets one registry list all three without guessing
    name stable, addressable name — distinct from instance_id, which stays per-process
    tenant_id the company this identity exists in; None = shared platform infrastructure, not admin
    existing fields capabilities, health_status, last_heartbeat — now shared by all three kinds
    Existing representation Maps to
    AgentOrgNode (models/agent_org.py:31) COMPANY_OS, name = its restart-stable agent_id slug (unique, :64), tenant_id = company_id
    AI-stack AgentIdentity / agent_type AI_STACK, name = agent_id, tenant_id = None — these are process-global singletons serving every tenant
    AgentTerminalSession (services/agent_terminal/models.py:32) SESSION, name = session_id; stable for the session's lifetime. No tenant field exists today

    2. Naming and addressing

    • name is unique within (kind, tenant_id); platform-wide when tenant_id is None. On collision, reject and require a rename — reusing the -<hex8> suffix convention Company OS hires already use (models/agent_org.py:44-45). Never silently overwrite a live entry: that is how a rogue process would squat a trusted name.
    • Role address (kind, name) resolves through presence to whichever live instance serves that name — "the RAG agent".
    • Instance address (SESSION, name, instance_id) is exact, with no fan-out — two people's sessions must never collide.

    3. The originator, and the authorization rule

    MessageHeader gains originator — set once, on first send, and never overwritten by a relay (unlike sender, which legitimately changes each hop at :413) — plus chain, the hop names, for audit and cycle detection.

    Rule: effective permission at hop i is the intersection of every hop from the originator through i.

    effective(hop_i) = permission(originator) ∩ permission(hop_1) ∩ … ∩ permission(hop_i)
    

    Not the originator's alone: that would let a low-privilege relay act as a transparent pass-through for a high-privilege originator. Intersection is the only form under which every hop's own refusal survives the whole chain — which is what "relaying can never widen what is permitted" requires.

    This is the mechanism of the #16950 bug, confirmed in code: BaseAgent._handle_communication_request (agents/base_agent.py:419) builds its AgentRequest from message.payload.content and never reads message.header.sender — the identity is dropped at exactly the point authorization needs it. Downstream, hold_scopes claims in the recipient's name (base_agent.py:249), and A2A claims every admitted peer's scopes as agent_id="a2a-executor" (a2a/task_executor.py:114), even though the real peer_id is in scope in the same function (:75). The fix threads originator from the header into the request and makes hold_scopes' identity the intersection-computed one. A2A's external wire format does not change: X-A2A-Agent-Id stays; only the internal attribution does.

    A peer's message is never human approval. Peer entries in the inbox (#16948) are typed distinctly from human steering entries, and re-enter the ordinary sensitive-tool gate (agent_loop/loop.py:81-83, :952-953) against the intersection — or are refused when no human is present to approve.

    Review correction — intersection needs one permission vocabulary. The three kinds express permission three ways today: AI-stack agents as allowed_work/forbidden_work profiles, Company OS agents as org roles, external A2A peers as Capability. You cannot intersect a role with a work list. Scopes — the currency hold_scopes already uses — are the common vocabulary, and each kind maps its native form onto scopes at one place. #16950 owns that mapping; until it exists, intersection is undefined across kinds.

    4. Can the sender be forged?

    Today, yes. The Redis channel rpushes and blpops raw JSON (agent_communication.py:223-242), deserialised by StandardMessage.from_json (:175) with no signature, and register_agent (:632) checks only for a duplicate id — it does not ask whether the caller may hold that name.

    Gate registration. Each kind registers only from its authoritative source: Company OS from a real AgentOrgNode row, AI-stack from the fixed bootstrap list, sessions from an authenticated session-creation call — never from a caller-declared identity blob. Receivers reject a sender or originator that does not resolve to a live registered entry.

    Review correction — what the gate does and does not stop. The first draft rejected "trusted transport" because the threat is an in-network process impersonating another agent — and then chose a design that does not stop that either. Anyone who can write to Redis can publish a message whose sender is any registered agent and whose originator is any registered principal; the registry check passes, because those identities are registered.

    Separating the two threats makes the scope honest:

    5. Tenancy

    Review correction — scope by the originator's tenant, not the sender's. The first draft allowed a message only if sender.tenant == recipient.tenant or the recipient was shared. That rejects replies from shared infrastructure: a shared RAG agent (tenant_id=None) answering a Company OS agent in tenant X fails both tests.

    Rule: a message belongs to its originator's tenant. It may reach a recipient in that tenant, or a shared recipient (tenant_id=None); anything else is a hard reject, not filter-and-continue. A shared agent acting inside a chain inherits the originator's tenant for everything it does in that chain — so it can reply into tenant X for a request that began in X, and never into Y.

    Today enforcement is nowhere uniform: Company OS filters only at its API layer (llc/api/agents.py:41-50) and does not use agent_communication at all; AI-stack agents and sessions carry no tenant field. The pattern to copy already exists in resources: AgentBudgetState keys on company_id:agent_id (llc/services/agent_budget_tracker.py:125), and LLCWorkspaceLease carries company_id (llc/models/workspace_lease.py:44).

    6. What it consolidates

    Six places hold agent identity or state: AgentHealthRegistry (agents/agent_client.py:70), AgentCapabilityRegistry (orchestration/agent_registry.py:213), AgentRegistryService over models.agent.Agent, DistributedAgentManager (agents/agent_orchestration/distributed_management.py:40), AgentOrgNode, and SessionManager's in-memory session map.

    The four that answer "is this agent alive and busy right now" collapse into one live presence registry keyed by (kind, tenant_id, name), fed from each source rather than reimplemented. The static capability catalogue and the durable identity rows stay — they answer different questions (what can it do, does it exist) and are looked up by the same name. Six overlapping places become three distinct ones.

    7. What each child consumes

    Child Consumes
    #16947 presence the whole identity type; owns the consolidation in §6. Build first
    #16948 next-turn delivery name-based addressing (§2); a peer inbox entry type distinct from human steering (agent_loop/types.py:345, which has no sender field today)
    #16949 idle notices presence's live busy/idle, scoped by name and tenant; transport is the existing event bus
    #16950 trust boundary §3 and §5, plus the scope-vocabulary mapping; replaces a2a-executor
    #16951 resources (name, tenant_id) as the claim key. Shares #16950's root cause: both hold_scopes call sites attribute claims to the executor, never the originator, so "agents sharing one API key back off together" is mis-attributed until the originator is threaded

    Could not be determined from code

    • Whether shared AI-stack agents ever need per-tenant isolation beyond request-scoped context.
    • What tenant a terminal session belongs to — no field exists, and it may be implied at the API layer.

    Decisions for the owner

    Asked separately; recorded here once answered.

  6. mrveiss commented on Sep 18, 2026

    @mrveiss
    OwnerAuthor

    Correction to the identity model above — authority is not carried by hold_scopes

    Investigation for #16950 showed that §3 of the design, and one of my own review corrections, targeted the wrong mechanism. Build against this correction, not against §3 as posted.

    agents/scope_enforcement.py is work-claim mutual exclusion — a lease — not authorization. "Agent B holds scope S" means B holds a claim on a resource, so that two agents do not write the same thing at once. B writing under its own claim is that design working correctly. So:

    The authority surfaces the originator must actually carry are:

    • the approval gates — enforce_work_item_approval and requires_approval_before
    • governed-identity boundaries — build_governed_identity
    • auth_role
    • A2A's trust-level capabilities

    The intersection rule in §3 still stands; it applies to these, not to claims. The open question of a common vocabulary is therefore still open — mapping approval categories, identity boundaries, roles and A2A capabilities onto one comparable form is #16950's to design.

    What the laundering on main actually is

    Not the claim path the issue first described — delegation drops the parent's approval gates.

    _handle_delegate_tool (chat_workflow/tool_handler.py:2316) passes the child only parent_agent_id (used for a log line) and auth_role. The child is built from build_governed_identity({"agent_id": agent_type}, …), so its requires_approval_before is empty and the work_item_id is gone. A parent whose write_file is held behind the work item's "writing files" gate delegates to documentation_agent, and the child's write_file runs unapproved. auth_role is carried deliberately (#13821); the gates are not.

    It sits behind AUTOBOT_DELEGATION_ENABLED, which is off by default — which limits the exposure, not the defect.

    Other findings from the same investigation

    • A2A enforces one capability of four. Only SUBMIT_TASKS is checked (api/a2a.py:227). QUERY_MEMORY, DEFINE_AGENTS and DISCOVERY are defined in the trust matrix and checked nowhere, so a trust level that denies them denies nothing. Being filed separately.
    • A2A's a2a-executor flattening does not break claim exclusion — holding the same claim requires the same agent and the same task (work_claims.py:307), so two peers' tasks still refuse each other. What is lost is attribution and authority, not exclusion.
    • The internal peer channel holds no claims at all. _handle_communication_request calls process_request directly and so skips execute_with_tracking. Relevant to feat(agents): coordinate agents competing for shared resources — API limits, CPU time, queues #16951.
    • An A2A peer cannot pin an agent identity on the paths traced so far. build_governed_identity has three callers, all inside chat_workflow (manager.py:3598, graph.py:415, delegation.py:137), and the A2A orchestrator's chat route (agents/chat_agent.py:process_chat_message) does not enter it. The distributed-processing path is not yet traced and is held as unknown, not as safe.

    Test approach for #16950, decided

    The failing test drives the real delegate handler with delegation enabled, from a parent held by an approval gate, and asserts the real enforce_work_item_approval holds the child's write_file — the outcome, not a field. Because the repository's pre-push hook blocks a push containing a failing test, and bypassing hooks is forbidden, it lands as xfail(strict=True, reason="#16950: …"). That reports XFAIL on main only because the gap is real, and turns into a failure the moment a fix makes it pass — so the marker is removed in the same commit that closes the hole, and cannot be forgotten.

  7. 2 remaining items

  8. mrveiss commented on Sep 18, 2026

    @mrveiss
    OwnerAuthor

    Owner decisions, 2026-09-18 — the identity model is final

    With these four settled, the model above (as corrected) is the one every child builds against.

    1. Permission across a chain: the intersection of every hop.

    effective(hop_i) = permission(originator) ∩ permission(hop_1) ∩ … ∩ permission(hop_i)
    

    Relaying can never widen what is permitted. The accepted cost: an agent that delegates must itself hold every permission the delegated work needs, so delegating agents need broad enough grants, or chains stay short. This applies to the authority surfaces named in the correction — approval gates, governed-identity boundaries, auth_role, A2A capabilities — not to work claims.

    2. An autonomous Company OS run's originator is the org agent itself. A scheduled heartbeat with no human in the loop is bounded by that agent's own configured org role, for everything it does and everything it delegates. No separate system principal.

    3. External A2A peers get identity for attribution and authorization only — they do not appear in discovery. An admitted peer's own identity replaces the shared a2a-executor for attribution and for the intersection rule, but external callers can neither see nor be discovered alongside internal agents. A2A stays a complement to internal coordination, not a way into it. The EXTERNAL identity kind therefore exists for attribution and is excluded from presence (#16947).

    4. Bus forgery: ship the confused-deputy fix now, with the boundary stated; harden next. #16950 fixes the laundering by well-behaved agents. Until #16962 lands, the trust boundary of the agent bus is Redis write access — a process with it can still forge a message as any registered agent. #16962 adds a per-agent key issued at registration and an HMAC on every message, covering originator so a relay cannot alter it.

    Build now

    Child Starts from
    #16947 presence the identity type (§1), naming (§2), and consolidation of the live-status registries (§6). EXTERNAL excluded
    #16950 trust the intersection rule over the four authority surfaces; delegation carrying the parent's approval gates and work_item_id; the admitted peer replacing a2a-executor. Negative control is #16958 (strict xfail)
    #16951 resources extend the claim waitlist rather than the unwritten workspace lease; wire the existing QuotaHeadroomStore read side. Claims attributed to the originator
    #16948, #16949 wait on #16947 (blocked_by)
    #16957 enforce the three A2A capabilities that are defined and checked nowhere
  9. mrveiss commented on Sep 18, 2026

    @mrveiss
    OwnerAuthor

    Correction to the rulings above: "A peer's message is never human approval" says peer inbox entries re-enter the ordinary sensitive-tool gate at agent_loop/loop.py:81-83, :952-953. That module is not wired in production (#11221: AgentLoop has no production caller; its package docstring says so). The ruling stands. The gate it means is the live tool seam: chat_workflow/tool_handler.py ToolHandlerMixin._dispatch_tool_call, with enforce_work_item_approval / _approval_category_for. Surfaced by #16948's implementer before any code shipped into the dead module.

  10. mrveiss commented on Sep 28, 2026

    @mrveiss
    OwnerAuthor

    Classification (for work assignment — not a fix proposal)


    Generated by Claude Code

  11. mrveiss commented on Sep 28, 2026

    @mrveiss
    OwnerAuthor

    How the three children actually relate — established by audit against origin/main @ 7adaca8c29, 2026-09-28

    Recording this on the umbrella because it currently exists only in a session's working memory, and a relationship that lives only in a transcript is the next stale premise.

    Three children, three distinct classes. None is a peer of another; none is a sub-umbrella. All three are native sub-issues of this issue and have no children of their own.

    Child Class State on main
    #16948 delivery timing — when a peer's message reaches the recipient Mechanism built and good (protocols/peer_inbox.py, queue drained at each kind's own turn boundary, fail-closed on unaddressable names, presence visibility doubling as the authorization check). Wired for AI_STACK chat and SESSION. 1 of 4 ACs met
    #16950 authority on relay — whose permissions a relayed request runs with Identity propagation is real and structural (protocols/message_origin.py, ContextVar-based so relays are not left to each agent's discipline); the authority algebra exists (security/authority.py meet()); enforcement runs on the A2A path only. 3 of 5 ACs met
    #16951 resource contention — who gets a limited resource when several want it 0 of 5. Three of four primitives exist with their semantics designed and no call site. Note at the top of that issue: PR#16964, which would have wired the quota read side, is CLOSED, not merged

    The one thing they share is a seam, not a class

    #16948 and #16950 have the same residual cause: BaseAgent.send_message_to_agent (base_agent.py:474-483) still routes to the immediate MessageType.REQUEST path (agent_communication.py:622-627, dispatched at :489-491 to the handler registered at base_agent.py:411). That is why turn-boundary delivery does not apply there and why no authority check runs there. Converting that one path satisfies parts of both — cross-linked on each with the detail. They stay separate issues: different properties, different tests, either deliverable without the other.

    #16951 shares nothing mechanical with the other two. Its AC1 is, however, an instance of a class tracked elsewhere: a declared mechanism with no consumer (#15826). user_rate_limiter, LLCWorkspaceLease (#16818) and the quota store's read side are three witnesses to it, not three new defects.

    What this means for the umbrella's own goal

    The design rule stated here — "a message from a peer is that peer's request, never the human's approval" — now holds on the converted path, structurally: peer_inbox.py:23-27 makes a drained message context, so it cannot invoke anything, and a tool call its content prompts still passes that kind's ordinary sensitive-tool gate. The second half of the rule — "an agent refused an action must not be able to get it done by asking another agent" — does not yet hold on the unconverted path, and #16950's AC4 test is currently the hierarchical case (a child inheriting a parent's hold) rather than the peer case it names.

    Open dependencies across the three: #16997 (v0.9.0, AI_STACK roles beyond chat), #16992 (v0.10.0, COMPANY_OS delivery), #16962 (v0.10.0, per-agent keys + MAC — until it lands, the origin chain is attributable but unauthenticated), #16818 (v0.9.0, the lease #16951 needs), #17334 is unrelated to this tree.

  12. modified the milestones: v0.9.0, v0.9-umbrellas on Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions