Repository navigation
Umbrella: agent-to-agent coordination alongside A2A — discovery, messaging, idle notice, shared resources #16946
Description
Activity
- added sub-issues
on Sep 18, 2026 Owner direction, 2026-09-18: agent coordination is top priority. Moved with all five children into v0.9.0 and marked
priority: critical.Build order, because three children share one foundation. Presence (#16947), next-turn messaging (#16948) and the trust fix (#16950) all need the same model of what an agent's identity is, across Company OS agents, AI-stack role agents and live sessions. Built independently they would produce three incompatible identity schemes — which is precisely how this codebase came to have four separate agent registries. So:
- First — the shared identity and addressing model, published on this issue. Every child consumes it; none invents its own.
- In parallel, the parts that do not depend on it: the inventory of existing resource primitives for feat(agents): coordinate agents competing for shared resources — API limits, CPU time, queues #16951 (leases, budget tracker, the unwired rate limiters), and a failing test for security(agents): an agent refused an action can get another agent to do it — no per-peer identity #16950 that demonstrates the permission laundering on
maintoday. - Then feat(agents): live presence — one registry of named agents with busy/idle state, across all three kinds #16947 builds presence on the model, and feat(agents): deliver peer messages at the recipient's next turn, not mid-task #16948 and feat(agents): subscribe once to be told when a peer is next idle #16949 follow it (both are
blocked_byfeat(agents): live presence — one registry of named agents with busy/idle state, across all three kinds #16947).
Shared agent identity and addressing model
The foundation every child of this umbrella builds on. Presence (#16947), next-turn messaging (#16948), idle notices (#16949), the trust fix (#16950) and resource coordination (#16951) all consume it. None invents its own identity scheme — this codebase already has six overlapping places holding agent identity, and the point is to reduce that, not add a seventh.
Code claims below were verified against
main. Sections marked Review correction changed the first draft.1. One identity type — extend
AgentIdentity, don't add a new oneprotocols/agent_communication.py:102-113AgentIdentityis already the typeMessageHeader.senderuses, andsend_messagealready stamps it from the sender's bound identity rather than caller data (:413). Extend it additively:Field Meaning kindCOMPANY_OS,AI_STACKorSESSION— new; lets one registry list all three without guessingnamestable, addressable name — distinct from instance_id, which stays per-processtenant_idthe company this identity exists in; None= shared platform infrastructure, not adminexisting fields capabilities,health_status,last_heartbeat— now shared by all three kindsExisting representation Maps to AgentOrgNode(models/agent_org.py:31)COMPANY_OS,name= its restart-stableagent_idslug (unique,:64),tenant_id=company_idAI-stack AgentIdentity/agent_typeAI_STACK,name=agent_id,tenant_id=None— these are process-global singletons serving every tenantAgentTerminalSession(services/agent_terminal/models.py:32)SESSION,name=session_id; stable for the session's lifetime. No tenant field exists today2. Naming and addressing
nameis unique within(kind, tenant_id); platform-wide whentenant_idisNone. On collision, reject and require a rename — reusing the-<hex8>suffix convention Company OS hires already use (models/agent_org.py:44-45). Never silently overwrite a live entry: that is how a rogue process would squat a trusted name.- Role address
(kind, name)resolves through presence to whichever live instance serves that name — "the RAG agent". - Instance address
(SESSION, name, instance_id)is exact, with no fan-out — two people's sessions must never collide.
3. The originator, and the authorization rule
MessageHeadergainsoriginator— set once, on first send, and never overwritten by a relay (unlikesender, which legitimately changes each hop at:413) — pluschain, the hop names, for audit and cycle detection.Rule: effective permission at hop i is the intersection of every hop from the originator through i.
effective(hop_i) = permission(originator) ∩ permission(hop_1) ∩ … ∩ permission(hop_i)Not the originator's alone: that would let a low-privilege relay act as a transparent pass-through for a high-privilege originator. Intersection is the only form under which every hop's own refusal survives the whole chain — which is what "relaying can never widen what is permitted" requires.
This is the mechanism of the #16950 bug, confirmed in code:
BaseAgent._handle_communication_request(agents/base_agent.py:419) builds itsAgentRequestfrommessage.payload.contentand never readsmessage.header.sender— the identity is dropped at exactly the point authorization needs it. Downstream,hold_scopesclaims in the recipient's name (base_agent.py:249), and A2A claims every admitted peer's scopes asagent_id="a2a-executor"(a2a/task_executor.py:114), even though the realpeer_idis in scope in the same function (:75). The fix threadsoriginatorfrom the header into the request and makeshold_scopes' identity the intersection-computed one. A2A's external wire format does not change:X-A2A-Agent-Idstays; only the internal attribution does.A peer's message is never human approval. Peer entries in the inbox (#16948) are typed distinctly from human steering entries, and re-enter the ordinary sensitive-tool gate (
agent_loop/loop.py:81-83,:952-953) against the intersection — or are refused when no human is present to approve.Review correction — intersection needs one permission vocabulary. The three kinds express permission three ways today: AI-stack agents as
allowed_work/forbidden_workprofiles, Company OS agents as org roles, external A2A peers asCapability. You cannot intersect a role with a work list. Scopes — the currencyhold_scopesalready uses — are the common vocabulary, and each kind maps its native form onto scopes at one place. #16950 owns that mapping; until it exists, intersection is undefined across kinds.4. Can the sender be forged?
Today, yes. The Redis channel
rpushes andblpops raw JSON (agent_communication.py:223-242), deserialised byStandardMessage.from_json(:175) with no signature, andregister_agent(:632) checks only for a duplicate id — it does not ask whether the caller may hold that name.Gate registration. Each kind registers only from its authoritative source: Company OS from a real
AgentOrgNoderow, AI-stack from the fixed bootstrap list, sessions from an authenticated session-creation call — never from a caller-declared identity blob. Receivers reject asenderororiginatorthat does not resolve to a live registered entry.Review correction — what the gate does and does not stop. The first draft rejected "trusted transport" because the threat is an in-network process impersonating another agent — and then chose a design that does not stop that either. Anyone who can write to Redis can publish a message whose
senderis any registered agent and whoseoriginatoris any registered principal; the registry check passes, because those identities are registered.Separating the two threats makes the scope honest:
- The confused deputy — the security(agents): an agent refused an action can get another agent to do it — no per-peer identity #16950 bug — is fixed by threading the originator and intersecting permissions. Well-behaved agents all send through
send_message, which stamps the sender and will preserve the originator, so an agent can no longer act with its own permissions on another's behalf. - Raw-bus forgery is not fixed by this design. Its trust boundary is Redis write access. Closing it needs a per-agent credential; the light form is a symmetric key issued at the registration gate, with each message carrying an HMAC over header and payload that the registry verifies — no PKI, and rotation is re-registration. That is a separate decision (below), not part of security(agents): an agent refused an action can get another agent to do it — no per-peer identity #16950's core.
5. Tenancy
Review correction — scope by the originator's tenant, not the sender's. The first draft allowed a message only if
sender.tenant == recipient.tenantor the recipient was shared. That rejects replies from shared infrastructure: a shared RAG agent (tenant_id=None) answering a Company OS agent in tenant X fails both tests.Rule: a message belongs to its originator's tenant. It may reach a recipient in that tenant, or a shared recipient (
tenant_id=None); anything else is a hard reject, not filter-and-continue. A shared agent acting inside a chain inherits the originator's tenant for everything it does in that chain — so it can reply into tenant X for a request that began in X, and never into Y.Today enforcement is nowhere uniform: Company OS filters only at its API layer (
llc/api/agents.py:41-50) and does not useagent_communicationat all; AI-stack agents and sessions carry no tenant field. The pattern to copy already exists in resources:AgentBudgetStatekeys oncompany_id:agent_id(llc/services/agent_budget_tracker.py:125), andLLCWorkspaceLeasecarriescompany_id(llc/models/workspace_lease.py:44).6. What it consolidates
Six places hold agent identity or state:
AgentHealthRegistry(agents/agent_client.py:70),AgentCapabilityRegistry(orchestration/agent_registry.py:213),AgentRegistryServiceovermodels.agent.Agent,DistributedAgentManager(agents/agent_orchestration/distributed_management.py:40),AgentOrgNode, andSessionManager's in-memory session map.The four that answer "is this agent alive and busy right now" collapse into one live presence registry keyed by
(kind, tenant_id, name), fed from each source rather than reimplemented. The static capability catalogue and the durable identity rows stay — they answer different questions (what can it do, does it exist) and are looked up by the samename. Six overlapping places become three distinct ones.7. What each child consumes
Child Consumes #16947 presence the whole identity type; owns the consolidation in §6. Build first #16948 next-turn delivery name-based addressing (§2); a peer inbox entry type distinct from human steering ( agent_loop/types.py:345, which has no sender field today)#16949 idle notices presence's live busy/idle, scoped by nameand tenant; transport is the existing event bus#16950 trust boundary §3 and §5, plus the scope-vocabulary mapping; replaces a2a-executor#16951 resources (name, tenant_id)as the claim key. Shares #16950's root cause: bothhold_scopescall sites attribute claims to the executor, never the originator, so "agents sharing one API key back off together" is mis-attributed until the originator is threadedCould not be determined from code
- Whether shared AI-stack agents ever need per-tenant isolation beyond request-scoped context.
- What tenant a terminal session belongs to — no field exists, and it may be implied at the API layer.
Decisions for the owner
Asked separately; recorded here once answered.
Correction to the identity model above — authority is not carried by
hold_scopesInvestigation for #16950 showed that §3 of the design, and one of my own review corrections, targeted the wrong mechanism. Build against this correction, not against §3 as posted.
agents/scope_enforcement.pyis work-claim mutual exclusion — a lease — not authorization. "Agent B holds scope S" means B holds a claim on a resource, so that two agents do not write the same thing at once. B writing under its own claim is that design working correctly. So:- §3's plan to make
hold_scopes' identity "the intersection-computed one" does not implement authorization. Claims belong to feat(agents): coordinate agents competing for shared resources — API limits, CPU time, queues #16951 (resource coordination), where the originator question is about attribution, not authority. - My review note that "scopes are the common permission vocabulary" was wrong for the same reason. Scopes are the vocabulary of what is being claimed, not of what is permitted.
The authority surfaces the originator must actually carry are:
- the approval gates —
enforce_work_item_approvalandrequires_approval_before - governed-identity boundaries —
build_governed_identity auth_role- A2A's trust-level capabilities
The intersection rule in §3 still stands; it applies to these, not to claims. The open question of a common vocabulary is therefore still open — mapping approval categories, identity boundaries, roles and A2A capabilities onto one comparable form is #16950's to design.
What the laundering on
mainactually isNot the claim path the issue first described — delegation drops the parent's approval gates.
_handle_delegate_tool(chat_workflow/tool_handler.py:2316) passes the child onlyparent_agent_id(used for a log line) andauth_role. The child is built frombuild_governed_identity({"agent_id": agent_type}, …), so itsrequires_approval_beforeis empty and thework_item_idis gone. A parent whosewrite_fileis held behind the work item's "writing files" gate delegates todocumentation_agent, and the child'swrite_fileruns unapproved.auth_roleis carried deliberately (#13821); the gates are not.It sits behind
AUTOBOT_DELEGATION_ENABLED, which is off by default — which limits the exposure, not the defect.Other findings from the same investigation
- A2A enforces one capability of four. Only
SUBMIT_TASKSis checked (api/a2a.py:227).QUERY_MEMORY,DEFINE_AGENTSandDISCOVERYare defined in the trust matrix and checked nowhere, so a trust level that denies them denies nothing. Being filed separately. - A2A's
a2a-executorflattening does not break claim exclusion — holding the same claim requires the same agent and the same task (work_claims.py:307), so two peers' tasks still refuse each other. What is lost is attribution and authority, not exclusion. - The internal peer channel holds no claims at all.
_handle_communication_requestcallsprocess_requestdirectly and so skipsexecute_with_tracking. Relevant to feat(agents): coordinate agents competing for shared resources — API limits, CPU time, queues #16951. - An A2A peer cannot pin an agent identity on the paths traced so far.
build_governed_identityhas three callers, all insidechat_workflow(manager.py:3598,graph.py:415,delegation.py:137), and the A2A orchestrator's chat route (agents/chat_agent.py:process_chat_message) does not enter it. The distributed-processing path is not yet traced and is held as unknown, not as safe.
Test approach for #16950, decided
The failing test drives the real
delegatehandler with delegation enabled, from a parent held by an approval gate, and asserts the realenforce_work_item_approvalholds the child'swrite_file— the outcome, not a field. Because the repository's pre-push hook blocks a push containing a failing test, and bypassing hooks is forbidden, it lands asxfail(strict=True, reason="#16950: …"). That reports XFAIL onmainonly because the gap is real, and turns into a failure the moment a fix makes it pass — so the marker is removed in the same commit that closes the hole, and cannot be forgotten.- §3's plan to make
2 remaining items
Owner decisions, 2026-09-18 — the identity model is final
With these four settled, the model above (as corrected) is the one every child builds against.
1. Permission across a chain: the intersection of every hop.
effective(hop_i) = permission(originator) ∩ permission(hop_1) ∩ … ∩ permission(hop_i)Relaying can never widen what is permitted. The accepted cost: an agent that delegates must itself hold every permission the delegated work needs, so delegating agents need broad enough grants, or chains stay short. This applies to the authority surfaces named in the correction — approval gates, governed-identity boundaries,
auth_role, A2A capabilities — not to work claims.2. An autonomous Company OS run's originator is the org agent itself. A scheduled heartbeat with no human in the loop is bounded by that agent's own configured org role, for everything it does and everything it delegates. No separate system principal.
3. External A2A peers get identity for attribution and authorization only — they do not appear in discovery. An admitted peer's own identity replaces the shared
a2a-executorfor attribution and for the intersection rule, but external callers can neither see nor be discovered alongside internal agents. A2A stays a complement to internal coordination, not a way into it. TheEXTERNALidentity kind therefore exists for attribution and is excluded from presence (#16947).4. Bus forgery: ship the confused-deputy fix now, with the boundary stated; harden next. #16950 fixes the laundering by well-behaved agents. Until #16962 lands, the trust boundary of the agent bus is Redis write access — a process with it can still forge a message as any registered agent. #16962 adds a per-agent key issued at registration and an HMAC on every message, covering
originatorso a relay cannot alter it.Build now
Child Starts from #16947 presence the identity type (§1), naming (§2), and consolidation of the live-status registries (§6). EXTERNALexcluded#16950 trust the intersection rule over the four authority surfaces; delegation carrying the parent's approval gates and work_item_id; the admitted peer replacinga2a-executor. Negative control is #16958 (strict xfail)#16951 resources extend the claim waitlist rather than the unwritten workspace lease; wire the existing QuotaHeadroomStoreread side. Claims attributed to the originator#16948, #16949 wait on #16947 ( blocked_by)#16957 enforce the three A2A capabilities that are defined and checked nowhere - added sub-issues
on Sep 18, 2026 Correction to the rulings above: "A peer's message is never human approval" says peer inbox entries re-enter the ordinary sensitive-tool gate at
agent_loop/loop.py:81-83, :952-953. That module is not wired in production (#11221:AgentLoophas no production caller; its package docstring says so). The ruling stands. The gate it means is the live tool seam:chat_workflow/tool_handler.pyToolHandlerMixin._dispatch_tool_call, withenforce_work_item_approval/_approval_category_for. Surfaced by #16948's implementer before any code shipped into the dead module.- added sub-issues
on Sep 18, 2026 Classification (for work assignment — not a fix proposal)
- Scope:
agents(secondarycoordination): internal agent-to-agent discovery, messaging and trust. - Primary files/dirs:
autobot-backend/protocols/agent_communication.pyandautobot-backend/agents/base_agent.py(verified on main)autobot-backend/a2a/,autobot-backend/acp/andautobot-backend/llc/api/agents.py(verified on main)
- Umbrella: yes, a container, not work. Children:
- open, v0.9.0: feat(agents): deliver peer messages at the recipient's next turn, not mid-task #16948, security(agents): an agent refused an action can get another agent to do it — no per-peer identity #16950, feat(agents): coordinate agents competing for shared resources — API limits, CPU time, queues #16951
- open, v0.10.0: security(agents): authenticate agent-bus messages — a per-agent key issued at registration, HMAC on every message #16962, decision(agents): should a peer message wake an idle AI-stack agent to start a new run? #16991, feat(agents): deliver a peer message at the start of a Company OS agent's next heartbeat run #16992
- open, no milestone: security(a2a): a peer's research task writes peer-chosen content into the knowledge base, and no capability covers writes #16967, security(agents): the sentiment agent reads and writes any session's working memory named in the request context #16968, AI_STACK peer-message drain only wired for the 'chat' role; 'rag'/'system_commands' silently lose messages #16997
- closed: feat(agents): live presence — one registry of named agents with busy/idle state, across all three kinds #16947, feat(agents): subscribe once to be told when a peer is next idle #16949, security(a2a): three of four trust-matrix capabilities are enforced nowhere — a level that denies them denies nothing #16957, feat(agents): capture a terminal session's tenant at creation — it cannot be recovered afterwards #16975, agents: the peer channel never delivers to the recipient — a request lands in the sender's own inbox and the sender answers itself #16986
- Blocked by:
- The owner decisions of 2026-09-18 (the identity model is final; the intersection rule; A2A peers excluded from discovery; bus forgery deferred to security(agents): authenticate agent-bus messages — a per-agent key issued at registration, HMAC on every message #16962). These unblock the children rather than block them.
- decision(agents): should a peer message wake an idle AI-stack agent to start a new run? #16991 is an open owner decision: whether a peer message should wake an idle agent.
- The children's individual blockers were not read.
- Read vs inferred: Read the issue body, its 5 comments, and the sub-issue list with states and milestones. The children's bodies were not read.
- Undetermined: the current blocked-by state of each open child.
Generated by Claude Code
- Scope:
How the three children actually relate — established by audit against
origin/main@7adaca8c29, 2026-09-28Recording this on the umbrella because it currently exists only in a session's working memory, and a relationship that lives only in a transcript is the next stale premise.
Three children, three distinct classes. None is a peer of another; none is a sub-umbrella. All three are native sub-issues of this issue and have no children of their own.
Child Class State on main#16948 delivery timing — when a peer's message reaches the recipient Mechanism built and good ( protocols/peer_inbox.py, queue drained at each kind's own turn boundary, fail-closed on unaddressable names, presence visibility doubling as the authorization check). Wired for AI_STACKchatand SESSION. 1 of 4 ACs met#16950 authority on relay — whose permissions a relayed request runs with Identity propagation is real and structural ( protocols/message_origin.py,ContextVar-based so relays are not left to each agent's discipline); the authority algebra exists (security/authority.pymeet()); enforcement runs on the A2A path only. 3 of 5 ACs met#16951 resource contention — who gets a limited resource when several want it 0 of 5. Three of four primitives exist with their semantics designed and no call site. Note at the top of that issue: PR#16964, which would have wired the quota read side, is CLOSED, not merged The one thing they share is a seam, not a class
#16948 and #16950 have the same residual cause:
BaseAgent.send_message_to_agent(base_agent.py:474-483) still routes to the immediateMessageType.REQUESTpath (agent_communication.py:622-627, dispatched at:489-491to the handler registered atbase_agent.py:411). That is why turn-boundary delivery does not apply there and why no authority check runs there. Converting that one path satisfies parts of both — cross-linked on each with the detail. They stay separate issues: different properties, different tests, either deliverable without the other.#16951 shares nothing mechanical with the other two. Its AC1 is, however, an instance of a class tracked elsewhere: a declared mechanism with no consumer (#15826).
user_rate_limiter,LLCWorkspaceLease(#16818) and the quota store's read side are three witnesses to it, not three new defects.What this means for the umbrella's own goal
The design rule stated here — "a message from a peer is that peer's request, never the human's approval" — now holds on the converted path, structurally:
peer_inbox.py:23-27makes a drained message context, so it cannot invoke anything, and a tool call its content prompts still passes that kind's ordinary sensitive-tool gate. The second half of the rule — "an agent refused an action must not be able to get it done by asking another agent" — does not yet hold on the unconverted path, and #16950's AC4 test is currently the hierarchical case (a child inheriting a parent's hold) rather than the peer case it names.Open dependencies across the three: #16997 (v0.9.0, AI_STACK roles beyond
chat), #16992 (v0.10.0, COMPANY_OS delivery), #16962 (v0.10.0, per-agent keys + MAC — until it lands, the origin chain is attributable but unauthenticated), #16818 (v0.9.0, the lease #16951 needs), #17334 is unrelated to this tree.
Goal
Agents inside AutoBot can find each other, message each other by name, learn when a peer is free, and coordinate access to resources they compete for — without any agent being able to borrow another's permissions.
It complements A2A, it does not replace it. A2A (
autobot-backend/a2a/) is an external, task-shaped federation protocol: an outside system submits one task and polls for one result. ACP (autobot-backend/acp/) is one editor driving one agent over stdio. Neither is a way for AutoBot's own agents to talk to each other.Scope — decided by the owner, 2026-09-18
Which agents: all three kinds, which today have three separate, unconnected representations:
protocols/agent_communication.pyaddresses by static IDPlus resource contention: when several agents compete for the same limited resource — API rate limits, CPU time, work queues — they coordinate rather than collide. This is what turns messaging into coordination.
What exists today (audited against
main)protocols/agent_communication.py:116-125—MessageHeadercarriessender,reply_to,correlation_idllc/api/agents.py:41lists Company OS agents with their last run status — a poll, not live, and Company OS only. Four separate registries exist, none of them live named presenceBaseAgent.send_message_to_agentdid not work: both channels delivered to the sender, which answered itself (#16986, fixed by #17001). Recipients are static IDs, and delivery runs the handler immediately, interrupting the recipient mid-task (#16948)agent_loop/loop.py:203-206drains a steering inbox only at the top of each iteration — built for human steering, never for peerstask_executor.py:114,agent_id="a2a-executor"), and the internal channel has no trust gate at allRecommended home: extend
protocols/agent_communication.py— it already has the envelope, the addressed request/response path and pluggable transports. Not A2A (an external, spec-governed protocol) and not ACP (single client, single agent).Design rule that must hold throughout
A message from a peer is that peer's request, never the human's approval. A peer cannot grant another peer a permission it does not hold, and an agent refused an action must not be able to get it done by asking another agent to do it. Today that second property does not hold — see the trust-boundary child.
Children
Tracked as native sub-issues below.