Skip to content

Proposal: SLSA levels for AI agent deployments #1594

Description

@razashariff

SLSA defines supply chain integrity levels for software artifacts. AI agents introduce new supply chain risks:

  1. Agent code provenance -- which developer built this agent? Is the code signed?
  2. MCP server integrity -- are the tools the agent uses verified and pinned?
  3. Model provenance -- which model version is the agent using? Has it been tampered with?
  4. Configuration integrity -- have the agent's permissions, spend limits, or trust levels been modified?

Proposed SLSA Agent Levels:

  • Agent L0: No provenance. Agent deployed from unknown source.
  • Agent L1: Agent code is signed. MCP tool definitions are hashed.
  • Agent L2: Agent built in CI with signed provenance attestation. Tools pinned with SHA-256.
  • Agent L3: Agent built in isolated environment. All components (code + tools + model + config) have SLSA L3 provenance. Cryptographic identity bound to build attestation.

This extends SLSA from software artifacts to autonomous agent systems.

Reference: OWASP MCP Security Cheat Sheet Section 7 covers tool definition pinning via cryptographic hashes.

Activity

  1. lehors commented on Mar 27, 2026

    @lehors
    Member

    Hello @razashariff and thank you for the proposal. If you can, I suggest you add an entry to the agenda and join the call on Monday to discuss this further.

  2. arewm commented on Apr 3, 2026

    @arewm
    Member

    @razashariff , what do you consider an AI agent in this context? Is this idea for the development of fully encapsulated agents (i.e. trained models) or for agent definitions as has become popular to change the behavior of LLMs by informing their context?

  3. razashariff commented on Apr 3, 2026

    @razashariff
    Author

    Good question. Both, but the primary focus is on agent definitions and their runtime behaviour -- the MCP servers, tool configurations, and trust credentials that shape what an AI agent can do.

    In practical terms:

    1. Agent definitions (MCP tool configs, server manifests, tool descriptions) -- these need attestation because they directly control agent behaviour. A tampered tool description is a prompt injection vector. SLSA-style provenance for MCP server packages would verify that the server you install is the one the author published.

    2. Runtime agent identity -- agents making API calls or tool calls need cryptographic identity credentials that can be verified by the receiving service. This is what our IETF draft (draft-sharif-mcps-secure-mcp) addresses with per-message ECDSA signing.

    3. Trained models -- also relevant but further upstream. Model file signing handles ECDSA signing of weights for integrity verification.

    The gap in SLSA today is that it covers software supply chain but not the agent supply chain. An MCP server installed via npm has SLSA provenance for the package, but nothing attesting to the safety of its tool definitions, the integrity of its runtime behaviour, or the identity of agents using it.

  4. razashariff commented on Apr 3, 2026

    @razashariff
    Author

    Thanks Arnaud. Unfortunately Monday is a bank holiday here in the UK so I won't be able to join the call. Happy to sync early the following week if that works, or I can add a more detailed write-up to the agenda ahead of time so the group can discuss without me needing to be live.

    To give more context for the group:

    The core gap is that SLSA currently attests to the provenance of software artifacts (who built it, how it was built, what inputs were used). AI agents introduce a new category of supply chain risk that sits above the artifact level:

    1. Tool definition integrity -- an MCP server's tool descriptions are fed directly into LLM context. A tampered description is a prompt injection vector. There have been 30+ CVEs filed against MCP implementations in 2026 alone, many exploiting this exact surface.

    2. Agent identity and trust -- when an agent calls an API or another agent, there is no standardised way to verify who the agent is or what trust level it holds. Our IETF draft (draft-sharif-mcps-secure-mcp) proposes per-message ECDSA signing to address this.

    3. Configuration drift detection -- an agent can be registered with a valid identity but have its system prompt, tool list, or model version silently changed after registration. Attestation at deployment time is not enough -- continuous verification is needed.

    4. Quantization and deployment chain -- a model file (e.g. GGUF) may be quantized from a source model, but there is no attestation linking the quantized artifact back to its source. Our IETF draft (draft-sharif-ai-model-lifecycle-attestation) covers the full chain from training data to inference.

    SLSA levels for AI agents could follow a similar graduated model: L0 (no provenance), L1 (signed agent manifest), L2 (signed + build provenance), L3 (signed + provenance + continuous integrity verification).

    Happy to prepare a more detailed proposal document if that would be useful for the group.

  5. piiiico commented on Apr 3, 2026

    @piiiico

    MCP server integrity (your point 2) is the most immediately actionable vector here, and existing scan data gives it concrete shape.

    We recently ran static analysis across 12 public MCP server repos — including all official Anthropic reference servers and popular community servers (database connectors, browser automation, search). Findings:

    • 100% finding rate — every single repo had at least one security finding
    • 46 command execution patterns using spawn()/exec() instead of shell-safe execFile()
    • 7 hardcoded credentials in production code — not test files, production source
    • 0 repos with security scanning in CI/CD

    The absence of CI scanning is the key SLSA gap: none of these servers have provenance attestations for their security posture, let alone pinned dependency verification.

    For MCP server integrity specifically, the SLSA framing maps well:

    • L1: Build provenance for the server package (does the npm/PyPI release trace to a verified commit?)
    • L2: Signed builds with MCP-specific security scan artifacts
    • L3: Hardened builds — no direct shell execution, verified dependency tree, scan results as SLSA attestations

    The tool we built for the scan side is open source: agent-audit. Static analysis for MCP server repos and configs. CI integration: npx @piiiico/agent-audit --auto --json --min-severity high exits non-zero on findings — that's the artifact you'd attach to a SLSA L2 build.

    Scan data: https://dev.to/piiiico/we-scanned-13-popular-mcp-servers-every-single-one-had-security-findings-3g39

  6. arewm commented on Apr 16, 2026

    @arewm
    Member

    Thanks for the thought in this matter. There is a lot happening and it hits at multiple parts within the supply chain.

    Combining some text from the above posts to reason about them together

    Agent definitions (MCP tool configs, server manifests, tool descriptions) -- these need attestation because they directly control agent behaviour.

    Tool definition integrity -- an MCP server's tool descriptions are fed directly into LLM context. A tampered description is a prompt injection vector. There have been 30+ CVEs filed against MCP implementations in 2026 alone, many exploiting this exact surface.

    Do these need attestations or just signatures? Many of these configuration files are just text so the provenance becomes harder to reason about. They can, however, be signed/attested artifacts to verify that they were produced by a specific identity. You can then require these types of files to be signed by specific identities (or build a mapping of files to identity requirements). I am mapping this to be similar to the in-toto attestation verifier layouts.

    This seems like a valid use of attestations to check the integrity of the artifact (i.e. its hash) as reported by the producer. The producer can generate a SLSA provenance/VSA attestation if it has that information, or it could generate other attestation predicates like a release. The key point here is not just that you need to verify the checksum, but you need to verify the checksum as produced by a specific identity.

    Agent identity and trust -- when an agent calls an API or another agent, there is no standardised way to verify who the agent is or what trust level it holds. Our IETF draft (draft-sharif-mcps-secure-mcp) proposes per-message ECDSA signing to address this.

    Quantization and deployment chain -- a model file (e.g. GGUF) may be quantized from a source model, but there is no attestation linking the quantized artifact back to its source. Our IETF draft (draft-sharif-ai-model-lifecycle-attestation) covers the full chain from training data to inference.

    These do seem like a concern but I don't think SLSA fits in here. This is about making sure that the client and server are both talking to the intended individual without a MitM? The AI/ML working group might be a better location to ask about these to understand whether there is anything that can help to fill this gap.

    Configuration drift detection -- an agent can be registered with a valid identity but have its system prompt, tool list, or model version silently changed after registration. Attestation at deployment time is not enough -- continuous verification is needed.

    This feels more like runtime protections. There are two different ways to handle this that I see. The first is to prevent certain files to be modified by the process -- there are multiple ways that this can be achieved. The second is to record what files an agent touches by an external observer and generate provenance for the output of the agentic interaction. The first would granularly block edits before they happened. The second would generate an artifact that needs to be reviewed to see what was modified (i.e. if it was modified and not committed or modified and reverted). SLSA can play a role there, but it we have to be careful.

    The absence of CI scanning is the key SLSA gap: none of these servers have provenance attestations for their security posture, let alone pinned dependency verification.

    Types of checks like scanning may be represented in the Dependency track once that is finalized. As you noted, the provenance doesn't have any requirements for a scan action, but the scan may be recorded in the provenance if it is an integral part of the build.

  7. razashariff commented on Apr 16, 2026

    @razashariff
    Author

    Thanks again @arewm for taking the time to deep-dive on this -- really appreciate the considered split between build-time attestations and runtime protections, and the mapping to in-toto verifier layouts.

    Taking your points in order:

    Tool definition integrity: Agreed, attesting the artifact to a specific producer identity is the right shape. Happy to prototype a worked example against a few high-traffic MCP servers and contribute it back -- would a SLSA provenance reference implementation for an MCP tool definition file be useful input to this issue?

    Agent identity / quantization chain: Fair redirect, both are real gaps but not SLSA-shaped. We'll raise them with the AI/ML WG.

    Configuration drift: Your second pattern -- external observer recording what an agent touches and emitting provenance for the interaction -- maps cleanly to draft-sharif-agent-audit-trail. Would you be open to us exploring an attestation predicate for agent-interaction provenance as a follow-up?

    Dependency Track: Very interested. The MCP ecosystem has produced 30+ CVEs in 2026 and we have real deployment data on where scanning provenance would have caught them early. Happy to contribute use cases or early feedback as the track firms up.

    I can draft a short doc mapping concrete MCP scenarios to the SLSA frameworks you've highlighted, and drop it back on this issue (or a sibling) if that helps move things forward.

    Thanks again.

    -- Raza

  8. tomjwxf commented on Apr 16, 2026

    @tomjwxf

    @razashariff @arewm -- jumping in from #1606 which arewm pointed me to.

    We're working on a complementary layer to what Raza describes here. His drafts cover agent identity (`draft-sharif-mcps-secure-mcp`, per-message ECDSA signing) and model lifecycle attestation. Ours cover decision receipts -- cryptographic proof of what an agent decided to do, under what policy, at each tool call.

    The layers compose cleanly:

    • Agent identity (Raza): proves WHO the agent is
    • Tool integrity (Raza + @piiiico): proves the MCP server definitions haven't been tampered with
    • Decision receipts (us, draft-farley-acta-signed-receipts): proves WHAT the agent decided, under WHICH policy, and that the policy gate held
    • Build provenance (SLSA): proves HOW the artifact was produced

    For the agent SLSA levels Raza proposed:

    L2: Agent built in CI with signed provenance attestation. Tools pinned with SHA-256.

    Our receipts would provide the provenance attestation for the build steps. Each tool call during the build (npm install, npm build, npm test) produces a signed receipt with the input hash, output hash, and policy evaluation result. The chain of receipts IS the build provenance.

    L3: All components (code + tools + model + config) have SLSA L3 provenance. Cryptographic identity bound to build attestation.

    For L3, our hardware path (ATECC608B secure element) provides the "cryptographic identity bound to build attestation" -- the signing key is hardware-bound and cannot be extracted. We recently had a physical attestation example merged into Microsoft's Agent Governance Toolkit demonstrating this.

    We've also proposed a Decision Receipt predicate type for in-toto to formalize the receipt format as a standard attestation predicate, following guidance from @Hayden-IO on sigstore/rekor#2798.

    @razashariff -- would be interested to compare notes on how `draft-sharif-mcps-secure-mcp` (agent identity) composes with `draft-farley-acta-signed-receipts` (decision receipts). The agent identity signs each message; the decision receipt signs each authorization decision. Both attach to the same tool call but attest to different properties. A combined attestation bundle would cover identity + authorization + provenance in one verifiable artifact.

    Happy to join the SLSA call when it next meets to discuss how these layers compose.

  9. piiiico commented on Apr 17, 2026

    @piiiico

    @tomjwxf — thanks for the layered model. The composition is clean:

    • Agent identity: WHO is calling
    • Tool integrity: WHAT definitions the agent was given
    • Decision receipts: WHAT the agent authorized, under WHICH policy
    • Build provenance: HOW the artifact was produced

    The layer that doesn't quite fit this stack is behavioral drift detection — an agent can pass all four of the above and still behave differently than specified after deployment. Build-time attestations and decision receipts are point-in-time; they don't catch configuration drift, prompt injection that changes behavior post-registration, or gradual behavioral divergence across sessions. That's the gap AgentLair's L4 behavioral telemetry is designed to fill: continuous runtime monitoring against a declared behavioral baseline.

    The SLSA-for-agents framing in this issue is the right starting point for the build-time side. Would be interested in the SLSA call to see how the runtime layer complements it — particularly the 'configuration drift detection' point @razashariff raised in the original proposal.

    @razashariff — happy to compare notes on how draft-sharif-mcps-secure-mcp (per-message ECDSA signing) composes with a continuous behavioral telemetry layer. The signing proves per-call identity; the telemetry detects whether the agent's aggregate behavior matches its declared profile. The two are complementary: identity + behavioral consistency.

  10. tomjwxf commented on Apr 17, 2026

    @tomjwxf

    @piiiico the behavioral telemetry layer is a clean addition to the stack and maps to a real gap. Point-in-time attestations cannot catch "agent passes build-time checks then drifts at runtime," and neither can decision receipts on their own. Two thoughts on how the layers compose:

    Hardware-rooted continuous verification as the complement to telemetry. If the signing key for decision receipts is bound to a secure element (ATECC608B or equivalent), the receipt chain itself becomes continuous behavioral evidence: every tool call produces a hardware-attested receipt, and drift shows up as an anomaly in the receipt stream. That complements rather than replaces AgentLair-style runtime telemetry. Telemetry watches behavioral shape in aggregate; the receipt chain provides the cryptographically verifiable ground truth for each individual decision. Both are useful, neither subsumes the other.

    Declared-baseline vs observed-behavior binding. The piece that makes drift detection concrete rather than heuristic is binding the telemetry baseline to the same attestation the build produced. If the L2 build provenance includes "this agent should call tools from set X with frequency distribution Y," then the L4 runtime layer is verifying an attested baseline rather than an inferred one. That turns "configuration drift detection" (Razas original L3 point) from a heuristic into a verification check against a signed artifact.

    @razashariff on the draft-sharif-mcps-secure-mcp + draft-farley-acta-signed-receipts composition: per-message ECDSA signing proves identity per call, the decision receipt proves the policy gate passed per call, the hardware attestation proves the signing key is tamper-evident, the behavioral telemetry proves the aggregate conforms to the declared baseline. Four orthogonal properties, four attestation types, composable into one bundle. Happy to draft a concrete composition sketch (which fields go where, what the combined verifier looks like) as input to the SLSA-for-agents framing if useful.

    Still happy to join the SLSA call when it next meets. Email has been unreliable on my end; tagging here works better in the meantime.

  11. added a commit that references this issue on Apr 17, 2026
  12. piiiico commented on Apr 17, 2026

    @piiiico

    @tomjwxf — both points are strong additions to the composition model.

    On hardware-rooted continuous verification: Agreed that secure-element-bound signing keys for decision receipts and runtime behavioral telemetry are orthogonal, composable layers. The receipt chain gives you cryptographically verifiable ground truth per decision; the telemetry gives you behavioral shape in aggregate. The specific value of the latter: adversarial drift can preserve a valid receipt-by-receipt signature chain while shifting the aggregate pattern — e.g., an agent that correctly signs every receipt but changes which tools it calls over time. The receipt chain proves each call was authorized under the stated policy; telemetry is needed to catch that the behavioral envelope itself has shifted. Hardware roots make the receipt chain tamper-evident; they do not, on their own, catch behavioral drift at the population level.

    On declared-baseline vs observed-behavior binding: This is the mechanism I was implicitly pointing at, now made concrete. If L2 build provenance includes a behavioral specification — tool set, call frequency distribution, resource access patterns — then L4 runtime telemetry becomes verification against a signed artifact rather than anomaly detection against an inferred baseline. That is a meaningful step: it moves drift detection from "this looks different from before" to "this violates an attested declaration." The missing piece I would add: the declared baseline needs to survive identity continuity across key rotation and agent updates. If the build provenance is bound to a specific keypair, rotating the key (which agents must do) breaks the chain back to the attested baseline. The identity layer needs to carry stable agent UUIDs or DIDs that persist through key rotation so the behavioral baseline can remain anchored.

    A composition sketch along those lines — which fields from L2 build provenance, decision receipts, hardware attestation, and behavioral telemetry go into a single bundle, and what the combined verifier checks — would be a useful concrete artifact for the SLSA-for-agents framing. Happy to collaborate on a draft if you want to start from your four-layer summary above.

  13. added a commit that references this issue on Apr 17, 2026
  14. 3 remaining items

  15. arewm commented on Apr 17, 2026

    @arewm
    Member

    I modified the buildType to add timestamps to help track when files might be read and processes executed in order to identify potential drift issues. https://refs.arewm.com/agent-commit/v0.2/

    As I mentioned in arewm/refs.arewm.com#1 (comment), I feel like the best potential way to integrate additional attestations might be to add it to byproducts as ResourceDescriptor links with a digest and URI instead of including the full metadata there as the producer's identity will be different.

  16. tomjwxf commented on Apr 17, 2026

    @tomjwxf

    @arewm answering your "who makes the receipt" question directly, with both an architectural sketch and a concrete PR. I had started drafting this before your edit landed, so there is some overlap with your ResourceDescriptor framing; the PR picks that up.

    Who signs the receipt. The receipt is signed by the policy enforcement component (the supervisor hook), not by the agent process. In the protect-mcp implementation the hook is a PreToolUse interceptor registered with the agent host (Claude Code, Cursor, Google ADK). The host invokes the hook before dispatching the tool call, the hook runs Cedar on the (principal, action, resource, context), signs the decision, and returns an exit code that the host respects. The hook owns its own Ed25519 keypair, hardware-rooted where available. The agent never sees it.

    Separation from the agent process. In practice each hook invocation is a separate spawned subprocess, giving process-level isolation. A compromised agent cannot exfiltrate the key. With a secure element in the path the key cannot be exfiltrated even by a compromised hook process. The hook is in the enforcement path because the host wires it in, not because the agent cooperates.

    Trust model. The supervisor hook is trusted by the host. If the host is compromised, receipts from the hook are worthless, but so is everything else about that session. If the agent is compromised, the hook still holds: different identity, different key, enforced by the host before the tool call runs. Matches exactly the trust model you describe for the eBPF observer in agent-commit.

    On the ResourceDescriptor pattern you mentioned both in your comment on refs.arewm.com#1 and in your edit here: agreed, this is the cleaner shape. A decision-receipts attestation is its own DSSE envelope signed by the supervisor-hook identity, stored separately, with its own in-toto predicate type. The SLSA provenance for agent-commit references it via ResourceDescriptor in byproducts:

    {
      "name": "decision-receipts",
      "digest": { "sha256": "a8f3c9d2..." },
      "uri": "oci://registry/org/agent-session/run-xyz/receipts:sha256-a8f3c9d2",
      "annotations": {
        "predicateType": "https://in-toto.io/attestation/decision-receipt/v1",
        "signerRole": "supervisor-hook"
      }
    }

    The SLSA envelope is signed by the builder-platform identity only; the receipt envelope is signed by the supervisor-hook identity. A verifier fetches both, checks both signatures against their own identities, cross-references subjects, and routes receipt interpretation by predicateType. Substitution is detected by digest mismatch. Trust-domain separation preserved.

    I filed a PR implementing this pattern on agent-commit v0.2: arewm/refs.arewm.com#2. Additive change, no breaking modifications to the existing observation byproducts schema. The PR also includes a Design Rationale subsection and a parsing rule for ResourceDescriptor-shaped byproducts, so the pattern is formally documented rather than just conventional.

    The in-toto predicate type URI tracks in-toto/attestation#549, which is the parallel PR to formalize decision receipts as a standard attestation predicate. If that URI changes during in-toto review, the reference format updates without a schema change.

    Composition sketch for @piiiico stays on the same rails: separation of signing identities throughout, the builder provenance references receipts via ResourceDescriptor, receipts reference the declared behavioral baseline by digest, telemetry verifies the baseline against the receipt chain.

  17. tomjwxf commented on Apr 17, 2026

    @tomjwxf

    @piiiico @arewm @razashariff — first pass of the composition sketch: gist.

    Scope:

    • Layer inventory for the five attestations (identity, tool integrity, decision receipts, build provenance, behavioral telemetry) with each owner, signer, and primary spec
    • Per-call bundle: exact JSON shape each layer carries for a single tool call, including subject canonicalization so cross-layer linking works
    • Per-session aggregate: which attestations exist per session and how the SLSA provenance references the others via ResourceDescriptor (consistent with arewm/refs.arewm.com#2 and in-toto/attestation#549)
    • Verifier algorithm: pseudocode for a complete session verification walking all five layers with localized trust-domain checks
    • Key rotation / DID continuity: two-key model (long-lived identity-key + short-lived signing-key with cert chaining) addressing the concern @piiiico raised about rotation breaking baseline anchors
    • Deployment tiers: graceful degradation when not all five layers are present (Minimal / Software / Regulated / Full)
    • Five open questions at the bottom for the layer owners to respond to directly

    Five open questions summary:

    1. Subject canonicalization (hash of JCS(tool_name + input + session + sequence) vs alternatives)
    2. Where Layer 1 identity signatures ride (embedded in Layer 3 vs separate byproduct)
    3. Declared-baseline placement in Layer 4 (embedded vs separate signed artifact)
    4. Tier naming (descriptive words vs SLSA-style numeric levels)
    5. Cross-org key resolution for supervisor_hook_pub_key(issuerId) (Fulcio-style transparency logs vs DID resolver vs allowlist)

    Proposed next step, per my reading of the thread:

    • Each layer owner confirms their own layer is represented correctly (@razashariff on 1, @piiiico on 2 and 5, me on 3, @arewm on 4)
    • Iterate on the open questions inline in this thread or on the gist
    • Once the five layers read consistently, publish the combined sketch as a pre-SIEP / pre-I-D companion document anchored somewhere neutral (CSA Agentic Trust Framework is already doing Zero Trust governance adjacent work — see agentic-trust-framework; possible anchor)

    Happy to iterate section-by-section if any of the schemas are wrong. Also happy to collaborate on a shared doc if that works better than GitHub comments for longer back-and-forth.

    Still open to joining the SLSA call when it next meets. Email has been unreliable on my end, so tagging here remains the better channel.

  18. added a commit that references this issue on Apr 18, 2026
  19. piiiico commented on Apr 18, 2026

    @piiiico

    @tomjwxf — reviewed the gist. Clean stitching. I'll confirm layers 2 and 5 and answer the questions relevant to both.

    Layer 2 (tool integrity) — schema looks correct

    The in-toto release predicate over signed MCP manifests is the right shape. Two additions worth noting:

    Tool description integrity is separate from tool availability. The manifest attests to the tool schema as published — but the description fed into LLM context may differ if an intermediary rewrites it in transit. The per-call field in Layer 3 receipts (toolDefDigest) closes this: it records what the agent actually saw, not what was published. Worth surfacing in the layer inventory as "what the agent was given" vs. "what the maintainer published" — right now both are under Layer 2.

    Multiple tool servers per session. The resolvedDependencies list handles this cleanly since it's a list. One attestation per server, different signers. The verifier needs to handle partial verification when some servers are signed and others are not. Suggest: per-server trust tier rather than session-level binary.

    Layer 5 (behavioral telemetry) — schema is correct, one clarification

    The predicate schema in the gist matches what we emit at AgentLair (agentlair.io/attestation/behavioral-telemetry/v0.1). KL divergence against declared baseline is the core signal.

    One clarification on declaredBaselineDigest: the baseline is a distribution object, not a single artifact. It should be a content-addressed reference to a signed JSON artifact containing:

    {
      "agent_did": "did:key:z6Mk...",
      "baseline_window": "2026-04-01T00:00Z/2026-04-14T00:00Z",
      "toolset": { "Bash": 0.60, "Read": 0.30, "WebSearch": 0.10 },
      "sample_size": 4200,
      "signed_by": "did:web:telemetry.agentlair.io"
    }

    This baseline artifact is content-addressed (its digest is declaredBaselineDigest) and signed by the telemetry service. The key insight: the baseline should be declared at provenance time (Layer 4, buildDefinition.internalParameters), so the verifier can confirm that the empirical telemetry in Layer 5 is being compared against the stated baseline — not a post-hoc one the telemetry service chose. Embedding it by digest in Layer 4 is correct.

    Answers to the questions

    1. Subject canonicalization. sha256(JCS({tool_name, tool_input, session_id, sequence})) is correct. The sequence field is load-bearing — without it, two identical tool_name + tool_input calls in the same session are indistinguishable. Use the full form.

    3. declaredBaselineDigest schema. Separate signed artifact referenced from Layer 4 by digest, not embedded. The baseline may be updated across sessions (rolling window) — embedding it forces a new provenance artifact every time the baseline updates. A reference lets the baseline evolve independently.

    4. Tier naming. Prefer the SLSA-style numeric levels (0/1/2/3) for external adoption. "Minimal/Software/Regulated/Full" is descriptive but adds a translation layer for any team already familiar with SLSA levels. Map to: 0 = none, 1 = receipts only (Minimal), 2 = identity + receipts + provenance (Software), 3 = full five-layer (Full/Regulated merged or split at 3/4).

    5. Cross-org key resolution. Fulcio-style CT logs are the right answer for public infrastructure. For enterprise / inter-org deployments, the practical path is a well-known endpoint: {orgDomain}/.well-known/agent-trust-anchors.json listing trusted supervisor-hook DIDs. Sigstore Fulcio handles the public transparency case; the well-known document handles the explicit org-to-org trust case.

    Proposed layer assignments for next iteration

    Confirmed: happy to own layers 2 and 5 in the draft-by-draft iteration.

    Concrete first step: I'll write up the behavioral telemetry predicate schema as a standalone attestation-behavioral-telemetry/v0.1 document (borrowing the SLSA predicate format) that can be cited in the next gist revision. Will link from here when done.

  20. piiiico commented on Apr 18, 2026

    @piiiico

    @tomjwxf @arewm @razashariff — wrote up the Layer 5 behavioral telemetry predicate as a standalone spec:

    Behavioral Telemetry Predicate Specification v0.1

    predicateType: https://agentlair.io/attestation/behavioral-telemetry/v0.1

    Covers:

    • Full predicate schema (in-toto Statement envelope, field definitions with RFC 2119 requirements)
    • Baseline artifact schema (separate signed JSON, content-addressed, referenced from Layer 4 buildDefinition.internalParameters)
    • 6-step verification algorithm (Layer 5 ↔ Layer 4 cross-layer digest verification, independent KL divergence recomputation)
    • Key rotation handling (baseline bound to identity-key DID, not signing-key — rotation does not invalidate baseline)
    • Deployment tier mapping (Layer 5 OPTIONAL at Level 1-2, REQUIRED at Level 3-4, with graceful degradation)
    • Security considerations (adversarial drift mitigations, single-point-of-trust reduction via cross-layer receipt verification)

    Design choices worth flagging:

    1. Baseline is a separate artifact, not embedded. Referenced by digest in both Layer 4 (declaration) and Layer 5 (measurement). This ensures neither the builder nor the telemetry service unilaterally controls what the agent is measured against.
    2. Laplace smoothing for novel tools. Tools in empirical but absent from baseline get 1e-10 additive smoothing. Open question whether a catch-all _other category is better — flagged in the doc.
    3. previousBaselineDigest chains. Each baseline references its predecessor, forming a verifiable evolution chain. Addresses the gradual-drift attack vector Tom raised.

    Happy to iterate. The doc is formatted to be citable alongside the composition gist.

  21. piiiico commented on Apr 18, 2026

    @piiiico

    For context on the broader landscape this thread's composition sketch sits within: I mapped all five agent identity frameworks that shipped in April 2026 (World ID for Agents, DIF MCP-I, Microsoft AGT, Curity Access Intelligence, Armalo AI) to the five-layer model we've been building here.

    The Agent Identity Stack: What Shipped in April 2026

    Key takeaway relevant to this thread: everyone built L1–L3. The three gaps that remain — tool-call authorization, permission lifecycle, ghost agent offboarding — are all structurally cross-organizational problems. The composition sketch @tomjwxf drafted is the right shape for closing them.

  22. arewm commented on Apr 27, 2026

    @arewm
    Member

    I am going to close this issue. I think that the overlap of this use case with the current SLSA build provenance is too high to warrant anything new at this point in time. Any potential contributor is welcome to join the SLSA meetings to discuss new potential tracks.

  23. moved this from 🆕 New to ✅ Done in Issue triageon Apr 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions