Skip to content

Map: Live Claude Code verification for the SPAWNED subagent-nesting inference rule #288

Description

@antejavor

Destination

Redrawn 2026-08-17 — this map's original destination (below, reached) is joined by a second one, brought back in from Out of scope at the user's explicit request rather than spun into a separate map:

Original destination (reached): A repeatable way to verify, against a real Claude Code session (not synthetic e2e data), that the three-tier SPAWNED subagent-nesting inference rule (Map: Harness-grounded graph model across the Context Graph family, ticket #277) actually produces the correct graph shape when a real model spawns exactly one subagent (Tier 1 only). Shipped as scripts/dev-memgraph.sh verify-nesting (since expanded and renamed test-graph-model) and a manual-trigger CI job.

New destination: Reduce test-mocking across the context-graph family where it gives false confidence — specifically, tests that assert on a Cypher query string via a mocked Memgraph client rather than proving the query actually executes correctly (the exact class of gap that let both of #290's real bugs through undetected). Convert the ones with no existing e2e coverage; delete the ones that are already redundantly proven by a real test_e2e.py; document the resulting policy in one shared place so future tests don't re-litigate it. Concretely:

  1. skills-graph's test_skill_graph.py: delete the ~half already behaviorally proven by test_e2e.py's 6 real tests (add_skill, get_skill, update_skill, delete_skill, list_skills, search_by_name, dependency management); convert the unique half (record_skill_usage's branching — create_missing, container-vs-session, agent-fallback; setup/drop) from string-matching to real e2e assertions on the resulting graph shape.
  2. unstructured2graph's test_memgraph.py: same delete-redundant/convert-unique split against its own test_e2e.py — needs its own file-by-file audit first (the skills-graph split above is already fully worked out; this one isn't yet).
  3. sessions-graph's test_reconciliation.py: swap the hand-rolled _fake_actions_graph stub for a real e2e-backed ActionsGraph instance with real recorded actions, keep the LLM mocked. Gives schema-drift protection between sessions-graph and actions-graph without needing OPENAI_API_KEY (unlike test_e2e_reconciliation.py, which already does this for real but only runs when that key is set).
  4. scripts/dev-memgraph.sh test-graph-model: add assertions for sessions-graph's own output ((:User)-[:HAD_SESSION]->(:Session), Session.reconciliation_status) — it's already wired into that live run and produces this data today, just unchecked. Same session, no new LLM cost.
  5. Policy: extend context-graph/CONTEXT-MAP.md with the mock-vs-e2e reasoning above (what stays mocked and why: hooks/wiring translation tests, pure model validation; what goes e2e and why: anything asserting real Cypher correctness) so it's written down once, not re-decided per test file.

End state: all five land, CONTEXT-MAP.md documents why, and the map closes again.

Notes

Domain: context-graph in memgraph/ai-toolkit — skills-graph/tests/, unstructured2graph/tests/, sessions-graph/tests/, scripts/dev-memgraph.sh, context-graph/CONTEXT-MAP.md.

This map carries execution (unchanged from the original charting) — the five items above are fully decided via /grilling; building them is the natural next step, not a separate handoff. No genuine blocker remains this time (unlike the original destination's ANTHROPIC_API_KEY provisioning task) — everything is ready to execute directly.

Explicitly considered and rejected during grilling:

  • Deleting all of test_skill_graph.py's mocked tests wholesale — rejected because roughly half (the record_skill_usage branching logic) has no e2e coverage anywhere else; deleting those would be a real regression in coverage, not a cleanup.
  • Leaving test_reconciliation.py's ActionsGraph fake as-is — initially agreed, then reversed once it became clear test_e2e_reconciliation.py's real coverage of the same risk is gated behind OPENAI_API_KEY and so doesn't protect a contributor who lacks that key locally.
  • A dedicated new docs/agents/testing-strategy.md for the policy — rejected in favor of extending the existing family-wide CONTEXT-MAP.md, since the reasoning is identical across all five packages, not package-specific domain knowledge.

Decisions so far

Not yet specified

(none — the new destination's scope is fully decided; the one remaining unknown, unstructured2graph's exact delete/convert split, is a mechanical audit folded into ticket #291 itself, not a separate open question.)

Out of scope

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions