Destination
Redrawn 2026-08-17 — this map's original destination (below, reached) is joined by a second one, brought back in from Out of scope at the user's explicit request rather than spun into a separate map:
Original destination (reached): A repeatable way to verify, against a real Claude Code session (not synthetic e2e data), that the three-tier SPAWNED subagent-nesting inference rule (Map: Harness-grounded graph model across the Context Graph family, ticket #277) actually produces the correct graph shape when a real model spawns exactly one subagent (Tier 1 only). Shipped as scripts/dev-memgraph.sh verify-nesting (since expanded and renamed test-graph-model) and a manual-trigger CI job.
New destination: Reduce test-mocking across the context-graph family where it gives false confidence — specifically, tests that assert on a Cypher query string via a mocked Memgraph client rather than proving the query actually executes correctly (the exact class of gap that let both of #290's real bugs through undetected). Convert the ones with no existing e2e coverage; delete the ones that are already redundantly proven by a real test_e2e.py; document the resulting policy in one shared place so future tests don't re-litigate it. Concretely:
skills-graph's test_skill_graph.py: delete the ~half already behaviorally proven by test_e2e.py's 6 real tests (add_skill, get_skill, update_skill, delete_skill, list_skills, search_by_name, dependency management); convert the unique half (record_skill_usage's branching — create_missing, container-vs-session, agent-fallback; setup/drop) from string-matching to real e2e assertions on the resulting graph shape.
unstructured2graph's test_memgraph.py: same delete-redundant/convert-unique split against its own test_e2e.py — needs its own file-by-file audit first (the skills-graph split above is already fully worked out; this one isn't yet).
sessions-graph's test_reconciliation.py: swap the hand-rolled _fake_actions_graph stub for a real e2e-backed ActionsGraph instance with real recorded actions, keep the LLM mocked. Gives schema-drift protection between sessions-graph and actions-graph without needing OPENAI_API_KEY (unlike test_e2e_reconciliation.py, which already does this for real but only runs when that key is set).
scripts/dev-memgraph.sh test-graph-model: add assertions for sessions-graph's own output ((:User)-[:HAD_SESSION]->(:Session), Session.reconciliation_status) — it's already wired into that live run and produces this data today, just unchecked. Same session, no new LLM cost.
- Policy: extend
context-graph/CONTEXT-MAP.md with the mock-vs-e2e reasoning above (what stays mocked and why: hooks/wiring translation tests, pure model validation; what goes e2e and why: anything asserting real Cypher correctness) so it's written down once, not re-decided per test file.
End state: all five land, CONTEXT-MAP.md documents why, and the map closes again.
Notes
Domain: context-graph in memgraph/ai-toolkit — skills-graph/tests/, unstructured2graph/tests/, sessions-graph/tests/, scripts/dev-memgraph.sh, context-graph/CONTEXT-MAP.md.
This map carries execution (unchanged from the original charting) — the five items above are fully decided via /grilling; building them is the natural next step, not a separate handoff. No genuine blocker remains this time (unlike the original destination's ANTHROPIC_API_KEY provisioning task) — everything is ready to execute directly.
Explicitly considered and rejected during grilling:
- Deleting all of
test_skill_graph.py's mocked tests wholesale — rejected because roughly half (the record_skill_usage branching logic) has no e2e coverage anywhere else; deleting those would be a real regression in coverage, not a cleanup.
- Leaving
test_reconciliation.py's ActionsGraph fake as-is — initially agreed, then reversed once it became clear test_e2e_reconciliation.py's real coverage of the same risk is gated behind OPENAI_API_KEY and so doesn't protect a contributor who lacks that key locally.
- A dedicated new
docs/agents/testing-strategy.md for the policy — rejected in favor of extending the existing family-wide CONTEXT-MAP.md, since the reasoning is identical across all five packages, not package-specific domain knowledge.
Decisions so far
Not yet specified
(none — the new destination's scope is fully decided; the one remaining unknown, unstructured2graph's exact delete/convert split, is a mechanical audit folded into ticket #291 itself, not a separate open question.)
Out of scope
Destination
Redrawn 2026-08-17 — this map's original destination (below, reached) is joined by a second one, brought back in from Out of scope at the user's explicit request rather than spun into a separate map:
Original destination (reached): A repeatable way to verify, against a real Claude Code session (not synthetic e2e data), that the three-tier
SPAWNEDsubagent-nesting inference rule (Map: Harness-grounded graph model across the Context Graph family, ticket #277) actually produces the correct graph shape when a real model spawns exactly one subagent (Tier 1 only). Shipped asscripts/dev-memgraph.sh verify-nesting(since expanded and renamedtest-graph-model) and a manual-trigger CI job.New destination: Reduce test-mocking across the context-graph family where it gives false confidence — specifically, tests that assert on a Cypher query string via a mocked Memgraph client rather than proving the query actually executes correctly (the exact class of gap that let both of #290's real bugs through undetected). Convert the ones with no existing e2e coverage; delete the ones that are already redundantly proven by a real
test_e2e.py; document the resulting policy in one shared place so future tests don't re-litigate it. Concretely:skills-graph'stest_skill_graph.py: delete the ~half already behaviorally proven bytest_e2e.py's 6 real tests (add_skill,get_skill,update_skill,delete_skill,list_skills,search_by_name, dependency management); convert the unique half (record_skill_usage's branching —create_missing, container-vs-session, agent-fallback;setup/drop) from string-matching to real e2e assertions on the resulting graph shape.unstructured2graph'stest_memgraph.py: same delete-redundant/convert-unique split against its owntest_e2e.py— needs its own file-by-file audit first (the skills-graph split above is already fully worked out; this one isn't yet).sessions-graph'stest_reconciliation.py: swap the hand-rolled_fake_actions_graphstub for a real e2e-backedActionsGraphinstance with real recorded actions, keep the LLM mocked. Gives schema-drift protection between sessions-graph and actions-graph without needingOPENAI_API_KEY(unliketest_e2e_reconciliation.py, which already does this for real but only runs when that key is set).scripts/dev-memgraph.sh test-graph-model: add assertions forsessions-graph's own output ((:User)-[:HAD_SESSION]->(:Session),Session.reconciliation_status) — it's already wired into that live run and produces this data today, just unchecked. Same session, no new LLM cost.context-graph/CONTEXT-MAP.mdwith the mock-vs-e2e reasoning above (what stays mocked and why: hooks/wiring translation tests, pure model validation; what goes e2e and why: anything asserting real Cypher correctness) so it's written down once, not re-decided per test file.End state: all five land,
CONTEXT-MAP.mddocuments why, and the map closes again.Notes
Domain:
context-graphinmemgraph/ai-toolkit—skills-graph/tests/,unstructured2graph/tests/,sessions-graph/tests/,scripts/dev-memgraph.sh,context-graph/CONTEXT-MAP.md.This map carries execution (unchanged from the original charting) — the five items above are fully decided via
/grilling; building them is the natural next step, not a separate handoff. No genuine blocker remains this time (unlike the original destination'sANTHROPIC_API_KEYprovisioning task) — everything is ready to execute directly.Explicitly considered and rejected during grilling:
test_skill_graph.py's mocked tests wholesale — rejected because roughly half (therecord_skill_usagebranching logic) has no e2e coverage anywhere else; deleting those would be a real regression in coverage, not a cleanup.test_reconciliation.py'sActionsGraphfake as-is — initially agreed, then reversed once it became cleartest_e2e_reconciliation.py's real coverage of the same risk is gated behindOPENAI_API_KEYand so doesn't protect a contributor who lacks that key locally.docs/agents/testing-strategy.mdfor the policy — rejected in favor of extending the existing family-wideCONTEXT-MAP.md, since the reasoning is identical across all five packages, not package-specific domain knowledge.Decisions so far
gh secret list. Unblocked the manual-trigger CI job piece of the original destination.agent-context-graph hook run claude-codeneeding explicit--connectorflags;actions-graph'sagent_spawning_tool_namesdefaulting to{"Task"}when real payloads report"Agent"— correction posted to Grilling: formalize the subagent-nesting parent-inference rule (including parallel/batch tool calls) #277).reduce-test-mocking). unstructured2graph's split turned out bigger than scoped: the entirepromote_entity_types_to_labels/promote_all_entity_types_to_labelsfamily had zero live coverage anywhere, not just a redundant mocked tier. Full suite green, livetest-graph-modelrun confirms all 13 checks (including the 2 new sessions-graph ones) pass against a real session.Not yet specified
(none — the new destination's scope is fully decided; the one remaining unknown,
unstructured2graph's exact delete/convert split, is a mechanical audit folded into ticket #291 itself, not a separate open question.)Out of scope