Repository navigation
context-graph: test-graph-model — verify the graph model against a real Claude Code session (map #288) - #290
Merged
Merged
Conversation
…ssion (map #288) Adds scripts/dev-memgraph.sh verify-nesting and a manual-trigger CI workflow that drive a real, non-interactive claude -p session restricted to the Task tool (forcing genuine subagent delegation), then assert the Tier 1 graph shape (HAS_AGENT, SPAWNED, HAS_ACTION) that #277's three-tier inference rule is supposed to produce -- the live-session verification that ticket flagged as required but never done, only checked against synthetic e2e data until now. Two real bugs surfaced running this against an actual session, not doc research: - `agent-context-graph hook run claude-code` silently no-ops without explicit --connector flags, undocumented outside the installed marketplace plugin's own generated command. - actions-graph's agent_spawning_tool_names defaulted to {"Task"}, but real hook payloads report the spawning tool as tool_name "Agent" -- "Task" is only Claude Code's CLI/UI-facing name for it. Tier 1 of #277's rule had never actually matched anything against a real session as a result. Fixed to {"Agent", "Task"}, with a regression test locking in the real name. Verified three consecutive times against a live Claude Code session before landing this.
Expands the check into a broader assertion of the actions-graph/skills-graph model against the same real, live Claude Code session, not just the one Agent/SPAWNED mechanism from before: - FOLLOWED_BY: the subagent's own tool-call chain is correctly ordered, and never crosses into the top-level session's own separate chain (#278). - PARENT_OF: ToolCall -> ToolResult correlation holds at both the top level (the spawning call itself) and inside the subagent. - USED_TOOL: the subagent's tool calls link to real Tool nodes. - USED_SKILL: attaches to the Agent, not the Session, when the skill read happens inside a subagent (#280) -- verified live for the first time by folding a real SKILL.md read into the same subagent call, so this still costs exactly one LLM session rather than a second one. All 11 checks passed against a real live Claude Code session before landing this (on top of the three earlier clean runs of the narrower check).
Comments should stand on their own without needing an external tracker lookup -- ticket numbers belong in the PR description and commit history, not inline. No behavior change.
4 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Implements the destination of wayfinder map #288: a repeatable way to verify, against a real Claude Code session (not synthetic e2e data), that a broad slice of the actions-graph/skills-graph model actually holds up.
scripts/dev-memgraph.sh test-graph-model— drives a scripted, non-interactiveclaude -psession restricted to the Task tool (forcing genuine subagent delegation), where the one spawned subagent both reads a realSKILL.mdfile and searches the repo — one LLM session, eleven assertions:Session -[:HAS_AGENT]-> Agent,Action -[:SPAWNED]-> Agent(the real spawning tool call),Agent -[:HAS_ACTION]-> Action(Grilling: formalize the subagent-nesting parent-inference rule (including parallel/batch tool calls) #277/Grilling: does a subagent instance become a first-class Agent node, or stay a flat Action pair with a populated PARENT_OF? #278)FOLLOWED_BY: the subagent's own chain is correctly ordered, and never crosses into the top-level session's own separate chain (Grilling: does a subagent instance become a first-class Agent node, or stay a flat Action pair with a populated PARENT_OF? #278)PARENT_OF:ToolCall→ToolResultcorrelation at both the top level and inside the subagentUSED_TOOL: the subagent's tool calls link to realToolnodesUSED_SKILL: attaches to theAgent, not theSession, when the skill read happens inside a subagent (Grilling: does Skill Usage need to attach to the specific nested Agent, not just the flat Session? #280) — verified live for the first time.github/workflows/test-live-graph-model.yaml— manualworkflow_dispatchjob wrapping the same check. Deliberately not on every push/PR: it drives a real, billed LLM session, and subagent-spawning is inherently non-deterministic — worth having on tap for deliberate use when changing something heavily in this area, not a PR gate.Two real bugs found, not just tooling built
Running this against an actual live session (rather than reasoning from docs) surfaced two genuine, previously-unknown gaps:
agent-context-graph hook run claude-codesilently does nothing without explicit--connector skills-graph --connector actions-graph --connector sessions-graphflags — undocumented outside the installed marketplace plugin's own generated command.actions-graph'sagent_spawning_tool_namesdefaulted to{"Task"}, but realPreToolUse/PostToolUsehook payloads report the spawning tool astool_name: "Agent"—"Task"is only how Claude Code's CLI/UI refers to it. This meant Tier 1 of Grilling: formalize the subagent-nesting parent-inference rule (including parallel/batch tool calls) #277's inference rule had never actually matched anything against a real session. Fixed to{"Agent", "Task"}, with a regression test locking in the real name. Correction posted directly to Grilling: formalize the subagent-nesting parent-inference rule (including parallel/batch tool calls) #277, since it invalidates that ticket's doc-derived assumption.Also worked around a PATH ambiguity discovered along the way: a stale, globally
uv tool installedagent-context-graphcan shadow this repo's own workspace source. The hook command now always resolves viauv run --package agent-context-graph, deterministically, regardless of what else is installed.Test plan
scripts/dev-memgraph.sh test-graph-modelrun against a real Claude Code session — all 11 checks passing (plus 3 earlier clean runs of the narrower pre-expansion check)scripts/dev-memgraph.sh testruff check/ruff format --checkcleantest-live-graph-model.yaml) needs its first liveworkflow_dispatchtrigger post-merge to fully validate the fresh-runner path (install steps, PATH wiring) — validated locally against every piece it depends on, but a GitHub-hosted runner run is the real proof