Skip to content

context-graph: test-graph-model — verify the graph model against a real Claude Code session (map #288) - #290

Merged
antejavor merged 3 commits into
mainfrom
live-nesting-verification
Aug 17, 2026
Merged

antejavor merged 3 commits into
mainfrom
live-nesting-verification

Conversation

@antejavor

@antejavor antejavor commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Implements the destination of wayfinder map #288: a repeatable way to verify, against a real Claude Code session (not synthetic e2e data), that a broad slice of the actions-graph/skills-graph model actually holds up.

Two real bugs found, not just tooling built

Running this against an actual live session (rather than reasoning from docs) surfaced two genuine, previously-unknown gaps:

  1. agent-context-graph hook run claude-code silently does nothing without explicit --connector skills-graph --connector actions-graph --connector sessions-graph flags — undocumented outside the installed marketplace plugin's own generated command.
  2. actions-graph's agent_spawning_tool_names defaulted to {"Task"}, but real PreToolUse/PostToolUse hook payloads report the spawning tool as tool_name: "Agent" — "Task" is only how Claude Code's CLI/UI refers to it. This meant Tier 1 of Grilling: formalize the subagent-nesting parent-inference rule (including parallel/batch tool calls) #277's inference rule had never actually matched anything against a real session. Fixed to {"Agent", "Task"}, with a regression test locking in the real name. Correction posted directly to Grilling: formalize the subagent-nesting parent-inference rule (including parallel/batch tool calls) #277, since it invalidates that ticket's doc-derived assumption.

Also worked around a PATH ambiguity discovered along the way: a stale, globally uv tool installed agent-context-graph can shadow this repo's own workspace source. The hook command now always resolves via uv run --package agent-context-graph, deterministically, regardless of what else is installed.

Test plan

  • scripts/dev-memgraph.sh test-graph-model run against a real Claude Code session — all 11 checks passing (plus 3 earlier clean runs of the narrower pre-expansion check)
  • Full suite (55 tests) green via scripts/dev-memgraph.sh test
  • ruff check / ruff format --check clean
  • CI workflow itself (test-live-graph-model.yaml) needs its first live workflow_dispatch trigger post-merge to fully validate the fresh-runner path (install steps, PATH wiring) — validated locally against every piece it depends on, but a GitHub-hosted runner run is the real proof

…ssion (map #288)

Adds scripts/dev-memgraph.sh verify-nesting and a manual-trigger CI workflow
that drive a real, non-interactive claude -p session restricted to the Task
tool (forcing genuine subagent delegation), then assert the Tier 1 graph
shape (HAS_AGENT, SPAWNED, HAS_ACTION) that #277's three-tier inference rule
is supposed to produce -- the live-session verification that ticket flagged
as required but never done, only checked against synthetic e2e data until now.

Two real bugs surfaced running this against an actual session, not doc
research:

- `agent-context-graph hook run claude-code` silently no-ops without explicit
  --connector flags, undocumented outside the installed marketplace plugin's
  own generated command.
- actions-graph's agent_spawning_tool_names defaulted to {"Task"}, but real
  hook payloads report the spawning tool as tool_name "Agent" -- "Task" is
  only Claude Code's CLI/UI-facing name for it. Tier 1 of #277's rule had
  never actually matched anything against a real session as a result. Fixed
  to {"Agent", "Task"}, with a regression test locking in the real name.

Verified three consecutive times against a live Claude Code session before
landing this.
Expands the check into a broader assertion of the actions-graph/skills-graph
model against the same real, live Claude Code session, not just the one
Agent/SPAWNED mechanism from before:

- FOLLOWED_BY: the subagent's own tool-call chain is correctly ordered, and
  never crosses into the top-level session's own separate chain (#278).
- PARENT_OF: ToolCall -> ToolResult correlation holds at both the top level
  (the spawning call itself) and inside the subagent.
- USED_TOOL: the subagent's tool calls link to real Tool nodes.
- USED_SKILL: attaches to the Agent, not the Session, when the skill read
  happens inside a subagent (#280) -- verified live for the first time by
  folding a real SKILL.md read into the same subagent call, so this still
  costs exactly one LLM session rather than a second one.

All 11 checks passed against a real live Claude Code session before landing
this (on top of the three earlier clean runs of the narrower check).
@antejavor antejavor changed the title context-graph: verify SPAWNED inference against a real Claude Code session (map #288) context-graph: test-graph-model — verify the graph model against a real Claude Code session (map #288) Aug 17, 2026
Comments should stand on their own without needing an external tracker
lookup -- ticket numbers belong in the PR description and commit history,
not inline. No behavior change.
@antejavor
antejavor merged commit 663b82d into main Aug 17, 2026
16 checks passed
@antejavor
antejavor deleted the live-nesting-verification branch August 17, 2026 21:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant