Skip to content

Research: how GitHub resource reads appear in harness hook payloads #455

Description

@antejavor

Part of #454

Question

How does a GitHub resource read show up in harness activity today? For each harness we support (Claude Code first, then Codex, Copilot CLI, Cursor, OpenCode, Antigravity CLI, Grok Build as far as docs allow): which tool calls fetch GitHub content — built-in web fetch, gh/curl via shell, GitHub MCP server tools, file reads of a cloned repo — and what do their hook payloads carry (tool name, input URL/args, full vs truncated/summarised result, size limits)? What reaches our Event Protocol and Actions Graph today? Can we distinguish a resource the user handed the agent (pasted URL, attached text in the prompt) from one the agent fetched itself?

Sources: this repo's adapters in context-graph/agent-context-graph + real Action nodes from scripts/dev-memgraph.sh exploration data, plus each harness's hook/tool docs.

Activity

self-assigned this
on Oct 7, 2026

antejavor commented on Oct 7, 2026

@antejavor
ContributorAuthor

Answer

The pre-tool input reliably tells us which resource was read in every harness except Antigravity, whose hooks carry no prompt and no output. The post-tool output is not a reliable copy of what the resource contains.

Key findings

  • Four ways GitHub content arrives:

    • the web-fetch tool (WebFetch / web_fetch / webfetch / read_url_content);
    • a shell command (gh issue view …, gh api …, curl …);
    • a GitHub MCP tool (e.g. mcp__github__*; Copilot CLI has github-mcp-server built in and on by default);
    • a file read of a local clone. A local clone gives no GitHub provenance unless a clone in the same session links the path to a repo.
  • Claude Code WebFetch returns {bytes, code, codeText, result, durationMs, url}. For most fetches result is a small model's answer to the agent's prompt, not the page. In our exploration Memgraph, a 524 KB page came back as about 500 characters. Some fetches skip that summarising step: one Markdown docs page came back whole. Which fetches skip it is undocumented.

  • Shell and MCP output does reach the hooks, but it may be truncated:

    • Claude Code Bash keeps about 30K characters inline;
    • Claude Code MCP is capped at 25K tokens by default;
    • Codex truncates shell output by policy but passes MCP results whole;
    • OpenCode cuts at 2,000 lines / 50 KB;
    • Grok Bash cuts at 20 KB.

    Codex's hosted web search fires no hook at all. Antigravity's PostToolUse has no output field.

  • Today every read becomes a ToolCall → ToolResult Action pair:

    • Input and output are stored as JSON strings in properties, and Claude Code also copies the whole tool_response into metadata.
    • A WebFetch result is stored as a Python repr of the dict.
    • Nothing parses the URL or gh argv into an identity.
    • Reconciliation cuts each result to 8,000 characters before chunking.
  • User-handed resources can be told apart from agent fetches:

    • A pasted URL or pasted body reaches us only through the user-prompt hook (as a UserMessage), never as a tool call. Cursor also sends structured attachments, and OpenCode sends parts.
    • We can link a pasted URL to a later agent fetch of the same URL within the session, but nothing does that today.
    • Caveats: Claude Code's UserPromptSubmit also fires for scheduled tasks, subagent reports and cross-session messages, so not every prompt comes from a human. Antigravity has no prompt hook.

Implications

  • Capture model: take a Resource's identity from the tool input (URL, parsed gh/curl argv, GitHub MCP args) and from URLs in the user prompt. Record reads as pointers ("session S read R via tool T, with the agent's fetch prompt, provenance FETCHED or HANDED_BY_USER"). Fetch canonical content ourselves, keyed on that identity, instead of trusting hook output, which is summarised or truncated.
  • Cache hit: a cached WebFetch answer cannot be replayed, because it is specific to the prompt. Hooks cannot swap tool output across harnesses (additionalContext is capped at 10K characters), so memory has to serve re-reads through recall over our stored content, not by intercepting the fetch. The harness's own caches are short-lived (WebFetch keeps results for 15 minutes), so our cache only adds value across sessions and users.

Open questions

  • Once Claude Code's inline output limit is exceeded, does PostToolUse.tool_response hold the full output or the persisted-file stub?
  • Which WebFetch calls skip summarisation?
  • Grok's real output payload is undocumented.
  • Should the Claude Code adapter unwrap WebFetch.result and stop duplicating tool_response into metadata?
  • Should Antigravity capture read transcriptPath?

Context: docs/research/harness-github-resource-reads.md on branch research/harness-github-resource-reads.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions