Skip to content

Map: Context-graph emergence pipeline — reconciliation, retrieval, eval #297

Description

@antejavor

Destination

A specced direction for context-graph's emergence pipeline — reconciliation/decay, retrieval+efficiency, and the eval loop driving both — under a reframed mission (organizational-knowledge emergence from agents, for agent/harness consumption — not session/tool storage). Ready to implement, not yet built.

Notes

Domain: context-graph family, memgraph/ai-toolkit. See context-graph/CONTEXT-MAP.md, reframed by this map's first resolved ticket.

Skills every session should consult: /grilling + /domain-modeling for HITL tickets; a /research subagent for research tickets — but constrain fan-out: #308's first attempt spawned 8 concurrent research subagents, exhausted the API session quota before writing anything, and had to be redone manually with sequential lookups and authenticated gh api.

This map carries execution (amended 2026-08-24, same override map #288 used). The eval-loop decisions are concrete enough to build against, so implementing them is the natural next step rather than a /to-spec handoff. Specifically, Tier 1 of the eval loop is fully unblocked: corpus format (#302), benchmark choice and fetch-and-convert model (#308), tier structure (#303), metrics and judge (#304), efficiency and isolation (#309), and the retrieval baseline (#300) are all settled. The buildable vertical slice is: fetch/convert LongMemEval v1 → inject into a dedicated eval database → reconcile → retrieval agent queries via run_cypher_query → score with ContextualRecall + one GEval + the deterministic payload-token count → report.

The three remaining open tickets do not block that slice, and two are better answered after it runs:

  • #305 (comparison report) is a prototype ticket — prototypes need real output to react to; designing a report format before any score exists is guessing at the shape of data nobody has seen.
  • #306 (cadence) depends on what a run actually costs and how long it takes — both facts obtained by running it.
  • #307 (gold-slice scripting) blocks the gold slice only, not Tier 1 or the injected part of Tier 2.

Standing preferences for this effort:

Decisions so far

  • Grilling: what should the eval loop score, and how? — blended end-to-end score (mostly recall via retrieval, structural checks lighter-weighted); synthetic-first unified corpus, org-signal-question framed; hybrid task-gen (bulk skip-harness + small real-harness gold slice) now, harden later; graded rubric (coverage+efficiency) via LLM-as-judge; human-gated promotion for this effort, full-auto is longer-term direction only.
  • Grilling: what's the retrieval baseline for eval v1 to run against? — existing run_cypher_query MCP tool, zero new build; dedicated retrieval layer deferred, to be shaped by where this baseline fails.
  • Domain-modeling: reframe context-graph's mission in CONTEXT-MAP.md — Context Graph's definition reworded to center raw→emerge (Collection Tier → Memory Tier) into durable organizational knowledge, for agents/harness to consume. Kept "avoid: Knowledge graph" (mechanism, not mission); rejected minting a new "Knowledge Lake" term as redundant with Collection/Memory Tier.
  • Grilling: what format does the eval corpus have, and where does it live? — deepeval's Golden as-is (no custom schema), serialized JSONL, committed to git. Memgraph-as-store rejected ("keep the ruler outside the thing being measured" — a schema migration could silently corrupt the answer key); Confident AI cloud rejected on the project's own owned-IP thesis. Upstream benchmark content is fetch-and-convert against a pinned release, committing only converted output, not vendored blobs.
  • Grilling: how does the LLM-as-judge scoring mechanism get built? — rubric splits by mechanism: LLM judges quality, plain code counts cost (refines Grilling: what should the eval loop score, and how? #299's "all via LLM-judge"). Quality = a minimal pair, built-in ContextualRecallMetric + one custom GEval coverage rubric; Faithfulness/AnswerRelevancy skipped at v1 on judge-call cost. Judge is Anthropic pinned to a dated model ID (decorrelates from the OpenAI-backed extraction pipeline; a moving alias would break cross-version comparability). Calibration = two small occasional checks, repeat-and-compare for noise plus ~25 human-graded items for bias.
  • Grilling: how is the deterministic efficiency metric instrumented, and how do eval runs avoid polluting the graph under test? — efficiency is retrieval payload size (tokens handed back to answer), not agent-side burn: a deterministic count over Golden.retrieval_context, no adapter work. Coverage is a hard gate, efficiency ranks within it (a weighted composite would let coverage be traded away invisibly). Isolation = one Memgraph database per eval batch, reset before fixture injection, using the existing MEMGRAPH_DATABASE config; batch-shared beats per-question because distractors are what make precision/efficiency meaningful. Markers inside a batch are provenance-only — never the mechanism keeping eval-agent traces out. Corrected during implementation (2026-08-25): per-batch databases need Memgraph Enterprise (multi-tenancy), so isolation is instead a dedicated eval instance cleared before each batch — same guarantee, community licence, no Enterprise secret in CI.
  • Research: which memory benchmark should the eval corpus derive from, and what does each actually exercise? — no benchmark ships both agent-harness activity and a gold-answer key; the field splits. Adopt LongMemEval v1 (MIT) for the text path and LongMemEval-V2 (Apache-2.0) for the action path; decline LoCoMo (CC BY-NC — excludes commercial use); hold the CC BY 4.0 SWE trajectory corpora in reserve as content for authored questions. Residual gap, all synthetic: skills-graph and subagent nesting are uncovered by every candidate, and the organizational/multiplayer question framing appears in none. Findings: docs/research/2026-08-memory-benchmarks.md on branch research/memory-benchmarks.
  • Grilling: who/what authors the eval corpus, and how? — two tiers scored separately, never blended: Tier 1 adopted from upstream (~951 questions, converted not authored, asks "does recall work mechanically"), Tier 2 hand-authored org-signal questions (asks "does it work for what we're building", and is what promotion hangs on). Both start small, scale only when the signal is too noisy. Tier 2 authorship splits by judgment: human writes questions + gold answers, fixture content comes from real traces rather than LLM prose (tidy synthetic fixtures would flatter retrieval). The real-harness gold slice is a subset of Tier 2, sized by component coverage — post-Research: which memory benchmark should the eval corpus derive from, and what does each actually exercise? #308 it is the only eval coverage skills-graph and subagent nesting get at all.
  • Grilling: what triggers an eval run, and how often? — two tiers. The fast, LLM-free tests run on every PR with no key provided (they skip themselves), staying free and deterministic. Real eval batches run on workflow_dispatch only, never on push/PR: cost (~2 LLM calls per session, ~900 sessions per 20-question batch), non-determinism (Grilling: how does the LLM-as-judge scoring mechanism get built? #304 handles that with calibration, not a fixed threshold), and decisively — a CI gate on a judged score is the automatic promotion Grilling: what should the eval loop score, and how? #299 ruled out in favour of a human reading a report. Not scheduled either: nothing consumes the output automatically, so a timer would buy reports nobody asked for. Implemented in PR #311.
  • Prototype: what does the human-gated comparison report look like? — verdict-first with an explicit noise floor, chosen over a neutral table and a regressions-first layout. Two behaviours carry it: the report refuses to compare runs pinned to a different corpus/judge/tokenizer (a delta across pins measures the pin change as if it were the change under test), and it will not call any delta real without calibration — coverage_is_real is three-valued, since "not real" and "cannot tell" are different answers to a decision-maker. A real coverage regression outranks any efficiency gain; efficiency alone never declares an improvement. Surfaced a sizing caveat for Grilling: how does the LLM-as-judge scoring mechanism get built? #304's calibration work: at 20 questions one flip is 5pp, so coverage granularity is coarser than a plausible noise floor.
  • Grilling: how does the gold-slice real-harness task get scripted? — facts planted in the session prompt (file-carried facts an explicit later expansion); the planting prompt lives on the Golden itself as additional_metadata["session_prompt"], not a joined fixture file that could drift. Sized by carrier — where the fact physically lives in the graph — rather than by component, since test-graph-model already verifies shape and this verifies recall. Starts with exactly one question: a fact stated only inside a subagent, because top-level recall is already covered by Tier 1 and the nested carrier has a demonstrated silent-failure mode (Grilling: does the Episode/Memory model need any change given the nesting redesign? #281's single-hop HAS_ACTION). Reuses test-graph-model's hard-won techniques, including avoiding literal SKILL.md paths in the prompt (skills-graph: SkillGraphConnector can record a false-positive Skill Usage from a prompt merely mentioning a SKILL.md path #293).

Not yet specified

  • Building the gold slice's one question — #307 decided its shape (a fact stated only inside a subagent, planting prompt on the Golden, reusing test-graph-model's techniques). Writing the runner that drives that live session and checks recall is implementation, not a further decision.
  • Reconciliation/emergence/decay design itself — what triggers decay, where the collection-tier/memory-tier boundary sits under the new organizational-knowledge framing, how existing sessions-graph reconciliation and skills-graph procedural mining fold into it. Waits on eval v1 producing a first score to design against.
  • Dedicated retrieval v2 (beyond the crude run_cypher_query baseline) — waits on where the baseline actually fails. Known prior art/debt to fold in: lightrag-memgraph's open vector-query bugs (perf(lightrag-memgraph): vector query() full label scan defeats vector index; closed: lightrag-memgraph: vector query() silently drops dead candidates with no signal).
  • "Efficiency" as its own axis, distinct from eval's coverage/efficiency answer-rubric — not yet pinned down as a separate concern (e.g. query cost/latency at scale) or folded into the same rubric.

Out of scope

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions