You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A specced direction for context-graph's emergence pipeline — reconciliation/decay, retrieval+efficiency, and the eval loop driving both — under a reframed mission (organizational-knowledge emergence from agents, for agent/harness consumption — not session/tool storage). Ready to implement, not yet built.
Notes
Domain: context-graph family, memgraph/ai-toolkit. See context-graph/CONTEXT-MAP.md, reframed by this map's first resolved ticket.
Skills every session should consult: /grilling + /domain-modeling for HITL tickets; a /research subagent for research tickets — but constrain fan-out: #308's first attempt spawned 8 concurrent research subagents, exhausted the API session quota before writing anything, and had to be redone manually with sequential lookups and authenticated gh api.
This map carries execution (amended 2026-08-24, same override map #288 used). The eval-loop decisions are concrete enough to build against, so implementing them is the natural next step rather than a /to-spec handoff. Specifically, Tier 1 of the eval loop is fully unblocked: corpus format (#302), benchmark choice and fetch-and-convert model (#308), tier structure (#303), metrics and judge (#304), efficiency and isolation (#309), and the retrieval baseline (#300) are all settled. The buildable vertical slice is: fetch/convert LongMemEval v1 → inject into a dedicated eval database → reconcile → retrieval agent queries via run_cypher_query → score with ContextualRecall + one GEval + the deterministic payload-token count → report.
The three remaining open tickets do not block that slice, and two are better answered after it runs:
#305 (comparison report) is a prototype ticket — prototypes need real output to react to; designing a report format before any score exists is guessing at the shape of data nobody has seen.
#306 (cadence) depends on what a run actually costs and how long it takes — both facts obtained by running it.
#307 (gold-slice scripting) blocks the gold slice only, not Tier 1 or the injected part of Tier 2.
Standing preferences for this effort:
Eval drives reconciliation/decay design, not the reverse — don't design component 3 (reconciliation/emergence/decay) before component 5 (eval) has a score to design against.
Retrieval starts crude: the existing run_cypher_query MCP tool (integrations/mcp-memgraph), zero new build for v1. A dedicated retrieval layer is deferred until the crude baseline's failures shape what it actually needs to be.
memgraph-toolbox (a dependency of every context-graph package) already ships a deepeval-based evaluations extra with one custom metric built (CoherenceEmbeddingsBasedMetric, evals/coherence.py). deepeval is this repo's established eval framework — the eval-build tickets below extend it rather than introduce a competing one (promptfoo/braintrust/ragas confirmed unused anywhere in the repo).
Decisions so far
Grilling: what should the eval loop score, and how? — blended end-to-end score (mostly recall via retrieval, structural checks lighter-weighted); synthetic-first unified corpus, org-signal-question framed; hybrid task-gen (bulk skip-harness + small real-harness gold slice) now, harden later; graded rubric (coverage+efficiency) via LLM-as-judge; human-gated promotion for this effort, full-auto is longer-term direction only.
Domain-modeling: reframe context-graph's mission in CONTEXT-MAP.md — Context Graph's definition reworded to center raw→emerge (Collection Tier → Memory Tier) into durable organizational knowledge, for agents/harness to consume. Kept "avoid: Knowledge graph" (mechanism, not mission); rejected minting a new "Knowledge Lake" term as redundant with Collection/Memory Tier.
Grilling: what format does the eval corpus have, and where does it live? — deepeval's Golden as-is (no custom schema), serialized JSONL, committed to git. Memgraph-as-store rejected ("keep the ruler outside the thing being measured" — a schema migration could silently corrupt the answer key); Confident AI cloud rejected on the project's own owned-IP thesis. Upstream benchmark content is fetch-and-convert against a pinned release, committing only converted output, not vendored blobs.
Grilling: how does the LLM-as-judge scoring mechanism get built? — rubric splits by mechanism: LLM judges quality, plain code counts cost (refines Grilling: what should the eval loop score, and how? #299's "all via LLM-judge"). Quality = a minimal pair, built-in ContextualRecallMetric + one custom GEval coverage rubric; Faithfulness/AnswerRelevancy skipped at v1 on judge-call cost. Judge is Anthropic pinned to a dated model ID (decorrelates from the OpenAI-backed extraction pipeline; a moving alias would break cross-version comparability). Calibration = two small occasional checks, repeat-and-compare for noise plus ~25 human-graded items for bias.
Grilling: how is the deterministic efficiency metric instrumented, and how do eval runs avoid polluting the graph under test? — efficiency is retrieval payload size (tokens handed back to answer), not agent-side burn: a deterministic count over Golden.retrieval_context, no adapter work. Coverage is a hard gate, efficiency ranks within it (a weighted composite would let coverage be traded away invisibly). Isolation = one Memgraph database per eval batch, reset before fixture injection, using the existing MEMGRAPH_DATABASE config; batch-shared beats per-question because distractors are what make precision/efficiency meaningful. Markers inside a batch are provenance-only — never the mechanism keeping eval-agent traces out. Corrected during implementation (2026-08-25): per-batch databases need Memgraph Enterprise (multi-tenancy), so isolation is instead a dedicated eval instance cleared before each batch — same guarantee, community licence, no Enterprise secret in CI.
Research: which memory benchmark should the eval corpus derive from, and what does each actually exercise? — no benchmark ships both agent-harness activity and a gold-answer key; the field splits. Adopt LongMemEval v1 (MIT) for the text path and LongMemEval-V2 (Apache-2.0) for the action path; decline LoCoMo (CC BY-NC — excludes commercial use); hold the CC BY 4.0 SWE trajectory corpora in reserve as content for authored questions. Residual gap, all synthetic: skills-graph and subagent nesting are uncovered by every candidate, and the organizational/multiplayer question framing appears in none. Findings: docs/research/2026-08-memory-benchmarks.md on branch research/memory-benchmarks.
Grilling: who/what authors the eval corpus, and how? — two tiers scored separately, never blended: Tier 1 adopted from upstream (~951 questions, converted not authored, asks "does recall work mechanically"), Tier 2 hand-authored org-signal questions (asks "does it work for what we're building", and is what promotion hangs on). Both start small, scale only when the signal is too noisy. Tier 2 authorship splits by judgment: human writes questions + gold answers, fixture content comes from real traces rather than LLM prose (tidy synthetic fixtures would flatter retrieval). The real-harness gold slice is a subset of Tier 2, sized by component coverage — post-Research: which memory benchmark should the eval corpus derive from, and what does each actually exercise? #308 it is the only eval coverage skills-graph and subagent nesting get at all.
Grilling: what triggers an eval run, and how often? — two tiers. The fast, LLM-free tests run on every PR with no key provided (they skip themselves), staying free and deterministic. Real eval batches run on workflow_dispatch only, never on push/PR: cost (~2 LLM calls per session, ~900 sessions per 20-question batch), non-determinism (Grilling: how does the LLM-as-judge scoring mechanism get built? #304 handles that with calibration, not a fixed threshold), and decisively — a CI gate on a judged score is the automatic promotion Grilling: what should the eval loop score, and how? #299 ruled out in favour of a human reading a report. Not scheduled either: nothing consumes the output automatically, so a timer would buy reports nobody asked for. Implemented in PR #311.
Prototype: what does the human-gated comparison report look like? — verdict-first with an explicit noise floor, chosen over a neutral table and a regressions-first layout. Two behaviours carry it: the report refuses to compare runs pinned to a different corpus/judge/tokenizer (a delta across pins measures the pin change as if it were the change under test), and it will not call any delta real without calibration — coverage_is_real is three-valued, since "not real" and "cannot tell" are different answers to a decision-maker. A real coverage regression outranks any efficiency gain; efficiency alone never declares an improvement. Surfaced a sizing caveat for Grilling: how does the LLM-as-judge scoring mechanism get built? #304's calibration work: at 20 questions one flip is 5pp, so coverage granularity is coarser than a plausible noise floor.
Building the gold slice's one question — #307 decided its shape (a fact stated only inside a subagent, planting prompt on the Golden, reusing test-graph-model's techniques). Writing the runner that drives that live session and checks recall is implementation, not a further decision.
Reconciliation/emergence/decay design itself — what triggers decay, where the collection-tier/memory-tier boundary sits under the new organizational-knowledge framing, how existing sessions-graph reconciliation and skills-graph procedural mining fold into it. Waits on eval v1 producing a first score to design against.
"Efficiency" as its own axis, distinct from eval's coverage/efficiency answer-rubric — not yet pinned down as a separate concern (e.g. query cost/latency at scale) or folded into the same rubric.
Destination
A specced direction for context-graph's emergence pipeline — reconciliation/decay, retrieval+efficiency, and the eval loop driving both — under a reframed mission (organizational-knowledge emergence from agents, for agent/harness consumption — not session/tool storage). Ready to implement, not yet built.
Notes
Domain:
context-graphfamily,memgraph/ai-toolkit. Seecontext-graph/CONTEXT-MAP.md, reframed by this map's first resolved ticket.Skills every session should consult:
/grilling+/domain-modelingfor HITL tickets; a/researchsubagent for research tickets — but constrain fan-out: #308's first attempt spawned 8 concurrent research subagents, exhausted the API session quota before writing anything, and had to be redone manually with sequential lookups and authenticatedgh api.This map carries execution (amended 2026-08-24, same override map #288 used). The eval-loop decisions are concrete enough to build against, so implementing them is the natural next step rather than a
/to-spechandoff. Specifically, Tier 1 of the eval loop is fully unblocked: corpus format (#302), benchmark choice and fetch-and-convert model (#308), tier structure (#303), metrics and judge (#304), efficiency and isolation (#309), and the retrieval baseline (#300) are all settled. The buildable vertical slice is: fetch/convert LongMemEval v1 → inject into a dedicated eval database → reconcile → retrieval agent queries viarun_cypher_query→ score withContextualRecall+ oneGEval+ the deterministic payload-token count → report.The three remaining open tickets do not block that slice, and two are better answered after it runs:
Standing preferences for this effort:
run_cypher_queryMCP tool (integrations/mcp-memgraph), zero new build for v1. A dedicated retrieval layer is deferred until the crude baseline's failures shape what it actually needs to be.actions-graphhas two independent write paths for subagent nesting (theagent-context-graphconnector path, and a separate directActionTrackerhooks path) — surfaced in Grilling: does a subagent instance become a first-class Agent node, or stay a flat Action pair with a populated PARENT_OF? #278's correction. Out of this map unless it blocks a ticket here.memgraph-toolbox(a dependency of every context-graph package) already ships adeepeval-basedevaluationsextra with one custom metric built (CoherenceEmbeddingsBasedMetric,evals/coherence.py).deepevalis this repo's established eval framework — the eval-build tickets below extend it rather than introduce a competing one (promptfoo/braintrust/ragas confirmed unused anywhere in the repo).Decisions so far
run_cypher_queryMCP tool, zero new build; dedicated retrieval layer deferred, to be shaped by where this baseline fails.Context Graph's definition reworded to center raw→emerge (Collection Tier → Memory Tier) into durable organizational knowledge, for agents/harness to consume. Kept "avoid: Knowledge graph" (mechanism, not mission); rejected minting a new "Knowledge Lake" term as redundant with Collection/Memory Tier.Goldenas-is (no custom schema), serialized JSONL, committed to git. Memgraph-as-store rejected ("keep the ruler outside the thing being measured" — a schema migration could silently corrupt the answer key); Confident AI cloud rejected on the project's own owned-IP thesis. Upstream benchmark content is fetch-and-convert against a pinned release, committing only converted output, not vendored blobs.ContextualRecallMetric+ one customGEvalcoverage rubric;Faithfulness/AnswerRelevancyskipped at v1 on judge-call cost. Judge is Anthropic pinned to a dated model ID (decorrelates from the OpenAI-backed extraction pipeline; a moving alias would break cross-version comparability). Calibration = two small occasional checks, repeat-and-compare for noise plus ~25 human-graded items for bias.Golden.retrieval_context, no adapter work. Coverage is a hard gate, efficiency ranks within it (a weighted composite would let coverage be traded away invisibly). Isolation = one Memgraph database per eval batch, reset before fixture injection, using the existingMEMGRAPH_DATABASEconfig; batch-shared beats per-question because distractors are what make precision/efficiency meaningful. Markers inside a batch are provenance-only — never the mechanism keeping eval-agent traces out. Corrected during implementation (2026-08-25): per-batch databases need Memgraph Enterprise (multi-tenancy), so isolation is instead a dedicated eval instance cleared before each batch — same guarantee, community licence, no Enterprise secret in CI.skills-graphand subagent nesting are uncovered by every candidate, and the organizational/multiplayer question framing appears in none. Findings:docs/research/2026-08-memory-benchmarks.mdon branchresearch/memory-benchmarks.skills-graphand subagent nesting get at all.workflow_dispatchonly, never on push/PR: cost (~2 LLM calls per session, ~900 sessions per 20-question batch), non-determinism (Grilling: how does the LLM-as-judge scoring mechanism get built? #304 handles that with calibration, not a fixed threshold), and decisively — a CI gate on a judged score is the automatic promotion Grilling: what should the eval loop score, and how? #299 ruled out in favour of a human reading a report. Not scheduled either: nothing consumes the output automatically, so a timer would buy reports nobody asked for. Implemented in PR #311.coverage_is_realis three-valued, since "not real" and "cannot tell" are different answers to a decision-maker. A real coverage regression outranks any efficiency gain; efficiency alone never declares an improvement. Surfaced a sizing caveat for Grilling: how does the LLM-as-judge scoring mechanism get built? #304's calibration work: at 20 questions one flip is 5pp, so coverage granularity is coarser than a plausible noise floor.additional_metadata["session_prompt"], not a joined fixture file that could drift. Sized by carrier — where the fact physically lives in the graph — rather than by component, sincetest-graph-modelalready verifies shape and this verifies recall. Starts with exactly one question: a fact stated only inside a subagent, because top-level recall is already covered by Tier 1 and the nested carrier has a demonstrated silent-failure mode (Grilling: does the Episode/Memory model need any change given the nesting redesign? #281's single-hopHAS_ACTION). Reusestest-graph-model's hard-won techniques, including avoiding literalSKILL.mdpaths in the prompt (skills-graph: SkillGraphConnector can record a false-positive Skill Usage from a prompt merely mentioning a SKILL.md path #293).Not yet specified
test-graph-model's techniques). Writing the runner that drives that live session and checks recall is implementation, not a further decision.sessions-graphreconciliation andskills-graphprocedural mining fold into it. Waits on eval v1 producing a first score to design against.run_cypher_querybaseline) — waits on where the baseline actually fails. Known prior art/debt to fold in:lightrag-memgraph's open vector-query bugs (perf(lightrag-memgraph): vector query() full label scan defeats vector index; closed: lightrag-memgraph: vector query() silently drops dead candidates with no signal).Out of scope