Part of #297
Question
Two coupled questions, surfaced resolving Grilling: how does the LLM-as-judge scoring mechanism get built?:
1. Where do the efficiency counts come from? #304 decided efficiency (hops / tool calls / tokens) is measured deterministically in code, not LLM-judged — a hand-written non-LLM BaseMetric, the memgraph-toolbox evals/coherence.py pattern. But the counts have to be sourced. Candidates:
- deepeval's own
LLMTestCase.tools_called / TOOLS_CALLED param, populated by whatever runs the retrieval agent;
- the agent SDK's own token/usage accounting;
- context-graph's own captured data — the eval agent's retrieval session is itself a harness session, so
actions-graph already records its tool calls and agent-context-graph already emits the events. Self-instrumenting, and arguably the most honest measure of "what did answering this actually cost."
2. How does an eval run avoid polluting the graph it measures? The third option above exposes the real problem, and it applies regardless of which option is chosen: if the eval agent runs as a real harness session with hooks active, its own activity is captured into the same graph the eval is scoring. Successive eval runs would then accumulate their own sessions as evaluable content, and a corpus question could be answered from a previous eval run's traces rather than from the planted fixture data. Candidates to grill through: hooks disabled for eval runs (loses the free instrumentation from option 3); a separate Memgraph database/instance per eval run; a marker on eval-origin nodes plus filtering at retrieval time; accepting the pollution as immaterial (needs an argument, not an assumption).
Note the tension with #302's reasoning: the corpus went into git specifically to keep the ruler outside the thing being measured. The same principle plausibly applies to the eval run's own traces.
Related and worth reading first: #307 (gold-slice real-harness scripting) hits the same "eval runs inside a real harness session" territory from the fixture-planting side.
Part of #297
Question
Two coupled questions, surfaced resolving Grilling: how does the LLM-as-judge scoring mechanism get built?:
1. Where do the efficiency counts come from? #304 decided efficiency (hops / tool calls / tokens) is measured deterministically in code, not LLM-judged — a hand-written non-LLM
BaseMetric, thememgraph-toolboxevals/coherence.pypattern. But the counts have to be sourced. Candidates:LLMTestCase.tools_called/TOOLS_CALLEDparam, populated by whatever runs the retrieval agent;actions-graphalready records its tool calls andagent-context-graphalready emits the events. Self-instrumenting, and arguably the most honest measure of "what did answering this actually cost."2. How does an eval run avoid polluting the graph it measures? The third option above exposes the real problem, and it applies regardless of which option is chosen: if the eval agent runs as a real harness session with hooks active, its own activity is captured into the same graph the eval is scoring. Successive eval runs would then accumulate their own sessions as evaluable content, and a corpus question could be answered from a previous eval run's traces rather than from the planted fixture data. Candidates to grill through: hooks disabled for eval runs (loses the free instrumentation from option 3); a separate Memgraph database/instance per eval run; a marker on eval-origin nodes plus filtering at retrieval time; accepting the pollution as immaterial (needs an argument, not an assumption).
Note the tension with #302's reasoning: the corpus went into git specifically to keep the ruler outside the thing being measured. The same principle plausibly applies to the eval run's own traces.
Related and worth reading first: #307 (gold-slice real-harness scripting) hits the same "eval runs inside a real harness session" territory from the fixture-planting side.