Skip to content

Grilling: how is the deterministic efficiency metric instrumented, and how do eval runs avoid polluting the graph under test? #309

Description

@antejavor

Part of #297

Question

Two coupled questions, surfaced resolving Grilling: how does the LLM-as-judge scoring mechanism get built?:

1. Where do the efficiency counts come from? #304 decided efficiency (hops / tool calls / tokens) is measured deterministically in code, not LLM-judged — a hand-written non-LLM BaseMetric, the memgraph-toolbox evals/coherence.py pattern. But the counts have to be sourced. Candidates:

  • deepeval's own LLMTestCase.tools_called / TOOLS_CALLED param, populated by whatever runs the retrieval agent;
  • the agent SDK's own token/usage accounting;
  • context-graph's own captured data — the eval agent's retrieval session is itself a harness session, so actions-graph already records its tool calls and agent-context-graph already emits the events. Self-instrumenting, and arguably the most honest measure of "what did answering this actually cost."

2. How does an eval run avoid polluting the graph it measures? The third option above exposes the real problem, and it applies regardless of which option is chosen: if the eval agent runs as a real harness session with hooks active, its own activity is captured into the same graph the eval is scoring. Successive eval runs would then accumulate their own sessions as evaluable content, and a corpus question could be answered from a previous eval run's traces rather than from the planted fixture data. Candidates to grill through: hooks disabled for eval runs (loses the free instrumentation from option 3); a separate Memgraph database/instance per eval run; a marker on eval-origin nodes plus filtering at retrieval time; accepting the pollution as immaterial (needs an argument, not an assumption).

Note the tension with #302's reasoning: the corpus went into git specifically to keep the ruler outside the thing being measured. The same principle plausibly applies to the eval run's own traces.

Related and worth reading first: #307 (gold-slice real-harness scripting) hits the same "eval runs inside a real harness session" territory from the fixture-planting side.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions