Part of #297
Question
How does the LLM-as-judge scoring mechanism (coverage + efficiency rubric, per #299) actually get built? deepeval (this repo's established eval framework — see the corpus-format ticket) offers two paths: its built-in GEval metric (custom natural-language criteria, no code) or a hand-written BaseMetric subclass, the same pattern evals/coherence.py already uses. Resolve which path, what the criteria/prompt actually says for "coverage" vs. "efficiency," which judge model, and whether any consistency/calibration check (e.g. repeat-and-compare) is needed before trusting a single judge run.
Depends on the corpus format ticket (need the LLMTestCase-or-otherwise shape settled first).
Part of #297
Question
How does the LLM-as-judge scoring mechanism (coverage + efficiency rubric, per #299) actually get built?
deepeval(this repo's established eval framework — see the corpus-format ticket) offers two paths: its built-inGEvalmetric (custom natural-language criteria, no code) or a hand-writtenBaseMetricsubclass, the same patternevals/coherence.pyalready uses. Resolve which path, what the criteria/prompt actually says for "coverage" vs. "efficiency," which judge model, and whether any consistency/calibration check (e.g. repeat-and-compare) is needed before trusting a single judge run.Depends on the corpus format ticket (need the
LLMTestCase-or-otherwise shape settled first).