Skip to content

Grilling: what format does the eval corpus have, and where does it live? #302

Description

@antejavor

Part of #297

Question

What format does the eval corpus (component 5) actually take, and where does it live?

Prior art found: memgraph-toolbox — a dependency every context-graph package already has — ships a deepeval-based evaluations extra (pyproject.toml: deepeval>=3.5.2), with one custom metric already built (CoherenceEmbeddingsBasedMetric in evals/coherence.py, extending deepeval.metrics.BaseMetric). deepeval is this repo's established eval framework; no other eval framework (promptfoo/braintrust/ragas) appears anywhere in the repo. deepeval's LLMTestCase (input / actual_output / expected_output / context) is a natural fit for the corpus shape already decided in #299 (question + ground-truth answer + retrieved context).

Resolve: does the corpus get stored as a deepeval-shaped dataset (and if so, where — a file in the family, a new memgraph-toolbox submodule, something else), or is there a reason to deviate from LLMTestCase's shape? This decision blocks the authorship-mechanics and judge-harness tickets, which need the shape settled first.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions