Skip to content

Add a multi-turn scenario/persona eval framework with LLM-judge verdict caching #16814

Description

@mrveiss

Problem

AutoBot's eval framework can only replay and score a single-turn golden-trajectory
fixture. There is no way to express a multi-turn, goal-directed conversation test
(scripted turns with expected events, or an LLM-played persona with a goal), no
yes/no/continue judge against an arbitrary natural-language criterion, and no caching
of judge verdicts across reruns.

Location

  • autobot-backend/eval/store.py:5-19 — a "golden trajectory" is a captured real run
    stored as a fixed JSON fixture (eval/golden/*.json); single-turn only.
  • autobot-backend/eval/runner.py:70-110 TrajectoryReplayer.replay_one — replays one
    golden's single-turn input, checks tools deterministically, scores the reply.
  • autobot-backend/rlm/evaluator.py:53-80 ResponseQualityEvaluator.evaluate() — LLM-as
    -self-judge, but returns a 0.0-1.0 numeric score parsed from a fixed
    SCORE/CRITIQUE/HINT prompt (_EVAL_PROMPT, line 25), not a verdict against an
    arbitrary criterion.
  • No verdict caching: grepped cache/hash in eval/*.py and rlm/evaluator.py — no
    matches, every replay re-pays for a fresh judge call.
  • No scenario/persona concept: grepped persona/scenario across the repo's YAML —
    no test-scenario hits.

Impact

Behavioral regressions in multi-turn, goal-directed agent conversations (the common
case for AutoBot's chat and voice agents) cannot be captured as regression tests today
— only single-turn fixture replay is possible. Every eval run also re-pays LLM cost
for judge calls that a previous run already answered identically.

Acceptance Criteria

  • A declarative scenario format (or equivalent) supports both a scripted
    multi-turn conversation and an LLM-played persona with a goal, layered over
    the existing golden-trajectory fixture harness rather than replacing it.
  • A judge can answer yes/no/continue (or equivalent) against an arbitrary
    natural-language criterion, distinct from the existing numeric
    ResponseQualityEvaluator self-score, which stays available for what it
    already covers.
  • Judge verdicts are cached by (criterion, conversation) with an explicit
    invalidation story (criterion or model change busts the cache) rather than
    caching silently going stale.
  • Verified against current origin/main code, with file:line evidence in the
    closing comment.

Discovered During

Research comparison in docs/research/streaming-voice-agent-pipeline-architecture.md.

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions