Problem
AutoBot's eval framework can only replay and score a single-turn golden-trajectory
fixture. There is no way to express a multi-turn, goal-directed conversation test
(scripted turns with expected events, or an LLM-played persona with a goal), no
yes/no/continue judge against an arbitrary natural-language criterion, and no caching
of judge verdicts across reruns.
Location
autobot-backend/eval/store.py:5-19 — a "golden trajectory" is a captured real run
stored as a fixed JSON fixture (eval/golden/*.json); single-turn only.
autobot-backend/eval/runner.py:70-110 TrajectoryReplayer.replay_one — replays one
golden's single-turn input, checks tools deterministically, scores the reply.
autobot-backend/rlm/evaluator.py:53-80 ResponseQualityEvaluator.evaluate() — LLM-as
-self-judge, but returns a 0.0-1.0 numeric score parsed from a fixed
SCORE/CRITIQUE/HINT prompt (_EVAL_PROMPT, line 25), not a verdict against an
arbitrary criterion.
- No verdict caching: grepped
cache/hash in eval/*.py and rlm/evaluator.py — no
matches, every replay re-pays for a fresh judge call.
- No scenario/persona concept: grepped
persona/scenario across the repo's YAML —
no test-scenario hits.
Impact
Behavioral regressions in multi-turn, goal-directed agent conversations (the common
case for AutoBot's chat and voice agents) cannot be captured as regression tests today
— only single-turn fixture replay is possible. Every eval run also re-pays LLM cost
for judge calls that a previous run already answered identically.
Acceptance Criteria
Discovered During
Research comparison in docs/research/streaming-voice-agent-pipeline-architecture.md.
Problem
AutoBot's eval framework can only replay and score a single-turn golden-trajectory
fixture. There is no way to express a multi-turn, goal-directed conversation test
(scripted turns with expected events, or an LLM-played persona with a goal), no
yes/no/continue judge against an arbitrary natural-language criterion, and no caching
of judge verdicts across reruns.
Location
autobot-backend/eval/store.py:5-19— a "golden trajectory" is a captured real runstored as a fixed JSON fixture (
eval/golden/*.json); single-turn only.autobot-backend/eval/runner.py:70-110TrajectoryReplayer.replay_one— replays onegolden's single-turn input, checks tools deterministically, scores the reply.
autobot-backend/rlm/evaluator.py:53-80ResponseQualityEvaluator.evaluate()— LLM-as-self-judge, but returns a 0.0-1.0 numeric score parsed from a fixed
SCORE/CRITIQUE/HINTprompt (_EVAL_PROMPT, line 25), not a verdict against anarbitrary criterion.
cache/hashineval/*.pyandrlm/evaluator.py— nomatches, every replay re-pays for a fresh judge call.
persona/scenarioacross the repo's YAML —no test-scenario hits.
Impact
Behavioral regressions in multi-turn, goal-directed agent conversations (the common
case for AutoBot's chat and voice agents) cannot be captured as regression tests today
— only single-turn fixture replay is possible. Every eval run also re-pays LLM cost
for judge calls that a previous run already answered identically.
Acceptance Criteria
multi-turn conversation and an LLM-played persona with a goal, layered over
the existing golden-trajectory fixture harness rather than replacing it.
natural-language criterion, distinct from the existing numeric
ResponseQualityEvaluatorself-score, which stays available for what italready covers.
invalidation story (criterion or model change busts the cache) rather than
caching silently going stale.
origin/maincode, with file:line evidence in theclosing comment.
Discovered During
Research comparison in
docs/research/streaming-voice-agent-pipeline-architecture.md.