You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Design proactivity and interaction-quality task slice #126
Related to #60 and #111. Follows the interactive-task direction reserved by #18 / ADR 0008.
Current behavior / design
As of this issue being written, Eval Lab tasks are non-interactive by default, and the project has reserved a future interactive-task contract for bounded follow-up questions. It does not yet have a task slice or rubric for evaluating whether a coding agent asks, warns, proceeds, or stays quiet at the right moment.
Why change
Recent agent-eval sources frame useful coding agents as proactive rather than merely autonomous. A coding agent can produce a correct patch while still being noisy, asking low-value questions, failing to warn about a real risk, or proceeding through a consequential ambiguity without asking. Eval Lab should be able to evaluate helpful intervention quality separately from final patch correctness.
Design a small post-starter task slice for interaction quality. The slice should define one or more tasks where the expected behavior depends on a well-timed question, warning, or abstention from noisy commentary. It should state how scripted answers, transcript evidence, final patch evidence, and human review fit together.
Candidate rubric dimensions:
question or warning was relevant to the task outcome;
timing was early enough to avoid wasted or unsafe work;
question was specific and answerable;
agent did not block progress unnecessarily;
agent avoided ambient commentary when no intervention was useful;
final handoff accurately reports what was clarified, assumed, checked, and left uncertain.
Acceptance criteria
The design names at least one concrete task shape where proactivity changes the expected evaluation outcome.
The design distinguishes helpful proactivity from noise, indecision, or excessive clarification.
The task evidence captures questions/warnings in order, scripted answers when used, final patch state, grader results, and final handoff text.
Parent
Related to #60 and #111. Follows the interactive-task direction reserved by #18 / ADR 0008.
Current behavior / design
As of this issue being written, Eval Lab tasks are non-interactive by default, and the project has reserved a future interactive-task contract for bounded follow-up questions. It does not yet have a task slice or rubric for evaluating whether a coding agent asks, warns, proceeds, or stays quiet at the right moment.
Why change
Recent agent-eval sources frame useful coding agents as proactive rather than merely autonomous. A coding agent can produce a correct patch while still being noisy, asking low-value questions, failing to warn about a real risk, or proceeding through a consequential ambiguity without asking. Eval Lab should be able to evaluate helpful intervention quality separately from final patch correctness.
Sources:
What to build
Design a small post-starter task slice for interaction quality. The slice should define one or more tasks where the expected behavior depends on a well-timed question, warning, or abstention from noisy commentary. It should state how scripted answers, transcript evidence, final patch evidence, and human review fit together.
Candidate rubric dimensions:
Acceptance criteria
Blocked by
None - design can start immediately.