Skip to content

Design proactivity and interaction-quality task slice #126

Description

@Jordak

Parent

Related to #60 and #111. Follows the interactive-task direction reserved by #18 / ADR 0008.

Current behavior / design

As of this issue being written, Eval Lab tasks are non-interactive by default, and the project has reserved a future interactive-task contract for bounded follow-up questions. It does not yet have a task slice or rubric for evaluating whether a coding agent asks, warns, proceeds, or stays quiet at the right moment.

Why change

Recent agent-eval sources frame useful coding agents as proactive rather than merely autonomous. A coding agent can produce a correct patch while still being noisy, asking low-value questions, failing to warn about a real risk, or proceeding through a consequential ambiguity without asking. Eval Lab should be able to evaluate helpful intervention quality separately from final patch correctness.

Sources:

What to build

Design a small post-starter task slice for interaction quality. The slice should define one or more tasks where the expected behavior depends on a well-timed question, warning, or abstention from noisy commentary. It should state how scripted answers, transcript evidence, final patch evidence, and human review fit together.

Candidate rubric dimensions:

  • question or warning was relevant to the task outcome;
  • timing was early enough to avoid wasted or unsafe work;
  • question was specific and answerable;
  • agent did not block progress unnecessarily;
  • agent avoided ambient commentary when no intervention was useful;
  • final handoff accurately reports what was clarified, assumed, checked, and left uncertain.

Acceptance criteria

  • The design names at least one concrete task shape where proactivity changes the expected evaluation outcome.
  • The design distinguishes helpful proactivity from noise, indecision, or excessive clarification.
  • The task evidence captures questions/warnings in order, scripted answers when used, final patch state, grader results, and final handoff text.
  • The proposal states which checks are deterministic, which are human-review rubric items, and which might later feed Prototype automated process judge for trial artifacts #77's process judge.
  • The issue explicitly states whether the slice belongs under Draft PRD for harder post-starter evaluation tasks #60, Define task-prep and process-readiness rubric #111, or a standalone interactive-task implementation issue.
  • Non-interactive suite summaries must remain separate from any interactive/proactivity trials.

Blocked by

None - design can start immediately.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestneeds-triageMaintainer needs to evaluate this issue

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions