Skip to content

feat(review): add requirement-to-output verification during agent review #216

Description

@CallMeR

Problem / Motivation

A major challenge in the current AI coding lifecycle is that agents can often produce useful output, but the execution process is still hard to audit, coordinate, and trace. ORGII is already addressing this problem with replayable sessions, agent trajectories, reviewability, and traceability.

However, there is another related problem that is often under-addressed: proving that the final output actually satisfies the user's original task requirements.

When users provide long or multi-part requirements, the request can usually be decomposed into several smaller requirement items. For example:

  1. Implement feature A.
  2. Adjust UI behavior B.
  3. Use specific copy, color, layout, or interaction behavior C.

Due to known LLM limitations such as "lost in the middle", an agent may complete the task in a way that appears broadly useful, but still fails to strictly satisfy part of the original request. This can happen in subtle ways:

  • A feature is partially implemented but misses an edge case.
  • A button exists, but the label, color, or placement does not match the request.
  • A UI element is present, but the interaction behavior is incomplete.
  • A requested requirement is implemented in code but not wired into the visible product.
  • A bug remains because the agent optimized for the general goal rather than the full requirement set.
  • The final answer says the task is complete, but there is no structured evidence that each user requirement was checked against the actual output.

In other words, the agent's work may be traceable, but the final deliverable is not necessarily requirement-verifiable.

For an agentic IDE focused on reviewability and traceability, it would be valuable to make the review process compare the user's original requirements against the produced artifact, not only review the code diff or agent trajectory.

Proposed Solution

Add a requirement-to-output verification step during the review workflow.

The core idea is to extract or derive a structured checklist from the user's original request, then compare that checklist against the final output, code changes, UI state, tests, screenshots, browser state, or other available artifacts.

A possible workflow:

  1. Requirement extraction

    • When a user submits a task, ORGII derives a structured list of requirement items from the original prompt.
    • Each item should preserve traceability back to the original user request.
    • The extracted checklist should include functional requirements, UI details, copy requirements, behavior requirements, constraints, and explicit non-goals where possible.
  2. Implementation phase

    • The coding agent works normally.
    • The requirement checklist remains attached to the session as review context.
  3. Review / verification phase

    • Before marking the task as complete, ORGII runs a verification pass that compares each requirement item against the produced output.

    • The verifier can inspect relevant evidence such as git diff, changed files, tests, screenshots, browser DOM, terminal output, app state, or replayed agent trajectory.

    • Each requirement receives a status such as:

      • satisfied
      • partially satisfied
      • not satisfied
      • uncertain / needs human review
  4. Evidence mapping

    • For each requirement, the review result should include supporting evidence where possible.

    • Examples:

      • File path and line range.
      • Screenshot reference.
      • DOM element or UI state.
      • Test result.
      • Terminal command output.
      • Agent step or session event reference.
  5. Repair loop

    • If any requirement is not satisfied, partially satisfied, or uncertain, ORGII should allow the agent to re-enter a repair loop.
    • The repair loop should focus specifically on the missing or incorrectly implemented requirement items.
    • After repair, the verification pass should run again.
  6. Review UI

    • Add a review panel that shows:

      • Original user request.
      • Extracted requirement checklist.
      • Verification status per requirement.
      • Evidence for each requirement.
      • Remaining gaps.
      • Optional "ask agent to fix unmet requirements" action.

This would turn the review process from "the agent produced something and we can inspect the trace" into "the agent produced something and we can verify it against the user's original intent."

Alternatives Considered

  1. Rely only on normal code review

    Traditional code review can catch some issues, but it often focuses on code quality, architecture, and obvious correctness. It does not guarantee that every part of the user's original request was preserved and checked.

  2. Rely only on tests

    Tests are useful, but they may not cover UI details, copy, layout, interaction behavior, or implicit user constraints. They also require the agent to know what should be tested, which is exactly where requirement loss can occur.

  3. Ask the agent to summarize what it did

    A final summary is helpful, but it is self-reported and not enough as evidence. The agent may claim a requirement is complete without independently checking the output.

  4. Let the user manually compare the result against the prompt

    This works for small tasks, but it does not scale well for long, multi-part requests. It also weakens the value of an agentic IDE that already has access to session history, code changes, browser state, terminal output, and replayable traces.

Acceptance Criteria

  • ORGII can derive a structured requirement checklist from the user's original task prompt.
  • Each checklist item preserves a reference to the original user requirement or prompt segment.
  • The review workflow can compare the final output against the checklist.
  • Each requirement item can be marked as satisfied, partially satisfied, not satisfied, or uncertain.
  • The review result can include evidence for each requirement, such as files, diffs, tests, screenshots, DOM state, terminal output, or session events.
  • The user can see unmet or uncertain requirements in a review UI.
  • The user can trigger a focused repair loop for unmet or partially satisfied requirements.
  • After repair, ORGII can re-run the requirement verification pass.
  • The feature should work even when the original user request is long and contains multiple independent requirement items.
  • The feature should not replace human review; it should provide structured evidence and highlight gaps for the user.

Additional Context

ORGII already emphasizes agent reviewability, traceability, session replay, and controllability. This feature would extend that direction from "what did the agent do?" to "did the agent actually satisfy what the user asked for?"

This is especially important for UI-heavy and product-oriented tasks, where small deviations can matter:

  • Wrong button label.
  • Incorrect color or spacing.
  • Missing loading state.
  • Incorrect interaction behavior.
  • Feature implemented but not exposed in the UI.
  • Layout close to correct but not matching the user's explicit requirement.
  • A multi-part request where one or two items are silently skipped.

The proposed feature could be implemented as a lightweight first version by generating a checklist and running an LLM-based review over the diff and available session artifacts. A more advanced version could integrate screenshots, browser state, DOM inspection, tests, and replayed session events as stronger evidence.

The goal is not to make the agent infallible. The goal is to make requirement drift visible, reviewable, and fixable.

Activity

  1. assigned and unassigned on Jul 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions