You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat(review): add requirement-to-output verification during agent review #216
A major challenge in the current AI coding lifecycle is that agents can often produce useful output, but the execution process is still hard to audit, coordinate, and trace. ORGII is already addressing this problem with replayable sessions, agent trajectories, reviewability, and traceability.
However, there is another related problem that is often under-addressed: proving that the final output actually satisfies the user's original task requirements.
When users provide long or multi-part requirements, the request can usually be decomposed into several smaller requirement items. For example:
Implement feature A.
Adjust UI behavior B.
Use specific copy, color, layout, or interaction behavior C.
Due to known LLM limitations such as "lost in the middle", an agent may complete the task in a way that appears broadly useful, but still fails to strictly satisfy part of the original request. This can happen in subtle ways:
A feature is partially implemented but misses an edge case.
A button exists, but the label, color, or placement does not match the request.
A UI element is present, but the interaction behavior is incomplete.
A requested requirement is implemented in code but not wired into the visible product.
A bug remains because the agent optimized for the general goal rather than the full requirement set.
The final answer says the task is complete, but there is no structured evidence that each user requirement was checked against the actual output.
In other words, the agent's work may be traceable, but the final deliverable is not necessarily requirement-verifiable.
For an agentic IDE focused on reviewability and traceability, it would be valuable to make the review process compare the user's original requirements against the produced artifact, not only review the code diff or agent trajectory.
Proposed Solution
Add a requirement-to-output verification step during the review workflow.
The core idea is to extract or derive a structured checklist from the user's original request, then compare that checklist against the final output, code changes, UI state, tests, screenshots, browser state, or other available artifacts.
A possible workflow:
Requirement extraction
When a user submits a task, ORGII derives a structured list of requirement items from the original prompt.
Each item should preserve traceability back to the original user request.
The extracted checklist should include functional requirements, UI details, copy requirements, behavior requirements, constraints, and explicit non-goals where possible.
Implementation phase
The coding agent works normally.
The requirement checklist remains attached to the session as review context.
Review / verification phase
Before marking the task as complete, ORGII runs a verification pass that compares each requirement item against the produced output.
The verifier can inspect relevant evidence such as git diff, changed files, tests, screenshots, browser DOM, terminal output, app state, or replayed agent trajectory.
Each requirement receives a status such as:
satisfied
partially satisfied
not satisfied
uncertain / needs human review
Evidence mapping
For each requirement, the review result should include supporting evidence where possible.
Examples:
File path and line range.
Screenshot reference.
DOM element or UI state.
Test result.
Terminal command output.
Agent step or session event reference.
Repair loop
If any requirement is not satisfied, partially satisfied, or uncertain, ORGII should allow the agent to re-enter a repair loop.
The repair loop should focus specifically on the missing or incorrectly implemented requirement items.
After repair, the verification pass should run again.
Review UI
Add a review panel that shows:
Original user request.
Extracted requirement checklist.
Verification status per requirement.
Evidence for each requirement.
Remaining gaps.
Optional "ask agent to fix unmet requirements" action.
This would turn the review process from "the agent produced something and we can inspect the trace" into "the agent produced something and we can verify it against the user's original intent."
Alternatives Considered
Rely only on normal code review
Traditional code review can catch some issues, but it often focuses on code quality, architecture, and obvious correctness. It does not guarantee that every part of the user's original request was preserved and checked.
Rely only on tests
Tests are useful, but they may not cover UI details, copy, layout, interaction behavior, or implicit user constraints. They also require the agent to know what should be tested, which is exactly where requirement loss can occur.
Ask the agent to summarize what it did
A final summary is helpful, but it is self-reported and not enough as evidence. The agent may claim a requirement is complete without independently checking the output.
Let the user manually compare the result against the prompt
This works for small tasks, but it does not scale well for long, multi-part requests. It also weakens the value of an agentic IDE that already has access to session history, code changes, browser state, terminal output, and replayable traces.
Acceptance Criteria
ORGII can derive a structured requirement checklist from the user's original task prompt.
Each checklist item preserves a reference to the original user requirement or prompt segment.
The review workflow can compare the final output against the checklist.
Each requirement item can be marked as satisfied, partially satisfied, not satisfied, or uncertain.
The review result can include evidence for each requirement, such as files, diffs, tests, screenshots, DOM state, terminal output, or session events.
The user can see unmet or uncertain requirements in a review UI.
The user can trigger a focused repair loop for unmet or partially satisfied requirements.
After repair, ORGII can re-run the requirement verification pass.
The feature should work even when the original user request is long and contains multiple independent requirement items.
The feature should not replace human review; it should provide structured evidence and highlight gaps for the user.
Additional Context
ORGII already emphasizes agent reviewability, traceability, session replay, and controllability. This feature would extend that direction from "what did the agent do?" to "did the agent actually satisfy what the user asked for?"
This is especially important for UI-heavy and product-oriented tasks, where small deviations can matter:
Wrong button label.
Incorrect color or spacing.
Missing loading state.
Incorrect interaction behavior.
Feature implemented but not exposed in the UI.
Layout close to correct but not matching the user's explicit requirement.
A multi-part request where one or two items are silently skipped.
The proposed feature could be implemented as a lightweight first version by generating a checklist and running an LLM-based review over the diff and available session artifacts. A more advanced version could integrate screenshots, browser state, DOM inspection, tests, and replayed session events as stronger evidence.
The goal is not to make the agent infallible. The goal is to make requirement drift visible, reviewable, and fixable.
Problem / Motivation
A major challenge in the current AI coding lifecycle is that agents can often produce useful output, but the execution process is still hard to audit, coordinate, and trace. ORGII is already addressing this problem with replayable sessions, agent trajectories, reviewability, and traceability.
However, there is another related problem that is often under-addressed: proving that the final output actually satisfies the user's original task requirements.
When users provide long or multi-part requirements, the request can usually be decomposed into several smaller requirement items. For example:
Due to known LLM limitations such as "lost in the middle", an agent may complete the task in a way that appears broadly useful, but still fails to strictly satisfy part of the original request. This can happen in subtle ways:
In other words, the agent's work may be traceable, but the final deliverable is not necessarily requirement-verifiable.
For an agentic IDE focused on reviewability and traceability, it would be valuable to make the review process compare the user's original requirements against the produced artifact, not only review the code diff or agent trajectory.
Proposed Solution
Add a requirement-to-output verification step during the review workflow.
The core idea is to extract or derive a structured checklist from the user's original request, then compare that checklist against the final output, code changes, UI state, tests, screenshots, browser state, or other available artifacts.
A possible workflow:
Requirement extraction
Implementation phase
Review / verification phase
Before marking the task as complete, ORGII runs a verification pass that compares each requirement item against the produced output.
The verifier can inspect relevant evidence such as git diff, changed files, tests, screenshots, browser DOM, terminal output, app state, or replayed agent trajectory.
Each requirement receives a status such as:
satisfiedpartially satisfiednot satisfieduncertain / needs human reviewEvidence mapping
For each requirement, the review result should include supporting evidence where possible.
Examples:
Repair loop
not satisfied,partially satisfied, oruncertain, ORGII should allow the agent to re-enter a repair loop.Review UI
Add a review panel that shows:
This would turn the review process from "the agent produced something and we can inspect the trace" into "the agent produced something and we can verify it against the user's original intent."
Alternatives Considered
Rely only on normal code review
Traditional code review can catch some issues, but it often focuses on code quality, architecture, and obvious correctness. It does not guarantee that every part of the user's original request was preserved and checked.
Rely only on tests
Tests are useful, but they may not cover UI details, copy, layout, interaction behavior, or implicit user constraints. They also require the agent to know what should be tested, which is exactly where requirement loss can occur.
Ask the agent to summarize what it did
A final summary is helpful, but it is self-reported and not enough as evidence. The agent may claim a requirement is complete without independently checking the output.
Let the user manually compare the result against the prompt
This works for small tasks, but it does not scale well for long, multi-part requests. It also weakens the value of an agentic IDE that already has access to session history, code changes, browser state, terminal output, and replayable traces.
Acceptance Criteria
satisfied,partially satisfied,not satisfied, oruncertain.Additional Context
ORGII already emphasizes agent reviewability, traceability, session replay, and controllability. This feature would extend that direction from "what did the agent do?" to "did the agent actually satisfy what the user asked for?"
This is especially important for UI-heavy and product-oriented tasks, where small deviations can matter:
The proposed feature could be implemented as a lightweight first version by generating a checklist and running an LLM-based review over the diff and available session artifacts. A more advanced version could integrate screenshots, browser state, DOM inspection, tests, and replayed session events as stronger evidence.
The goal is not to make the agent infallible. The goal is to make requirement drift visible, reviewable, and fixable.