Repository navigation
test: add representative coding-agent evidence fixture - #258
harshitethic wants to merge 8 commits into
Conversation
| {"event_type":"file_read","timestamp":1789300802.0,"event_id":"ev-read","session_id":"fixture-review-001","data":{"path":"src/example.py","source":"repository"}} | ||
| {"event_type":"tool_call","timestamp":1789300803.0,"event_id":"tool-test-1","session_id":"fixture-review-001","data":{"tool":"shell","command":"python -m pytest -q tests/test_example.py"}} | ||
| {"event_type":"tool_result","timestamp":1789300804.0,"event_id":"result-test-1","session_id":"fixture-review-001","parent_id":"tool-test-1","duration_ms":620.0,"data":{"exit_code":1,"output":"1 failed"}} | ||
| {"event_type":"error","timestamp":1789300804.1,"event_id":"ev-error","session_id":"fixture-review-001","parent_id":"tool-test-1","data":{"message":"example test failed","recoverable":true}} |
There was a problem hiding this comment.
result-test-1 and this ERROR are both terminal children of tool-test-1. The evidence-health core in #256 classifies the fixture as partial with duplicate_tool_outcome, in addition to the intended orphan gap. Current hook capture also emits one terminal ERROR for a failed tool call, not both events. Model the failure with one terminal event, or declare and assert this second defect explicitly.
|
|
||
| def test_fixture_uses_one_session_and_expected_count(self): | ||
| self.assertEqual(len(self.events), self.expected["expected_event_count"]) | ||
| self.assertEqual({event.session_id for event in self.events}, {self.expected["session_id"]}) |
There was a problem hiding this comment.
expected.json declares both provider and fixture_version, but the tests never assert either value against the session-start event. Removing or changing the source attribution/version therefore leaves CI green. Assert both fields here so the versioned fixture contract cannot drift silently.
|
Addressed both fixture review points: removed the duplicate |
Summary
Adds a versioned, synthetic representative session fixture as a reusable regression input for #250.
Scenario covered
The fixture contains:
Artifacts
tests/fixtures/representative_session/events.ndjson— canonicalTraceEventinputtests/fixtures/representative_session/annotations.jsonl— human review overlaytests/fixtures/representative_session/expected.json— deterministic IDs/count/gap expectationstests/test_representative_fixture.py— schema + scenario regression coveragedocs/representative-fixture.md— trust boundary and refresh procedureWhy the gap is intentional
The fixture includes an orphaned tool result on purpose. Review UI/evidence-health work must surface missing relationships instead of inventing a parent and presenting a falsely complete timeline.
Scope
This is a focused foundation for #250 rather than a claim to close every consumer listed there. Follow-up PRs can wire the same fixture into evidence health, share/export, OTLP mapping, and the local review UI without each subsystem inventing different sample data.
Validation
The test parses the fixture through the repository's own
TraceEvent.from_json()andAnnotation.from_json()models and asserts the required failure/retry/recovery, redaction, annotation, relationship, and intentional-gap invariants.I could not execute the suite locally because the connected development machine is offline; upstream CI is the verification gate for this branch.
Refs #250.