Skip to content

feat: add deterministic evidence-health assessment core - #256

Open
harshitethic wants to merge 12 commits into
Siddhant-K-code:mainfrom
harshitethic:feat/evidence-health-core
Open

harshitethic wants to merge 12 commits into
Siddhant-K-code:mainfrom
harshitethic:feat/evidence-health-core

Conversation

@harshitethic

@harshitethic harshitethic commented Sep 5, 2026 •

Copy link
Copy Markdown

Summary

Adds a dependency-free, versioned evidence-health core for captured agent sessions, plus focused tests and user-facing semantics documentation.

This is a scoped foundation for #240 rather than claiming the entire issue is complete. The PR intentionally implements the deterministic calculation layer first; wiring the result into every CLI/API/dashboard surface can build on this without duplicating health rules.

What the core detects

  • missing / duplicate session start and end markers
  • events recorded outside the session boundary
  • mixed session IDs
  • missing / duplicate event IDs
  • non-finite timestamps and timestamp regressions
  • unpaired tool calls/results
  • unpaired LLM requests/responses
  • duplicate and out-of-order terminal outcomes
  • failed tool/LLM calls represented by parent-linked ERROR terminal events
  • observable export failures
  • explicitly declared provider blind spots
  • active sessions vs finalized sessions with a missing end marker

The result is versioned and machine-readable with healthy, partial, unknown, and invalid states plus stable reason codes.

Trust boundary

A key constraint from #240 is that a clean timeline must not be presented as proof of complete capture. This implementation therefore never guesses provider limitations from absent events. Provider blind spots must be supplied explicitly by the capture adapter / future capture matrix work (#239), and they are kept distinct from observed structural defects.

The empty-evidence path also preserves any observable export failures and declared provider blind spots instead of returning early and discarding those signals; the overall state remains unknown because there is still no event stream to assess.

That keeps the core local and dependency-free, matching the repository's architecture constraints.

Tests

tests/test_evidence_health.py covers the healthy path, empty/active/finalized sessions, malformed boundaries and IDs, timestamp issues, tool and LLM relationship failures, duplicate/out-of-order outcomes, failed calls terminating in ERROR, provider blind spots, export failures, and versioned serialization.

Documentation

docs/evidence-health.md documents statuses, reason semantics, the distinction between observed defects and provider limitations, and what a healthy result does not claim.

Scope / follow-ups

This PR does not claim to close #240 yet. Remaining integration work includes exposing the same result through the CLI/replay/API/dashboard and sourcing provider limitations from the capture matrix rather than ad-hoc callers. Keeping those integrations separate avoids inventing provider metadata before #239 defines it and avoids coupling the calculation engine to the dashboard work in #242.

Validation

Fresh-clone focused validation on PR head 91dbcf5dcac074cb8890b76068d6888fea04cbad:

PYTHONPATH=src python -m unittest discover -s tests -p 'test_evidence_health.py' -v

Result: 17 tests passed.

A full unittest discovery was also started on the available Windows host, but unrelated platform-sensitive tests are not clean there, so I am not claiming full-suite success from that run. Upstream GitHub Actions is still action_required with zero jobs created, so upstream CI itself has not executed yet.

No runtime dependencies or storage-format changes are introduced.

Comment thread src/agent_trace/evidence_health.py Outdated
Comment thread src/agent_trace/evidence_health.py Outdated
Comment thread src/agent_trace/evidence_health.py

Copy link
Copy Markdown
Author

Validation update: I reproduced the PR head (91dbcf5dcac074cb8890b76068d6888fea04cbad) from a fresh clone and ran the focused evidence-health suite with the repository's src/ layout on PYTHONPATH:

PYTHONPATH=src python -m unittest discover -s tests -p 'test_evidence_health.py' -v

Result: 17 tests passed.

I also started the full unittest discovery. It exercises many platform-specific tests that are not clean on this Windows host, so I am not treating that run as authoritative full-suite validation. Upstream Actions still shows action_required with zero jobs, so CI itself has not run.

Copy link
Copy Markdown
Author

@Siddhant-K-code this one is ready for maintainer review. It keeps the scope narrow to the deterministic evidence-health core for #240, and I ran the focused suite on the PR head (17 tests passed). I deliberately left CLI/dashboard wiring and provider-matrix integration out so the core can be reviewed independently. Happy to adjust the API/reason codes if you want a different shape before merge.

and event.event_type == EventType.ERROR
and event.parent_id.strip() in requests
)
if not is_result and not is_terminal_error:

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ERROR events with a missing or mistyped parent_id are skipped here because both predicates are false. A stream containing start, ERROR(parent_id="missing"), and end is therefore reported healthy, which hides a broken relationship. Surface parent-linked orphan errors as a non-healthy reason and retain the error event as provenance.

reasons: tuple[EvidenceHealthReason, ...] = field(default_factory=tuple)
schema_version: int = EVIDENCE_HEALTH_SCHEMA_VERSION
provider: str = ""
capture_method: str = ""

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The result records provider and capture_method but not the applicable capture-matrix version. That loses the provenance needed to interpret provider_blind_spot reasons after the matrix changes, and #240 explicitly requires this version in the health result. Add a matrix version or stable matrix reference to the schema and serialization.

@Siddhant-K-code

Copy link
Copy Markdown
Owner

@harshitethic thanks for the work on these six PRs. They're close. What's blocking merge is the review I left on September 20, so here's everything still open in one place.

#256 evidence-health core

#257 capture matrix

  • Claude rollback: setup --cli claude only prints the hook JSON. It doesn't write ~/.claude/settings.json. (thread)
  • Codex rollback: setup --cli codex overwrites hooks.json, so removing our commands can't restore earlier hooks. Documenting that accurately is enough for this PR; the merge-or-backup fix in setup can be a follow-up. (thread)
  • Codex stop is captured (handle_stop, covered by test_codex_hooks.py). Keep session_end unavailable. (thread)
  • Validate verification_status against an allowed set, then require the metadata only for verified. (thread)
  • Have the Markdown/JSON consistency test check coverage and status values, or generate the table from the JSON. (thread)

#258 representative fixture

  • tool-test-1 has two terminal children (result-test-1 and the ERROR), so feat: add deterministic evidence-health assessment core #256 classifies the fixture as partial with duplicate_tool_outcome on top of the intended orphan gap. Use one terminal event, or declare and assert the second defect. (thread)
  • Assert provider and fixture_version from expected.json against the session-start event. (thread)

#259 disclosure preview

  • Hook TOOL_CALL events keep commands and paths under data["arguments"], so normal hook traces report zero of each. Read the nested mapping (see policy.py or mcp_scan.py) and add a hook-shaped regression test. (thread)

#260 review projection

  • Parent resolution ignores parent type and duplicate IDs. Build an ID-to-events index, validate allowed parent types, and surface missing or ambiguous links as gaps. (thread)
  • Nothing produces is_test, is_command, recovery_of, or recovered, so those filters are empty for real traces. Derive them from recorded fields or add a normalization step, and test a hook-shaped trace. (thread)

#261 v1 workflow doc

  • Add the copy-or-merge step into ~/.claude/settings.json, or link to the setup instructions, before the verification step. (thread)

Suggested order is #256 → #257 → #258 → #259 → #260 → #261, since #258 and #261 depend on how the earlier ones settle. Once you push, I'll approve the CI runs and squash-merge each PR as it goes green.

If you don't have time for some of these, tell me which ones and I'll fix them on your branches with your authorship kept.

@harshitethic

Copy link
Copy Markdown
Author

Addressed the two outstanding review points:

  • Orphaned ERROR events with missing/unknown parent_id now produce non-healthy reasons (unlinked_error / orphan_error) with their source event IDs, while correctly paired failed tool/LLM requests remain valid terminal outcomes.
  • Evidence-health results now accept and serialize capture_matrix_version, which callers can populate from the provider capture matrix (product: add verified setup diagnostics and a provider capture matrix #239); the docs explain the provenance boundary.
  • Added regressions for unknown-parent and missing-parent errors and schema serialization.

Validation on updated head: PYTHONPATH=src python3 -m unittest discover -s tests -p 'test_evidence_health.py' -v — 19 passed, plus git diff --check.

This update was AI-assisted (ChatGPT GPT-6) and the focused tests were executed on my Mac. Please let me know if you want a different orphan-error reason taxonomy.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

product: calculate and surface per-session evidence health

2 participants