Skip to content

testplan_show_test_results_from_build_id still omits detailed fields after #1011 #1682

Description

Summary

testplan_show_test_results_from_build_id still returns lightweight test-result rows without testCaseTitle, errorMessage, stackTrace, automatedTestName, or automatedTestStorage in @azure-devops/mcp 2.10.0 and current main.

This appears to be a regression introduced during review of #1011, which closed #887 and #985.

Current behavior

For a build with published test results, the tool returns rows shaped like:

{
  "id": 100000,
  "outcome": "Passed",
  "durationInMs": 6753386,
  "runId": "1706006238"
}

The tool description promises test titles, error messages, and stack traces, but those properties are absent from the serialized response.

Root cause analysis

The first commit in #1011 changed the implementation from getTestResultDetailsForBuild to:

  1. getTestRuns(project, buildUri)
  2. getTestResults(project, runId, ..., outcomes) for each run

That matches the PR description, which states that getTestResultDetailsForBuild returns lightweight entries without the detailed fields, while per-run getTestResults returns full TestCaseResult objects.

A later review commit in the same PR switched the implementation back to:

getTestResultDetailsForBuild(
  project,
  buildid,
  undefined,
  undefined,
  outcomeFilter,
  undefined,
  true
)

Current main and published 2.10.0 still use that API. The formatter then reads fields such as r.testCaseTitle, but they are undefined on lightweight rows, so JSON.stringify silently omits them.

There was no later revert after merge: merge commit 67b9aa68169f90fc311b648ad470d5f95c2724b8 remains an ancestor of main. The ineffective API choice was included in the final merged form of #1011.

Test gap

The unit tests mock getTestResultDetailsForBuild as returning fully populated result objects. That does not match the observed production API response, so the tests validate the formatter but not the API contract that supplies its input.

Expected behavior

Each returned row should include, when available:

  • testCaseTitle
  • errorMessage
  • stackTrace
  • automatedTestName
  • automatedTestStorage
  • outcome, duration, result ID, and run ID

Outcome filtering should continue to work without requiring callers to retrieve every passed result.

Suggested fix

Restore the per-run detailed-results path from the first revision of #1011:

  1. Query runs for the build.
  2. Query getTestResults for each run with server-side outcome filtering.
  3. Bound request concurrency for builds with many runs.
  4. Add an integration/contract test proving the selected Azure DevOps API actually supplies testCaseTitle, errorMessage, and stackTrace; do not mock those fields onto the lightweight response type.

References

Activity

  1. danhellem commented on Oct 7, 2026

    @danhellem
    Collaborator

    Have you tried using the remote MCP Server instead?

  2. Paraphern commented on Oct 7, 2026

    @Paraphern

    Contract-level bisect, in case it helps narrow the window. I pin tool contracts per release and diff hashes (method at the bottom). For testplan_show_test_results_from_build_id:

    version contract hash tools pinned
    2.5.0 f5be27eb260e60b1 86
    2.6.0 a3440170985a743d 90
    2.7.0 a3440170985a743d 90
    2.8.0 a3440170985a743d 90
    2.9.0 a3440170985a743d 40
    2.10.0 a3440170985a743d 40

    Two takeaways:

    1. This tool's contract changed exactly once, 2.5.0 -> 2.6.0, and never after. Support outcome filtering option in test results #1011 ("Support outcome filtering option in test results") merged 2026-03-18 - the same day 2.6.0 shipped. If the review commits that switched the implementation back to the lightweight fetch are the regression, they rode 2.6.0, and 2.6.0 is the version boundary to diff against. The 2.9.0 consolidation (90 -> 40 tools) never touched this tool's contract.
    2. The description never changed across the whole range - the promised fields have been advertised continuously while the serialized output dropped them, so nothing at the contract surface flags the regression. That's exactly the gap contract pinning can't see through either; response-shape assertions would be the next layer.

    Method: pin-and-diff with an open-source zero-dependency tool, 4-command repro in the audit report: https://github.com/Paraphern/rugsnare/blob/main/audits/npm-top-mcp-drift-2026-10.md (Section 1 also has the full 2.5.0 -> 2.10.0 map: 75 removed / 29 added / 6 rewritten tool names - happy to post the lists here or as JSON if a migration note would use them).

  3. danhellem commented on Oct 7, 2026

    @danhellem
    Collaborator

    Please take a look at #1684 and let me know if this solves your issues. I would love to have more than myself testing it before we merge

  4. Paraphern commented on Oct 7, 2026

    @Paraphern

    Tested the branch at the contract level (pinned the built server and diffed against the 2.10.0 release). What I can confirm from here:

    • Startup/tool surface: clean. 44 tools pinned vs 40 in the 2.10.0 release - the 4 extra come from main being ahead (repo_pull_request_org, repo_file_write, wit_work_item_attachment_upload, wit_work_item_attachment_link), nothing removed.
    • testplan_show_test_results_from_build_id contract: description unchanged; input schema has exactly one change - outcomes went from free-form string array to an enum of the 15 valid outcome values (Unspecified ... NotImpacted). That matches the PR's "stricter schema validation for outcome filters" precisely.
    • What I can't verify without a real org + published test runs: the response fields themselves (testCaseTitle, errorMessage, stackTrace). The paging-over-runs approach in the code is exactly what the testplan_show_test_results_from_build_id still omits detailed fields after #1011 #1682 root-cause analysis pointed at (per-run results return full TestCaseResult objects), so it looks right - but the runtime confirmation needs someone with a live instance.

    One small note worth a changelog line when this ships: the outcomes enum is technically a client-facing input-schema change - free-form strings (e.g. a lowercase failed) that 2.10.0 accepted now fail validation. Probably exactly what you want, but agents/skills passing lowercase values will notice.

    For what it's worth, this package is on our weekly drift watch (pinned baselines, diffed on every release) - once the fix lands in a published version, the watcher will pick it up automatically.

  5. DavidS-msft commented on Oct 8, 2026

    @DavidS-msft
    MemberAuthor

    Confirmed with a live organization and published test runs.

    Baseline @azure-devops/mcp 2.10.0 returned only id, outcome, durationInMs, and runId for the reproducing failed result. I built PR #1684 at c6357be, ran the full suite (1,351/1,351 passed), tool validation, and the TypeScript build, then invoked that exact branch against the same build/result. The PR returned testCaseTitle, the full errorMessage, automatedTestName, and automatedTestStorage, so it resolves #1682 for this live case. Outcome filtering with Failed also worked correctly.

    I also tested the hosted remote MCP server directly. It already returns those detailed fields for the same result. That server is different from my current setup: I am currently using the local stdio package (npx -y @azure-devops/mcp@latest, currently 2.10.0), while the remote server is the Azure DevOps-hosted streamable-HTTP service and receives updates separately/earlier.

  6. Paraphern commented on Oct 8, 2026

    @Paraphern

    That completes the stack nicely - your live-org run covers exactly the gap we flagged (response fields need a real org; we could only verify the contract surface). Between the maintainer's suite, the contract-level check and your live verification, #1684 looks merge-ready from the outside.

    One observation worth a docs line for anyone pinning this package: the local/remote split you describe is itself a drift surface. The npm stdio channel (2.10.0, fields missing) and the hosted streamable-HTTP channel (fields present) are versioned separately - same product, different contracts per channel, and the remote one moves out-of-band from npm releases. Consumers who pin and diff only the npm package get no signal when the remote contract changes; worth stating explicitly whenever the remote server's versioning gets documented.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions