Skip to content

[release/13.6] Fix remaining Outerloop failures in DCP, CLI, and Dashboard tests - #20664

Merged
Jose Perez Rodriguez (joperezr) merged 9 commits into
release/13.6from
backport/pr-20466-to-release/13.6
Oct 2, 2026
Merged

Jose Perez Rodriguez (joperezr) merged 9 commits into
release/13.6from
backport/pr-20466-to-release/13.6

Conversation

@aspire-repo-bot

@aspire-repo-bot aspire-repo-bot Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Backport of #20466 to release/13.6

/cc David Negstad (@danegsta) Ankit Jain (@radical)

Customer Impact

Customers running Aspire 13.6.0 can see a critical Watch task over Kubernetes ContainerExec resources terminated unexpectedly error approximately one minute after AppHost startup. The application otherwise continues running, but the terminated watch can prevent subsequent container exec/terminal state and log updates from being observed.

Testing

The source PR validated KubernetesServiceTests and PeriodicRestartAsyncEnumerableTests with 14 passing tests, including 10 repeated stalled-watch runs and 20 repeated independent-cancellation runs. It also passed 61 CLI install strategy tests, 13 dashboard interaction tests, 10 repeated smooth-scroll tests, and builds of the deployment E2E and template test projects.

Risk

Low. The only shipping behavior change is localized to DCP Kubernetes watch connection retry and cancellation handling; the remaining changes affect test infrastructure. DCP watches already reconnect periodically, and focused tests cover stalled connections and unrelated cancellation propagation.

Regression?

Yes — this regressed in Aspire 13.6.0 following #19641. The reporter confirmed the failure occurs consistently on 13.6.0 and did not occur on 13.5.3.

Ankit Jain (radical) and others added 9 commits October 1, 2026 21:31
LocalArchive E2E installs copied built packages into the local hive but
left the CLI on its baked channel, so aspire add could not resolve the
14.0.0-ci AppHost SDK.

DCP watches also inherited the one-minute non-streaming API timeout even
though the Kubernetes client buffers their long-lived response body.
The dashboard smooth-scroll test separately depended on observing an
intermediate compositor frame that Windows headless Chromium can skip.

Select the local hive after LocalArchive installs, keep watch requests
alive until their periodic restart, and drive the scrollend contract
deterministically in the Playwright test.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
LocalArchive installs can auto-detect a PR package suffix and place packages
in a pr-N hive, which conflicts with configuring the installed CLI to use
the local hive.

Long-lived watch factories can also be canceled by their periodic restart
token before returning an enumerable. That cancellation currently escapes
and terminates the outer watch instead of recreating it.

Force LocalArchive installs into the local hive and restart factories whose
restart interval expires. Cover both LocalArchive command paths and both
periodic-enumerable overloads.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The periodic factory restart handler currently treats every cancellation as
a restart while the outer token is active. A factory canceled for another
reason would be silently retried instead of propagating its failure.

Handle cancellation only when the periodic restart token fired, and verify
that unrelated cancellation propagates for both enumerable overloads.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
PR archive tests install pre-downloaded packages through LocalArchive.
Forcing those packages into the local hive breaks the CLI's baked pr-N
channel, so generated AppHosts cannot resolve the PR AppHost SDK.

Allow the installer to infer the hive from the package version and only
configure the global local channel for non-PR archives. Cover both PR
and CI package identities with focused tests.

Periodic restart handling also retried independent cancellation in the
struct factory and class enumerator paths. Require the periodic token to
be canceled before restarting, and verify unrelated cancellation is
propagated without retrying the factory.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
LocalArchive CI builds retain a daily assembly identity even after user
configuration selects the local channel. Generated projects therefore use
daily feeds and cannot resolve CI SDK packages from the local hive.

Kubernetes watches can also stall while establishing their initial HTTP
response. Giving one attempt the periodic restart token prevents retries
for five minutes and leaves template tests without resource updates.

Override the CLI identity for non-PR local archives. Give each watch
connection attempt the API timeout and retry timed-out attempts, while
keeping established streams under the periodic restart token.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Template outerloop tests locate resource rows by an exact Fluent UI class
string. Current rows render with the stable resource-row class, so visible
resources are ignored until the tests time out.

Use the dashboard's existing resource-row selector and scope grid-cell
lookups to each matched row.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Fluent data-grid cells expose the gridcell role on custom elements rather
than native td elements. The tag-qualified selector finds no cells after
the resource-row locator succeeds.

Select grid cells by role within each matched resource row.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Fluent data-grid copies the configured row class onto each plain grid
cell. Matching resource-row alone therefore returns both tr rows and td
cells, and cell matches have no descendant grid cells.

Restrict the resource locator to tr elements while keeping cell lookups
scoped to each row.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Periodic restart filters treated any OperationCanceledException observed
after the restart interval as restart cancellation, even when the
exception belonged to an independent token. Require the exception token
to match the restart token, and cover both overloads during factory
creation and enumeration with deterministic cancellation synchronization.

Match local archive PR package suffixes to the lowercase [0-9a-g]+
contract used by both installer scripts so hive selection behaves
consistently.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

🚀 Dogfood this PR with:

⚠️ WARNING: Do not do this without first carefully reviewing the code of this PR to satisfy yourself it is safe.

curl -fsSL https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.sh | bash -s -- 20664

Or

  • Run remotely in PowerShell:
iex "& { $(irm https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.ps1) } 20664"

@github-actions github-actions Bot added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Oct 1, 2026
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Tests selector

52 / 99 PR test projects · 4 PR jobs, from 12 changed files.

Selected PR test projects (52 / 99)

Aspire.Cli.EndToEnd.Tests, Aspire.Dashboard.Tests, Aspire.Hosting.Analyzers.Tests, Aspire.Hosting.Azure.Kubernetes.Tests, Aspire.Hosting.Azure.Kusto.Tests, Aspire.Hosting.Azure.Provisioning.Tests, Aspire.Hosting.Azure.Tests, Aspire.Hosting.Blazor.Tests, Aspire.Hosting.Browsers.Tests, Aspire.Hosting.CodeGeneration.Go.Tests, Aspire.Hosting.CodeGeneration.Java.Tests, Aspire.Hosting.CodeGeneration.Python.Tests, Aspire.Hosting.CodeGeneration.Rust.Tests, Aspire.Hosting.CodeGeneration.TypeScript.Tests, Aspire.Hosting.Containers.Tests, Aspire.Hosting.DevTunnels.Tests, Aspire.Hosting.Docker.Tests, Aspire.Hosting.Dotnet.Tests, Aspire.Hosting.DotnetTool.Tests, Aspire.Hosting.EntityFrameworkCore.Tests, Aspire.Hosting.Foundry.Tests, Aspire.Hosting.Garnet.Tests, Aspire.Hosting.Go.Tests, Aspire.Hosting.Java.Tests, Aspire.Hosting.JavaScript.Tests, Aspire.Hosting.Kafka.Tests, Aspire.Hosting.Keycloak.Tests, Aspire.Hosting.Kubernetes.Tests, Aspire.Hosting.Maui.Tests, Aspire.Hosting.Milvus.Tests, Aspire.Hosting.MongoDB.Tests, Aspire.Hosting.MySql.Tests, Aspire.Hosting.Nats.Tests, Aspire.Hosting.OpenAI.Tests, Aspire.Hosting.Oracle.Tests, Aspire.Hosting.Orleans.Tests, Aspire.Hosting.PostgreSQL.Tests, Aspire.Hosting.Python.Tests, Aspire.Hosting.Qdrant.Tests, Aspire.Hosting.RabbitMQ.Tests, Aspire.Hosting.Radius.Tests, Aspire.Hosting.Redis.Tests, Aspire.Hosting.RemoteHost.Tests, Aspire.Hosting.Rust.Tests, Aspire.Hosting.Seq.Tests, Aspire.Hosting.SqlServer.Tests, Aspire.Hosting.Testing.Tests, Aspire.Hosting.Tests, Aspire.Hosting.Valkey.Tests, Aspire.Hosting.Yarp.Tests, Aspire.Playground.Tests, Aspire.Templates.Tests

Selected PR jobs (4)

cli-starter-validation, extension-e2e, polyglot, typescript-api-compat


How these were chosen — grouped by what changed

⚠️ 45 of the 52 selected test projects come from a single change — src/Aspire.Hosting/Dcp/KubernetesService.cs.

🔧 src/Aspire.Hosting/Dcp/KubernetesService.cs (changed source)
→ 45 via the project graph

show 45

Aspire.Hosting.Analyzers.Tests (2 hops), Aspire.Hosting.Azure.Kubernetes.Tests (2 hops), Aspire.Hosting.Azure.Kusto.Tests (2 hops), Aspire.Hosting.Azure.Provisioning.Tests (3 hops), Aspire.Hosting.Azure.Tests, Aspire.Hosting.Browsers.Tests (2 hops), Aspire.Hosting.CodeGeneration.Go.Tests, Aspire.Hosting.CodeGeneration.Java.Tests, Aspire.Hosting.CodeGeneration.Python.Tests, Aspire.Hosting.CodeGeneration.Rust.Tests, Aspire.Hosting.CodeGeneration.TypeScript.Tests, Aspire.Hosting.Containers.Tests (2 hops), Aspire.Hosting.DevTunnels.Tests (2 hops), Aspire.Hosting.Docker.Tests (2 hops), Aspire.Hosting.DotnetTool.Tests (2 hops), Aspire.Hosting.EntityFrameworkCore.Tests (2 hops), Aspire.Hosting.Foundry.Tests (2 hops), Aspire.Hosting.Garnet.Tests (2 hops), Aspire.Hosting.Go.Tests (2 hops), Aspire.Hosting.Java.Tests (2 hops), Aspire.Hosting.JavaScript.Tests (2 hops), Aspire.Hosting.Kafka.Tests (2 hops), Aspire.Hosting.Keycloak.Tests (2 hops), Aspire.Hosting.Kubernetes.Tests (2 hops), Aspire.Hosting.Maui.Tests, Aspire.Hosting.Milvus.Tests (2 hops), Aspire.Hosting.MongoDB.Tests (2 hops), Aspire.Hosting.MySql.Tests (2 hops), Aspire.Hosting.Nats.Tests (2 hops), Aspire.Hosting.OpenAI.Tests (2 hops), Aspire.Hosting.Oracle.Tests (2 hops), Aspire.Hosting.Orleans.Tests (2 hops), Aspire.Hosting.PostgreSQL.Tests (2 hops), Aspire.Hosting.Python.Tests (2 hops), Aspire.Hosting.Qdrant.Tests (2 hops), Aspire.Hosting.RabbitMQ.Tests (2 hops), Aspire.Hosting.Redis.Tests (2 hops), Aspire.Hosting.RemoteHost.Tests, Aspire.Hosting.Rust.Tests (2 hops), Aspire.Hosting.Seq.Tests (2 hops), Aspire.Hosting.SqlServer.Tests (2 hops), Aspire.Hosting.Testing.Tests (2 hops), Aspire.Hosting.Valkey.Tests (2 hops), Aspire.Hosting.Yarp.Tests (2 hops), Aspire.Playground.Tests

🧪 tests/Aspire.Hosting.Tests/Dcp/KubernetesServiceTests.cs (changed test)
→ 1 directly: Aspire.Hosting.Tests
→ 3 via the project graph: Aspire.Hosting.Blazor.Tests, Aspire.Hosting.Dotnet.Tests, Aspire.Hosting.Radius.Tests

📦 affected project Aspire.Hosting
→ 1 test: Aspire.Cli.EndToEnd.Tests

🧪 tests/Aspire.Cli.EndToEnd.Tests/Helpers/CliE2EAutomatorHelpers.cs (changed test)
→ 1 directly: Aspire.Cli.EndToEnd.Tests

🧪 tests/Aspire.Cli.EndToEnd.Tests/Helpers/CliInstallStrategyTests.cs (changed test)
→ 1 directly: Aspire.Cli.EndToEnd.Tests

🧪 tests/Aspire.Cli.EndToEnd.Tests/NewChannelNuGetConfigTests.cs (changed test)
→ 1 directly: Aspire.Cli.EndToEnd.Tests

🧪 tests/Aspire.Dashboard.Tests/Integration/Playwright/DashboardInteractionsTests.cs (changed test)
→ 1 directly: Aspire.Dashboard.Tests

🧪 tests/Aspire.Hosting.Tests/Utils/PeriodicRestartAsyncEnumerableTests.cs (changed test)
→ 1 directly: Aspire.Hosting.Tests

🧪 tests/Aspire.Templates.Tests/TemplateTestsBase.cs (changed test)
→ 1 directly: Aspire.Templates.Tests

Job reasons

Job Triggered by
cli-starter-validation affected project Aspire.Hosting.JavaScript
extension-e2e • src/Aspire.Hosting/Dcp/KubernetesService.cs, src/Aspire.Hosting/Utils/PeriodicRestartAsyncEnumerable.cs
• affected project Aspire.Hosting
polyglot • affected project Aspire.Hosting
• affected project Aspire.Hosting.Azure.Provisioning
typescript-api-compat affected project Aspire.Hosting

Selection computed for commit 30b542f.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Retrying the failed CI jobs for this pull request from the CI run attempt. The rerun is being tracked in the rerun attempt.

@bart-vmware

Copy link
Copy Markdown
Contributor

We independently ran into the problem this PR fixes (#20609), with severe impact, and can confirm the backport fixes it. Details below in case they help with prioritizing it for 13.6.1.

Impact is larger than "ContainerExec / terminal updates". The Executable and Endpoint watches die the same way. When the Executable watch is dead, the AppHost never learns that DCP started project resources: DCP runs them (Running/Healthy in its logs), but the AppHost keeps them in Waiting, never runs their health checks and never requests their logs. Anything that WaitFors a container taking longer than ~60 seconds to become healthy hangs. In our case that is CI on macOS (Docker in Lima/QEMU, image pull plus database init take 5-7 minutes), where every test using DistributedApplicationTestingBuilder with such a container times out. Same tests pass on 13.5.x and on Windows/Linux. Full analysis and a minimal repro are in #20609.

Verification. We built v13.6.0 plus only the two src/Aspire.Hosting hunks of this PR (Aspire.Hosting built with AssemblyVersion=13.6.0.0, swapped into the test output), with the official assembly as a control, using a container whose health is delayed on purpose:

Assembly Health delay Result
official 13.6.0 90 s fails: Watch task over Kubernetes … terminated unexpectedly, project stuck in Waiting
13.6.0 + this PR 90 s passes, no critical messages
13.6.0 + this PR 390 s (crosses the 5-minute periodic restart) passes, no critical messages

Both hunks matter in practice: without the PeriodicRestartAsyncEnumerable change, a watch still waiting for its first event is killed when the 5-minute restart fires. Several of our macOS runs created the first executable 5 to 7 minutes after startup. We tested on Windows only, and Aspire.Hosting was built with -p:DefaultTargetFramework=net10.0 because the repo requires an SDK 11 preview.

Request. This PR has no milestone yet, and release/13.6 is already versioned 13.6.1. Could it be merged for 13.6.1? We are carrying a reflection-based workaround for it in our test fixture and would like to remove it.

We also found two other problems that this PR does not cover. They are reported separately and are not a reason to hold this one back: #20691 and microsoft/dcp#286.

@joperezr
Jose Perez Rodriguez (joperezr) merged commit 07f7de7 into release/13.6 Oct 2, 2026
646 of 649 checks passed
@aspire-repo-bot

Copy link
Copy Markdown
Contributor Author

✅ No documentation update needed.

Step 5 branch taken: excluded → base_branch_is_release, head_branch_is_backport, title_release_prefix, body_backport_marker

This PR is a backport and is out of scope for docs generation per the exclusion rule, which overrides recommendation regardless of value.

  • Triggered signals: no signals triggered (signal_count: 0, triggered_signals: []).
  • Exclusion reasons (from signals.json):
    1. base_branch_is_release — base ref is release/13.6 (weak/supporting signal only).
    2. head_branch_is_backport — head branch follows the backport-bot naming convention.
    3. title_release_prefix — PR title is prefixed [release/13.6].
    4. body_backport_marker — PR body contains a backport marker referencing the original PR.

Since this is a backport of an already-merged change, its user-facing documentation (if any was needed) should already have been authored against the original forward PR on the default branch. Drafting a second docs PR here would be duplicate noise. No docs PR was created.

This was referenced Oct 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Aspire 13.6.0 always errors with "Watch task over Kubernetes ContainerExec resources terminated unexpectedly"

4 participants