Repository navigation
[release/13.6] Fix remaining Outerloop failures in DCP, CLI, and Dashboard tests - #20664
Jose Perez Rodriguez (joperezr) merged 9 commits into
Conversation
LocalArchive E2E installs copied built packages into the local hive but left the CLI on its baked channel, so aspire add could not resolve the 14.0.0-ci AppHost SDK. DCP watches also inherited the one-minute non-streaming API timeout even though the Kubernetes client buffers their long-lived response body. The dashboard smooth-scroll test separately depended on observing an intermediate compositor frame that Windows headless Chromium can skip. Select the local hive after LocalArchive installs, keep watch requests alive until their periodic restart, and drive the scrollend contract deterministically in the Playwright test. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
LocalArchive installs can auto-detect a PR package suffix and place packages in a pr-N hive, which conflicts with configuring the installed CLI to use the local hive. Long-lived watch factories can also be canceled by their periodic restart token before returning an enumerable. That cancellation currently escapes and terminates the outer watch instead of recreating it. Force LocalArchive installs into the local hive and restart factories whose restart interval expires. Cover both LocalArchive command paths and both periodic-enumerable overloads. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The periodic factory restart handler currently treats every cancellation as a restart while the outer token is active. A factory canceled for another reason would be silently retried instead of propagating its failure. Handle cancellation only when the periodic restart token fired, and verify that unrelated cancellation propagates for both enumerable overloads. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
PR archive tests install pre-downloaded packages through LocalArchive. Forcing those packages into the local hive breaks the CLI's baked pr-N channel, so generated AppHosts cannot resolve the PR AppHost SDK. Allow the installer to infer the hive from the package version and only configure the global local channel for non-PR archives. Cover both PR and CI package identities with focused tests. Periodic restart handling also retried independent cancellation in the struct factory and class enumerator paths. Require the periodic token to be canceled before restarting, and verify unrelated cancellation is propagated without retrying the factory. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
LocalArchive CI builds retain a daily assembly identity even after user configuration selects the local channel. Generated projects therefore use daily feeds and cannot resolve CI SDK packages from the local hive. Kubernetes watches can also stall while establishing their initial HTTP response. Giving one attempt the periodic restart token prevents retries for five minutes and leaves template tests without resource updates. Override the CLI identity for non-PR local archives. Give each watch connection attempt the API timeout and retry timed-out attempts, while keeping established streams under the periodic restart token. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Template outerloop tests locate resource rows by an exact Fluent UI class string. Current rows render with the stable resource-row class, so visible resources are ignored until the tests time out. Use the dashboard's existing resource-row selector and scope grid-cell lookups to each matched row. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Fluent data-grid cells expose the gridcell role on custom elements rather than native td elements. The tag-qualified selector finds no cells after the resource-row locator succeeds. Select grid cells by role within each matched resource row. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Fluent data-grid copies the configured row class onto each plain grid cell. Matching resource-row alone therefore returns both tr rows and td cells, and cell matches have no descendant grid cells. Restrict the resource locator to tr elements while keeping cell lookups scoped to each row. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Periodic restart filters treated any OperationCanceledException observed after the restart interval as restart cancellation, even when the exception belonged to an independent token. Require the exception token to match the restart token, and cover both overloads during factory creation and enumeration with deterministic cancellation synchronization. Match local archive PR package suffixes to the lowercase [0-9a-g]+ contract used by both installer scripts so hive selection behaves consistently. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
🚀 Dogfood this PR with:
curl -fsSL https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.sh | bash -s -- 20664Or
iex "& { $(irm https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.ps1) } 20664" |
Tests selector52 / 99 PR test projects · 4 PR jobs, from 12 changed files. Selected PR test projects (52 / 99)
Selected PR jobs (4)
How these were chosen — grouped by what changed
🔧 show 45
🧪 📦 affected project 🧪 🧪 🧪 🧪 🧪 🧪 Job reasons
Selection computed for commit |
|
Retrying the failed CI jobs for this pull request from the CI run attempt. The rerun is being tracked in the rerun attempt. |
|
We independently ran into the problem this PR fixes (#20609), with severe impact, and can confirm the backport fixes it. Details below in case they help with prioritizing it for 13.6.1. Impact is larger than "ContainerExec / terminal updates". The Verification. We built v13.6.0 plus only the two
Both hunks matter in practice: without the Request. This PR has no milestone yet, and release/13.6 is already versioned 13.6.1. Could it be merged for 13.6.1? We are carrying a reflection-based workaround for it in our test fixture and would like to remove it. We also found two other problems that this PR does not cover. They are reported separately and are not a reason to hold this one back: #20691 and microsoft/dcp#286. |
07f7de7
into
release/13.6
|
✅ No documentation update needed. Step 5 branch taken: This PR is a backport and is out of scope for docs generation per the exclusion rule, which overrides
Since this is a backport of an already-merged change, its user-facing documentation (if any was needed) should already have been authored against the original forward PR on the default branch. Drafting a second docs PR here would be duplicate noise. No docs PR was created. |
Backport of #20466 to release/13.6
/cc David Negstad (@danegsta) Ankit Jain (@radical)
Customer Impact
Customers running Aspire 13.6.0 can see a critical
Watch task over Kubernetes ContainerExec resources terminated unexpectedlyerror approximately one minute after AppHost startup. The application otherwise continues running, but the terminated watch can prevent subsequent container exec/terminal state and log updates from being observed.Testing
The source PR validated
KubernetesServiceTestsandPeriodicRestartAsyncEnumerableTestswith 14 passing tests, including 10 repeated stalled-watch runs and 20 repeated independent-cancellation runs. It also passed 61 CLI install strategy tests, 13 dashboard interaction tests, 10 repeated smooth-scroll tests, and builds of the deployment E2E and template test projects.Risk
Low. The only shipping behavior change is localized to DCP Kubernetes watch connection retry and cancellation handling; the remaining changes affect test infrastructure. DCP watches already reconnect periodically, and focused tests cover stalled connections and unrelated cancellation propagation.
Regression?
Yes — this regressed in Aspire 13.6.0 following #19641. The reporter confirmed the failure occurs consistently on 13.6.0 and did not occur on 13.5.3.