Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -229,6 +229,13 @@ repos:
files: '^scripts/dev/((install_)?merge_train_guard\.py|tests/test_(install_)?merge_train_guard\.py)$'
pass_filenames: false

- id: rc3-home-gpu-retest-contract
name: RC3 home GPU retest kit regressions (ADR-1386)
entry: python3 -B -m unittest discover -s scripts/dev/tests -p test_rc3_home_gpu_retest.py
language: system
files: '^(scripts/dev/(rc3-home-gpu-retest\.sh|rc3_retest_helpers\.py|tests/test_rc3_home_gpu_retest\.py)|docs/state\.md)$'
pass_filenames: false

- id: test-safe-subprocess-boundary
name: Bounded process execution and fail-soft consumer regressions
entry: python3 -m unittest scripts.lib.test_safe_subprocess scripts.lib.test_backlog_tracker scripts.ci.tests.test_agent_eligibility_precheck
Expand Down
9 changes: 9 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,15 @@
`backend_used`, so a run that mixes device twins and CPU extractors says so.


- **RC3 home GPU retest kit** (`scripts/dev/rc3-home-gpu-retest.sh`): runs the
verify-and-time commands that the `docs/state.md` rows carry for the RTX 4090
(CUDA), the Arc A380 (SYCL) and the gfx1036 iGPU (HIP), one entry per row,
with every device run under that device's lock. It writes a log and the JSON
of every run per row plus a summary table, and `--baseline DIR` compares a
pull request's run with an earlier run on `master`. See
[the retest guide](docs/development/rc3-home-gpu-retest.md) (ADR-1386).


### Changed

- **`vmaf` reads its two inputs ahead of scoring, on one thread each
Expand Down
7 changes: 7 additions & 0 deletions changelog.d/added/rc3-home-gpu-retest-kit.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
- **RC3 home GPU retest kit** (`scripts/dev/rc3-home-gpu-retest.sh`): runs the
verify-and-time commands that the `docs/state.md` rows carry for the RTX 4090
(CUDA), the Arc A380 (SYCL) and the gfx1036 iGPU (HIP), one entry per row,
with every device run under that device's lock. It writes a log and the JSON
of every run per row plus a summary table, and `--baseline DIR` compares a
pull request's run with an earlier run on `master`. See
[the retest guide](docs/development/rc3-home-gpu-retest.md) (ADR-1386).
1 change: 1 addition & 0 deletions core/test/test_meson_secret_env_sanitization.py
Original file line number Diff line number Diff line change
Expand Up @@ -118,6 +118,7 @@
Path(".zed/tasks.json"): ("/workspace/scripts/ci/run_meson_test.py",),
Path(".claude/skills/bisect-regression/scaffold.sh"): ("scripts/ci/run_meson_test.py",),
Path("scripts/dev/preflight.sh"): ("scripts/ci/run_meson_test.py",) * 3,
Path("scripts/dev/rc3-home-gpu-retest.sh"): ("scripts/ci/run_meson_test.py",),
Path("scripts/setup/ubuntu.sh"): ("scripts/ci/run_meson_test.py",),
Path("scripts/sync-pelorus-interop.sh"): ("scripts/ci/run_meson_test.py",),
}
Expand Down
44 changes: 44 additions & 0 deletions docs/adr/1386-rc3-home-gpu-retest-kit.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
<!-- markdownlint-disable MD013 MD060 -->
# ADR-1386: One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks

- **Status**: Accepted
- **Date**: 2026-09-30
- **Deciders**: lusoris
- **Tags**: tooling, testing, verification, rc3, cuda, hip, sycl, fork-local

## Context

RC3 work moved from the office box (Arc B580, UHD 770) to `ryzen-4090-arc` (RTX 4090, Arc A380, gfx1036 iGPU). Fifteen open rows of [`docs/state.md`](../state.md) carry verify-and-time commands for that box, and one closed row still asks for an A380 check of the oneAPI release image. The rows were written by different changes over two days, so their commands differ in fixtures, frame counts, tolerances, and in whether they call `scripts/dev/speed_gpu_parity.py` or spell out `vmaf` runs and a Python one-liner. Run by hand, that is about seventy commands per backend. Several open pull requests (#1636 HIP parity, #1637 CUDA parity, #1639 CUDA cambi and SpEED, #1630 SYCL strict floating point) are to be measured against what `master` gives on these devices. The handoff issue #1641 names a script that runs them all as the first task at home.

The box is also shared: several agents build and run on its three GPUs at once. A timing taken while another job holds the GPU measures the other job, and two parity runs on one device can each slow the other past a timeout.

## Decision

`scripts/dev/rc3-home-gpu-retest.sh` holds one entry per `docs/state.md` row and backend. Each entry spells out its row's commands in bash (the JSON comparisons, medians and summaries live in `scripts/dev/rc3_retest_helpers.py`); the script never reads `docs/state.md` at run time. A pull request that adds or changes a row's commands for this box changes the entry in the same pull request. A test checks that every entry names a row that exists, and a pre-commit hook runs it when the kit or `docs/state.md` changes.

Every device run holds that device's `flock` (`cuda-4090.lock`, `hip-gfx1036.lock`, `sycl-a380.lock`), taken before a timing block's clock starts and held through its repetitions, and the kit never holds two device locks at once. CPU reference runs take no lock. Each entry keeps its JSON output, and `--baseline DIR` compares a later run's output with an earlier one, which is how the rows' "before" and "after" runs are done. Timings record the load average and, for CUDA, what the 4090 already had in use.

## Alternatives considered

| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| **Chosen**: explicit entries per row in one bash script, JSON logic in a small Python helper | each command is readable next to its row; a row's quirks (serial CPU for motion, 48 vs 22 frames, 2.2e-15 for cambi) stay exact; `--dry-run` shows what will run | a row edit needs a matching entry edit | — |
| Parse the commands out of `docs/state.md` at run time | never out of step with the rows | the rows are prose with inline code, shell loops and placeholders ("the ported build", "N = 2 and 22"); a parser would guess, and a wrong guess reports a wrong verdict | rejected: the rows are written for people, not for a parser |
| Extend `scripts/dev/speed_gpu_parity.py` to cover every row | one tool, already used by five rows | the other rows compare different keys with different bounds, need `feature_backends` and fallback-warning checks, meson tests, and a Docker image; the script would grow a row table of its own | rejected: keep it the per-feature parity tool it is, and call it from the kit where a row does |
| Run every entry in the `vmaf-dev-mcp` container | the container is the canonical environment | the rows are written for a host `build/` and the release-image row runs Docker itself; device locks are host files | rejected for this kit; the container remains the place for published numbers (ADR-1102) |
| No device locks; ask for an idle box | simpler | the box is shared by design; an unlocked timing measures whichever job ran alongside | rejected: correctness runs must not collide, and timings need to say what they shared |
| One lock for the whole kit run | simple | holds all three GPUs for an hour, blocking every other job | rejected: per-device locks, released between runs |

## Consequences

- **Positive**: one command gives a dated, per-row record of what `master` does on this box, and the same command with `--baseline` gives the before/after comparison the rows ask for. Findings that only show on real hardware (a row whose command errors on `master`, a build option that does not build) surface in one pass.
- **Negative**: the entries duplicate the rows' commands, so they can drift; the entry test only proves that each entry's row exists, not that its commands still match the row's text.
- **Neutral / follow-ups**: rows that the open pull requests add for this box (for example `T-CUDA-FP-CONTRACT-DEFAULT-2026-09-29` and `T-HIP-FP-CONTRACT-DEFAULT-2026-09-29` from #1630) get entries when they land. Whether a check should require an entry for every row that names `ryzen-4090-arc` is left open.

## References

- Handoff issue [#1641](https://github.com/VMAFx/vmafx/issues/1641), "Home retest kit" (req, verbatim): "Every open CUDA/HIP row in `docs/state.md` carries its own verify-and-time commands (grep `ryzen-4090-arc`); a `scripts/dev/rc3-home-gpu-retest.sh` that runs them all is the first home task."
- Paraphrased from the dispatching instruction for this change: encode each command explicitly with its row id instead of parsing the markdown, let the user pick backends, a build directory, `--list` and `--only`, select the A380 and the gfx1036 explicitly, hold the per-device lock for every device run, and write a log per row plus a summary table.
- [ADR-1185](1185-backend-perf-baseline-methodology.md): median-of-N timing and recording the load, which the kit follows.
- [ADR-0165](0165-state-md-bug-tracking.md): `docs/state.md` as the bug ledger the entries point at.
- [Research-1386](../research/1386-rc3-home-gpu-retest-master-baseline.md): the first `master` run on `ryzen-4090-arc`.
1 change: 1 addition & 0 deletions docs/adr/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -1173,4 +1173,5 @@ public authority; documentation never links into either local root.
| [ADR-1383](1383-state-md-three-way-conflict-resolver.md) | `scripts/dev/resolve-state-md-conflict.py` resolves a conflicted `docs/state.md` by a three-way merge of the index stages: rows and move tombstones keyed by bug id (state = text plus section), disposition rows keyed by label with their id lists merged as sets, other lines three-way by line; a record both sides changed differently stops with exit 1 and nothing written until `--take NAME=ours\|theirs`; the result must pass `check-state-md-rows.sh`. Replaces "master side wins". | Accepted | process, state-hygiene, tooling, git, fork-local |
| [ADR-1379](1379-cuda-cambi-device-resident-pipeline.md) | Run `cambi_cuda` entirely on the device (the ADR-1357 design on CUDA): device preprocessing, spatial mask, per-scale decimate / mode filter, column-histogram c-values and exact 128-bit top-K pooling, one 88-byte readback per frame; `cambi.c` shares its window guard and host constants with both device twins. | Accepted | cuda, gpu, cambi, performance, numerics, rc3, fork-local |
| [ADR-1380](1380-cuda-speed-device-resident-pipeline.md) | The CUDA SpEED twins run the whole per-frame chain on the device (the ADR-1358 chain on CUDA), 25x25 eigenvalues and QR included, with round-to-nearest intrinsics and `--fmad=false`; one 40-byte readback per frame; `speed_internal_gpu_configure()` sets up both the SYCL and CUDA pipelines | Accepted | cuda, speed, gpu, performance, numerics, rc3, fork-local |
| [ADR-1386](1386-rc3-home-gpu-retest-kit.md) | `scripts/dev/rc3-home-gpu-retest.sh` runs the verify-and-time commands the `docs/state.md` rows carry for `ryzen-4090-arc` (RTX 4090 CUDA, Arc A380 SYCL, gfx1036 HIP): one explicit entry per row and backend, never parsed from the markdown at run time; every device run holds that device's `flock` and a timing block takes it before its clock starts; each entry keeps its JSON so `--baseline` gives the rows' before/after comparison. | Accepted | tooling, testing, verification, rc3, cuda, hip, sycl, fork-local |
| [ADR-1247](1247-scorecard-exact-head-gates.md) | Bind Scorecard gates to their measured source and scope | Accepted | ci, security, supply-chain |
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
| [ADR-1386](1386-rc3-home-gpu-retest-kit.md) | `scripts/dev/rc3-home-gpu-retest.sh` runs the verify-and-time commands the `docs/state.md` rows carry for `ryzen-4090-arc` (RTX 4090 CUDA, Arc A380 SYCL, gfx1036 HIP): one explicit entry per row and backend, never parsed from the markdown at run time; every device run holds that device's `flock` and a timing block takes it before its clock starts; each entry keeps its JSON so `--baseline` gives the rows' before/after comparison. | Accepted | tooling, testing, verification, rc3, cuda, hip, sycl, fork-local |
1 change: 1 addition & 0 deletions docs/adr/_index_fragments/_order.txt
Original file line number Diff line number Diff line change
Expand Up @@ -1081,3 +1081,4 @@
1383-state-md-three-way-conflict-resolver
1379-cuda-cambi-device-resident-pipeline
1380-cuda-speed-device-resident-pipeline
1386-rc3-home-gpu-retest-kit
3 changes: 2 additions & 1 deletion docs/adr/by-tag/cuda.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@

Auto-generated by `scripts/docs/generate-adr-by-tag.sh`. Edit ADR `Tags:` lines to update.

141 ADR(s) carry this tag.
142 ADR(s) carry this tag.

| ID | Title |
|----|-------|
Expand Down Expand Up @@ -148,3 +148,4 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh`. Edit ADR `Tags:` lines
| [ADR-1336](../1336-cuda-context-owned-resource-teardown.md) | Tear down CUDA resources in their owning context |
| [ADR-1379](../1379-cuda-cambi-device-resident-pipeline.md) | Run the CUDA CAMBI extractor entirely on the device |
| [ADR-1380](../1380-cuda-speed-device-resident-pipeline.md) | The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic |
| [ADR-1386](../1386-rc3-home-gpu-retest-kit.md) | One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks |
3 changes: 2 additions & 1 deletion docs/adr/by-tag/fork-local.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@

Auto-generated by `scripts/docs/generate-adr-by-tag.sh`. Edit ADR `Tags:` lines to update.

447 ADR(s) carry this tag.
448 ADR(s) carry this tag.

| ID | Title |
|----|-------|
Expand Down Expand Up @@ -454,3 +454,4 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh`. Edit ADR `Tags:` lines
| [ADR-1379](../1379-cuda-cambi-device-resident-pipeline.md) | Run the CUDA CAMBI extractor entirely on the device |
| [ADR-1380](../1380-cuda-speed-device-resident-pipeline.md) | The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic |
| [ADR-1383](../1383-state-md-three-way-conflict-resolver.md) | Resolve `docs/state.md` rebase conflicts three-way, keyed by bug id |
| [ADR-1386](../1386-rc3-home-gpu-retest-kit.md) | One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks |
3 changes: 2 additions & 1 deletion docs/adr/by-tag/hip.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@

Auto-generated by `scripts/docs/generate-adr-by-tag.sh`. Edit ADR `Tags:` lines to update.

99 ADR(s) carry this tag.
100 ADR(s) carry this tag.

| ID | Title |
|----|-------|
Expand Down Expand Up @@ -106,3 +106,4 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh`. Edit ADR `Tags:` lines
| [ADR-1296](../1296-gpu-failure-path-stub-interposition.md) | GPU init failure paths are tested device-free, by compiling the backend TU against runtime stubs |
| [ADR-1320](../1320-cuda-hip-kernel-header-dependency-tracking.md) | CUDA fatbin and HIP HSACO kernel header dependency tracking |
| [ADR-1325](../1325-integer-adm-barten-fixed-point-normalization.md) | Normalize integer ADM Barten weights with one exponent per scale |
| [ADR-1386](../1386-rc3-home-gpu-retest-kit.md) | One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks |
16 changes: 8 additions & 8 deletions docs/adr/by-tag/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -146,7 +146,7 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh` from each ADR's `Tags:`
| [cross-backend](cross-backend.md) | 2 |
| [cross-backend-parity](cross-backend-parity.md) | 3 |
| [cross-surface](cross-surface.md) | 1 |
| [cuda](cuda.md) | 141 |
| [cuda](cuda.md) | 142 |
| [darwin](darwin.md) | 2 |
| [data](data.md) | 2 |
| [datasets](datasets.md) | 2 |
Expand Down Expand Up @@ -225,7 +225,7 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh` from each ADR's `Tags:`
| [float_moment](float_moment.md) | 1 |
| [floating-point](floating-point.md) | 2 |
| [fork-internal](fork-internal.md) | 2 |
| [fork-local](fork-local.md) | 447 |
| [fork-local](fork-local.md) | 448 |
| [fork-policy](fork-policy.md) | 1 |
| [fp64](fp64.md) | 1 |
| [fr](fr.md) | 1 |
Expand Down Expand Up @@ -265,7 +265,7 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh` from each ADR's `Tags:`
| [headers](headers.md) | 1 |
| [helm](helm.md) | 9 |
| [hfr](hfr.md) | 2 |
| [hip](hip.md) | 99 |
| [hip](hip.md) | 100 |
| [hooks](hooks.md) | 5 |
| [housekeeping](housekeeping.md) | 2 |
| [http](http.md) | 5 |
Expand Down Expand Up @@ -462,7 +462,7 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh` from each ADR's `Tags:`
| [rc](rc.md) | 3 |
| [rc-blocking](rc-blocking.md) | 1 |
| [rc1](rc1.md) | 1 |
| [rc3](rc3.md) | 6 |
| [rc3](rc3.md) | 7 |
| [rclone](rclone.md) | 4 |
| [readme](readme.md) | 1 |
| [rebase](rebase.md) | 2 |
Expand Down Expand Up @@ -545,7 +545,7 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh` from each ADR's `Tags:`
| [supply-chain](supply-chain.md) | 33 |
| [sve2](sve2.md) | 1 |
| [swagger](swagger.md) | 1 |
| [sycl](sycl.md) | 115 |
| [sycl](sycl.md) | 116 |
| [t3-15](t3-15.md) | 2 |
| [t3-15c](t3-15c.md) | 1 |
| [t6-3b](t6-3b.md) | 1 |
Expand All @@ -557,7 +557,7 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh` from each ADR's `Tags:`
| [temporal](temporal.md) | 2 |
| [test](test.md) | 21 |
| [testdata](testdata.md) | 4 |
| [testing](testing.md) | 70 |
| [testing](testing.md) | 71 |
| [tests](tests.md) | 8 |
| [thread-safety](thread-safety.md) | 3 |
| [threading](threading.md) | 13 |
Expand All @@ -566,7 +566,7 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh` from each ADR's `Tags:`
| [tiny-model](tiny-model.md) | 1 |
| [tinyai](tinyai.md) | 1 |
| [toolchain](toolchain.md) | 1 |
| [tooling](tooling.md) | 52 |
| [tooling](tooling.md) | 53 |
| [tools](tools.md) | 19 |
| [touched-file-rule](touched-file-rule.md) | 3 |
| [training](training.md) | 37 |
Expand All @@ -591,7 +591,7 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh` from each ADR's `Tags:`
| [validation](validation.md) | 3 |
| [vendored](vendored.md) | 5 |
| [vendoring](vendoring.md) | 3 |
| [verification](verification.md) | 1 |
| [verification](verification.md) | 2 |
| [video-saliency](video-saliency.md) | 1 |
| [vif](vif.md) | 16 |
| [vlm](vlm.md) | 1 |
Expand Down
3 changes: 2 additions & 1 deletion docs/adr/by-tag/rc3.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@

Auto-generated by `scripts/docs/generate-adr-by-tag.sh`. Edit ADR `Tags:` lines to update.

6 ADR(s) carry this tag.
7 ADR(s) carry this tag.

| ID | Title |
|----|-------|
Expand All @@ -13,3 +13,4 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh`. Edit ADR `Tags:` lines
| [ADR-1370](../1370-sycl-float-ssim-device-decimation.md) | float\_ssim\_sycl decimates on the device, bit-identical to the CPU |
| [ADR-1379](../1379-cuda-cambi-device-resident-pipeline.md) | Run the CUDA CAMBI extractor entirely on the device |
| [ADR-1380](../1380-cuda-speed-device-resident-pipeline.md) | The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic |
| [ADR-1386](../1386-rc3-home-gpu-retest-kit.md) | One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks |
3 changes: 2 additions & 1 deletion docs/adr/by-tag/sycl.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@

Auto-generated by `scripts/docs/generate-adr-by-tag.sh`. Edit ADR `Tags:` lines to update.

115 ADR(s) carry this tag.
116 ADR(s) carry this tag.

| ID | Title |
|----|-------|
Expand Down Expand Up @@ -122,3 +122,4 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh`. Edit ADR `Tags:` lines
| [ADR-1369](../1369-sycl-shared-planes-light-twins.md) | SYCL twins read the planes the state uploads once per frame; opt-in shared chroma planes and a device-side slot fence |
| [ADR-1370](../1370-sycl-float-ssim-device-decimation.md) | float\_ssim\_sycl decimates on the device, bit-identical to the CPU |
| [ADR-1371](../1371-sycl-motion-diff-first-pipeline.md) | SYCL motion differences the frames before the blur, in one shared kernel |
| [ADR-1386](../1386-rc3-home-gpu-retest-kit.md) | One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks |
3 changes: 2 additions & 1 deletion docs/adr/by-tag/testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@

Auto-generated by `scripts/docs/generate-adr-by-tag.sh`. Edit ADR `Tags:` lines to update.

70 ADR(s) carry this tag.
71 ADR(s) carry this tag.

| ID | Title |
|----|-------|
Expand Down Expand Up @@ -77,3 +77,4 @@ Auto-generated by `scripts/docs/generate-adr-by-tag.sh`. Edit ADR `Tags:` lines
| [ADR-1336](../1336-cuda-context-owned-resource-teardown.md) | Tear down CUDA resources in their owning context |
| [ADR-1342](../1342-rc1-external-tester-report-bundle.md) | RC1 external tester evidence bundle |
| [ADR-1361](../1361-psnr-hvs-area-scaled-parity-tolerance.md) | Scale the psnr\_hvs cross-backend tolerance with the CPU's float-sum length |
| [ADR-1386](../1386-rc3-home-gpu-retest-kit.md) | One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks |
Loading
Loading