Skip to content

test(ai/scripts): coverage round 3 — calibrate_phase_f_recipes + analyze_knob_sweep helpers - #755

Closed
lusoris wants to merge 1 commit into
masterfrom
test/ai-scripts-coverage-round3
Closed

lusoris wants to merge 1 commit into
masterfrom
test/ai-scripts-coverage-round3

Conversation

@lusoris

@lusoris lusoris commented Jun 6, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Adds ai/tests/test_calibrate_phase_f_recipes_unit.py (50 tests) covering the 7 pure helper functions in calibrate_phase_f_recipes.py left uncovered by the existing provenance-only smoke test: mos_to_vmaf_proxy, saliency_benefit_to_intensity, _iter_corpus_rows, _ugc_target_vmaf_offset, _ugc_tight_interval_width, _resolution_dominance, _ugc_saliency_benefit_fraction, and calibrate() integration.
  • Adds ai/tests/test_analyze_knob_sweep_unit.py (29 tests) covering 5 helpers in analyze_knob_sweep.py not exercised by the existing 3-test suite: _stable_knob_repr, _slug, _closest_bare_at_bitrate, write_slice_csv, write_summary_md.
  • All 79 tests run without GPU, corpus, or model downloads (<100 ms total).

Test plan

  • cd ai && python -m pytest tests/test_calibrate_phase_f_recipes_unit.py tests/test_analyze_knob_sweep_unit.py -v — 79 passed, 0 failed
  • Runs cleanly alongside existing test_calibrate_phase_f_recipes.py and test_knob_sweep_analysis.py (83 total)
  • ruff check passes on both new files

ADR-0108 deliverables

  • Research digest: no digest needed: tests only exercise existing pure functions, no new design decisions
  • Decision matrix: no alternatives: only-one-way fix (add missing unit tests)
  • AGENTS.md invariant note: no rebase-sensitive invariants added
  • Reproducer / smoke-test: cd ai && python -m pytest tests/test_calibrate_phase_f_recipes_unit.py tests/test_analyze_knob_sweep_unit.py -v
  • Changelog fragment: changelog.d/added/ai-scripts-coverage-round3.md
  • docs/rebase-notes.md entry: no rebase impact — test-only addition

🤖 Generated with Claude Code

…alyze_knob_sweep helpers (round 3)

The existing test_calibrate_phase_f_recipes.py covers only the main()
invocation (run-provenance smoke).  The existing test_knob_sweep_analysis.py
covers pareto_frontier, stratify, and detect_recipe_regressions via a
20-row synthetic fixture.

This commit adds coverage for the remaining pure helper functions:

* test_calibrate_phase_f_recipes_unit.py (50 tests):
  - mos_to_vmaf_proxy: clamping, boundary MOS values, string coercion
  - saliency_benefit_to_intensity: threshold boundaries for all three labels
  - _iter_corpus_rows: valid rows, blank-line skip, malformed-JSON skip,
    missing-field skip, optional duration_s default
  - _ugc_target_vmaf_offset: empty/small corpus, symmetric distribution,
    heavy-tail sign, clamp bounds, rounding
  - _ugc_tight_interval_width: empty corpus fallback, uniform/wide
    distributions, floor/cap enforcement, rounding
  - _resolution_dominance: empty, single-resolution, split, dominant bucket,
    portrait-vs-landscape distinction
  - _ugc_saliency_benefit_fraction: fallback, no-qualify, all-qualify,
    half-qualify, high-MOS exclusion, square-aspect treatment
  - calibrate() integration: all four recipe classes, UGC/proxy provenance
    tags, required recipe keys, saliency label validity

* test_analyze_knob_sweep_unit.py (29 tests):
  - _stable_knob_repr: empty dict, single entry, alphabetical sort, non-Mapping
    input, numeric values, insertion-order invariance
  - _slug: alphanumeric passthrough, hyphen/underscore preserved, space/slash
    replacement, empty-string fallback, all-special-chars
  - _closest_bare_at_bitrate: no-bare-rows, within tolerance, outside
    tolerance, picks closest, exact match, zero-tolerance
  - write_slice_csv: filename pattern, header row, data rows, slug-safe names,
    auto-creates output directory
  - write_summary_md: file creation, slice count, no-regression message,
    regression table, auto-creates output directory

All 79 tests run without GPU, corpus, or model downloads (<100 ms total).

no rebase impact: test-only addition; no existing golden assertion modified.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@lusoris

lusoris commented Jun 8, 2026

Copy link
Copy Markdown
Contributor Author

Superseded — bundled into PR #845 (chore/bundle-8-drafts-r1) to drain in one CI cycle. Merge SHA: a42e4ea.

@lusoris lusoris closed this Jun 8, 2026
@lusoris
lusoris deleted the test/ai-scripts-coverage-round3 branch June 8, 2026 00:51
lusoris added a commit that referenced this pull request Oct 1, 2026
…DR-1429) (#1753)

* docs(api): state what an index gap and a query before the flush do (ADR-1429)

vmaf_read_pictures() rejects a repeated or earlier index with -EINVAL and
accepts an index that skips values; after such a gap the motion extractors
write no motion2 / motion3 for the later pictures. A score asked for before
the flush returns the value or -EAGAIN. libvmaf.h and docs/api now say both,
and the state ledger records that Netflix/vmaf#910, #755 and #1180 do not
reproduce as wrong values on the fork.

* docs: regenerate the indexes and the citation map after rebasing
lusoris added a commit that referenced this pull request Oct 1, 2026
…23e8f2 (#1761)

* docs: record the fork's check of the upstream defects verified on 6ec23e8f2

Fifteen defects reproduced on Netflix master were run against the fork:
three reproduced and are fixed (#1305, #1420 as a hang, #1613), two are
documented (#910, #755 and #1180), ten are not affected. The dated section
in known-upstream-bugs.md and the Confirmed not-affected rows of the state
ledger carry the evidence; Netflix 8e7a1ac4e (revert of #1476) needs
nothing from the fork.

* docs: regenerate the indexes and the citation map after rebasing
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant