Skip to content

test(coverage): push critical-file coverage above ADR-0922 ratchet floors - #508

Merged
lusoris merged 1 commit into
masterfrom
test/coverage-push-pr469-floors
May 31, 2026
Merged

lusoris merged 1 commit into
masterfrom
test/coverage-push-pr469-floors

Conversation

@lusoris

@lusoris lusoris commented May 31, 2026

Copy link
Copy Markdown
Contributor

Summary

PR #469 ratcheted Coverage Gate floors (ADR-0922): OVERALL_MIN 37→70 %, CRITICAL_MIN 85→90 %, plus per-file floors on dnn_api.c / ort_backend.c (78→83 %) and tiny_extractor_template.h (10→15 %). Also added a per-PR delta gate (no touched file may drop > 0.5pp). The last successful master CI run measured two critical files below the new floor: read_json_model.c at 88.4 % vs new 90 %, and model_loader.c at 86.9 % vs new 90 %.

This PR adds 31 targeted unit tests that push both files above their new floors. No production code touched; no golden assertions or places= tolerances modified.

Type

  • test — test-only

Coverage delta (unit-test-only, container build matching CI flags)

File Before After Floor Delta Status
core/src/read_json_model.c 88.00 % 92.00 % 90 % +4.00 pp PASS
core/src/dnn/model_loader.c 87.20 % 90.00 % 90 % +2.80 pp PASS
core/src/dnn/dnn_api.c 83.40 % 83.40 % 83 % 0 PASS
core/src/dnn/dnn_attach_api.c 92.00 % 92.00 % 90 % 0 PASS
core/src/dnn/onnx_scan.c 93.40 % 93.40 % 90 % 0 PASS
core/src/dnn/op_allowlist.c 100 % 100 % 90 % 0 PASS
core/src/dnn/ort_backend.c 78.90 % 78.90 % 83 % 0 FAIL (*)
core/src/dnn/tensor_io.c 99.30 % 99.30 % 90 % 0 PASS
core/src/dnn/tiny_extractor_template.h 78.50 % 78.50 % 15 % 0 PASS
core/src/opt.c 100 % 100 % 90 % 0 PASS

(*) ort_backend.c remaining uncovered lines split into two structural categories on CPU-only ORT builds:

  1. Non-CPU execution-provider success branches (CUDA / OpenVINO / CoreML / ROCm) — the conditional sess->ep_name = "..." assignments inside the EP-selector switch never fire because each try_append_<ep> returns non-zero on a CPU-only build.
  2. ORT-API failure paths (GetTensorElementType, CreateTensorWithDataAsOrtValue, Run) — require either ORT-API mock injection or actual ORT misbehaviour, neither reachable from a happy-path unit test.

The structural ceiling on the current CPU-only ORT CI runner is approximately 79 %. Pushing above the 83 % floor needs (a) ORT-API mock injection via ort_backend_internal.h, or (b) a working CUDA / OpenVINO EP on the runner — both follow-up scope.

Tests added by file

core/test/test_model.c (+17 read_json_model.c branch tests):

  • 8 type-mismatch -EINVAL branches across parse_model_dict_array_key / parse_model_dict_chroma_correction / parse_model_dict_score_clip (slopes/intercepts/feature_names/feature_opts_dicts/model not array or string, score_clip min/max not number)
  • 3 parse_score_transform_poly branches (JSON_NULL disables, NUMBER sets value, STRING rejects)
  • parse_score_transform_knots_key NULL-disable and array-walk paths
  • parse_score_transform_bool_str non-string reject
  • parse_intercepts mid-loop type-mismatch (after first valid entry)
  • parse_model_dict_entry unrecognised-key skip path

core/test/dnn/test_model_loader.c (+13 model_loader.c branch tests):

  • extract_string_array ADR-0976 leak-free contract for encoder_vocab (third call site after output_names + feature_order)
  • extract_string_array empty-array happy path (line 211 — was unreachable through any existing fixture)
  • extract_string_array -ERANGE over-cap (encoder_vocab with 33 entries vs 32 cap)
  • extract_string_array trailing-junk -EINVAL (non-comma non-] after a valid entry)
  • extract_string_array non-string element -EINVAL (output_names: [42])
  • extract_int -ERANGE (onnx_opset: 99999999999999999999999 → default 0)
  • extract_float_array trailing-junk path via feature_mean: [1.5 Z]
  • 3 quant_mode literal-string branches: "static", "qat" (existing tests covered "dynamic" + unknown-fallback only)
  • kind: "filter" literal-string branch (existing tests covered "fr" + "nr" only)
  • 7 resolve_codec_alias branches: hevc/h265 → libx265, av1 → libsvtav1, vp9 → libvpx-vp9, vvc/h266 → libvvenc, avc → libx264 (existing tests covered h264 → libx264 only)
  • slower preset ordinal (the only preset name not in the existing preset-table grid)
  • vmaf_dnn_codec_block_fill NULL-vocab-entry skip branch

core/test/dnn/test_ort_internals.c (+1 ort_backend.c branch test):

  • vmaf_ort_run multi-output happy-path drives the per-output copy_output_tensor() + ReleaseValue() loops with n_outputs > 1 using the shipped model/tiny/smoke_multi_output_v0.onnx fixture

Reproducer

# Build with CI's exact coverage flags:
docker run --rm --entrypoint /bin/bash \
  -v "$PWD:/wt" -w /wt/core vmaf-dev-mcp:local -lc \
  'rm -rf build-coverage && meson setup build-coverage \
    --buildtype=debug -Db_coverage=true \
    -Denable_cuda=false -Denable_sycl=false \
    -Denable_float=true -Denable_avx512=true \
    -Denable_dnn=enabled \
    -Dc_args=-fprofile-update=atomic \
    -Dcpp_args=-fprofile-update=atomic && \
   ninja -C build-coverage && \
   meson test -C build-coverage --num-processes 1 && \
   pip install --quiet "gcovr>=8.0" && \
   /opt/vmaf-venv/bin/gcovr --root .. \
     --filter "src/.*" --exclude ".*/test/.*" \
     --exclude ".*/tests/.*" --exclude ".*/subprojects/.*" \
     --gcov-ignore-parse-errors=negative_hits.warn \
     --gcov-ignore-parse-errors=suspicious_hits.warn \
     --txt build-coverage/coverage.txt build-coverage'
# Verify the new tests pass + per-file coverage numbers
grep -E "model_loader|read_json_model|dnn_api|ort_backend" \
  core/build-coverage/coverage.txt

Bug-status hygiene

  • no state delta: REASON — pure test-only PR; no bug closed, opened, or ruled out.

Netflix golden-data gate

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.

Known follow-ups

  1. ort_backend.c 78.9 % vs 83 % floor — structural ceiling on CPU-only ORT builds. Requires either ORT-API mock injection (via existing ort_backend_internal.h hook surface) or a working non-CPU EP on the runner. Out of scope for a unit-test-only PR.
  2. OVERALL 47.6 % unit-test-only vs 70 % floor — the master CI's overall measurement (45.6 % at the last successful run, sha 21597ff8) comes from instrumenting both the C unit suite and the full Python test suite under the instrumented libvmaf.so / vmaf CLI. Closing the OVERALL gap by porting ~300 Python tests into native C tests is multi-PR scope. The 70 % floor will likely trip on master itself on the next successful CI run until either the floor is reconciled with the measured value or the Python-suite Coverage Gate path lands additional tests.

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: targeted-tests-only PR following the ratchet floors established by ADR-0922; the gap analysis is in the commit message + this PR body.
  • Decision matrix — no alternatives: only-one-way fix (adding unit tests is the canonical path; the production code is correct and the floors are correct).
  • AGENTS.md invariant note — no rebase-sensitive invariants (test-only additions to existing fork-added test files; no upstream-mirror surface touched).
  • Reproducer / smoke-test command — see "Reproducer" above.
  • CHANGELOG fragment — changelog.d/added/coverage-push-pr469-floors.md.
  • Rebase note — no rebase impact: fork-local test files only; no upstream Netflix path touched.

🤖 Generated with Claude Code

…oors

PR #469 ratcheted the Coverage Gate floors:
  - OVERALL_MIN: 37% -> 70%
  - CRITICAL_MIN: 85% -> 90%
  - PER_FILE_MIN[dnn_api.c]: 78% -> 83%
  - PER_FILE_MIN[ort_backend.c]: 78% -> 83%
  - PER_FILE_MIN[tiny_extractor_template.h]: 10% -> 15%
  + new per-PR delta gate (scripts/ci/coverage-delta-check.sh)

The last successful master CI run (21597ff, 2026-05-31 10:53) reported
read_json_model.c at 88.4% and model_loader.c at 86.9% — both now below
the new 90% critical floor. This PR adds 31 unit tests targeting the
under-floor paths.

After this PR (unit-test-only measurement, container build matches CI
flags: -Db_coverage=true -Denable_float=true -Denable_avx512=true
-Denable_dnn=enabled -Dc_args=-fprofile-update=atomic):

  core/src/read_json_model.c       88.00% -> 92.00%  (+4.00pp)  PASS
  core/src/dnn/model_loader.c      87.20% -> 90.00%  (+2.80pp)  PASS
  core/src/dnn/dnn_api.c           83.40%             (unchanged) PASS
  core/src/dnn/dnn_attach_api.c    92.00%             (unchanged) PASS
  core/src/dnn/ort_backend.c       78.90%             (unchanged) FAIL (*)
  core/src/dnn/onnx_scan.c         93.40%             (unchanged) PASS
  core/src/dnn/op_allowlist.c     100.00%             (unchanged) PASS
  core/src/dnn/tensor_io.c         99.30%             (unchanged) PASS
  core/src/opt.c                  100.00%             (unchanged) PASS

(*) ort_backend.c's remaining uncovered lines are ORT-API error-status
branches and non-CPU execution-provider success branches. On CPU-only
ORT builds (current CI runner) the EP try_append_{cuda,openvino,coreml,
rocm} calls all return non-zero and the conditional `sess->ep_name = ...`
assignments are unreachable; the GetTensorElementType-failure / calloc-
failure branches require ORT-API mock injection. The structural
ceiling on this configuration is ~79%. Pushing above 83% needs either:
  (a) ORT-API mocking via the ort_backend_internal.h test hook, or
  (b) a working CUDA / OpenVINO EP on the CI runner.

New tests by file:

  core/test/test_model.c (+17 tests):
    test_json_model_slopes_not_array
    test_json_model_intercepts_not_array
    test_json_model_feature_names_not_array
    test_json_model_feature_opts_dicts_not_array
    test_json_model_model_payload_not_string
    test_json_model_chroma_correction_not_number
    test_json_model_score_clip_min_not_number
    test_json_model_score_clip_max_not_number
    test_json_model_score_transform_poly_null_disables
    test_json_model_score_transform_knots_null_disables
    test_json_model_score_transform_bool_str_not_string
    test_json_model_unrecognised_model_dict_key
    test_json_model_intercepts_mid_not_number
    test_json_model_score_transform_poly_number_sets_value
    test_json_model_score_transform_poly_string_rejects
    test_json_model_score_transform_knots_array_walks

  core/test/dnn/test_model_loader.c (+13 tests):
    test_sidecar_encoder_vocab_malformed_no_leak
    test_sidecar_empty_arrays_are_valid
    test_sidecar_encoder_vocab_over_max_returns_erange
    test_sidecar_array_trailing_junk_wipes
    test_sidecar_array_non_string_element_wipes
    test_sidecar_opset_overflow_returns_default
    test_sidecar_feature_mean_trailing_junk
    test_sidecar_quant_mode_static
    test_sidecar_quant_mode_qat
    test_sidecar_kind_filter
    test_codec_block_fill_aliases_hevc_av1_vp9_vvc
    test_codec_block_fill_preset_slower
    test_codec_block_fill_null_vocab_entry_is_skipped

  core/test/dnn/test_ort_internals.c (+1 test):
    test_ort_run_multi_output_smoke

Verification:
  meson test -C core/build-coverage --num-processes 1
  -> 85/85 PASS (including all 31 new tests)

Each new test exercises at least one previously-uncovered branch
identified from the CI master-tip gcovr txt summary (run 26710610282,
sha 21597ff). No production code changed; no golden assertions or
places= tolerances touched.

This PR does NOT close the OVERALL 70% floor (current unit-test-only
measurement at 47.6%). CI's overall measurement is higher because the
Coverage Gate job also runs the full Python test suite under the
instrumented libvmaf.so / vmaf CLI. Pushing the unit-test-only overall
above 70% would require porting ~300 Python tests into native C tests
— out of scope for a single PR. The last successful master CI showed
overall at 45.6% — the new 70% floor will trip on master itself until
the per-file ratchet's measurement assumptions are reconciled.

Closes part of the PR #469 ratchet gap; ort_backend.c + overall remain
follow-ups.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@lusoris
lusoris marked this pull request as ready for review May 31, 2026 22:06
@lusoris
lusoris merged commit 705077b into master May 31, 2026
76 of 99 checks passed
@lusoris
lusoris deleted the test/coverage-push-pr469-floors branch May 31, 2026 22:06
@lusoris lusoris added this to the 1.0.0 — First release milestone Sep 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant