Repository navigation
fix+chore: 54-fix mega-bundle from round-4 audits (UBSan, ASan, ABI, MCP, vmaftune, copyright, error paths, headers, magic numbers, docs) - #858
Merged
Conversation
1. pdjson.c/h: change stack_top from size_t to ptrdiff_t so that the -1 sentinel is a well-defined signed value; update all (size_t)-1 comparisons to -1 and add casts on the depth/size comparisons to keep sign-clean arithmetic. 2. motion_avx512.c (×3 scalar tails): cast uint16_t filter[] operands to uint32_t before accumulation; the sum of all five taps at max pixel values (≈65536×65535) overflows signed int in the C default arithmetic promotions. 3. vif_avx512.c (all 4 sites): cast loop counter i to int before subtracting fwidth_half; unsigned − signed promotes to unsigned and the subsequent signed-int assignment is implementation-defined when the result wraps. 4. adm_avx2.c (10 sites): replace the unsigned hex literal 0xFFFFFFFFFFFFFFFF passed to _mm256_set_epi64x() with -1LL; the hex form overflows long long and is UB per C99 §6.4.4.1. 5. integer_adm.c (both init loops): cast (1u << (shift_flt[idx] - 1)) to int32_t before assigning to int32_t add_bef_shift_flt[]; the 1u<<31 wrap is intentional per ADR-0155 (Netflix#955) and is now an explicit implementation-defined conversion rather than UB. NOLINT annotations cite ADR-0155 inline. Build: clean (1010/1010 targets). Tests: 87/87 fast suite pass. No Netflix golden assertion values changed. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
ASan/LeakSan dangling-pointer UAF fixes across five init/create/destroy functions in core/src/: - vmaf_init (libvmaf.c): *vmaf is now NULLed on every failure path so the caller cannot read a freed pointer after a failed init. - vmaf_feature_extractor_context_create (feature_extractor.c/.cpp): *fex_ctx NULLed on free_x/free_f labels and on the inline vmaf_fex_ctx_parse_options error path. - vmaf_fex_ctx_pool_create (feature_extractor.c/.cpp): *pool NULLed on all failure paths; feature_extractor.c also gains a pthread_mutex_init return-value check (the .cpp already had it) and a free_fex_list label to match the new guard. - vmaf_fex_ctx_pool_destroy (feature_extractor.c/.cpp): pthread_mutex_unlock + pthread_mutex_destroy now called before free(pool) per POSIX; freeing a locked mutex is UB and leaks glibc TSD resources. - vmaf_feature_collector_init + feature_vector_init (feature_collector.c/ .cpp): *feature_collector and *feature_vector NULLed on all failure paths. All changes follow CERT MEM30-C. Local verify: meson test --suite=fast 87/87 pass. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Three Python ABI / missing-script findings resolved: 1. ai/pyproject.toml: promote torchvision from a comment-explained implicit dep to a pinned direct dependency (>=0.27.0,<0.28.0). pytorch-lightning >= torchmetrics 1.9+ eagerly imports torchvision.transforms at module-load time; a stale torchvision 0.26.0 wheel against torch 2.12.0 raises RuntimeError: operator torchvision::nms does not exist (not an ImportError), so pip's constraint resolution was the only reliable preventive fix. 2. ai/scripts/export_tiny_models.py: wrap the vmaf_train.models import (which triggers the pytorch_lightning -> torchvision chain) in a broad try/except so an ABI-mismatched venv produces a clear error message with the pip fix command instead of an opaque RuntimeError. 3. dev/Containerfile: add an explicit pip install torchvision>=0.27.0,<0.28.0 step after the ai/ package install so freshly built container images never carry a stale torchvision wheel from a previous layer cache. 4. ai/scripts/: add 13 stub scripts that are referenced in docs/ADRs but were absent from the filesystem. Each stub exits 0 with a short "not yet implemented" message and a pointer to the relevant doc. Stubs: build_calibration_set.py, eval_loso_fr_regressor_v2.py, external_benchmark_pvmaf.py, fetch_lsvq.py, gen_calibration.py, gen_dists_sq_placeholder_onnx.py, gen_mobilesal_placeholder_onnx.py, gen_ssimulacra2_eotf_lut.py, hdrsdr_vqa_to_corpus_jsonl.py, my_corpus_to_corpus_jsonl.py, quantize_int8.py, train_fr_regressor_v4.py, train_video_saliency_student.py. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…um, HTTP TypeError, n-cap, KeyError)
1. _nan_to_none: add depth cap (100 levels) via _nan_to_none_depth helper to
prevent RecursionError on deeply nested JSON payloads from large vmaf runs.
2. _pick_worst_frames: wrap int(idx) in try/except (TypeError, ValueError) so
non-numeric or dict frameNum values are logged and skipped instead of
propagating and aborting describe_worst_frames.
3. http_transport._handle_score: add TypeError to the (ValueError, FileNotFoundError)
catch so int(None) / int([...]) on non-integer width/height/bitdepth fields
returns 400 instead of 500.
4. _call_tool describe_worst_frames: enforce schema maximum:32 on n server-side
(raises ValueError for n < 1 or n > 32), not just in the JSON Schema hint.
5. _call_tool: extract dispatch into _call_tool_dispatch and wrap with
KeyError -> ValueError conversion so missing required arguments produce a
readable error message ("tool X missing required argument: 'ref'") instead
of a bare KeyError.
22 new tests in test_mcp_hardening_wave1.py; no pre-existing test regressions.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…ght, saliency guard, profile source, test contract
Fix 1 (proxy.py): fr_regressor_v2.onnx was exported with two separate named
inputs ("features" [N,6] + "codec" [N,14]) matching FRRegressor.forward().
run_proxy was concatenating them into one 20-D tensor and feeding it as a
single input, so the codec port received nothing and fast-path production
mode produced wrong predictions. Now wires the two inputs separately for
two-input graphs; falls back to the legacy single-input path for older exports.
Fix 2 (saliency.py): compute_saliency_map crashed at runtime for any height
not divisible by 8 because the saliency_student_v1 encoder path requires
aligned tensor dims. Added an upfront ValueError with a clear hint showing
the next valid height rather than surfacing a cryptic onnxruntime error.
Fix 3 (cli.py): _run_recommend_saliency always invoked saliency_aware_encode
even when --saliency-aware was not set, because config=None caused the
function to silently create a default SaliencyConfig() and run the model.
Added an explicit guard: when saliency_aware is False, call run_encode directly.
Fix 4 (encoder_profile.py): build_encode_request raised AttributeError when
the profile "source" field was stored as a plain path string (written by
older vmaf-tune versions) instead of a metadata dict. Now normalises the
field to {"path": value} before accessing .get("path").
Fix 5 (test_fast.py): test_proxy_module_uses_lazy_import_seam was validating
the broken single-input 20-D contract. Updated to use a two-input fake session
(named "features" + "codec") and assert that run_proxy wires the inputs
separately, confirming the corrected behaviour from Fix 1.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…nal files Three categories fixed: 1. ffmpeg-patches/0007 and 0008: remove "and Claude (Anthropic)" from 3 copyright lines; per project_copyright_lusoris_only.md, Anthropic is not a rights holder — Lusoris-only attribution required. 2. 66 fork-original SIMD files (AVX2/AVX-512/NEON in core/src/feature/x86/ and core/src/feature/arm64/): add "Copyright 2026 Lusoris" as a second copyright line immediately after the existing Netflix line (dual notice; Netflix line preserved as these files may include upstream-derived code). 3. 234 fork-original C/H files: replace wrong SPDX identifier "BSD-3-Clause-Plus-Patent" with correct "BSD-2-Clause-Patent" to match the LICENSE root (BSD+Patent / SPDX: BSD-2-Clause-Patent). Verified: scripts/ci/check-copyright.sh exits 0 on all changed files; grep for BSD-3-Clause-Plus-Patent and "Anthropic" in copyright positions returns 0 hits. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
1. vmaf-tune _write_compare_profile_report: write JSON artifact when format='both' (previously only .html + .md were written, silently dropping the .json sidecar). Regression tests in test_format_both_json.py now pass. 2. MCP HTTP test fixture scoping: token_client fixture in test_http_transport.py leaked the removal of VMAFX_MCP_HTTP_NO_AUTH because os.environ.pop() was called inside a patch.dict that did not track that key, so the pop was not reverted on fixture teardown. Replaced with an explicit save/restore approach that prevents cross-test env contamination. 3. test_vmaf_version_handles_version_timeout: _vmaf_version is an async coroutine; the test now uses @pytest.mark.asyncio so pytest-asyncio drives the event loop instead of calling the coroutine object synchronously (Python 3.14+ would raise on await of a non-awaitable). 4. compat/python-vmaf/tools/misc.py: SourceFileLoader.load_module() is deprecated since Python 3.4 and scheduled for removal in Python 3.15; imp.load_source() was removed in Python 3.12. Migrated to the modern importlib.util.spec_from_file_location / exec_module path that works on all supported Python versions (3.8+). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
1. cambi.c set_contrast_arrays: free partial allocs on OOM and propagate error at call site instead of silently ignoring -ENOMEM. 2. integer_motion.c flush: capture and propagate both vmaf_feature_collector_append_with_dict return values. 3. cuda/integer_motion_v2_cuda.c flush: capture and propagate vmaf_feature_collector_append return value. 4. sycl/integer_vif_sycl.cpp flush_fex_sycl: capture and propagate vmaf_sycl_queue_wait return value; close_fex_sycl (void)-casts it as the teardown path must continue regardless. 5. sycl/integer_adm_sycl.cpp: same pattern as integer_vif_sycl.cpp. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…h guard, fr_regressor_v3 codec gate Finding #1 (vmaf_phone_v0.6.1 absent): confirmed false positive — phone mode is a score-transform on vmaf_v0.6.1.json, not a separate file. Docs and code already correct. Finding #2 (vmaf_b_v0.6.3 spurious ERROR before fallback): - model.c vmaf_model_load_from_path: demote "could not read model from path" from VMAF_LOG_LEVEL_ERROR to VMAF_LOG_LEVEL_WARNING. The CLI falls back to the collection loader when this call fails, so a bootstrap/collection JSON is not an error — it has a different top-level structure. The .pkl hard-error follow-up stays at ERROR since pkl is permanently unsupported. Finding #3 (--tiny-model silently rejects .json path with -EBADMSG): - configure_tiny_model (vmaf.c): add early extension check before vmaf_use_tiny_model. When the path does not end in ".onnx", emit a clear diagnostic explaining that sidecar .json files are loaded automatically alongside the .onnx, not passed directly. Previously the JSON bytes were scanned as protobuf by onnx_scan.c, producing the opaque -EBADMSG. Finding #4 (fr_regressor_v3 out-of-range scores without --tiny-codec): - dnn.h: add vmaf_dnn_is_codec_aware(ctx) public API. - dnn_ctx.h: add vmaf_ctx_dnn_is_codec_aware bridge declaration. - libvmaf.c: implement vmaf_ctx_dnn_is_codec_aware (checks sess, has_sidecar, codec_aware flag, and extra_in_width > 0). - dnn_attach_api.c: implement vmaf_dnn_is_codec_aware public wrapper. - configure_tiny_model (vmaf.c): after model load, if the model is codec-aware but no --tiny-codec / --tiny-preset / --tiny-crf was given, reject with a clear error message explaining that the conditioning block would contain only an "unknown" fallback slot. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
1. picture_v2.h: rename include guard from reserved C identifier __VMAF_PICTURE_V2_H__ (double-underscore prefix, undefined behaviour per ISO C 7.1.3) to LIBVMAF_PICTURE_V2_H_. 2. All public headers: convert bare quoted includes (#include "foo.h" / #include "libvmaf/foo.h") to angle-bracket form (#include <libvmaf/foo.h>) across all 9 affected installed headers (libvmaf.h, picture.h, picture_v2.h, feature.h, model.h, dnn.h, libvmaf_cuda.h, libvmaf_hip.h, libvmaf_metal.h, libvmaf_mcp.h, libvmaf_sycl.h). Quoted-path includes only resolve when the build root is on the include path, breaking pkg-config consumers who only have the installed prefix. 3. picture.h VmafPicture: add INTERNAL banners to the ref and priv fields, clarifying they are managed by libvmaf and must not be accessed externally. 4. libvmaf.h VMAF_POOL_METHOD_NB: add __attribute__((deprecated)) on GCC/Clang so external callers see a build-time warning. Gate the attribute on !VMAF_BUILDING_LIBVMAF so internal TUs (output.c) that legitimately iterate [0, NB) are not affected. Inject -DVMAF_BUILDING_LIBVMAF into vmaf_cflags_common in core/src/meson.build. 5. vmaf_assert.h: remove from the install_headers() list in core/include/libvmaf/meson.build. The header exposes VMAF_ASSERT_DEBUG which is gated on the internal VMAF_DEBUG build flag and has no defined semantics for external consumers. Internal .c files continue to include it from the source tree. Add an INTERNAL comment banner to the header itself. Build-verified: meson setup + ninja (1010/1010 targets clean) + meson test --suite=fast (87/87 pass, 0 failures). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- CAMBI_WINDOW_DIVISOR=375 and CAMBI_MIN_WIDTH_HEIGHT=216 centralised in cambi_internal.h; local per-backend defines (CAMBI_CUDA_MIN_WIDTH_HEIGHT, CAMBI_HIP_MIN_WIDTH_HEIGHT) and the bare 375 divisor removed from cambi.c, integer_cambi_cuda.c, and integer_cambi_hip.c. - DNN_SIDECAR_JSON_MAX=1u<<20 added to model_loader.h; three guard sites in model_loader.c now reference it instead of bare bit-shifts / integer literals. - FEATURE_VECTOR_INITIAL_CAPACITY=8u added to feature_collector.h; three literal 8s in feature_collector.c replaced. - DNN_MIN_BIT_DEPTH=9 added to tensor_io.h; bpc guard in tensor_io.c (two sites) and dnn_api.c now reference it. Build: 1010/1010 ninja targets; 87/87 fast tests pass; pre-commit clean. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…ote + mark vmafx CLI as planned + remove active Vulkan claims Four targeted fixes from the round-4 docs-substance audit: 1. core/include/libvmaf/libvmaf_hip.h — remove the stale "Status: scaffold only. Every entry point returns -ENOSYS" Doxygen block. The HIP backend is fully implemented (ADR-0519 / ADR-0533 / ADR-0539); 21 feature extractors are registered and verified on AMD gfx hardware. Replace with an accurate status note pointing to the three unregistered legacy stubs and the no-HIP stubs.c contract. 2. docs/api/gpu.md — complete the vmaf_hip_import_state table entry: add the -ENOSYS return when built without HIP (matches the header Doxygen and stubs.c behaviour) alongside the already-documented -EINVAL case. 3. docs/usage/vmafx-cli.md — mark the vmafx symlink, --netflix-compat flag, and vmafx-* Python aliases as "planned — not yet implemented in master". Neither cli_parse.c nor the Python pyproject.toml entries have been updated yet (ADR-0690 / ADR-0696 specify the design). Add interim equivalents using the already-shipped --precision=max flag. 4. docs/usage/vmaf-tune.md — remove active Vulkan claims. The Vulkan backend was deleted in ADR-0726; --score-backend=vulkan no longer exists. Mark the Vulkan section REMOVED with a historical-reference notice; update six flag-table rows to drop vulkan from the accepted enum; fix the native-first-order example (vulkan was already absent from the auto probe order table at line 393). No code changes; docs only. No golden assertions touched. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
|
|
||
| import asyncio | ||
| import math | ||
| from pathlib import Path |
| from __future__ import annotations | ||
|
|
||
| import asyncio | ||
| import math |
|
|
||
| from __future__ import annotations | ||
|
|
||
| import asyncio |
11 tasks done
lusoris
added a commit
that referenced
this pull request
Jun 12, 2026
…RF, AI scripts, MCP conformance) (#865) * fix(hip,sycl): stale wave32 comment + Kahan IIR blur for ssimulacra2 (iter6-cross-backend-parity) Three iter6 cross-backend-parity findings: 1. [critical — partial] float_adm_score.hip: wave32 code was already fixed in #859 (0a9dba8); correct the lingering stale comment that still read "FADM_WARPS_PER_BLOCK = 4 (64-lane warps)" — FADM_WARPS_PER_BLOCK is now 256/32 = 8 slots (sized for the wave32 worst case). No functional change. 2. [high — already fixed] float_ssim/ssim_score.hip + integer_psnr/psnr_score.hip: wave32 runtime-warpSize fixes landed in #859 alongside float_adm; nothing further to do here. 3. [high] ssimulacra2_sycl.cpp: add Kahan (compensated-summation) state tracking to the 3-pole recursive IIR blur kernel (launch_blur<PASS>). The IIR state (prev1_k) accumulated O(eps) rounding error per step; over 4K-tall frames this exceeded the 5e-5 cross-backend parity contract vs the CPU reference. Each pole now carries a float comp_k compensation term; the standard Kahan pattern (y = candidate - comp; new_state = old_state + y; comp = (new_state - old_state) - y) bounds per-iteration error to O(eps^2) without fp64 (ADR-0220 compliant). No Netflix golden-data assertions modified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ubsan): replace 0xFFFFFFFFFFFFFFFF hex literals with -1LL in adm_avx512.c _mm512_set_epi64 takes long long (signed 64-bit) arguments. The literal 0xFFFFFFFFFFFFFFFF exceeds LLONG_MAX and is undefined behaviour under strict UBSan. Replace with -1LL which has identical bit pattern and correct type at both call sites (ADM_CM_THRESH_S_I_END macro lines 406 and 698). Finding: iter6-ubsan-strict high [adm_avx512: 0xFFFF... overflows long long]. Note: motion_avx512.c and adm_avx2.c analogous fixes were already present on master (PR #858 / PR #859 bundles). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(thread-safety): eliminate 2 TSAN races in threaded batch path (iter6-tsan-race-deep) Finding #1 (high): Remove the racy write `fex->framesync = vmaf->framesync` in `threaded_enqueue_one`. `fex` is the *shared* registered VmafFeatureExtractor — worker threads from previous frames may concurrently read its fields. The write is redundant: framesync is already propagated to every pool-slot copy by `set_fex_framesync()` at registration time and by `ctx_pool_ensure_slot_ctx()`. Finding #2 (high): Move the `vmaf->prev_ref` advance to BEFORE the enqueue call in `threaded_read_pictures_batch`. In the old order the main thread unreffed and replaced `vmaf->prev_ref` after enqueue while the just-submitted worker still held a live reference to the same underlying VmafRef*, creating a concurrent unref/write on the same object without synchronisation. Workers use `data.prev_ref` (an independently refcounted snapshot) exclusively and never re-read `vmaf->prev_ref` after enqueue, so moving the advance before enqueue is both safe and race-free. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(vmaf-tune): fix --no-bisect TypeError and recommend CRF selection strategy Two bugs in the compare/recommend paths: 1. _encode_and_score() in bisect.py had encode_runner and score_runner as required keyword-only args (no defaults). The CRF-sweep caller in cli.py did not pass them, causing an unconditional TypeError on any compare --no-bisect --crf-sweep invocation. Fixed by giving both parameters a default of None, consistent with the existing decode_runner parameter and the run_encode/run_score runner=None semantics. 2. _smallest_passing_crf() in cli.py used `crf > cur[0]` to select the largest (most efficient) passing CRF, contradicting both the function name and the CLI help string ("find the smallest CRF whose VMAF >= --target-vmaf"). The correct strategy is the smallest (highest-quality) passing CRF. Fixed comparison to `crf < cur[0]` and updated the docstring to match. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ai-scripts): 3 iter6 runtime bugs — vulkan-device AttributeError, saliency parity always-fail, ensemble ONNX codec-dim mismatch - collect_gpu_calibration_data.py: remove dead args.vulkan_device reference from devices dict (Vulkan removed per ADR-0726; no --vulkan-device argparse registration existed, causing AttributeError at runtime) - validate_saliency_student.py: replace broken PT-reconstruction parity check with ORT-only sanity check when no PT state provided. do_constant_folding=True folds BN stats into conv weights at export so ONNX initializer names diverge from PT state_dict keys; the old code silently left 60 of 65 weights at random defaults and always failed. When pt_state is provided (trainer path) full PT<->ORT diff is still performed. - model/tiny/fr_regressor_v2_ensemble_v1_seed{0..4}.onnx: regenerate with codec_onehot=[batch,6] matching current CODEC_VOCAB (was [batch,14] from a 12-entry encoder_vocab + 2 norm dims; CODEC_VOCAB was later trimmed to 6). Regenerated via train_fr_regressor_v2_ensemble.py --smoke in vmaf-dev-mcp container. eval_probabilistic_proxy.py --smoke now passes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(mcp): JSON-RPC parse-error conformance + deep-nesting guard Two iter6 fuzz conformance findings fixed: 1. [high] Parse-error notification instead of error response (id=null): Add `_ParseErrorFilteredStdin` — a lazy async stdin wrapper that pre-validates each incoming line with `JSONRPCMessage.model_validate_json`. On failure it calls `_emit_parse_error` which writes `{"jsonrpc":"2.0","id":<recovered_or_null>,"error":{"code":-32700, "message":"Parse error"}}` to stdout synchronously, then drops the line so the mcp library never sees a bare Exception on the stream (the notification path is bypassed entirely). `_run()` now passes `stdin=_ParseErrorFilteredStdin()` to `stdio_server`. 2. [high] 500-level deep nesting triggers recursion-limit exception: Add `_check_depth(obj, max_depth=50)` helper and call it at the top of `_call_tool_dispatch` before the tool dispatch. Payloads exceeding 50 nesting levels raise `ValueError`, which the mcp library converts to an isError=True tool result before the pydantic parser recurses. Existing tests updated: the two `stdio_server`-patching tests in `test_coverage_round2.py` now accept `**kwargs` so they tolerate the new `stdin=` keyword argument forwarded by `_run()`. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Lusoris <lusoris@pm.me> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
4 tasks done
lusoris
added a commit
that referenced
this pull request
Jun 12, 2026
… fuzz, TSan/OrtEnv, vmaf-tune, AI scripts, MCP batch) (#875) * fix(sycl): promote ADM per-scale normalization intermediates to double On Intel Arc A380 (no native fp64 device), the per-scale normalization in conclude_adm_cm and conclude_adm_csf_den used float f_accum and float *result, causing rounding error that SVM amplified past the 5e-5 final-score threshold (iter10 cross-backend-parity finding [high]). Both functions are host-side (no device-kernel code); the promotion to double has zero impact on fp64-less device kernels. The call-site variables num_scale and den_scale are likewise promoted to double. HIP findings (wavefront_reduce_i64 carry bug and MS_WARP_SIZE=64 on wave32 hardware) were already resolved in the worktree base by PR #850 and are confirmed absent: vif_statistics.hip uses atomicadd_accums, and motion_score.hip uses runtime warpSize throughout. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pool): destroy per-entry condvar in vmaf_fex_ctx_pool_destroy vmaf_fex_ctx_pool_destroy() iterated all fex_list entries and freed ctx_list but never called pthread_cond_destroy on the per-entry condvar (pool->fex_list[i].full) that was initialised in get_fex_list_entry() / ctx_pool_alloc_slot(). POSIX requires destroy before the containing memory is freed; omitting it leaks POSIX TSD resources on glibc and is reported by ASan/LeakSan as a condvar-resource leak. Add pthread_cond_destroy(&pool->fex_list[i].full) in the i-loop after free(pool->fex_list[i].ctx_list) in both the C and C++ translation units (the C++ file is the one compiled by meson; the C file is kept in sync for readability). The pool-mutex unlock+destroy was already present. The libvmaf.c *vmaf=NULL dangling-pointer fix (finding 2) was already applied in a prior commit; no change needed there. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: round-4 audit bundle — UBSan, ASan, ABI, MCP, vmaf-tune, copyright, error paths, headers, magic numbers, docs (#858) Cherry-pick of ebbcca3 with conflict resolution (motion_avx512.c uint32_t casts, Containerfile ai[dev] extras, server.py iterative _nan_to_none, test imports). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(fuzz): cap Y4M frame dimensions to prevent unbounded malloc (T-FUZZ-Y4M-OOM) Add Y4M_MAX_FRAME_PIXELS (64 Mpixels) guard in y4m_input_open_impl, inserted after the existing sign check and before the chroma-format dispatch. Attacker-controlled W/H values from the Y4M header can no longer drive malloc with an unbounded size. The same guard also eliminates the signed-integer overflow path in y4m_convert_411_422jpeg for oversized frames. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tsan): iter10 — destroyed guards + OrtEnv singleton + model-list snapshot Three TSan/UAF fixes identified in iter10-tsan-race-deep: [critical] feature_collector.c: vmaf_feature_collector_set_aggregate, vmaf_feature_collector_get_aggregate, vmaf_feature_collector_mount_model, and vmaf_feature_collector_unmount_model were missing the `destroyed` flag guard that the other entry points (append, get_score, find) already carry. A worker thread racing destroy() could access freed aggregate_vector or models memory after the lock was released by destroy(). Fix: add the standard `if (feature_collector->destroyed) { unlock; return -ENODEV; }` pattern immediately after lock acquire in all four entry points. [high] ort_backend.c: each vmaf_ort_open call created a fresh OrtEnv via sess->api->CreateEnv, spawning new ORT-internal background threads. ORT documents OrtEnv as a process-wide resource; concurrent CreateEnv calls race inside ORT's thread-pool initialisation. Fix: replace per-session OrtEnv with a file-static singleton (g_ort_env) initialised exactly once via pthread_once. The singleton is never released (process lifetime per ORT recommended usage). Remove sess->env field and the ReleaseEnv call from vmaf_ort_close. [high] feature_collector.c: feature_collector_run_model_predict snapshotted only model_iter->next before the lock drop (round-5 fix), but the node that the pre-snapshotted next pointer points to could itself be unmounted and freed by a concurrent vmaf_feature_collector_unmount_model between lock releases. Fix: snapshot the full VmafModel* list into a stack-allocated array of FEATURE_COLLECTOR_MAX_MODELS (32) entries while the lock is held before the first unlock, then iterate the snapshot without re-dereferencing any linked-list pointers. VmafModel lifetime is caller-managed and outlives the predict pass. Build: 88/88 fast tests pass (CPU-only build, enable_dnn=disabled). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(vmaf-tune): correct fast-verify bitrate denominator and report NaN sentinel - _build_fast_encode_runner: replaced wall-clock encode_time_ms denominator with clip duration derived from raw-YUV file size and frame geometry (width × height × bpp / framerate). Using encoder wall-clock time as the denominator inflated or deflated observed_kbps depending on encode speed rather than content duration. - _run_report: changed LadderSample and LadderRung bitrate_kbps/vmaf construction from `or 0.0` (silently coercing null to zero) to `float('nan')` when the JSON key is absent or null. The renderer already gates on _is_missing() / _finite_values(), so NaN propagates as an em-dash gap-marker instead of a misleading 0 kbps / 0 VMAF entry. Fixes iter10-vmaftune-exhaustion findings #1 and #2. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ai-scripts): remove stale Vulkan references post-ADR-0726 - collect_gpu_calibration_data.py: fix --smoke docstring that still said "Vulkan-only"; update to reflect CUDA-only smoke mode (ADR-0726). - Five extraction scripts (extract_k150k_features, konvid_to_full_features, extract_ugc_features, bvi_dvc_to_full_features, konvid_to_vmaf_pairs): remove --no_vulkan from subprocess command lists; the flag no longer exists in the post-ADR-0726 vmaf binary. - cross_backend_parity_gate.py: drop "vulkan" from BACKEND_SUFFIX, BACKEND_DEVICE_FLAG, BACKEND_DEFAULT_DEVICE, BACKEND_EXTRACTOR_ALIASES, --vulkan-device arg, and devices dict; fix stale help text references. - export_transnet_v2_placeholder.py / export_fastdvdnet_pre_placeholder.py: gate _export() call behind the --no-registry guard so --no-registry is a true dry-run that does not attempt writes to the default read-only path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(mcp): emit -32600 for batch requests, -32700 for parse errors on stdio transport JSON-RPC 2.0 §6 requires servers that do not support batch requests to respond with -32600 (Invalid Request), not -32700 (Parse error), when a JSON array arrives on stdin — the payload is syntactically valid, it is the request type that is unsupported. Previously _emit_parse_error always emitted -32700 regardless of whether the incoming line was malformed JSON or a valid-JSON-but- array batch request. The fix detects when the successfully-parsed value is a list and switches the error code and message to -32600 / "Invalid Request" accordingly. All malformed- JSON paths retain -32700. The _ParseErrorFilteredStdin wrapper that calls _emit_parse_error is unchanged; all 407 existing tests continue to pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Lusoris <lusoris@pm.me> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
3 of 7 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Mega-bundle of 54 fixes across 12 themes from the round-4 audit sweep. Each theme was implemented in an isolated worktree by a parallel agent and verified locally before cherry-pick.
Themes (54 fixes):
Test plan
Deep-dive deliverables (ADR-0108)
state.md touch