Skip to content

Software-mode slow capture ends in FFmpeg exit 255 (SIGTERM): kill-reason observability + encoder exit-state propagation request #3744

Description

@aikdna

Summary

On a production-scale, DOM-heavy composition under forced software raster (--no-browser-gpu, SwiftShader), the render ends with Streaming encode failed: FFmpeg exited with code 255 whose stderr tail reads Exiting normally, received signal 15. — twice, ~19.7 min and ~13.6 min end-to-end, no usable output. We have not reproduced this with synthetic compositions, and we do not attribute the terminating signal source. This is an observability/robustness request, not a claimed root-cause report.

Environment (pinned)

component version
hyperframes CLI 0.6.52 (npm, local node_modules/.bin/hyperframes)
Node.js 24.19.0
OS macOS 26.2 (25C56), Apple Silicon
Chrome (puppeteer-launched) 152.0.7977.76 (system Chrome via PUPPETEER_EXECUTABLE_PATH)
ffmpeg 8.1.2

Observed

Composition: single-document assembly (root + 3 bare-id scene wrappers), 1920x1080@30, 19 s = 570 frames, ~60 KB inline DOM. Forced software raster, two attempts, both exit 1:

  • attempt 1: end-to-end ≈19.7 min; last ffmpeg progress line frame=265 fps=0.5 ... elapsed=0:08:13 inside the error's stderr tail; final error Streaming encode failed: FFmpeg exited with code 255, tail ends Exiting normally, received signal 15.
  • attempt 2: end-to-end ≈13.6 min; identical signature.
  • External 10 s process sampling around both attempts: the Chrome browser process was alive in every sample until HyperFrames tore down; the process that received SIGTERM is the ffmpeg streaming-encoder child.
  • GPU default mode on the same page: completes (~89 s incl. a 45 s player-ready wait — a separate topic, not chained causally here; the same warning preceded five successful renders and one no-warning control).

What isolation achieved: a canvas-driven, 570-frame, DOM-light repro under --no-browser-gpu finishes end-to-end in ≈115 s (≈5 s compile + 45 s unrelated wait + ≈65–70 s capture+encode, ~0.12 s/frame). So the abort appears tied to slow per-frame capture of a DOM-heavy page under software raster, not to frame count alone — and we could not shrink it further into a synthetic repro.

Unverified code reading, for triage only (not our diagnosis): in dist/cli.js (0.6.52), spawnStreamingEncoder (~L32025) attaches two SIGTERM paths to the ffmpeg child: an idle watchdog ffmpegStreamingTimeout (default 600 000 ms, env FFMPEG_STREAMING_TIMEOUT_MS) that is reset on writes Node reports as accepted, and the job-level abortSignal handler. Both could produce exactly "signal 15"; we cannot tell which one fired from outside the process. For completeness: the capture loop ignores writeFrame's boolean return — but in Node that return is Writable backpressure semantics (buffer threshold), not encoder liveness, and we deliberately do not propose treating false as fatal.

Asks

  1. Record the kill reason in the surfaced error (which watchdog/abort path fired, frames accepted vs attempted at kill time) so consumers don't have to read the bundle to enumerate suspects.
  2. Propagate encoder-child exit state (an explicit encoder error/dead status), so the capture loop can stop early instead of spending the remaining capture budget against a terminated encoder — failure at ~19.7 min cost us the full render twice.
  3. Optional guidance: is there a recommended producer configuration for pages whose software-mode capture rate is ≤1 fps (e.g. the enableChunkedEncode path that ships disabled by default)?

Package hygiene: redacted log excerpts available on request (consumer wrapper lines removed; scene ids replaced by host-a/b/c; terminal control sequences stripped, verified by scan with 0 remaining control bytes; no consumer-project paths or internal designations).

Activity

  1. Ha1baraA11 commented on Sep 7, 2026

    @Ha1baraA11

    This sounds like the renderer is treating the FFmpeg termination as a generic capture failure, so the useful part of the error gets lost.

    Before looking at the encoder path, could you share the exact command and the FFmpeg stderr from one failing run? The exit code would help too, especially since the issue mentions SIGTERM/255.

    Also, does the same input fail in both software and hardware capture modes? If it only happens in software mode, that should narrow this down to the capture/encoder lifecycle rather than the composition itself.

  2. jrusso1020 commented on Sep 7, 2026

    @jrusso1020
    Collaborator

    @aikdna You are on quite an old version of HyperFrames and I think we have since fixed quite a few issues around this area of the codebase

    Would you mind updating and seeing if this issue still occurs?

  3. jrusso1020 commented on Sep 7, 2026

    @jrusso1020
    Collaborator

    @aikdna Following up with the specific changes behind the upgrade request: #1372, included in v0.6.95 and later, changed streaming writes to wait for backpressure to drain, refresh the inactivity timer after drain, and stop capture when an encoder write fails. In v0.6.52, a buffered write did not refresh the timer when it later drained, and capture could continue after the encoder had exited.

    I checked unmodified current main / v0.8.31: the relevant tests pass, and a real FFmpeg run with large frames completed beyond a shortened inactivity budget while frames continued arriving. That validates the current mechanism; it does not reproduce your DOM-heavy composition or establish who sent SIGTERM in the original runs.

    Please retry on the latest release and confirm the version reported by the same CLI binary your pipeline invokes. If it still fails, please share the command, stderr tail, and a minimal reproducer if possible. The fuller kill-reason/frame-count diagnostics are not all present today; let's first establish whether the render failure survives the upgrade.

  4. aikdna commented on Sep 8, 2026

    @aikdna
    Author

    Retested on the current release as requested — hyperframes@0.8.31 (self-reported
    via --version from the same CLI binary we rendered with; Node 24.19.0, system
    Chrome via PUPPETEER_EXECUTABLE_PATH, macOS 26.2, ffmpeg 8.1.2). Summary:
    the failure no longer reproduces on 0.8.31, and we now have a synthetic page
    that does reproduce the original signature on 0.6.52.
    That is exactly the
    "count as fixed on our side" condition we proposed, so thank you — details below.

    Background on our earlier gap: the compositions in the original package only hit
    ~2 fps software-capture, which was fast enough to stay out of the failure regime
    (it passed on 0.6.52 too). Building on your note that #1372 made streaming
    writes wait for back-pressure drain (refreshing the inactivity timer after drain,
    and stopping capture when an encoder write fails), we scaled the page until
    software capture dropped into the production regime (~0.2 fps).

    Isomorphic synthetic (same registration shape as our production page: root
    timeline registered + 3 bare-id section wrappers that never register; 1920x1080
    @30; raw requestAnimationFrame driver; 324 backdrop-filter + large-shadow
    cards plus a blurred blended overlay; 3 s/90-frame calibration variant and full
    19 s/570-frame variant). Command identical in all cells:

    hyperframes render --format mp4 --quality standard --fps 30 --workers 1 --no-browser-gpu --output out.mp4
    
    page 0.6.52 0.8.31
    90f software fails: rc=1, e2e 626 s; capture iterates all 90 frames while ffmpeg's last progress line reads frame=84 fps=0.2 ... elapsed=0:08:00; error Streaming encode failed: FFmpeg exited with code 255, stderr tail ends Exiting normally, received signal 15. — same signature as our two production attempts passes: rc=0, e2e 113 s, ffprobe 90 frames / 3.000 s
    570f software (= production duration) (not rerun — the 90-frame failure is duration-independent; the kill tracks the last accepted write, not frame count) passes: rc=0, e2e 423 s, ffprobe 570 frames / 19.000 s, blackdetect finds no black intervals
    GPU (default) mode passes (rc=0, 64 s) passes (rc=0, 64 s)

    Two observations that line up with #1372:

    1. On 0.6.52, capture kept running past the encoder's death (90 frames iterated
      against 84 ever accepted) — the "buffered writes don't refresh the timer on
      later drain" behavior you described. On 0.8.31 the drain-aware write path
      both keeps the encoder alive and paces the capture: wall time for the same
      heavy page dropped from a stall/kill at ~10 min to a clean ~2 min finish.
    2. The data-no-timeline lesson from Bare data-composition-id hosts (static sections) burn the full player-ready budget on every render — contract question + API improvement request #3743 also applies here: on 0.8.31 the
      same page still pays the 45 s Topic-1 wait (we left hosts unmarked on
      purpose, to isolate this variable) but finishes correctly.

    As asked, the failure artifacts from the 0.6.52 cell (exact command above, full
    stderr tail incl. the libx264 exit block, exit code 1) are ready to share; say
    the word if a minimal-repro HTML for the passing/failing pair would still be
    useful for a regression fixture. The kill-reason diagnostics from our original
    ask #1 are no longer needed for us to diagnose this — with back-pressure handled,
    we'd deprioritize it to nice-to-have.

    From our side this can be closed as fixed by #1372 / present since v0.6.95.
    Thank you for tracking down the exact change — it saved us a long hunt, and the
    care you put into the answer shows.

  5. miga-heygen commented on Sep 8, 2026

    @miga-heygen
    Contributor

    Not reproducible on v0.8.30+ — underlying issue (render timeout SIGTERM) was fixed in earlier releases. Reporter's stale 0.6.52 is no longer supported.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions