Skip to content

Merge dev and finish offline strokes in the normal stroke path - #1

Merged
daghack merged 11 commits into
daghack:offline-stroke-pathfrom
darkly-art:resolve-conflicts
Oct 6, 2026
Merged

daghack merged 11 commits into
daghack:offline-stroke-pathfrom
darkly-art:resolve-conflicts

Conversation

@TheTechromancer

@TheTechromancer TheTechromancer commented Oct 6, 2026 •

Copy link
Copy Markdown

Hi Talon, thanks for this. This PR is based on your profiling/benchmark, and I kept your original commit.

Basically while your PR was open, dev changed how strokes render. So I resolved the conflicts and integrated your feature directly into the normal stroke path. That includes the batching idea which is adapted to the new stabilizer. Your recorded stroke now takes 26.9 ms through plain begin_stroke / stroke_to / end_stroke, against 22.2 ms for stroke_path on the same machine. So stroke_path itself is gone: just keep calling that loop, with no render between samples.

For offline work, set input.predictionHorizon to 0, or strokes overshoot their last sample by a few pixels.

Merging this updates darkly-art#135, which we'll then review and merge into dev.


For your agent

This PR merges darkly-art/darkly dev (8e2bef0, including darkly-art#134) into offline-stroke-path, then adds one commit. Net change against dev, relative to 3f99aea:

Removed (superseded by dev's design):

  • DarklyEngine::stroke_path, BrushRender, StrokeOp::pen_sample, StrokeEngine::stabilize_all, and the stroke_path_* tests. After Stabilization Improvements darkly-art/darkly#134, stroke_to only records the sample and flush_stroke runs once per frame and once from end_stroke, so a stroke with no render between samples is stabilized and rendered once, at pen-up.
  • The from-scratch LaplacianStabilizer::push_all. Dev's Laplacian is windowed (edges trajectories, chord_t, pull), so the override no longer compiles or applies.

Kept, adapted to dev:

  • Dab-cap phase split in StrokeEngine::place_dab via BrushGpuContext::submit_and_continue. A headless stroke's single pen-up phase reaches MAX_DABS_PER_PHASE, so this is load-bearing. Tests: tests/preview_dab_cap.rs, headless_stroke_splits_phases_at_the_dab_cap.
  • DarklyEngine::set_stroke_seed(Option<u32>). Test: stroke_seed_pins_random_nodes.
  • StrokeEngine::render_whole, used by the brush preview.
  • The bench comparison, as stroke_replay_bench --no-frames (a flush per event against no frames between events).

Added:

  • flush_stroke(final_flush: bool): the pen-up flush saves no checkpoints and renders its range as one segment. A frameless stroke is 3 submissions (prologue, segment, commit) plus one per MAX_DABS_PER_PHASE dabs, at any stabilizer strength.
  • Batched stabilization: StabilizerAlgorithm::push_all is now the required method (push wraps it). StrokeEngine::take_divergence hands the stack every sample recorded since the last flush in one call. The resampler retracts its provisional tip once and pushes once. Prediction rebuilds once. The Laplacian relaxes a batch of M vertices once, at N x (M + N) updates, bit-identical to sequential pushes. MAX_COMMITS_PER_PUSH is renamed MAX_COMMITS_PER_SAMPLE.
  • Default input.predictionHorizon lowered from 15 to 10 ms. Prediction tails are still committed at pen-up, a known open issue, so pin it to 0 offline with darkly::config::set("input.predictionHorizon", ConfigValue::Float(0.0)).

Usage for offline rendering:

engine.set_stroke_seed(Some(seed));             // reproducible random nodes
engine.begin_stroke(layer)?;
for op in ops { engine.stroke_to(op); }          // no engine.render() in between
engine.end_stroke();                             // stabilizes and renders once

Measured (stroke_replay_bench --no-frames, recorded_curvy_stroke.json, 1024x512, release, prediction off; the no-frames figure includes a pen-up composite):

Brush flush per event no frames 3f99aea stroke_path
Ink Pen (0.6) 194 ms, 1290 submits 26.9 ms, 3 submits 22.2 ms
Calligraphy 230 ms 54.1 ms 49.0 ms
Rough Ink 190 ms 34.3 ms 33.9 ms

Regression tests, each shown failing before its fix: pen_up_flush_saves_no_checkpoints, live_pen_up_flush_renders_one_segment, a_batch_costs_one_relaxation, headless_stroke_relaxes_once, plus the dab-cap tests above. Equivalence: headless_stroke_matches_frame_by_frame, batches_are_bit_exact_with_from_scratch, a_batch_reaches_the_inner_as_one_push, batched_stack_matches_sequential.

TheTechromancer and others added 11 commits October 4, 2026 02:34
The Laplacian's pull per vertex is now the pen's speed past it over a
reference speed (500 CSS px/s), capped at 1. Fast motion is smoothed in
full, a slow pivot keeps its corner, and a stopped pen is exact. The
registry's `from_params` receives the stroke's view scale so the reference
speed is in CSS pixels at any zoom.

Two rendering defects fixed along the way:

- The divergence walk stopped at the first vertex within tolerance of its
  rendered position, assuming everything behind it was unchanged. Against
  rendered positions that assumption fails, because vertices are
  re-rendered at different moments: a freshly rendered vertex ended the
  walk while vertices behind it drifted unchecked, then snapped back later
  as a visible disconnect on quick wide curves. `find_divergence` now scans
  the whole window for the earliest vertex out of tolerance.
- `DIVERGENCE_EPSILON` drops from 0.5 to 0.1 CSS px. Neighbouring vertices
  rendered at different moments could each sit half a pixel off, giving the
  rendered stroke a 0.57 px ripple the smoothed polyline did not have.

Regression tests: slow_corner_keeps_its_shape,
moved_vertex_behind_an_unchanged_one_is_reported,
rendered_positions_stay_within_epsilon.
The Sponge lagged at full size and stabilization because every pen
sample ran a cost that grew with the stroke, several times per frame.

The Laplacian relaxed the whole polyline from its raw points on every
push, at 160 sweeps over a vertex every 6 CSS px. A push now relaxes
only the N + 2 vertices its influence bound says can move. The vertex
left of the window replays the per-sweep values recorded when it was
an earlier window's left edge, so the result is bit-identical to the
from-scratch relaxation at N x (N + 1) updates per push. Prediction
likewise copies only the inner polyline's changed tail.

Every coalesced pointer sample also ran the full diff, rewind, replay
and commit. stroke_to now only feeds the stabilizer, and flush_stroke
runs that cycle once per frame, from render and end_stroke.

- Divergence tracking moves from the stabilizer algorithms to the
  stroke engine (DivergenceDiff), whose walk starts from the tip as
  last rendered, so it holds however many events arrive between
  frames. The resampler's per-event window widening goes with it.
- The rewind and append paths share one segment loop. The unreachable
  no-stroke-buffer fallback and StrokeEngine::move_to are deleted.
- compute_segment_boundaries keeps the vi = 0 anchor when a stroke's
  first frame renders a single vertex.
- The [frame-perf] log gains a stroke sub-phase.
- The rewind counter is renamed test_stroke_rewinds.

The stroke_rewind.rs oracle tests had been comparing a run with
itself since unstabilized strokes stopped rewinding. They now paint
stabilized with a zero divergence epsilon (test_set_divergence_epsilon)
and assert that rewinds happened. Test drivers that relied on
per-sample rendering now render a frame per sample.

stroke_replay_bench now runs on bench_device, so 1080p cells fit.
The browser bench page passes dpr.

Native Sponge replay at 720p: cpu p95 per event falls from 12.9 to
7.9 ms, and non-submit cpu late in the stroke from 9.8 to 5.1 ms.
…om 2x

The present shader took one bilinear tap from a root composite allocated
with a single mip level. Below 50% zoom the taps skip texels, so sharp
edges hit or miss by phase and crawl while panning. Separately, the auto
pixel filter snapped to nearest from 100% zoom, where texels cannot tile
evenly, so edges staircased between 100% and 200%.

The root accumulator pair now carries a mip chain, regenerated lazily with
the existing box-pyramid rescale pass only when a minifying present follows
a changed composite (tracked by a tick stamp against composite_built, no
readback). The present reads one integer mip level, floor(log2(1/zoom)),
sampled bilinearly, as GEGL and Krita's High Quality mode do; the 50% to
100% range is unchanged. The sampling policy moves out of the shader into
view.rs as a unit-tested PixelFilter enum, with auto switching to nearest
at 2x (Krita, GEGL). The three view-uniform upload paths collapse into one,
and RescalePass builds its halve uniforms once instead of per level per
frame.

Regression tests in tests/present_zoom.rs fail against the unfixed present
(255 everywhere at 1/8 zoom where 223 is the box average; 128 vs 191 at
zoom 0.3; a pure-black column at 1.5x) and pass after, plus a laziness
counter test and view.rs unit tests for the policy.
Merge dev into offline-stroke-path

dev's per-frame stroke flush (#134) already renders a stroke no frame
ran during in one pass at pen-up, so stroke_path, BrushRender,
StrokeOp::pen_sample and StrokeEngine::stabilize_all are dropped, along
with the stroke_path tests. push_all no longer fits dev's windowed
Laplacian and is dropped here; batching returns in a later change.

Kept from this branch and adapted to dev:

- the dab-cap phase split in place_dab and submit_and_continue, now
  reachable from a headless stroke's single pen-up phase; regression
  tests preview_dab_cap.rs and headless_stroke_splits_phases_at_the_dab_cap
- set_stroke_seed, with stroke_seed_pins_random_nodes
- render_whole for the brush preview
- the bench comparison, as stroke_replay_bench --no-frames (a flush per
  event against no frames)
A stroke's events were each pushed through the stabilizer as they
arrived, and the Laplacian relaxed an N + 2 vertex window per resampled
vertex. A stroke no frame ran during, as a headless embedder paints,
computed every intermediate polyline and rendered only the last: at
stabilize 0.6 that was 27.5 ms of a 52.5 ms stroke.

stroke_to now only records the event. take_divergence hands the
stabilizer every event since the last flush through push_all, now the
required trait method (push wraps it). The resampler retracts its
provisional tip once and hands the algorithm one batch; prediction
rebuilds its tail once. The Laplacian relaxes a batch once, from the
recorded edge to the tip, at N x (M + N) vertex updates for M vertices:
never more than one vertex at a time, and bit-identical to it. It keeps
one edge trajectory, vertex len - N - 2, which no later point or tip
retract changes; retract_tip drops the weight that read the retracted
point. MAX_COMMITS_PER_PUSH is renamed MAX_COMMITS_PER_SAMPLE.

The headless recorded stroke at 0.6 drops from 52.5 to 26.9 ms
(stroke_path in PR #135: 22.2 ms). Live painting relaxes once per frame
instead of once per vertex.

Tests: a_batch_costs_one_relaxation and headless_stroke_relaxes_once
fail on the per-vertex path; batches_are_bit_exact_with_from_scratch,
a_batch_reaches_the_inner_as_one_push, batched_stack_matches_sequential.

Based on the batching in PR #135 by daghack.
@daghack
daghack merged commit f822c6e into daghack:offline-stroke-path Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants