Repository navigation
Merge dev and finish offline strokes in the normal stroke path - #1
Merged
daghack merged 11 commits intoOct 6, 2026
Merged
Conversation
The Laplacian's pull per vertex is now the pen's speed past it over a reference speed (500 CSS px/s), capped at 1. Fast motion is smoothed in full, a slow pivot keeps its corner, and a stopped pen is exact. The registry's `from_params` receives the stroke's view scale so the reference speed is in CSS pixels at any zoom. Two rendering defects fixed along the way: - The divergence walk stopped at the first vertex within tolerance of its rendered position, assuming everything behind it was unchanged. Against rendered positions that assumption fails, because vertices are re-rendered at different moments: a freshly rendered vertex ended the walk while vertices behind it drifted unchecked, then snapped back later as a visible disconnect on quick wide curves. `find_divergence` now scans the whole window for the earliest vertex out of tolerance. - `DIVERGENCE_EPSILON` drops from 0.5 to 0.1 CSS px. Neighbouring vertices rendered at different moments could each sit half a pixel off, giving the rendered stroke a 0.57 px ripple the smoothed polyline did not have. Regression tests: slow_corner_keeps_its_shape, moved_vertex_behind_an_unchanged_one_is_reported, rendered_positions_stay_within_epsilon.
The Sponge lagged at full size and stabilization because every pen sample ran a cost that grew with the stroke, several times per frame. The Laplacian relaxed the whole polyline from its raw points on every push, at 160 sweeps over a vertex every 6 CSS px. A push now relaxes only the N + 2 vertices its influence bound says can move. The vertex left of the window replays the per-sweep values recorded when it was an earlier window's left edge, so the result is bit-identical to the from-scratch relaxation at N x (N + 1) updates per push. Prediction likewise copies only the inner polyline's changed tail. Every coalesced pointer sample also ran the full diff, rewind, replay and commit. stroke_to now only feeds the stabilizer, and flush_stroke runs that cycle once per frame, from render and end_stroke. - Divergence tracking moves from the stabilizer algorithms to the stroke engine (DivergenceDiff), whose walk starts from the tip as last rendered, so it holds however many events arrive between frames. The resampler's per-event window widening goes with it. - The rewind and append paths share one segment loop. The unreachable no-stroke-buffer fallback and StrokeEngine::move_to are deleted. - compute_segment_boundaries keeps the vi = 0 anchor when a stroke's first frame renders a single vertex. - The [frame-perf] log gains a stroke sub-phase. - The rewind counter is renamed test_stroke_rewinds. The stroke_rewind.rs oracle tests had been comparing a run with itself since unstabilized strokes stopped rewinding. They now paint stabilized with a zero divergence epsilon (test_set_divergence_epsilon) and assert that rewinds happened. Test drivers that relied on per-sample rendering now render a frame per sample. stroke_replay_bench now runs on bench_device, so 1080p cells fit. The browser bench page passes dpr. Native Sponge replay at 720p: cpu p95 per event falls from 12.9 to 7.9 ms, and non-submit cpu late in the stroke from 9.8 to 5.1 ms.
…om 2x The present shader took one bilinear tap from a root composite allocated with a single mip level. Below 50% zoom the taps skip texels, so sharp edges hit or miss by phase and crawl while panning. Separately, the auto pixel filter snapped to nearest from 100% zoom, where texels cannot tile evenly, so edges staircased between 100% and 200%. The root accumulator pair now carries a mip chain, regenerated lazily with the existing box-pyramid rescale pass only when a minifying present follows a changed composite (tracked by a tick stamp against composite_built, no readback). The present reads one integer mip level, floor(log2(1/zoom)), sampled bilinearly, as GEGL and Krita's High Quality mode do; the 50% to 100% range is unchanged. The sampling policy moves out of the shader into view.rs as a unit-tested PixelFilter enum, with auto switching to nearest at 2x (Krita, GEGL). The three view-uniform upload paths collapse into one, and RescalePass builds its halve uniforms once instead of per level per frame. Regression tests in tests/present_zoom.rs fail against the unfixed present (255 everywhere at 1/8 zoom where 223 is the box average; 128 vs 191 at zoom 0.3; a pure-black column at 1.5x) and pass after, plus a laziness counter test and view.rs unit tests for the policy.
Stabilization Improvements
Merge dev into offline-stroke-path dev's per-frame stroke flush (#134) already renders a stroke no frame ran during in one pass at pen-up, so stroke_path, BrushRender, StrokeOp::pen_sample and StrokeEngine::stabilize_all are dropped, along with the stroke_path tests. push_all no longer fits dev's windowed Laplacian and is dropped here; batching returns in a later change. Kept from this branch and adapted to dev: - the dab-cap phase split in place_dab and submit_and_continue, now reachable from a headless stroke's single pen-up phase; regression tests preview_dab_cap.rs and headless_stroke_splits_phases_at_the_dab_cap - set_stroke_seed, with stroke_seed_pins_random_nodes - render_whole for the brush preview - the bench comparison, as stroke_replay_bench --no-frames (a flush per event against no frames)
A stroke's events were each pushed through the stabilizer as they arrived, and the Laplacian relaxed an N + 2 vertex window per resampled vertex. A stroke no frame ran during, as a headless embedder paints, computed every intermediate polyline and rendered only the last: at stabilize 0.6 that was 27.5 ms of a 52.5 ms stroke. stroke_to now only records the event. take_divergence hands the stabilizer every event since the last flush through push_all, now the required trait method (push wraps it). The resampler retracts its provisional tip once and hands the algorithm one batch; prediction rebuilds its tail once. The Laplacian relaxes a batch once, from the recorded edge to the tip, at N x (M + N) vertex updates for M vertices: never more than one vertex at a time, and bit-identical to it. It keeps one edge trajectory, vertex len - N - 2, which no later point or tip retract changes; retract_tip drops the weight that read the retracted point. MAX_COMMITS_PER_PUSH is renamed MAX_COMMITS_PER_SAMPLE. The headless recorded stroke at 0.6 drops from 52.5 to 26.9 ms (stroke_path in PR #135: 22.2 ms). Live painting relaxes once per frame instead of once per vertex. Tests: a_batch_costs_one_relaxation and headless_stroke_relaxes_once fail on the per-vertex path; batches_are_bit_exact_with_from_scratch, a_batch_reaches_the_inner_as_one_push, batched_stack_matches_sequential. Based on the batching in PR #135 by daghack.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Hi Talon, thanks for this. This PR is based on your profiling/benchmark, and I kept your original commit.
Basically while your PR was open, dev changed how strokes render. So I resolved the conflicts and integrated your feature directly into the normal stroke path. That includes the batching idea which is adapted to the new stabilizer. Your recorded stroke now takes 26.9 ms through plain
begin_stroke/stroke_to/end_stroke, against 22.2 ms forstroke_pathon the same machine. Sostroke_pathitself is gone: just keep calling that loop, with norenderbetween samples.For offline work, set
input.predictionHorizonto 0, or strokes overshoot their last sample by a few pixels.Merging this updates darkly-art#135, which we'll then review and merge into dev.
For your agent
This PR merges
darkly-art/darklydev (8e2bef0, including darkly-art#134) intooffline-stroke-path, then adds one commit. Net change against dev, relative to 3f99aea:Removed (superseded by dev's design):
DarklyEngine::stroke_path,BrushRender,StrokeOp::pen_sample,StrokeEngine::stabilize_all, and thestroke_path_*tests. After Stabilization Improvements darkly-art/darkly#134,stroke_toonly records the sample andflush_strokeruns once per frame and once fromend_stroke, so a stroke with norenderbetween samples is stabilized and rendered once, at pen-up.LaplacianStabilizer::push_all. Dev's Laplacian is windowed (edgestrajectories,chord_t,pull), so the override no longer compiles or applies.Kept, adapted to dev:
StrokeEngine::place_dabviaBrushGpuContext::submit_and_continue. A headless stroke's single pen-up phase reachesMAX_DABS_PER_PHASE, so this is load-bearing. Tests:tests/preview_dab_cap.rs,headless_stroke_splits_phases_at_the_dab_cap.DarklyEngine::set_stroke_seed(Option<u32>). Test:stroke_seed_pins_random_nodes.StrokeEngine::render_whole, used by the brush preview.stroke_replay_bench --no-frames(a flush per event against no frames between events).Added:
flush_stroke(final_flush: bool): the pen-up flush saves no checkpoints and renders its range as one segment. A frameless stroke is 3 submissions (prologue, segment, commit) plus one perMAX_DABS_PER_PHASEdabs, at any stabilizer strength.StabilizerAlgorithm::push_allis now the required method (pushwraps it).StrokeEngine::take_divergencehands the stack every sample recorded since the last flush in one call. The resampler retracts its provisional tip once and pushes once. Prediction rebuilds once. The Laplacian relaxes a batch ofMvertices once, atN x (M + N)updates, bit-identical to sequential pushes.MAX_COMMITS_PER_PUSHis renamedMAX_COMMITS_PER_SAMPLE.input.predictionHorizonlowered from 15 to 10 ms. Prediction tails are still committed at pen-up, a known open issue, so pin it to 0 offline withdarkly::config::set("input.predictionHorizon", ConfigValue::Float(0.0)).Usage for offline rendering:
Measured (
stroke_replay_bench --no-frames, recorded_curvy_stroke.json, 1024x512, release, prediction off; the no-frames figure includes a pen-up composite):stroke_pathRegression tests, each shown failing before its fix:
pen_up_flush_saves_no_checkpoints,live_pen_up_flush_renders_one_segment,a_batch_costs_one_relaxation,headless_stroke_relaxes_once, plus the dab-cap tests above. Equivalence:headless_stroke_matches_frame_by_frame,batches_are_bit_exact_with_from_scratch,a_batch_reaches_the_inner_as_one_push,batched_stack_matches_sequential.