Repository navigation
Add stroke_path for whole-path offline strokes - #135
Merged
Merged
Conversation
DarklyEngine::stroke_path(layer, &[StrokeOp]) paints a stroke whose samples are all known up front. It stabilizes them in one batch and renders the final polyline once (one submission per dab phase), instead of the live path's per-event rewind, re-render and full-layer commit. Output equals the live path's full re-render at every stabilizer strength, and the default live path at strength 0. Undo is recorded as for any stroke. - StrokeEngine::render_whole, shared with the brush preview renderer - StabilizerAlgorithm::push_all; the Laplacian override relaxes once (O(N*L) instead of O(N*L^2)) with a bit-identical result - place_dab flushes and submits when a phase reaches MAX_DABS_PER_PHASE; a longer phase overflowed the dab buffer (regression test: tests/preview_dab_cap.rs) - DarklyEngine::set_stroke_seed for reproducible random-node brushes - stroke_replay_bench --whole-path compares live against stroke_path
Member
|
@daghack thanks for your work on this. It's a good feature to have. I just did a bunch of work on the stabilizer including some performance improvements which should get us about halfway there. I'll make a PR to your PR and you can let me know if it looks good. |
Merge dev into offline-stroke-path dev's per-frame stroke flush (darkly-art#134) already renders a stroke no frame ran during in one pass at pen-up, so stroke_path, BrushRender, StrokeOp::pen_sample and StrokeEngine::stabilize_all are dropped, along with the stroke_path tests. push_all no longer fits dev's windowed Laplacian and is dropped here; batching returns in a later change. Kept from this branch and adapted to dev: - the dab-cap phase split in place_dab and submit_and_continue, now reachable from a headless stroke's single pen-up phase; regression tests preview_dab_cap.rs and headless_stroke_splits_phases_at_the_dab_cap - set_stroke_seed, with stroke_seed_pins_random_nodes - render_whole for the brush preview - the bench comparison, as stroke_replay_bench --no-frames (a flush per event against no frames)
A stroke's events were each pushed through the stabilizer as they arrived, and the Laplacian relaxed an N + 2 vertex window per resampled vertex. A stroke no frame ran during, as a headless embedder paints, computed every intermediate polyline and rendered only the last: at stabilize 0.6 that was 27.5 ms of a 52.5 ms stroke. stroke_to now only records the event. take_divergence hands the stabilizer every event since the last flush through push_all, now the required trait method (push wraps it). The resampler retracts its provisional tip once and hands the algorithm one batch; prediction rebuilds its tail once. The Laplacian relaxes a batch once, from the recorded edge to the tip, at N x (M + N) vertex updates for M vertices: never more than one vertex at a time, and bit-identical to it. It keeps one edge trajectory, vertex len - N - 2, which no later point or tip retract changes; retract_tip drops the weight that read the retracted point. MAX_COMMITS_PER_PUSH is renamed MAX_COMMITS_PER_SAMPLE. The headless recorded stroke at 0.6 drops from 52.5 to 26.9 ms (stroke_path in PR darkly-art#135: 22.2 ms). Live painting relaxes once per frame instead of once per vertex. Tests: a_batch_costs_one_relaxation and headless_stroke_relaxes_once fail on the per-vertex path; batches_are_bit_exact_with_from_scratch, a_batch_reaches_the_inner_as_one_push, batched_stack_matches_sequential. Based on the batching in PR darkly-art#135 by daghack.
Member
Merge dev and finish offline strokes in the normal stroke path
Contributor
Author
|
@TheTechromancer Excellent! I've merged and updated the PR. |
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
TheTechromancer
approved these changes
Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
I use Darkly headless, as a library, to render journal pages offline from strokes generated in code. Every stroke is fully known before it is drawn, but the only way to paint it was to replay it point by point through
begin_stroke/stroke_to/end_stroke, as if a pen were drawing it live. Profiling my renderer showed 94% of the main thread insidestroke_to, with nearly half of that blocked waiting on Metal command buffers: the time was going to per-event submissions, rewinds and full-layer commits, not to drawing.This adds a way to hand Darkly a whole stroke at once and have it drawn in a single pass. It should help anyone driving Darkly without a live pen, such as offline or batch rendering, stroke replay, or generated art, and it keeps output reproducible by letting the caller pin the random seed.
Summary
Adds
DarklyEngine::stroke_path(layer, &[StrokeOp]) -> Result<(), String>, which paints a brush stroke whose samples are all known up front (offline rendering, replay) in one pass. The live path treats everystroke_toas a pen event: stabilize, rewind to a checkpoint, re-render the tip with its new lookahead, and composite the whole layer, which is about 3 submissions per event and up to about 10 when the stabilizer diverges.stroke_pathstabilizes all samples in one batch, renders the final polyline once from the stroke-start state, and commits once: one submission per dab phase. Undo is recorded as for any stroke.Also fixes a crash that any dab phase over
MAX_DABS_PER_PHASE(16384) hits today, reachable now through the brush preview.Changes
stroke_path(engine/painting.rs): refuses if a stroke is open, or if any op is notStrokeOp::BrushStroke, before touching state. Opens withbegin_stroke, runs every op's growth and undo prelude first (extracted fromgpu_stroke_toasprepare_stroke_op), so the existing lazy init creates the stroke buffer once at the final extent. Then it renders throughbrush_stroke_towith a privateBrushRender::Path, and closes withend_stroke.brush_stroke_totakesBrushRender::Event | Pathin place of eight raw pen fields; thePatharm returns early, so the per-event body is not re-indented. Prediction is never applied to a path.StrokeEngine::render_whole:begin_stroke, every dab of the stabilized polyline,commit, into one context. The brush preview renderer uses it in place of its three hand-written phases, so preview and whole-path strokes share one render.StabilizerAlgorithm::push_allwith a looping default.LaplacianStabilizeroverrides it to relax once: O(NL) instead of O(NL^2) for an L-point path, with a bit-identical polyline.StrokeEngine::stabilize_allfeeds it.StrokeEngine::place_dab): when a dab-batching terminal's queue reachesMAX_DABS_PER_PHASE, it flushes the terminals and submits throughBrushGpuContext::submit_and_continue, which is factored out offlush_if_needed. Over-cap phases used to tripqueue_dab's debug assert, and in release they fail wgpu validation (a panic natively, a dropped stroke on the web).DarklyEngine::set_stroke_seed(Option<u32>): a session setting that seedsrandomnodes for reproducible renders. The defaultNonekeeps the wall-clock seed.StrokeOpderivesClone, Copyand gainspen_sample(), the one conversion toPaintInformationfor both paths.stroke_replay_bench --whole-path: times the recording live against onestroke_pathcall, through GPU completion.docs/brush/architecture.md; whole-path strokes in the per-frame flow indocs/brush/stabilization.md.Output equivalence
stroke_pathis byte-identical to the live path forced through a full re-render on every event (test_set_full_rerender), at every stabilizer strength, with prediction off.Against the default live path it is byte-identical at
stabilize = 0when no dab clips at the layer edge before a later growth. Tested for every builtin brush.At
stabilize > 0the default live path keeps segments drawn against lookahead points that later moved, so output differs. Measured on a 204-event recorded stroke at 1024x512:Tests
tests/preview_dab_cap.rs(regression): a 16k+ dab zig-zag throughBrushStrokePreviewRenderer::render_stroke. It fails ondevwith thequeue_daboverflow assert and passes with the split.tests/stroke_rewind.rs: the harness is generalized (per-side oracle flag, prediction pinned off, stroke seed pinned). New tests:1 + dabs / MAX_DABS_PER_PHASE;push_allmatches sequentialpushbit for bit.Performance
stroke_replay_bench --whole-path, recorded 204-event stroke, 1024x512, release, Apple GPU, timed through GPU completion:stroke_pathDenser input (for example 1 px samples) gains more, since the live cost is per event.
Notes
stroke_pathis a plainpub fn, not a#[handler]; there is no frontend caller yet.tests/watercolor.rs::watercolor_mark_is_invariant_to_flush_groupingfails identically ondevon the machine this was tested on.🤖 Generated with Claude Code