Repository navigation
Explore further dashboard performance optimizations #718
Description
Activity
Profiling results: Bokeh property validation skip
Summary
Tested
@without_property_validationonSessionUpdater._periodic_update_impl(the entire periodic update cycle). No measurable improvement in overall cycle time (~206ms avg in both cases), but the investigation revealed useful information.The decorator works, but validation is a small fraction of total cost
Expanding the pyinstrument session files to show all hidden frames:
Metric Baseline With decorator Bokeh prepare_value/validateaggregate1.152s/10s 0.211s/10s Total Pipe.send7.354s/10s 6.962s/10s Total cycle 9.858s/10s 9.653s/10s The Bokeh validation sum is inflated by pyinstrument tree nesting (children counted in parents). Real savings are ~3-5% of cycle time — within run-to-run noise.
Why Dask's 50% doesn't apply to us
Dask updates
ColumnDataSource.datadirectly — property validation dominates their update path. Our updates go through the full HoloViews pipeline (Pipe.send→ DynamicMap → element construction → range computation → model sync), so Bokeh validation is a small slice at the end. The outerfreeze()batching already collapses model sync to one recompute, further reducing the number of property sets.New finding: HoloViews
datetime_typesisinstance checkExpanding the
_update_rangessubtree (2.15s of 9.86s total) revealed a surprising cost:_update_ranges (2.149s) └─ get_extents (2.047s) └─ _get_subplot_extents (1.820s) └─ RectanglesPlot.get_extents (1.542s) └─ max_range (1.097s) ├─ datetime_types.__instancecheck__ (0.589s) ← !! ├─ numpy.nanmin/nanmax (0.198s) └─ filterwarnings (0.097s)0.59s per 10s cycle (6%) is spent checking
isinstance(value, datetime_types)inside HoloViews'max_rangeutility. This uses lazy module imports (_LazyModule) that repeatedly resolvepandas_datetime_typesandcftime_typeson every call. This is pure overhead for our use case (we never use datetime axes).The hard-range short-circuit (setting explicit dimension ranges from the Autoscaler) would eliminate this entire
_update_ranges→get_extentspath, making it the highest-impact optimization to try next.Verdict
Not worth adding now — the savings are real but too small to measure reliably (~3-5%). Worth revisiting once the larger costs (range computation, element construction) are addressed, as validation would then become a proportionally larger share.
Profile data is in
.scratch/profiles/baseline/and.scratch/profiles/without-property-validation/on theprofilerbranch.Profiling results: Hard dimension ranges from Autoscaler
Summary
Implemented
_apply_hard_ranges()onPlotter— forwards autoscaler bounds to HoloViews viaredim.range()so_compute_group_rangeshort-circuits thenanmin/nanmaxdata scan. Profiled before/after on the DREAM dashboard.Correction to previous comment
The previous analysis predicted hard ranges would "eliminate the entire
_update_ranges→get_extentspath." This turned out to be wrong.compute_rangesandget_extentsare two separate range computation paths withinupdate_frame:update_frame (5.68s) ├─ _get_frame 2.08s DynamicMap callback + element cloning ├─ _update_ranges 1.81s Compute axis extents from pre-computed ranges ├─ compute_ranges 0.60s Compute data ranges (short-circuit target) ← helped ├─ wrapper (subplots) 0.51s Update child Bokeh plots └─ _update_plot 0.49s Sync Bokeh model stateHard ranges short-circuit
compute_rangesonly (the data scan). The_update_ranges→get_extentspath is a downstream consumer — it reads the pre-computed ranges dict and combines them across all overlay subplots. It always runs regardless of hard ranges.Measured impact
Session update (GUI thread,
Pipe.sendpath):Metric Baseline Hard-ranges Change Pipe.send7.35s/10s 5.75s/10s -1.60s (-22%) compute_ranges0.94s 0.60s -0.34s (-36%) _compute_group_range0.91s 0.58s -0.33s (-37%) get_extents(top)2.05s 1.75s -0.30s (-15%) max_range(total)~1.4s ~1.1s -0.3s (-21%) Avg call 206 ms 175 ms -31 ms (-15%) Orchestrator compute (background thread):
Metric Baseline Hard-ranges Change ImagePlotter.plot1.92s 2.57s +0.66s _apply_hard_ranges— 0.68s new cost redim.range()deep-copies the entire HoloViews element (including data arrays), adding 0.68s to the compute path. Net savings across both threads: ~0.9s/30s.Why
get_extentsis still expensive (1.75s)The cost breakdown inside
get_extents:GeomMixin.get_extents(Rectangles ROI overlay): 1.35s — iterates 4 kdims, callsmax_rangefor each pair of start/end dimensionsmax_range: 0.97s total — each call doesimport pandas,warnings.filterwarnings(), numpy array creation,np.nanmin/np.nanmax. Called per dimension per subplotdimension_range: 0.48s — combines hard/soft/data ranges with padding per axis
This is
O(subplots × dimensions)work per frame. The per-call overhead inmax_range(pandas import guard, warning context manager, numpy array allocation,datetime_typesisinstance check from previous comment) accumulates across the many calls. Hard ranges don't help here — these functions always run to combine extents from all overlay children.Verdict
The
compute_rangesshort-circuit works as designed (-37%), but the overall impact is modest becauseget_extentsis the larger cost and can't be bypassed from our side. Theredim.range()deep-copy overhead (0.68s) partially offsets the savings. Further range optimization would require changes in HoloViews itself (e.g., cachingmax_rangeresults, removing the datetime isinstance checks for non-datetime data, or avoiding deep copies inredim.range()).Profile data is in
.scratch/profiles/baseline/and.scratch/profiles/hard-ranges/on theprofilerbranch.Status re-evaluated against main (what landed, what was rejected with profiling reasons, what remains — CDS hot-path spike and websocket compression): #1096 (comment)
I think this is likely superseded by #1055, and anyway a new analysis is probably more useful since this is so old.
Context
The
multi-session-optimizationbranch addressed the top-3 dashboard performance costs (outerdoc.models.freeze()batching, skippipe.send()for hidden tabs, version-based polling). Deep research into the Panel/Bokeh/HoloViews stack and the broader scientific dashboard community identified several further optimization opportunities worth exploring.Pre-optimization cost breakdown (from profiling before the branch's freeze batching):
recompute(model graph rebuild)freeze())compute_ranges+get_extents)After the outer
freeze()batching, range computation is likely the dominant remaining cost. Fresh profiling would confirm the new breakdown.Optimization opportunities
Disable Bokeh property validation during updates
Bokeh validates property types on every model/CDS update. Dask's distributed dashboard disabled this and reported ~50% reduction in Bokeh overhead. The mechanism:
This is safe when data shapes are controlled (as ours are by backend services). Could be applied to the
_batched_update()body. Low effort, measurable impact, zero risk.Reference: bokeh/bokeh#6042, bokeh/bokeh#7072 (quadratic
stream()performance caused by validation).Set explicit hard ranges on HoloViews dimensions
HoloViews forces
framewise=Truefor all DynamicMaps (hardcodedself.dynamicoverride incompute_ranges()), causing range recomputation on every frame. However, range computation short-circuits when dimensions have explicit hard ranges:Our
Autoscaleralready computes bounds. If those bounds are applied via.redim.range()on the elements before sending through the pipe, the expensivenanmin/nanmaxdata scans could be eliminated. Medium effort, high impact — likely the biggest remaining win given that freeze batching already addressed the recompute cost.Reference: holoviz/holoviews#4909 (framewise not respected on composite plots).
Direct Bokeh CDS updates for the streaming hot path
Each
pipe.send()triggers the full HoloViews overhead chain: DynamicMap callback, element construction, range computation, model diffing, property validation. For plots where only data changes (not structure), updating the BokehColumnDataSourcedirectly would bypass all of this:A November 2025 Discourse discussion confirms DynamicMap is "significantly slower than using Bokeh directly." This is a larger refactor but has the highest performance ceiling. A hybrid approach is possible: use HoloViews for initial plot construction and interactive features (ROI editing, linked tools), but update data via direct CDS manipulation in the streaming path. High effort, highest impact.
WebSocket compression
Panel supports per-message deflate compression:
Worth benchmarking for histogram/image data. Trade-off is ~300KB memory per connection and some CPU. Low effort, impact depends on data sizes.
Broader context
--num-procs(doesn't help with shared state) orhold/batching (already done).