Skip to content

Explore further dashboard performance optimizations #718

Description

@SimonHeybrock

Context

The multi-session-optimization branch addressed the top-3 dashboard performance costs (outer doc.models.freeze() batching, skip pipe.send() for hidden tabs, version-based polling). Deep research into the Panel/Bokeh/HoloViews stack and the broader scientific dashboard community identified several further optimization opportunities worth exploring.

Pre-optimization cost breakdown (from profiling before the branch's freeze batching):

Component % of cycle (before) Status after branch
Bokeh recompute (model graph rebuild) ~44% Reduced (N recomputes collapsed to 1 via outer freeze())
HoloViews range computation (compute_ranges + get_extents) ~28% Unchanged — not addressed by current optimizations
Subplot glyph updates ~11% Unchanged
Other ~17% Unchanged

After the outer freeze() batching, range computation is likely the dominant remaining cost. Fresh profiling would confirm the new breakdown.

Optimization opportunities

Disable Bokeh property validation during updates

Bokeh validates property types on every model/CDS update. Dask's distributed dashboard disabled this and reported ~50% reduction in Bokeh overhead. The mechanism:

from bokeh.core.properties import without_property_validation

@without_property_validation
def update_callback():
    source.data = new_data

This is safe when data shapes are controlled (as ours are by backend services). Could be applied to the _batched_update() body. Low effort, measurable impact, zero risk.

Reference: bokeh/bokeh#6042, bokeh/bokeh#7072 (quadratic stream() performance caused by validation).

Set explicit hard ranges on HoloViews dimensions

HoloViews forces framewise=True for all DynamicMaps (hardcoded self.dynamic override in compute_ranges()), causing range recomputation on every frame. However, range computation short-circuits when dimensions have explicit hard ranges:

# In HoloViews _compute_group_range:
if all(util.isfinite(r) for r in el_dim.range):
    data_range = (None, None)  # Skip data scan entirely

Our Autoscaler already computes bounds. If those bounds are applied via .redim.range() on the elements before sending through the pipe, the expensive nanmin/nanmax data scans could be eliminated. Medium effort, high impact — likely the biggest remaining win given that freeze batching already addressed the recompute cost.

Reference: holoviz/holoviews#4909 (framewise not respected on composite plots).

Direct Bokeh CDS updates for the streaming hot path

Each pipe.send() triggers the full HoloViews overhead chain: DynamicMap callback, element construction, range computation, model diffing, property validation. For plots where only data changes (not structure), updating the Bokeh ColumnDataSource directly would bypass all of this:

bokeh_source = plot_handle.handles['source']
bokeh_source.data = {'x': new_x, 'y': new_y}

A November 2025 Discourse discussion confirms DynamicMap is "significantly slower than using Bokeh directly." This is a larger refactor but has the highest performance ceiling. A hybrid approach is possible: use HoloViews for initial plot construction and interactive features (ROI editing, linked tools), but update data via direct CDS manipulation in the streaming path. High effort, highest impact.

WebSocket compression

Panel supports per-message deflate compression:

panel serve --websocket-compression-level 6 --websocket-compression-mem-level 8

Worth benchmarking for histogram/image data. Trade-off is ~300KB memory per connection and some CPU. Low effort, impact depends on data sizes.

Broader context

  • ESSlivedata is one of the most advanced web-based scientific live data dashboards in the neutron/synchrotron community. Most facilities (ORNL, ISIS, ESRF, European XFEL) still use desktop Qt/Java GUIs.
  • CHESS (Cornell) tried Bokeh, then Panel, then switched to Plotly Express, but for a much simpler use case (10s refresh, no interactive features).
  • Nobody in the Panel community has reported solving the multi-session + high-frequency update problem. The standard advice is --num-procs (doesn't help with shared state) or hold/batching (already done).
  • Bokeh memory leaks from incomplete session cleanup are a long-standing issue worth monitoring in long-running deployments.
  • No realistic alternative framework solves the fundamental Bokeh per-session model graph overhead without losing capabilities we depend on (ROI editing, linked tools, Datashader, HoloViews declarative API).

Activity

  1. SimonHeybrock commented on Mar 2, 2026

    @SimonHeybrock
    MemberAuthor

    Profiling results: Bokeh property validation skip

    Summary

    Tested @without_property_validation on SessionUpdater._periodic_update_impl (the entire periodic update cycle). No measurable improvement in overall cycle time (~206ms avg in both cases), but the investigation revealed useful information.

    The decorator works, but validation is a small fraction of total cost

    Expanding the pyinstrument session files to show all hidden frames:

    Metric Baseline With decorator
    Bokeh prepare_value/validate aggregate 1.152s/10s 0.211s/10s
    Total Pipe.send 7.354s/10s 6.962s/10s
    Total cycle 9.858s/10s 9.653s/10s

    The Bokeh validation sum is inflated by pyinstrument tree nesting (children counted in parents). Real savings are ~3-5% of cycle time — within run-to-run noise.

    Why Dask's 50% doesn't apply to us

    Dask updates ColumnDataSource.data directly — property validation dominates their update path. Our updates go through the full HoloViews pipeline (Pipe.send → DynamicMap → element construction → range computation → model sync), so Bokeh validation is a small slice at the end. The outer freeze() batching already collapses model sync to one recompute, further reducing the number of property sets.

    New finding: HoloViews datetime_types isinstance check

    Expanding the _update_ranges subtree (2.15s of 9.86s total) revealed a surprising cost:

    _update_ranges (2.149s)
    └─ get_extents (2.047s)
       └─ _get_subplot_extents (1.820s)
          └─ RectanglesPlot.get_extents (1.542s)
             └─ max_range (1.097s)
                ├─ datetime_types.__instancecheck__ (0.589s)  ← !!
                ├─ numpy.nanmin/nanmax (0.198s)
                └─ filterwarnings (0.097s)
    

    0.59s per 10s cycle (6%) is spent checking isinstance(value, datetime_types) inside HoloViews' max_range utility. This uses lazy module imports (_LazyModule) that repeatedly resolve pandas_datetime_types and cftime_types on every call. This is pure overhead for our use case (we never use datetime axes).

    The hard-range short-circuit (setting explicit dimension ranges from the Autoscaler) would eliminate this entire _update_ranges → get_extents path, making it the highest-impact optimization to try next.

    Verdict

    Not worth adding now — the savings are real but too small to measure reliably (~3-5%). Worth revisiting once the larger costs (range computation, element construction) are addressed, as validation would then become a proportionally larger share.

    Profile data is in .scratch/profiles/baseline/ and .scratch/profiles/without-property-validation/ on the profiler branch.

  2. SimonHeybrock commented on Mar 2, 2026

    @SimonHeybrock
    MemberAuthor

    Profiling results: Hard dimension ranges from Autoscaler

    Summary

    Implemented _apply_hard_ranges() on Plotter — forwards autoscaler bounds to HoloViews via redim.range() so _compute_group_range short-circuits the nanmin/nanmax data scan. Profiled before/after on the DREAM dashboard.

    Correction to previous comment

    The previous analysis predicted hard ranges would "eliminate the entire _update_ranges → get_extents path." This turned out to be wrong. compute_ranges and get_extents are two separate range computation paths within update_frame:

    update_frame (5.68s)
    ├─ _get_frame          2.08s  DynamicMap callback + element cloning
    ├─ _update_ranges      1.81s  Compute axis extents from pre-computed ranges
    ├─ compute_ranges      0.60s  Compute data ranges (short-circuit target) ← helped
    ├─ wrapper (subplots)  0.51s  Update child Bokeh plots
    └─ _update_plot        0.49s  Sync Bokeh model state
    

    Hard ranges short-circuit compute_ranges only (the data scan). The _update_ranges → get_extents path is a downstream consumer — it reads the pre-computed ranges dict and combines them across all overlay subplots. It always runs regardless of hard ranges.

    Measured impact

    Session update (GUI thread, Pipe.send path):

    Metric Baseline Hard-ranges Change
    Pipe.send 7.35s/10s 5.75s/10s -1.60s (-22%)
    compute_ranges 0.94s 0.60s -0.34s (-36%)
    _compute_group_range 0.91s 0.58s -0.33s (-37%)
    get_extents (top) 2.05s 1.75s -0.30s (-15%)
    max_range (total) ~1.4s ~1.1s -0.3s (-21%)
    Avg call 206 ms 175 ms -31 ms (-15%)

    Orchestrator compute (background thread):

    Metric Baseline Hard-ranges Change
    ImagePlotter.plot 1.92s 2.57s +0.66s
    _apply_hard_ranges — 0.68s new cost

    redim.range() deep-copies the entire HoloViews element (including data arrays), adding 0.68s to the compute path. Net savings across both threads: ~0.9s/30s.

    Why get_extents is still expensive (1.75s)

    The cost breakdown inside get_extents:

    • GeomMixin.get_extents (Rectangles ROI overlay): 1.35s — iterates 4 kdims, calls max_range for each pair of start/end dimensions
    • max_range: 0.97s total — each call does import pandas, warnings.filterwarnings(), numpy array creation, np.nanmin/np.nanmax. Called per dimension per subplot
    • dimension_range: 0.48s — combines hard/soft/data ranges with padding per axis

    This is O(subplots × dimensions) work per frame. The per-call overhead in max_range (pandas import guard, warning context manager, numpy array allocation, datetime_types isinstance check from previous comment) accumulates across the many calls. Hard ranges don't help here — these functions always run to combine extents from all overlay children.

    Verdict

    The compute_ranges short-circuit works as designed (-37%), but the overall impact is modest because get_extents is the larger cost and can't be bypassed from our side. The redim.range() deep-copy overhead (0.68s) partially offsets the savings. Further range optimization would require changes in HoloViews itself (e.g., caching max_range results, removing the datetime isinstance checks for non-datetime data, or avoiding deep copies in redim.range()).

    Profile data is in .scratch/profiles/baseline/ and .scratch/profiles/hard-ranges/ on the profiler branch.

  3. SimonHeybrock commented on Jul 21, 2026

    @SimonHeybrock
    MemberAuthor

    Status re-evaluated against main (what landed, what was rejected with profiling reasons, what remains — CDS hot-path spike and websocket compression): #1096 (comment)

  4. SimonHeybrock commented on Jul 21, 2026

    @SimonHeybrock
    MemberAuthor

    I think this is likely superseded by #1055, and anyway a new analysis is probably more useful since this is so old.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions