Skip to content

Surface out-of-range event diagnostics to users #716

Description

@SimonHeybrock

Context

DREAM sends ev44 messages where time_of_flight values fall outside the configured histogram bin range (e.g., all-negative TOF). These events are silently dropped by sc.hist(), causing the detector view to show zero counts with no indication of why. PR #713 attempted to fix this by adding overflow/underflow bins with ±inf sentinel edges, but revealed design concerns:

  • Brittle downstream handling: every plotter and the autoscaler must know to strip ±inf edges before rendering. This is an implicit protocol that's easy to get wrong when adding new consumers.
  • Unclear semantics: the EFU team confirmed negative TOF should be ignored, making it questionable whether these events belong in the histogram data at all.

PR #713 is closed in favor of a cleaner approach (to be determined here).

Approaches considered

Log and forget

Just log a warning when events fall outside the bin range. Cheap, no data model changes.

Downside: logs are invisible to instrument scientists watching the dashboard. If the detector view is blank, nothing tells them data is arriving but being dropped.

Separate overflow/underflow counts (extra coord or attr)

Store out-of-range event counts as scalar attributes alongside the histogram, not as bins.

Upside: clean data model, no plotter changes needed.
Downside: still needs a UI element to surface the info.

Job status heartbeat warning

Detect out-of-range events (cheaply, by comparing event count before vs histogram sum after sc.hist()) and surface the count as a warning_message on the JobStatus heartbeat. The existing JobStatusWidget already renders warnings — no new UI needed.

Upside: uses existing infrastructure end-to-end.
Downside: requires a way for workflows to communicate warnings back to the job. The current Workflow.finalize() protocol returns only dict[str, Any] with no diagnostic channel.

Extend Workflow protocol with WorkflowResult

Change finalize() to return a WorkflowResult(data=..., warnings=...) dataclass instead of a plain dict. StreamProcessorWorkflow is the natural place to separate data targets from diagnostic targets — it already maps Sciline keys to output names. The Sciline provider computes the warning, StreamProcessorWorkflow routes it, Job.get() consumes it generically.

Upside: generic mechanism any workflow can use, not special-cased to detector view.
Downside: protocol change that touches all Workflow implementations and Job.get().

Open questions

  • Is a generic diagnostic channel worth the protocol change, or is logging sufficient for now?
  • If we go with warnings on the heartbeat, should the Sciline provider produce a formatted string or a structured value that gets formatted elsewhere?
  • Are there other workflows that would benefit from a diagnostic channel, making the investment worthwhile?
  • Should this be a warning (transient, per-heartbeat) or a persistent indicator in the detector view UI?

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:backendServices, Kafka, message pipeline, preprocessors, job handlingenhancement

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions