Skip to content

Monitoring for errors #49

Description

@jeremyestein

Description

It would be nice to be able to inspect the system quickly and easily, certainly without having to fire up a shell.

Definition of Done

  • A basic script looking for errors or inconsistencies is available (only this is in scope for Waveform project).
  • We know within a reasonable amount of time that something (an upload or processing step) failed, so we can manually re-run, and caution users not to use the incomplete data
  • Perhaps showing a summary of data that has been received (can't really check for errors, as we don't know what data's expected)
  • Separately monitor compressed HL7 files vs 2 kinds of parquets vs csv files.

Dependencies

Comments

This could roll into the next project on open telemetry.
It would be useful to have shared error monitoring in SAFEHR rather than done separately for individual projects.

Suggested implementation

Mount the waveform-export and waveform-saved-messages directories read-only into the emap streamlit container. Remember that waveform-export is not a part of Emap and neither directory is guaranteed to have any contents.

Make a new streamlit page that has controls on it for date, hour, variable, maybe bed.

For HL7 monitoring

A query method with caching (careful not to cache current data that is fast changing...) that takes the above parameters and does a query along the lines of:

find waveform-save-messages/${date}*/${bed}/ -name '*.bz' perl -pe 's/\r/\n/g' | grep -Po 'NM\|\d+\|' | sort | uniq -c

(you could use pure python instead)

The purpose is to see at a glance what types of data were coming through at various times. I think a common use case is wanting to know what is coming through right now.

For export pipeline monitoring

Something in a similar spirit to the above but in the waveform-export directory.

To check for errors, you could compare the number of files in original-csv vs pseudonymised vs ftps-logs, or you could scan the snakemake log file.

Activity

  1. changed the title [-]Monitoring for errors etc[/-] [+]Monitoring for errors[/+] on Feb 4, 2026
  2. MonikaSvata commented on Feb 23, 2026

    @MonikaSvata
    Collaborator

    Parked until the project resumes.

  3. MonikaSvata commented on Aug 3, 2026

    @MonikaSvata
    Collaborator

    @jeremyestein to touchbase with @HChughtai and @p-j-smith to check what they have done in Telemetry for monitoring and if we can reuse it.

  4. skeating commented on Aug 24, 2026

    @skeating
    Collaborator

    Jeremy is going to split this into several issues each of which should produce 1 PR

  5. jeremyestein commented on Sep 1, 2026

    @jeremyestein
    CollaboratorAuthor

    A lot of this is now covered by reporting to OpenTelemetry. Any mention of scanning the disk to look for errors is likely better handled by sending metrics as soon as the error is identified. Eg. bad interchange message, FTPS upload failed.

  6. MonikaSvata commented on Sep 2, 2026

    @MonikaSvata
    Collaborator

    New tickets to be created for the items that are pending. Afterwards, this can be closed.

  7. jeremyestein commented on Sep 2, 2026

    @jeremyestein
    CollaboratorAuthor

    Much of this has been implemented already, in the form of telemetry that's sent when eg. an interchange message succeeds or fails, or an upload is made over FTPS.

    But it's a very wide issue, so it will be closed and superseded by:

  8. MonikaSvata commented on Sep 4, 2026

    @MonikaSvata
    Collaborator

    New tickets created and prioritised. Closing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions