Skip to content
tacular-omicsPublic

About

Self-contained, URL-safe mass spectrum interchange for Python and TypeScript

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

spectrl

CI PyPI Python DOI License

Put a mass spectrum directly in a URL.

Encodes one spectrum's peak arrays and modeled mzML metadata into a compact, URL-safe token. The encoded payload lives in the string. No backend is required.

spectrl.v3.<mode>.<base64url(payload)>.<checksum>

Try the browser demo · Read the format specification · See the changelog

A spectrl token embedded in a URL and decoded into a mass spectrum, with portable text, a shared format across implementations, and local decoding.

A spectrum travels as ordinary URL-safe text and decodes entirely client-side.

Why

Use spectrl to share a spectrum in a URL, QR code, notebook, paper, or application handoff. The token contains the spectrum itself, so decoding does not depend on an external service or file. A Universal Spectrum Identifier (USI) points to a spectrum in a repository. Spectrl embeds the spectrum. Use a USI when long-term repository lookup is the goal, and spectrl when a compact, self-contained handoff is more useful.

A local spectrum is encoded into a spectrl token, placed after the # of a URL, and decoded in the browser. The token has three parts: identifier, encoded spectrum data, and checksum.

What's included

Component Purpose
spectrl Python package Reference encoder/decoder, mzML bridge, URL helpers, and CLI
js/ Independent TypeScript implementation for browsers and Node
rust/ Independent Rust implementation (library and spectrl CLI), written from the specification and vectors
SPECIFICATION.md Normative spectrl.v3 wire-format specification
test-vectors/ Shared positive, negative, and cross-language conformance vectors

Install

pip install spectrl

Requires Python 3.12+. All core numeric encodings work in Python and Pyodide.

To install the unreleased development version:

pip install "spectrl @ git+https://github.com/tacular-omics/spectrl.git"

The TypeScript implementation is published as @spectrl-ms/spectrl:

npm install @spectrl-ms/spectrl

Rust

The Rust crate is published as spectrl on crates.io. It was written from SPECIFICATION.md, the registry and the shared vectors only, and writes the same tokens as Python and TypeScript byte for byte:

cargo add spectrl            # library
cargo install spectrl        # `spectrl encode|decode|inspect` CLI

It requires Rust 1.85+ and a C compiler. The crate forbids unsafe; zlib and Brotli come from the bundled reference C libraries, which is what makes z and b tokens byte-identical. Brotli is a default feature. See rust/README.md.

For service integration, see the producer and consumer examples, including precision policies, optional decoder budgets, HTTP limits, and worker guidance. JavaScript requires Node 22+ and also supports browser bundlers.

Quick start

Encode from mzmlpy

The mzML bridge is optional. Install it with pip install "spectrl[mzml]". Applications such as Spectacular that already depend on mzmlpy do not need the extra.

from mzmlpy.run import Mzml
from spectrl import encode_spectrum, from_mzmlpy

with Mzml("data.mzML") as mzml:
    spec = mzml.spectra[0]
    token = encode_spectrum(from_mzmlpy(spec))

print(token)
# spectrl.v3.z.…

Encode manually

import numpy as np
from spectrl import encode_spectrum
from spectrl.model import InlineSpectrum, SpectrlCvParam

spec = InlineSpectrum(
    default_array_length=3,
    mz=np.array([147.0, 175.1, 246.2]),
    intensity=np.array([1e5, 8e4, 3e4]),
    id="scan=42",
    params=[
        SpectrlCvParam(accession="MS:1000511", value=2),   # ms level
        SpectrlCvParam(accession="MS:1000130"),             # positive scan
        SpectrlCvParam(accession="MS:1000127"),             # centroid
    ],
)

token = encode_spectrum(spec)

Decode

from spectrl import decode_token

decoded = decode_token(token)
print(decoded.mz)        # numpy array
print(decoded.intensity) # numpy array
print(decoded.id)        # "scan=42"

URL bindings

from spectrl import to_fragment, to_query, to_data_uri, extract_token

# Embed in a URL fragment. This is recommended because it is never sent to the server.
url = to_fragment(token, "https://viewer.example.com/spectrum")
# https://viewer.example.com/spectrum#spectrl.v3.z.…

# Or as a query parameter
url = to_query(token, "https://viewer.example.com/spectrum")
# https://viewer.example.com/spectrum?d=spectrl.v3.z.…

# Or as a data URI
uri = to_data_uri(token)
# data:application/vnd.spectrl;v=3,spectrl.v3.z.…

# Extract token back from any of the above
token = extract_token(url)

Additional arrays

Beyond the dedicated m/z, intensity, and charge fields, attach any per-peak array by PSI-MS accession (a standard mzML binary array) or by a free-text name (a non-standard MS:1000786 array). Every ion-mobility variant lives here, so multiple distinct mobility arrays are preserved. int32/float32 dtypes are preserved. ArrayAccession provides readable string-valued enum keys while still allowing future PSI-MS accessions as plain strings.

import numpy as np
from spectrl import ArrayAccession, UnitAccession, encode_spectrum, decode_token
from spectrl.model import InlineSpectrum

spec = InlineSpectrum(
    default_array_length=3,
    mz=np.array([147.0, 175.1, 246.2]),
    intensity=np.array([1e5, 8e4, 3e4]),
    extra_arrays={
        ArrayAccession.RAW_INVERSE_REDUCED_ION_MOBILITY: np.array([0.82, 0.91, 1.05]),
        "MS:1000517": np.array([120.0, 80.0, 45.0]),         # signal-to-noise array (named CV)
        "iso_score": np.array([0.98, 0.91, 0.74], np.float32),  # non-standard (MS:1000786)
    },
    array_units={
        ArrayAccession.RAW_INVERSE_REDUCED_ION_MOBILITY:
            UnitAccession.VOLT_SECOND_PER_SQUARE_CENTIMETER,
    },
)
decoded = decode_token(encode_spectrum(spec))
decoded.extra_arrays["iso_score"]  # float32 array, round-tripped
decoded.mobility_arrays["MS:1003008"]  # accession-keyed filtered view

Auxiliary arrays retain their native numeric types and exact values in both profiles. Default lossy encoding picks, per array, the smallest of an exact encoding and bounded candidates: a 0.1 ppm grid for floating m/z, and for nonnegative floating intensity exact integer words for counts, 12-bit rounded floats, or a log1p grid. Ties and unsupported domains keep the exact encoding. JavaScript exposes the same arrays through extraArrays.

Override any array independently when needed:

token = encode_spectrum(spec, lossless=True, array_encodings={
    "mz": "modular-delta-shuffle",
    "intensity": "byte-shuffle",
    "iso_score": "raw",
})

Supported names are raw, byte-shuffle, modular-delta-shuffle, quantized, and rounded-float. Namespaced custom encodings can be registered explicitly. With lossless=True, every array remains exact and lossy overrides are rejected. The fixed lossless defaults are modular delta plus shuffle for m/z, byte shuffle for intensity, and raw words for other arrays. Both profiles compress the complete CBOR document once. There is no per-array compressor or PSI codec alias.

Unknown arrays remain exact unless the caller passes allow_unsafe_lossy_custom=True with an explicit lossy encoding. In JavaScript, enable Brotli with await installBrotli() from @spectrl-ms/spectrl/brotli to read or write Brotli payloads.

Payload compression

Fidelity and compression are independent options:

token = encode_spectrum(spec, lossless=True, compression="zlib")
token = encode_spectrum(spec, compression="auto")

The default is zlib level 6. Other choices are raw and brotli at quality 5. Install spectrl[brotli] to enable Brotli in Python. Explicit auto selects the shortest complete token among available backends, with ties preferring zlib, raw, then Brotli. The token stores the selected method. Explicitly requesting an unavailable backend raises an error.

The core array encodings are raw words (0), byte shuffle (1), modular delta plus shuffle (2), quantized words (3), and rounded floating-point words (4). Encoding 3 shares one unsigned integer layout for linear and log1p quantization, with optional first-order delta. Its parameters record scale, byte width, and optional log and delta flags. For example, an explicit linear intensity grid can be requested with:

token = encode_spectrum(spec, array_encodings={
    "intensity": {"encoding": [3, 1, {"scale": 100, "width": 4}]},
})

This rounds to increments of 0.01, with a checked absolute error bound of 0.005. Choose a width that fits the indices. A tighter tolerance requires a larger scale. The default log1p intensity candidate uses scale 3600, whose bound is (x + 1) * expm1(0.5 / 3600). It is not a strict relative bound near zero. When the smallest positive intensity m is below 1, as in normalized spectra, the scale grows to ceil(1800 * (m + 1) / m), so every positive intensity stays within about 0.028% of itself and none rounds to zero. The m/z default uses a logarithmic grid calibrated to a maximum error of 0.1 ppm relative to each source value. Its scale is derived from the smallest positive m/z, with a numerical margin and a check of the reconstructed values. Zero remains exactly zero. For example, the allowed error is 0.00001 at m/z 100 and 0.0001 at m/z 1000. All peaks remain present.

Encoding 4 keeps each float's sign, exponent and leading bits mantissa bits, so a normal value stays within a relative 2^-(bits+1) of itself and keeps its declared float32 or float64 type. Its parameters are bits and byte width:

token = encode_spectrum(spec, array_encodings={
    "intensity": {"encoding": [4, 1, {"bits": 12, "width": 4}]},
})

The default profile measures each candidate as encoded array bytes for raw payloads and as zlib level 6 output otherwise, so pass the same compression to encoding_plan that the token will use. See the specification for domains and tie order.

User params (free-text metadata)

For values with no CV term, attach named user parameters at the spectrum or scan level. They're omitted entirely when empty, so a spectrum without any is byte-identical to one produced before the feature existed.

from spectrl.model import InlineSpectrum, SpectrlUserParam

spec = InlineSpectrum(
    default_array_length=3, mz=mz, intensity=intensity,
    user_params=[
        SpectrlUserParam(name="Mascot score", value=42.7),
        SpectrlUserParam(name="reanalysis note", value="rerun semitryptic"),
    ],
)

from_mzmlpy reads spectrum- and scan-level userParams automatically. The JS implementation exposes the same via userParams. Values use their native scalar types; there is no separate type annotation. The mzML importer converts declared numeric user parameters into numbers, rejects invalid or out-of-range numeric values, and retains other values as text. Prefer a CV term whenever one exists. UserParams are heavier (no accession to compress) and uncontrolled.

Trim large spectra

from spectrl import top_n

# Keep the 50 most intense peaks before encoding
trimmed = top_n(spec, 50)
token = encode_spectrum(trimmed)

Quality, import, and sharing budgets

Use encoding_report() to measure error in the exact token it returns, fit_to_budget() to propose explicit peak or user-param removal for a URL byte budget, and parse_peak_list() / format_peak_list() for two-column text, CSV, or TSV. The TypeScript package provides the same workflows as encodingReport, fitToBudget, parsePeakList, and formatPeakList.

The browser demo supports importing your own peaks, downloading quality reports, and previewing a budget candidate before applying it. The Python conversion_report() API and spectrl convert-mzml command report observable mzML omissions alongside the converted spectrum.

See the workflow guide for runnable examples, JSON dtype preservation, CLI commands, and the limits of each report.

Lossless encoding

# Default is lossy quantization
# Use lossless=True for bit-exact native arrays
token = encode_spectrum(spec, lossless=True)

Token size

Token length grows with the number of points in the spectrum. The plot shows the 237 benchmark spectra from the spectrl manuscript, in the default and bit-exact encodings, against a cautious query-string limit and Chromium's URL limit. Large profile spectra belong in the URL fragment, not a query string.

Token length in characters against points per spectrum on log axes, for MS1, MS2, and MS3 centroid and profile spectra, default and bit-exact.

Token format

spectrl.v3.<mode>.<base64url(payload)>.<checksum>

The token text parts and the thirteen integer-keyed fields of the CBOR document: peak count, spectrum ID, CV parameters, scan list, precursors, products, peak arrays, free-text params, source, acquisition, processing, extensions, and ontology versions.

  • spectrl.v3: stable spectrl identifier + explicit v3 format version. The prefix is the version's only carrier.
  • Mode: z is the default zlib-compressed CBOR. r is raw CBOR and b is Brotli. All use unpadded base64url. Readers require r and z, with an optional b capability.
  • The payload contains one CBOR document (RFC 8949). Outer compression includes metadata and already encoded array blobs. Inspection expands this layer without decoding numeric arrays.
  • The required trailing checksum is CRC-32/ISO-HDLC over everything before the last ., encoded as eight lowercase hexadecimal characters. It detects accidental corruption without decoding the CBOR payload.
  • Header: a CBOR map with integer keys mirroring mzML structure: ms level, polarity, scan times, precursor isolation window, activation method, and collision energy.
  • Array blobs: one per array type (m/z, intensity, charge, and accession-keyed additional arrays, including every ion-mobility variant), each encoded through a versioned numeric encoding and embedded inline in the CBOR document as a byte string.

Validation

The shared conformance vectors test field-level Python/TypeScript interoperability in both directions, and the Rust crate passes every vector file. scripts/check_token_parity.py and scripts/check_adversarial_parity.py compare all three implementations (Rust is skipped with a message when cargo is absent). The test suites also cover malformed and adversarial inputs, canonicalization, URL bindings, mzML conversion, and all core numeric encodings.

# Python: lint, formatting check, and tests
just check

# TypeScript: install, typecheck, test, and build
cd js
npm ci
npm run typecheck
npm test
npm run build

# Rust: format, lint, and tests
cd rust
cargo fmt --check
cargo clippy --all-targets -- -D warnings
cargo test

Run just release-check from the repository root for the full Python, TypeScript, distribution, and demo release gate.

CLI

# Encode from JSON
echo '{"mz":[147.0,175.1],"intensity":[1e5,8e4]}' | spectrl encode

# Decode a token
echo "spectrl.v3.z.…" | spectrl decode

# Inspect the header as readable JSON
echo "spectrl.v3.z.…" | spectrl inspect

Demo

A browser demo encodes example spectra live, shows the shareable URL + QR, and decodes + plots them entirely client-side (no server). Use the hosted demo or launch it locally:

just demo   # → http://127.0.0.1:8000

See demo/ for details.

Design

  • mzML-aligned: modeled metadata uses mzML cvParam semantics and existing ontology accessions. A token is not an arbitrary mzML <spectrum> XML round-trip: run-level references, processing provenance, source-file links, and unmodeled XML structure are outside its scope.
  • CV binding: accession constants are generated into spectrl from its shared registry and validated against mzmlpy's StrEnum enums during development. Core encoding and decoding do not import an mzML parser.
  • Reproducible across implementations: Python and JavaScript produced identical complete tokens in all 1,836 comparisons, including 237 benchmark spectra under both encoding profiles with raw, zlib, and Brotli payloads; Rust matched both on all 414 targeted and seeded tokens. Matching metadata, array dtypes, and exact core encoding settings give portable token equality in raw payload mode. Compressed and lossy equality was verified for the tested runtime versions. Different settings can still produce different tokens for the same spectrum, and the CRC-32 is a corruption check rather than a spectrum hash. See reproducibility results and the specification.
  • Scope: represents measured spectra and acquisition context across mass spectrometry. Molecular identifications and fragment assignments are outside the format.

Scope and security

  • Decoding applies resource budgets by default, because a token usually arrives from somewhere untrusted. Tighten them per call with decode_token(token, limits=DecodeLimits(...)) or decodeToken(token, limits), or raise them to the format ceilings for a trusted producer with DecodeLimits.unlimited() and UNLIMITED_DECODE_LIMITS. See service integration for the default values, byte accounting, and examples.

  • URL lengths vary by browser and receiving system. Encoding warns above 8 KiB. Use top_n() or a repository identifier for spectra that are too large.

  • A bounded lossy profile is the default. Pass lossless=True when bit-exact arrays are required.

  • The trailing checksum detects accidental corruption. It does not authenticate the sender or make untrusted content safe.

  • spectrl preserves modeled spectrum-level metadata, not an entire mzML file or its run-level provenance. See the specification for the exact data model and decoder limits.

Specification

The normative token format is specified in SPECIFICATION.md (an open specification governed in this repository). This README is a tutorial. The specification is the contract. A machine-readable CV/codec/key registry lives in schema/registry.json.

Spectra convert both ways. spectrl.formats writes a decoded spectrum as mzML, MGF or MS2 and reads spectra back from any of them, reporting what a target format cannot represent rather than dropping it. mzML is the interchange format: a token written to mzML and read back is the same spectrum.

spectrl encode run.mzML --index 42 > token.txt   # one spectrum out of a run
spectrl decode token.txt --output spectrum.mzML  # and back again

spectrl.v3 carries measured spectra and acquisition context. Its header uses keys 0 through 12, with spectrum-level free-text parameters at key 7 and source-declared ontology versions at key 12. Molecular identifications belong in the surrounding application.

Contributing

See CONTRIBUTING.md and the Code of Conduct. Changes to the on-the-wire token format are governed more strictly. See the Format changes section of the contributing guide.

Bug reports and focused pull requests are welcome. Please report security problems privately as described in SECURITY.md.

Citation

If spectrl supports published work, cite the archived software release rather than the moving main branch. GitHub exposes the current metadata through CITATION.cff.

License

Licensed under the Apache License 2.0. If you use spectrl in research, please cite it via CITATION.cff. Third-party test-data attribution is recorded in NOTICE.

Related

  • mzmlpy: the mzML parser this library bridges from

V3 pipelines and selected-spectrum context

V3 separates numeric encoding from whole-document compression. No PSI codec term is needed. Earlier token versions and superseded development layouts are not supported. The trailing CRC32 protects the complete encoded token.

from spectrl import ArrayEncoding, encode_spectrum

token = encode_spectrum(spectrum, lossless=True, array_encodings={
    "mz": ArrayEncoding(encoding="modular-delta-shuffle"),
    "intensity": "byte-shuffle",
})

float32, float64, and int32 inputs retain their declared representation when encoded exactly, and encoding 4 keeps float32 or float64. Encoding 3 reconstructs float64. Lossy arrays record their error settings.

Use from_mzmlpy(spectrum, run=mzml) to include resolved source, instrument, software, and processing context. Repeated CV terms and nested user parameters are preserved. conversion_report(..., run=mzml) reports omissions. Trimming records selection history and scopes acquisition-derived summaries to their source.

register_encoding(id, Encoding(...), revision=1) installs trusted local callbacks. Numeric callbacks receive array/type/parameters for encoding and bytes/type/count/parameters for decoding. Each implementation supplies a parameter validator. Custom encoding registration is the only codec extension API. See the complete wire and callback contract. Inspection can read unknown operation descriptors without decoding the arrays. Full decoding fails clearly when a required operation or extension is unavailable.

Python JSON interchange preserves native dtypes with array_dtypes and extra_array_dtypes. Opaque byte strings use {"$spectrl":"bytes","hex":"..."}. Maps with non-string keys or a literal $spectrl key use a tagged map with an items list of key/value pairs, so extension data round-trips without ambiguity. These tags belong to the JSON API, not the CBOR wire format.

About

Self-contained, URL-safe mass spectrum interchange for Python and TypeScript

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages