Problem
Only the deflate backend has any decompressed-output cap
(max_message_size, and even that was broken — see child 01). The other three
WebSocket compression backends decompress a frame fully, unbounded, before
any protocol-level check can see the size:
compress_snappy.py:481-482 — self._decompressor.decompress(data) (no cap;
python-snappy StreamDecompressor.decompress has no output-length arg).
compress_bzip2.py:525-526 — self._decompressor.decompress(data) (the stdlib
BZ2Decompressor.decompress does accept max_length, but it is unused here).
compress_brotli.py:486-487 — self._decompressor.process(data) (no cap).
The base class mandates nothing — compress_base.py:60-63 is a bare class; the
decompress interface (start_decompress_message / decompress_message_data /
end_decompress_message) is duck-typed convention only, with no size-cap concept.
After child 02, the protocol layer enforces maxMessagePayloadSize on
uncompressed bytes post-decompress for these codecs — which stops the bypass,
but still lets a single frame inflate fully into memory first (bounded only by
maxFramePayloadSize on the wire). This issue closes that gap by making the
bounded-decompress guarantee uniform across all backends.
Fix (approach)
- Add a bounded variant to the base decompress API, e.g.
decompress_message_data(self, data, max_output_len=None), defaulting to
unbounded (backward compatible), with a documented contract: return at most
max_output_len bytes; if more output remained, signal "too large" (raise
PayloadExceededError) rather than truncate.
- Implement it precisely where the library supports it:
- deflate — via child 01's tail-draining loop (native zlib
max_length).
- bzip2 — via
BZ2Decompressor.decompress(data, max_length) + tail drain.
- Implement it as decompress-then-check where the library has no output cap:
- snappy, brotli — decompress the frame (already wire-bounded by
maxFramePayloadSize), then enforce max_output_len. Document this weaker
"bounded per-frame, not per-chunk" guarantee explicitly.
- The protocol layer (child 02) passes the remaining message budget as
max_output_len so inflation stops at the cap for the codecs that support it.
Red/green test plan (TDD, on this PR)
- Parametrized unit tests over each available backend (skip when the optional
dependency is absent — snappy/bzip2/brotli are behind optional imports):
- RED: with a
max_output_len set, an over-budget message must yield a
clean PayloadExceededError; today snappy/bzip2/brotli ignore the bound.
- Under-budget round-trips byte-exact.
- deflate/bzip2 fragmented-input tail-drain correctness.
- GREEN after the fix.
- Both backends import the same modules; one unit suite covers Twisted + asyncio.
Acceptance criteria
- Every compression backend honors
max_output_len with a clean typed rejection
on overflow; none truncates silently.
- Optional backends degrade gracefully (skipped when the dependency is missing).
- Base-API contract documented. Tests proven red-before / green-after in CI.
References
This work is being completed with AI assistance (Claude Code).
Problem
Only the deflate backend has any decompressed-output cap
(
max_message_size, and even that was broken — see child 01). The other threeWebSocket compression backends decompress a frame fully, unbounded, before
any protocol-level check can see the size:
compress_snappy.py:481-482—self._decompressor.decompress(data)(no cap;python-snappy
StreamDecompressor.decompresshas no output-length arg).compress_bzip2.py:525-526—self._decompressor.decompress(data)(the stdlibBZ2Decompressor.decompressdoes acceptmax_length, but it is unused here).compress_brotli.py:486-487—self._decompressor.process(data)(no cap).The base class mandates nothing —
compress_base.py:60-63is a bare class; thedecompress interface (
start_decompress_message/decompress_message_data/end_decompress_message) is duck-typed convention only, with no size-cap concept.After child 02, the protocol layer enforces
maxMessagePayloadSizeonuncompressed bytes post-decompress for these codecs — which stops the bypass,
but still lets a single frame inflate fully into memory first (bounded only by
maxFramePayloadSizeon the wire). This issue closes that gap by making thebounded-decompress guarantee uniform across all backends.
Fix (approach)
decompress_message_data(self, data, max_output_len=None), defaulting tounbounded (backward compatible), with a documented contract: return at most
max_output_lenbytes; if more output remained, signal "too large" (raisePayloadExceededError) rather than truncate.max_length).BZ2Decompressor.decompress(data, max_length)+ tail drain.maxFramePayloadSize), then enforcemax_output_len. Document this weaker"bounded per-frame, not per-chunk" guarantee explicitly.
max_output_lenso inflation stops at the cap for the codecs that support it.Red/green test plan (TDD, on this PR)
dependency is absent — snappy/bzip2/brotli are behind optional imports):
max_output_lenset, an over-budget message must yield aclean
PayloadExceededError; today snappy/bzip2/brotli ignore the bound.Acceptance criteria
max_output_lenwith a clean typed rejectionon overflow; none truncates silently.
References
This work is being completed with AI assistance (Claude Code).