Problem
PerMessageDeflate.decompress_message_data is the only place in the codebase
that applies a decompressed-output cap, via zlib's max_length argument:
# src/autobahn/websocket/compress_deflate.py:811-814
def decompress_message_data(self, data):
if self.max_message_size is not None:
return self._decompressor.decompress(data, self.max_message_size)
return self._decompressor.decompress(data)
Per Python's zlib semantics, when decompressobj().decompress(data, max_length)
is called with a nonzero max_length, it returns at most max_length bytes
and stores the remaining not-yet-produced output in
decompressor.unconsumed_tail, which must be fed back into decompress() in a
loop to retrieve the rest. Autobahn never reads unconsumed_tail (zero
occurrences in src/), so when max_message_size is set and a message exceeds
it, the excess is silently discarded rather than cleanly rejected.
Worse, the discarded tail leaves the decompressor mid-stream, so the subsequent
end_decompress_message (which feeds the \x00\x00\xff\xff sync-flush trailer)
throws on a corrupt stream.
Reproduction (plain zlib, mirrors Autobahn's calls)
import zlib
payload = b"a" * 2000
co = zlib.compressobj(zlib.Z_DEFAULT_COMPRESSION, zlib.DEFLATED, -15)
comp = (co.compress(payload) + co.flush(zlib.Z_SYNC_FLUSH))[:-4] # 18 bytes on the wire
do = zlib.decompressobj(-15)
out = do.decompress(comp, 1500) # cap = 1500
# out -> 1500 bytes (message was 2000)
# do.unconsumed_tail -> 5 bytes, DROPPED by Autobahn
do.decompress(b"\x00\x00\xff\xff") # end_decompress_message trailer
# raises: zlib error -3 "invalid distance code"
So the deflate cap does not produce a clean "message too large" outcome — it
produces truncation followed by a decode error. This is a latent
data-corruption bug on its own, and it is why Crossbar's test_cb_zip_bomb.py
"passes" (the connection drops) for the wrong reason: the drop comes from the
zlib -3 decode error, not from a size-limit rejection.
Fix (approach)
Make decompress_message_data drain unconsumed_tail in a bounded loop up to
the cap, and when output would exceed the cap, stop and signal "message too
large" cleanly (raise PayloadExceededError, matching how the transport already
reports oversize on send) rather than truncating + corrupting. Preserve the
uncapped fast path unchanged when max_message_size is None.
This makes deflate's bounded decompress a correct primitive that child issue 02
(protocol-level enforcement) and 03 (base-API generalization) build on.
Red/green test plan (TDD, on this PR)
- Unit level on
compress_deflate (no network, deterministic; runs locally
via just and in CI):
- RED: construct a
PerMessageDeflate(..., max_message_size=N), feed a
message whose inflated size is > N, assert a clean PayloadExceededError
(or the agreed signal) — today this either truncates silently or raises the
wrong error (zlib -3). Prove red in CI.
- Assert the under-cap case still round-trips byte-exact (no regression).
- Assert a multi-
decompress (fragmented input) case drains the tail correctly.
- GREEN: after the fix, all assertions pass.
- Matrix: this is deflate-specific; both Twisted/asyncio import the same module,
so one unit test covers both backends. (snappy/bzip2/brotli generalization is
child 03.)
Acceptance criteria
- Setting deflate
max_message_size and exceeding it yields a clean, typed
rejection, never truncated/corrupt data and never a zlib -3.
- Under-cap messages are byte-exact.
- New unit test present, proven red-before / green-after in CI on this PR.
References
This work is being completed with AI assistance (Claude Code).
Problem
PerMessageDeflate.decompress_message_datais the only place in the codebasethat applies a decompressed-output cap, via zlib's
max_lengthargument:Per Python's zlib semantics, when
decompressobj().decompress(data, max_length)is called with a nonzero
max_length, it returns at mostmax_lengthbytesand stores the remaining not-yet-produced output in
decompressor.unconsumed_tail, which must be fed back intodecompress()in aloop to retrieve the rest. Autobahn never reads
unconsumed_tail(zerooccurrences in
src/), so whenmax_message_sizeis set and a message exceedsit, the excess is silently discarded rather than cleanly rejected.
Worse, the discarded tail leaves the decompressor mid-stream, so the subsequent
end_decompress_message(which feeds the\x00\x00\xff\xffsync-flush trailer)throws on a corrupt stream.
Reproduction (plain zlib, mirrors Autobahn's calls)
So the deflate cap does not produce a clean "message too large" outcome — it
produces truncation followed by a decode error. This is a latent
data-corruption bug on its own, and it is why Crossbar's
test_cb_zip_bomb.py"passes" (the connection drops) for the wrong reason: the drop comes from the
zlib -3decode error, not from a size-limit rejection.Fix (approach)
Make
decompress_message_datadrainunconsumed_tailin a bounded loop up tothe cap, and when output would exceed the cap, stop and signal "message too
large" cleanly (raise
PayloadExceededError, matching how the transport alreadyreports oversize on send) rather than truncating + corrupting. Preserve the
uncapped fast path unchanged when
max_message_size is None.This makes deflate's bounded decompress a correct primitive that child issue 02
(protocol-level enforcement) and 03 (base-API generalization) build on.
Red/green test plan (TDD, on this PR)
compress_deflate(no network, deterministic; runs locallyvia
justand in CI):PerMessageDeflate(..., max_message_size=N), feed amessage whose inflated size is
> N, assert a cleanPayloadExceededError(or the agreed signal) — today this either truncates silently or raises the
wrong error (
zlib -3). Prove red in CI.decompress(fragmented input) case drains the tail correctly.so one unit test covers both backends. (snappy/bzip2/brotli generalization is
child 03.)
Acceptance criteria
max_message_sizeand exceeding it yields a clean, typedrejection, never truncated/corrupt data and never a
zlib -3.References
This work is being completed with AI assistance (Claude Code).