Repository navigation
Reject non-canonical Base32 input (non-zero trailing bits and malformed padding) #13
Description
Activity
- addedenhancementNew feature or requestNew feature or request
on Jul 21, 2026 Implementations MUST include appropriate pad characters at the end of encoded data unless the specification referring to this document explicitly states otherwise.For TOTP, the governing spec is the otpauth Key URI Format (Google Authenticator's de-facto standard), and that convention stores/accepts the base32 secret unpadded. That's why real secrets in the wild are unpadded — not cause 4648 permits it freely, but because the referring spec omits it. So "unpadded is fine for TOTP" is correct, via 3.2's delegation, not in general.
Since the applicable spec (otpauth) omits padding, the right call for this library is option (a) accept unpadded, but if any = is present, require the correct count/position. Requiring full canonical padding (option b) would reject legitimate otpauth secrets even in strict mode. (Let's pick option a)
§3.3 also validates #12. The very next paragraph says implementations MUST reject non-alphabet characters "unless the specification referring to this document explicitly states otherwise." So the old lenient skip-everything behavior actually violated 4648's default — which is precisely what the #12 strict mode restores. (Lenient stays as an opt-out for the MIME-style "be liberal" behavior §3.3 also mentions.)
Extended strict mode (issue #12's pStrict / EBase32DecodeError) to reject non-canonical input: (1) non-zero trailing bits per RFC 4648 section 3.5, caught via ExtractLastBits(vBuffer, vBitsInBuffer) after the decode loop -- e.g. 'MZ======' and 'MY======' both decode to 'f' leniently, but strict rejects the non-canonical 'MZ======'; and (2) a data character appearing after a pad character. Full padding count/length validation and the padded-vs-unpadded policy were deferred to new open ticket #19, because real TOTP secrets are frequently stored unpadded and requiring canonical padding needs a deliberate policy decision. Resolved in v1.0.39
Completed in v1.0.39. Padding count/length validation and the padded-vs-unpadded policy are deferred to #19.
- added 2 commits that reference this issue
on Jul 22, 2026
TBase32.Decodeaccepts non-canonical input: it discards leftover trailing bitswithout checking they are zero, and ignores padding (
=) entirely. RFC 4648section 3.5 says decoders SHOULD reject encodings where the pad/trailing bits are
non-zero, because they are not produced by a conforming encoder.
Consequences:
are silently thrown away), which masks copy/paste corruption of a secret.
TestLongerInput_LoremIpsum(Source/Tests/DUnit/radRTL.Base32Encoding.Tests.pas:83)already documents the defect -- it asserts that appending a stray
L:still decodes to the original text (line 85 shows that two extra chars
LLchange the result, so the tolerance is exactly one non-canonical trailing char).
Add opt-in validation that rejects non-canonical input. Like the strict-decode
work, this is off by default so current lenient behavior is preserved.
Implementation notes:
(
Add opt-in strict-decode mode to TBase32.Decode ...). Preferably fold bothunder a single strict/canonical switch (the same
pStrictparameter and thesame
EBase32DecodeErrorexception) so callers get one "be strict" flag ratherthan two overlapping ones. Decide this when the first of the two is implemented.
vBitsInBuffer. For a canonical encoding those residual bits must all be zero.In strict mode, raise if
ExtractLastBits(vBuffer, vBitsInBuffer) <> 0.=count and position are legalfor the data length -- padding may only appear as a contiguous run at the very
end, the run length must match
pDataLength mod 8(0/1/3/4/6 pad chars for thevalid quantum sizes), and no alphabet character may follow a pad char. Reject
otherwise. If padding validation proves large, split it into its own ticket and
land the trailing-bit check first (it is the higher-value RFC-3.5 fix).
pStrict = False) path must remain byte-for-byte unchanged; theexisting
TestLongerInput_LoremIpsumassertions stay valid for lenient mode.Acceptance criteria:
LOREM_BASE32+'L'tolerance case.EBase32DecodeError.LOREM_BASE32+'L'(a non-canonical trailing char) is rejected.