Repository navigation
Schema.isMaxLength uses UTF-16 code units but emits JSON Schema maxLength #6336
Description
Activity
Blast range: string length checks as Unicode code points
Suggested fix under analysis: change Effect string length checks to count Unicode
code points instead of JavaScript UTF-16 code units, while preserving current
array/object collection semantics.Direct runtime behavior
Schema.String.check(Schema.isMaxLength(n))- Astral characters such as
"💩"count as1instead of2. - Some values currently rejected become accepted.
- Astral characters such as
Schema.String.check(Schema.isMinLength(n))- Astral characters count as
1instead of2. - Some values currently accepted become rejected.
- Astral characters count as
Schema.String.check(Schema.isLengthBetween(min, max))- Same semantic change for both bounds.
Schema.Char- Would accept single-code-point astral characters such as
"💩". - This conflicts with its current docs, which explicitly describe
JavaScriptString.length.
- Would accept single-code-point astral characters such as
Schema.NonEmptyString/Schema.isNonEmpty()- No practical behavior change for strings, because only empty vs non-empty
matters.
- No practical behavior change for strings, because only empty vs non-empty
JSON Schema and representation
Schema.toJsonSchemaDocument- Existing
minLength/maxLengthoutput would become semantically aligned
with JSON Schema for strings.
- Existing
SchemaRepresentation.fromJsonSchemaDocument- Imported JSON Schema
minLength/maxLengthwould produce Effect schemas
that match JSON Schema string-length semantics.
- Imported JSON Schema
SchemaRepresentation.toCodeDocument- Generated code can stay structurally the same, but the meaning of emitted
Schema.isMinLength/Schema.isMaxLengthcalls changes.
- Generated code can stay structurally the same, but the meaning of emitted
- AI structured output helpers that consume Effect JSON Schema output
- JSON Schema shape is unchanged; runtime/string semantics become closer to
the schema sent to providers.
- JSON Schema shape is unchanged; runtime/string semantics become closer to
Arbitrary generation
- Length metadata currently feeds fast-check string constraints.
- fast-check string length is unit-based and defaults to ASCII-compatible units,
so most generated strings remain valid under code-point semantics. - fast-check can generate Unicode code-point units with
fc.string({ unit: "binary", minLength, maxLength }).- If Effect switches string checks to code-point semantics, string arbitrary
generation should likely useunit: "binary"so length constraints align
with runtime validation and JSON Schema. - If Effect keeps JavaScript
.lengthsemantics, usingunit: "binary"
would generate astral characters formaxLength(1)that the final filter
then rejects, increasing discard pressure.
- If Effect switches string checks to code-point semantics, string arbitrary
- Final schema filters still validate generated values, so correctness is
preserved. - Unicode-heavy arbitrary generation may need extra focused tests if we want to
guarantee efficient generation around astral characters.
Public compatibility risks
-
This is a breaking runtime behavior change for users relying on UTF-16 code
unit length. -
Relevant use cases include storage limits, protocol limits, database column
limits, and integrations where JavaScript.lengthis intentionally the
contract. -
Users needing old semantics would need a custom filter such as:
Schema.String.check(Schema.makeFilter((s) => s.length <= n))
Performance risks
- Current
.lengthchecks are O(1). - Code-point checks are O(n) for strings.
- The implementation should not use
Array.from(input).length, because it
allocates. - Prefer threshold-aware scans:
hasCodePointLengthAtMost(input, max)can stop aftermax + 1code points.hasCodePointLengthAtLeast(input, min)can stop aftermincode points.- exact/between checks should avoid full counting when a bound already fails.
Not affected
- Array length checks: still array element count.
- Object property count checks.
sizechecks forSet,Map, and other size-bearing values.- Numeric/date/property/unique checks.
- JSON Schema output shape for length checks: still
minLength,maxLength,
minItems, andmaxItems.
Required repo work
- Update
Schema.isMinLength,Schema.isMaxLength, andSchema.isLengthBetween. - Update docs/JSDoc for those checks.
- Update
Schema.Chardocs and tests. - Add regression tests for astral characters:
isMaxLength(1)accepts"💩".isMinLength(2)rejects"💩".isLengthBetween(1, 1)accepts"💩".Schema.Characcepts"💩"if we accept the new semantics.
- Add a changeset, because this is runtime behavior change.
Comparison context
- Zod v4 uses JavaScript
.lengthat runtime and emits JSON Schema
minLength/maxLength. - Valibot uses JavaScript
.lengthfor normallength/minLength/
maxLength; its companion JSON Schema package maps those to JSON Schema
minLength/maxLength. - Valibot also exposes separate byte and grapheme validators, which keeps its
different counting modes explicit.
Recommendation status
The fix is semantically clean if Effect wants JSON Schema alignment to be the
primary contract for string length checks. It is not drawback-free: it changes
runtime behavior and has performance cost unless implemented carefully.Relevant use cases include storage limits, protocol limits, database column
limits, and integrations where JavaScript .length is intentionally the
contract.I would argue that in most cases, Unicode code points are probably the better measure for database column limits. I believe that in Postgres, MySQL, and Oracle under default text encodings, VARCHAR columns are measured in terms of Unicode code points. So most code that relies on counting UTF-16 code units for compatibility with database column limits probably already has subtle bugs
Also for TEXT columns, functions like
length(col)andchar_length(col)in Postgres at least measure "characters" / code points, not UTF-16 code units. As far as I can tell, Postgres doesn't even support the UTF-16 character set
What version of Effect is running?
effect@4.0.0-beta.92What steps can reproduce the bug?
Run the minimal repro:
https://github.com/sbking/effect-schema-maxlength-repro
The repro defines:
and validates the astral Unicode string
"💩"with both Effect Schema and AJV using the generated JSON Schema.What is the expected behavior?
The Effect schema and the JSON Schema generated from it should validate the same string values.
In particular, if
Schema.toJsonSchemaDocumentemits JSON SchemamaxLength, the corresponding Effect string-length check should have JSON Schema-compatible semantics for strings, or the generated JSON Schema should avoid emittingmaxLengthfor a check with different runtime semantics.What do you see instead?
The generated JSON Schema accepts a value that the source Effect schema rejects.
Current repro output:
Additional information
Schema.isMaxLengthappears to validate JavaScript string length using UTF-16 code units. JSON SchemamaxLengthis defined in terms of Unicode code points.Those semantics diverge for astral Unicode characters:
"💩".length === 2Array.from("💩").length === 1"💩"for{ "type": "string", "maxLength": 1 }"💩"forSchema.String.check(Schema.isMaxLength(1))Pinned repro versions:
effect@4.0.0-beta.92ajv@8.20.0tsx@4.22.4typescript@6.0.3Related Discord discussion:
https://discord.com/channels/795981131316985866/1521281900428398745/1521341901041569823