Skip to content

Schema.isMaxLength uses UTF-16 code units but emits JSON Schema maxLength #6336

Description

@sbking

What version of Effect is running?

effect@4.0.0-beta.92

What steps can reproduce the bug?

Run the minimal repro:

https://github.com/sbking/effect-schema-maxlength-repro

npm install
npm run repro

The repro defines:

const schema = Schema.String.check(Schema.isMaxLength(1))
const jsonSchemaDocument = Schema.toJsonSchemaDocument(schema)

and validates the astral Unicode string "💩" with both Effect Schema and AJV using the generated JSON Schema.

What is the expected behavior?

The Effect schema and the JSON Schema generated from it should validate the same string values.

In particular, if Schema.toJsonSchemaDocument emits JSON Schema maxLength, the corresponding Effect string-length check should have JSON Schema-compatible semantics for strings, or the generated JSON Schema should avoid emitting maxLength for a check with different runtime semantics.

What do you see instead?

The generated JSON Schema accepts a value that the source Effect schema rejects.

Current repro output:

input: 💩
String.prototype.length / UTF-16 code units: 2
Array.from(input).length / Unicode code points: 1
Generated JSON Schema document: {
  "dialect": "draft-2020-12",
  "schema": {
    "type": "string",
    "allOf": [
      {
        "maxLength": 1
      }
    ]
  },
  "definitions": {}
}
AJV validates generated JSON Schema: true
Effect Schema validates original schema: false

Additional information

Schema.isMaxLength appears to validate JavaScript string length using UTF-16 code units. JSON Schema maxLength is defined in terms of Unicode code points.

Those semantics diverge for astral Unicode characters:

  • "💩".length === 2
  • Array.from("💩").length === 1
  • AJV accepts "💩" for { "type": "string", "maxLength": 1 }
  • Effect rejects "💩" for Schema.String.check(Schema.isMaxLength(1))

Pinned repro versions:

  • effect@4.0.0-beta.92
  • ajv@8.20.0
  • tsx@4.22.4
  • typescript@6.0.3

Related Discord discussion:

https://discord.com/channels/795981131316985866/1521281900428398745/1521341901041569823

Activity

  1. gcanti commented on Jun 30, 2026

    @gcanti
    Contributor

    Blast range: string length checks as Unicode code points

    Suggested fix under analysis: change Effect string length checks to count Unicode
    code points instead of JavaScript UTF-16 code units, while preserving current
    array/object collection semantics.

    Direct runtime behavior

    • Schema.String.check(Schema.isMaxLength(n))
      • Astral characters such as "💩" count as 1 instead of 2.
      • Some values currently rejected become accepted.
    • Schema.String.check(Schema.isMinLength(n))
      • Astral characters count as 1 instead of 2.
      • Some values currently accepted become rejected.
    • Schema.String.check(Schema.isLengthBetween(min, max))
      • Same semantic change for both bounds.
    • Schema.Char
      • Would accept single-code-point astral characters such as "💩".
      • This conflicts with its current docs, which explicitly describe
        JavaScript String.length.
    • Schema.NonEmptyString / Schema.isNonEmpty()
      • No practical behavior change for strings, because only empty vs non-empty
        matters.

    JSON Schema and representation

    • Schema.toJsonSchemaDocument
      • Existing minLength / maxLength output would become semantically aligned
        with JSON Schema for strings.
    • SchemaRepresentation.fromJsonSchemaDocument
      • Imported JSON Schema minLength / maxLength would produce Effect schemas
        that match JSON Schema string-length semantics.
    • SchemaRepresentation.toCodeDocument
      • Generated code can stay structurally the same, but the meaning of emitted
        Schema.isMinLength / Schema.isMaxLength calls changes.
    • AI structured output helpers that consume Effect JSON Schema output
      • JSON Schema shape is unchanged; runtime/string semantics become closer to
        the schema sent to providers.

    Arbitrary generation

    • Length metadata currently feeds fast-check string constraints.
    • fast-check string length is unit-based and defaults to ASCII-compatible units,
      so most generated strings remain valid under code-point semantics.
    • fast-check can generate Unicode code-point units with
      fc.string({ unit: "binary", minLength, maxLength }).
      • If Effect switches string checks to code-point semantics, string arbitrary
        generation should likely use unit: "binary" so length constraints align
        with runtime validation and JSON Schema.
      • If Effect keeps JavaScript .length semantics, using unit: "binary"
        would generate astral characters for maxLength(1) that the final filter
        then rejects, increasing discard pressure.
    • Final schema filters still validate generated values, so correctness is
      preserved.
    • Unicode-heavy arbitrary generation may need extra focused tests if we want to
      guarantee efficient generation around astral characters.

    Public compatibility risks

    • This is a breaking runtime behavior change for users relying on UTF-16 code
      unit length.

    • Relevant use cases include storage limits, protocol limits, database column
      limits, and integrations where JavaScript .length is intentionally the
      contract.

    • Users needing old semantics would need a custom filter such as:

      Schema.String.check(Schema.makeFilter((s) => s.length <= n))

    Performance risks

    • Current .length checks are O(1).
    • Code-point checks are O(n) for strings.
    • The implementation should not use Array.from(input).length, because it
      allocates.
    • Prefer threshold-aware scans:
      • hasCodePointLengthAtMost(input, max) can stop after max + 1 code points.
      • hasCodePointLengthAtLeast(input, min) can stop after min code points.
      • exact/between checks should avoid full counting when a bound already fails.

    Not affected

    • Array length checks: still array element count.
    • Object property count checks.
    • size checks for Set, Map, and other size-bearing values.
    • Numeric/date/property/unique checks.
    • JSON Schema output shape for length checks: still minLength, maxLength,
      minItems, and maxItems.

    Required repo work

    • Update Schema.isMinLength, Schema.isMaxLength, and Schema.isLengthBetween.
    • Update docs/JSDoc for those checks.
    • Update Schema.Char docs and tests.
    • Add regression tests for astral characters:
      • isMaxLength(1) accepts "💩".
      • isMinLength(2) rejects "💩".
      • isLengthBetween(1, 1) accepts "💩".
      • Schema.Char accepts "💩" if we accept the new semantics.
    • Add a changeset, because this is runtime behavior change.

    Comparison context

    • Zod v4 uses JavaScript .length at runtime and emits JSON Schema
      minLength / maxLength.
    • Valibot uses JavaScript .length for normal length / minLength /
      maxLength; its companion JSON Schema package maps those to JSON Schema
      minLength / maxLength.
    • Valibot also exposes separate byte and grapheme validators, which keeps its
      different counting modes explicit.

    Recommendation status

    The fix is semantically clean if Effect wants JSON Schema alignment to be the
    primary contract for string length checks. It is not drawback-free: it changes
    runtime behavior and has performance cost unless implemented carefully.

  2. sbking commented on Jun 30, 2026

    @sbking
    ContributorAuthor

    Relevant use cases include storage limits, protocol limits, database column
    limits, and integrations where JavaScript .length is intentionally the
    contract.

    I would argue that in most cases, Unicode code points are probably the better measure for database column limits. I believe that in Postgres, MySQL, and Oracle under default text encodings, VARCHAR columns are measured in terms of Unicode code points. So most code that relies on counting UTF-16 code units for compatibility with database column limits probably already has subtle bugs

    Also for TEXT columns, functions like length(col) and char_length(col) in Postgres at least measure "characters" / code points, not UTF-16 code units. As far as I can tell, Postgres doesn't even support the UTF-16 character set

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions