Skip to content

ext/standard: Add a fast path for ASCII in htmlspecialchars() - #24145

Open
ArtUkrainskiy wants to merge 1 commit into
php:masterfrom
ArtUkrainskiy:htmlspecialchars-plain-runs
Open

ArtUkrainskiy wants to merge 1 commit into
php:masterfrom
ArtUkrainskiy:htmlspecialchars-plain-runs

Conversation

@ArtUkrainskiy

@ArtUkrainskiy ArtUkrainskiy commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

This is the first part of #18126, split into a separate PR.

htmlspecialchars() currently calls get_next_char() for every input character and then checks the entity table. This is unnecessary for most ASCII input.

This changes the loop in php_escape_html_entities_ex() so that:

  • get_next_char() is only called for bytes >= 0x80. ASCII is a single byte in all supported charsets.
  • In htmlspecialchars() without ENT_DISALLOWED, ordinary ASCII bytes are copied directly until one of & " ' < > is found.
  • get_next_char() is marked zend_always_inline.

The ASCII copy path is not used with ENT_DISALLOWED, since control characters still have to be checked, nor by htmlentities(), since HTML5 has entities for other ASCII characters too. The exception is htmlentities() with ENT_XML1 or a multibyte charset other than UTF-8: LIMIT_ALL reduces those to the basic entities, so they take the copy path as well.

get_next_char() is large enough that compilers kept it out of line. Forcing it inline removes the call overhead for non-ASCII input as well. This also follows @bukka's suggestion in #18126 to optimize the existing decoder rather than add another one.

Benchmarks

Release builds, master vs this PR:

Each row groups several inputs of the htmlspecialchars suite; the per-input numbers are in the report linked below.

htmlspecialchars() input i7-13700H, GCC 13 Ryzen 9 5950X, GCC 15 instructions
short strings, 8–40 bytes 1.4–1.75× faster 1.4–2.5× faster 1.4–3.4× fewer
English text, 64 B – 64 KB 1.7–2.5× faster 2.7–3.8× faster 4.4–6.5× fewer
Markdown, JS, 4 KB 2.8× faster 3.7–4.1× faster 5.1–6.4× fewer
HTML, 1–64 KB 1.8× faster 2.5–2.9× faster 3.4–3.7× fewer
Russian (UTF-8), 1 KB and 64 KB 1.2–1.5× faster 1.15–1.2× faster 1.3× fewer
Chinese (UTF-8) 1–64 KB, Shift_JIS 1 KB, Big5 1 KB 1.2× faster same to 1.2× faster 1.2× fewer
1 KB of only & / only " 1.4× / 1.2× faster 1.3× faster / same 1.2–1.4× / 1.1× fewer
cp1251 Cyrillic, 1 KB and 64 KB same to 1.08× faster 1.3–1.7× faster 1.35–1.5× fewer

In Symfony Demo (benchmark/benchmark.php, callgrind), htmlspecialchars() and htmlentities() are called 221 times per request and account for about 1.3% of the instructions.

This patch reduces the instruction count for the whole request by 0.87% (39.48M → 39.14M), or 1.01% with the tracing JIT. The average htmlspecialchars() call goes from 2291 to 719 instructions.

The WordPress benchmark calls htmlspecialchars() only six times per request, so there is no measurable change there.

Other effects

htmlentities() shares the same loop. Apart from the cases above, it does not use the ASCII copy path, but avoiding the decoder for ASCII still makes it about 1.2–1.4× faster.

There is also a small improvement in json_encode() for non-ASCII strings. ext/json uses php_next_utf8_char(), which wraps get_next_char(cs_utf_8, ...). After inlining, the compiler can eliminate the charset switch and the unused decoders.

This increases the wrapper from 23 to 757 bytes and saves about 13–14 instructions per non-ASCII character:

json_encode() input i7-13700H, GCC 13 Ryzen 9 5950X, GCC 15 instructions
Russian text 1.06–1.13× faster same to 1.04× faster 1.1× fewer
Chinese text 1.10–1.13× faster 1.04–1.09× faster 1.1× fewer
ASCII text; json_decode() same same same

ASCII input does not use this decoder, and json_decode() has its own scanner.

utf8_decode() and ext/xml use the same wrapper, but I did not benchmark them.

Verification

I compared the output against master for 20,000 random inputs using 13 flag combinations, 14 charsets, both functions, and both values of double_encode. The output was identical.

Full results and reproduction commands:

https://github.com/ArtUkrainskiy/php-src-bench/tree/main/reports/htmlspecialchars-ascii-fast-path

The ASCII copy loop still has one branch per byte. A SIMD version using zend_simd.h will be a separate PR.

ASCII is a single byte in every supported charset, so get_next_char() is
only called for bytes >= 0x80, and is inlined there. htmlspecialchars()
never changes ASCII other than & " ' < >, so such bytes are copied in a
loop of their own instead of being looked up in the entity table one by one.

Output is unchanged. htmlspecialchars():

  8-40 byte strings             1.4-1.7x faster
  English text                  1.7-2.5x faster
  Markdown, JS                  2.8x faster
  HTML                          1.8x faster
  Russian, Chinese text         1.2-1.5x faster
  only & or "                   1.2-1.4x faster

htmlentities() is 1.3-1.4x faster. Symfony Demo needs 0.87% fewer
instructions per request.
@ArtUkrainskiy
ArtUkrainskiy marked this pull request as ready for review October 6, 2026 11:15
@ArtUkrainskiy
ArtUkrainskiy requested a review from bukka as a code owner October 6, 2026 11:15
@ArtUkrainskiy

Copy link
Copy Markdown
Contributor Author

cc @devnexen @LamentXU123 — you've reviewed and merged the recent ext/standard optimizations, in case this one is in your area too.

@staabm

staabm commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

//cc @nicolas-grekas as you also worked on performance of this function

@nicolas-grekas

Copy link
Copy Markdown
Contributor

Thanks for the ping @staabm

@ArtUkrainskiy see #23957 which looks complementary.

@iliaal iliaal left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fuzzed this against master under ASAN (14 charsets, all flag combinations, both functions, both double_encode values): identical output, no reports. One correction to the description: LIMIT_ALL clears all for htmlentities() with non-UTF-8 multibyte charsets or ENT_XML1, so those calls take the ASCII copy path too, which is fine.

Comment thread ext/standard/html.c

/* {{{ get_next_char */
static inline unsigned int get_next_char(
static zend_always_inline unsigned int get_next_char(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Forcing this inline also pulls it into the ambiguous-entity lookahead in find_entity_for_char(), which is cold. Do you have the overall html.o text size delta?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the fuzzing and the LIMIT_ALL catch, I'll fix the description.

html.o text grows by 2.2 KB (37,445 → 39,620): the big loop goes from 3.7 to 6.7 KB, the standalone get_next_char (1.5 KB) disappears, php_next_utf8_char grows to 757 bytes.

The lookahead copy is ~700 bytes of that. I tried keeping it out of line via a never_inline wrapper, but the wrapper alone is 1.5 KB, so the object only gets bigger, and that
build was slower across most of the suite (GCC reshuffles the registers in the big function). So I'd rather keep it as is.

Kocal added a commit to Kocal/Twig that referenced this pull request Oct 7, 2026
With the UTF-8 charset, the `html` strategy calls `htmlspecialchars()`, which decodes the string one character at a time: about 5 ns per byte, so 1.6 µs for the 280-character class list of a styled button, even when nothing needs escaping. On valid UTF-8, all it does is replace `&`, `"`, `'`, `<` and `>`: `ENT_SUBSTITUTE` only matters for invalid sequences.

Strings longer than 32 bytes are now checked for UTF-8 validity with `preg_match('//u')` and escaped with `strtr()` on those five characters. Invalid UTF-8 and shorter strings still go through `htmlspecialchars()`, whose fixed cost is lower below that length. The output is identical: compared against `htmlspecialchars()` on every Unicode code point (in three positions) and on invalid sequences (lone continuation bytes, overlong forms, surrogates, code points above U+10FFFF, truncated sequences), 3.3 million strings, no difference on PHP 8.4 and 8.5. The shortcut only applies on PHP < 8.7 (`PHP_VERSION_ID < 80700`), since php/php-src#23957 and php/php-src#24145 (both still open, targeting PHP 8.7) make `htmlspecialchars()` itself skip the work on input that needs no encoding.

Benchmark below, PHP 8.4, 3 interleaved runs each side, identical output, ns per escape:

| String | Before | After |
| --- | --- | --- |
| short id (`user-42`) | **155-165 ns** | **168-175 ns** (below the threshold; the length check costs ~10 ns) |
| 40-character title | **327-334 ns** | **269-285 ns** |
| 280-character CSS class list | **1590-1630 ns** | **498-511 ns** (3.2x) |
| paragraph with quotes and tags | **1634-1663 ns** | **758-766 ns** (2.2x) |
| accented paragraph | **1310-1321 ns** | **373-380 ns** (3.5x) |

End to end, on an admin dashboard built with Symfony UX components that escapes about 2,500 attribute values per request (mostly long Tailwind class lists), CPU time per request goes from **36.7-38.7 ms** to **35.7-36.4 ms** (minimum of 6 runs of 60 requests each side), about 3%.

```php
<?php

// Run from the repository root: php bench.php
require getcwd().'/vendor/autoload.php';

$escaper = new Twig\Runtime\EscaperRuntime();
$classes = 'inline-flex shrink-0 items-center justify-center rounded-lg border border-transparent bg-clip-padding text-sm font-medium whitespace-nowrap transition-all outline-none select-none focus-visible:border-ring focus-visible:ring-3 focus-visible:ring-ring/50 disabled:pointer-events-none disabled:opacity-50';
$cases = [
    'short (id)' => 'user-42',
    'title (40 chars)' => 'How we made our dashboard ten times faster',
    'CSS classes (280 chars)' => $classes,
    'paragraph with quotes' => str_repeat('Twig escapes "quotes" & <tags> in a sentence like this one. ', 5),
    'accented paragraph' => str_repeat('Une phrase accentuée, déjà très répétée. ', 6),
];
foreach ($cases as $label => $string) {
    $times = [];
    for ($run = 0; $run < 7; ++$run) {
        $start = hrtime(true);
        for ($i = 0; $i < 200000; ++$i) {
            $escaper->escape($string);
        }
        $times[] = (hrtime(true) - $start) / 200000;
    }
    sort($times);
    printf("%-24s %5.0f ns/escape | sha1 %s\n", $label, $times[3], sha1($escaper->escape($string)));
}
```

Blackfire profiles (Symfony UX dashboard):
- Before: https://app.blackfire.io/envs/5f4f9a62-eaa0-45ee-b7b3-a1b879f550e9/profiles/b344f602-a21f-41ef-a585-46fc865e0880/graph
- After: https://app.blackfire.io/envs/5f4f9a62-eaa0-45ee-b7b3-a1b879f550e9/profiles/e5e16b8b-8895-40e7-9952-43282e3bd0df/graph
- Comparison: https://app.blackfire.io/envs/5f4f9a62-eaa0-45ee-b7b3-a1b879f550e9/profiles/compare/b344f602-a21f-41ef-a585-46fc865e0880...e5e16b8b-8895-40e7-9952-43282e3bd0df/graph

Blackfire barely moves (**2.29 s** -> **2.30 s** overall) because its per-call instrumentation dominates calls this short, even though `htmlspecialchars()` drops out of the profile entirely (5,398 calls per request before).

This came out of profiling Symfony UX components on https://github.com/Kocal/sf-ux-perfs-twig-components.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants