Repository navigation
ext/standard: Add a fast path for ASCII in htmlspecialchars() - #24145
ArtUkrainskiy wants to merge 1 commit into
Conversation
ASCII is a single byte in every supported charset, so get_next_char() is only called for bytes >= 0x80, and is inlined there. htmlspecialchars() never changes ASCII other than & " ' < >, so such bytes are copied in a loop of their own instead of being looked up in the entity table one by one. Output is unchanged. htmlspecialchars(): 8-40 byte strings 1.4-1.7x faster English text 1.7-2.5x faster Markdown, JS 2.8x faster HTML 1.8x faster Russian, Chinese text 1.2-1.5x faster only & or " 1.2-1.4x faster htmlentities() is 1.3-1.4x faster. Symfony Demo needs 0.87% fewer instructions per request.
|
cc @devnexen @LamentXU123 — you've reviewed and merged the recent ext/standard optimizations, in case this one is in your area too. |
|
//cc @nicolas-grekas as you also worked on performance of this function |
|
Thanks for the ping @staabm @ArtUkrainskiy see #23957 which looks complementary. |
iliaal
left a comment
There was a problem hiding this comment.
Fuzzed this against master under ASAN (14 charsets, all flag combinations, both functions, both double_encode values): identical output, no reports. One correction to the description: LIMIT_ALL clears all for htmlentities() with non-UTF-8 multibyte charsets or ENT_XML1, so those calls take the ASCII copy path too, which is fine.
|
|
||
| /* {{{ get_next_char */ | ||
| static inline unsigned int get_next_char( | ||
| static zend_always_inline unsigned int get_next_char( |
There was a problem hiding this comment.
Forcing this inline also pulls it into the ambiguous-entity lookahead in find_entity_for_char(), which is cold. Do you have the overall html.o text size delta?
There was a problem hiding this comment.
Thanks for the fuzzing and the LIMIT_ALL catch, I'll fix the description.
html.o text grows by 2.2 KB (37,445 → 39,620): the big loop goes from 3.7 to 6.7 KB, the standalone get_next_char (1.5 KB) disappears, php_next_utf8_char grows to 757 bytes.
The lookahead copy is ~700 bytes of that. I tried keeping it out of line via a never_inline wrapper, but the wrapper alone is 1.5 KB, so the object only gets bigger, and that
build was slower across most of the suite (GCC reshuffles the registers in the big function). So I'd rather keep it as is.
With the UTF-8 charset, the `html` strategy calls `htmlspecialchars()`, which decodes the string one character at a time: about 5 ns per byte, so 1.6 µs for the 280-character class list of a styled button, even when nothing needs escaping. On valid UTF-8, all it does is replace `&`, `"`, `'`, `<` and `>`: `ENT_SUBSTITUTE` only matters for invalid sequences.
Strings longer than 32 bytes are now checked for UTF-8 validity with `preg_match('//u')` and escaped with `strtr()` on those five characters. Invalid UTF-8 and shorter strings still go through `htmlspecialchars()`, whose fixed cost is lower below that length. The output is identical: compared against `htmlspecialchars()` on every Unicode code point (in three positions) and on invalid sequences (lone continuation bytes, overlong forms, surrogates, code points above U+10FFFF, truncated sequences), 3.3 million strings, no difference on PHP 8.4 and 8.5. The shortcut only applies on PHP < 8.7 (`PHP_VERSION_ID < 80700`), since php/php-src#23957 and php/php-src#24145 (both still open, targeting PHP 8.7) make `htmlspecialchars()` itself skip the work on input that needs no encoding.
Benchmark below, PHP 8.4, 3 interleaved runs each side, identical output, ns per escape:
| String | Before | After |
| --- | --- | --- |
| short id (`user-42`) | **155-165 ns** | **168-175 ns** (below the threshold; the length check costs ~10 ns) |
| 40-character title | **327-334 ns** | **269-285 ns** |
| 280-character CSS class list | **1590-1630 ns** | **498-511 ns** (3.2x) |
| paragraph with quotes and tags | **1634-1663 ns** | **758-766 ns** (2.2x) |
| accented paragraph | **1310-1321 ns** | **373-380 ns** (3.5x) |
End to end, on an admin dashboard built with Symfony UX components that escapes about 2,500 attribute values per request (mostly long Tailwind class lists), CPU time per request goes from **36.7-38.7 ms** to **35.7-36.4 ms** (minimum of 6 runs of 60 requests each side), about 3%.
```php
<?php
// Run from the repository root: php bench.php
require getcwd().'/vendor/autoload.php';
$escaper = new Twig\Runtime\EscaperRuntime();
$classes = 'inline-flex shrink-0 items-center justify-center rounded-lg border border-transparent bg-clip-padding text-sm font-medium whitespace-nowrap transition-all outline-none select-none focus-visible:border-ring focus-visible:ring-3 focus-visible:ring-ring/50 disabled:pointer-events-none disabled:opacity-50';
$cases = [
'short (id)' => 'user-42',
'title (40 chars)' => 'How we made our dashboard ten times faster',
'CSS classes (280 chars)' => $classes,
'paragraph with quotes' => str_repeat('Twig escapes "quotes" & <tags> in a sentence like this one. ', 5),
'accented paragraph' => str_repeat('Une phrase accentuée, déjà très répétée. ', 6),
];
foreach ($cases as $label => $string) {
$times = [];
for ($run = 0; $run < 7; ++$run) {
$start = hrtime(true);
for ($i = 0; $i < 200000; ++$i) {
$escaper->escape($string);
}
$times[] = (hrtime(true) - $start) / 200000;
}
sort($times);
printf("%-24s %5.0f ns/escape | sha1 %s\n", $label, $times[3], sha1($escaper->escape($string)));
}
```
Blackfire profiles (Symfony UX dashboard):
- Before: https://app.blackfire.io/envs/5f4f9a62-eaa0-45ee-b7b3-a1b879f550e9/profiles/b344f602-a21f-41ef-a585-46fc865e0880/graph
- After: https://app.blackfire.io/envs/5f4f9a62-eaa0-45ee-b7b3-a1b879f550e9/profiles/e5e16b8b-8895-40e7-9952-43282e3bd0df/graph
- Comparison: https://app.blackfire.io/envs/5f4f9a62-eaa0-45ee-b7b3-a1b879f550e9/profiles/compare/b344f602-a21f-41ef-a585-46fc865e0880...e5e16b8b-8895-40e7-9952-43282e3bd0df/graph
Blackfire barely moves (**2.29 s** -> **2.30 s** overall) because its per-call instrumentation dominates calls this short, even though `htmlspecialchars()` drops out of the profile entirely (5,398 calls per request before).
This came out of profiling Symfony UX components on https://github.com/Kocal/sf-ux-perfs-twig-components.
This is the first part of #18126, split into a separate PR.
htmlspecialchars()currently callsget_next_char()for every input character and then checks the entity table. This is unnecessary for most ASCII input.This changes the loop in
php_escape_html_entities_ex()so that:get_next_char()is only called for bytes >=0x80. ASCII is a single byte in all supported charsets.htmlspecialchars()withoutENT_DISALLOWED, ordinary ASCII bytes are copied directly until one of& " ' < >is found.get_next_char()is markedzend_always_inline.The ASCII copy path is not used with
ENT_DISALLOWED, since control characters still have to be checked, nor byhtmlentities(), since HTML5 has entities for other ASCII characters too. The exception ishtmlentities()withENT_XML1or a multibyte charset other than UTF-8:LIMIT_ALLreduces those to the basic entities, so they take the copy path as well.get_next_char()is large enough that compilers kept it out of line. Forcing it inline removes the call overhead for non-ASCII input as well. This also follows @bukka's suggestion in #18126 to optimize the existing decoder rather than add another one.Benchmarks
Release builds, master vs this PR:
Each row groups several inputs of the
htmlspecialcharssuite; the per-input numbers are in the report linked below.htmlspecialchars()input&/ only"In Symfony Demo (
benchmark/benchmark.php, callgrind),htmlspecialchars()andhtmlentities()are called 221 times per request and account for about 1.3% of the instructions.This patch reduces the instruction count for the whole request by 0.87% (39.48M → 39.14M), or 1.01% with the tracing JIT. The average
htmlspecialchars()call goes from 2291 to 719 instructions.The WordPress benchmark calls
htmlspecialchars()only six times per request, so there is no measurable change there.Other effects
htmlentities()shares the same loop. Apart from the cases above, it does not use the ASCII copy path, but avoiding the decoder for ASCII still makes it about 1.2–1.4× faster.There is also a small improvement in
json_encode()for non-ASCII strings. ext/json usesphp_next_utf8_char(), which wrapsget_next_char(cs_utf_8, ...). After inlining, the compiler can eliminate the charset switch and the unused decoders.This increases the wrapper from 23 to 757 bytes and saves about 13–14 instructions per non-ASCII character:
json_encode()inputjson_decode()ASCII input does not use this decoder, and
json_decode()has its own scanner.utf8_decode()and ext/xml use the same wrapper, but I did not benchmark them.Verification
I compared the output against master for 20,000 random inputs using 13 flag combinations, 14 charsets, both functions, and both values of
double_encode. The output was identical.Full results and reproduction commands:
https://github.com/ArtUkrainskiy/php-src-bench/tree/main/reports/htmlspecialchars-ascii-fast-path
The ASCII copy loop still has one branch per byte. A SIMD version using
zend_simd.hwill be a separate PR.