Repository navigation
perf(query): parse with SLL prediction first, LL only on a syntax error - #4963
Conversation
Tracking
Standard development
CI Testing Labels
Documentation checklist
|
413293a to
6b15c6f
Compare
Benchmarks and detailsWhere LL pays. Parsing master's grammar in LL mode over the 2984 gql_behave statements, 17% make at least one full-context prediction:
Whole corpus (warm DFA, best of 3): 958 ms LL vs 21 ms SLL. Statements without a full-context prediction are 1.4x faster too. The Movies-graph Uncached EXPLAIN of 100 list comprehensions, master → branch:
gql_behave suite wall time (Release, 7 interleaved runs each, a fresh server per run, no concurrent build, identical pass/fail counts):
Test time is not the point of the change; the numbers show the saving is visible end to end on the larger suite. |
The parser ran in LL prediction mode. There, every decision SLL finds ambiguous runs a full-context simulation, and a truly ambiguous input pays for one every time: `[x IN xs WHERE x:A AND true | x]` parses 20x slower than the same comprehension without the conjunct. Parsing now starts in SLL mode with a bail-out error strategy, which resolves such a decision without the simulation. ANTLR guarantees SLL either builds the tree LL would or fails, so only a failure is parsed again in LL mode, which also words the syntax error as before. Uncached EXPLAIN of 100 comprehensions: `x:A AND true | x` 76.4 -> 3.9 ms, `reduce(... CASE WHEN x:A ...)` 56.7 -> 6.5 ms, a comprehension without a conflict 4.8 -> 3.6 ms. Error messages are unchanged.
The LL retry was only tested for a syntax error. Two valid shapes fail SLL, a variable-length pattern with a property map and reduce() over exists(). They must still parse, and the parse must pay for full-context prediction.
The SLL/LL tree guarantee holds because the grammar has no semantic predicates or actions; the comment now says so. FullContextPredictions() counts only LL-pass predictions, and the redundant tokens_.seek(0) is gone, since Parser::reset() seeks the stream itself. The fallback test swaps reduce(a = exists(()), ...), which the visitor rejects, for a `$param` var-length bound that is valid and bails in SLL on a no-viable-alternative error. The syntax-error test asserts the full message.
832b901 to
7896e51
Compare
The constructor now tries ParseSLL() and falls back to ParseLL(), in place of an early return from a try block and an empty catch. Each pass sets its own prediction mode and error strategy.
|



What. The Cypher parser predicts with SLL first and parses again with LL only when SLL fails. Queries that hit an ambiguous grammar decision parse 40–700x faster. The tree and every error message are unchanged.
How.
src/query/frontend/opencypher/parser.hpp: the first pass runsPredictionMode::SLLwithBailErrorStrategy. OnParseCancellationExceptionthe token stream is rewound and the query is parsed withPredictionMode::LL,DefaultErrorStrategyand the existing error listener, which words the syntax error as before.CypherParserTest.ValidQueryNeedsNoFullContextPredictionasserts that valid queries make no full-context prediction.Why. At each decision the parser first predicts with SLL, which ignores the call stack, so its result is cached. In the default LL mode, a decision SLL finds ambiguous falls back to LL: the lookahead is simulated again with the real call stack, and that result is not cached, so every parse of the query pays for it again. For a true ambiguity the alternatives never separate, so the scan runs long and still ends at the lowest alternative. SLL can resolve the conflict itself by taking the lowest viable alternative at once. LL picks that same alternative unless the call stack rules it out, and then the SLL parse fails and the LL retry runs.
This is ANTLR's documented fast path: The Definitive ANTLR 4 Reference, section 13.7 "Maximizing Parser Speed", and Adaptive LL(*) Parsing (Parr, Harwell, Fisher, OOPSLA 2014), section 3.2. Theorem 6.5 there proves SLL either builds the tree LL would or fails; the grammar has no semantic predicates or actions.
Benchmarks and details are in the comment below.