Skip to content

perf(query): parse with SLL prediction first, LL only on a syntax error - #4963

Merged
imilinovic merged 4 commits into
masterfrom
perf/two-stage-parsing
Sep 30, 2026
Merged

imilinovic merged 4 commits into
masterfrom
perf/two-stage-parsing

Conversation

@imilinovic

@imilinovic imilinovic commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

What. The Cypher parser predicts with SLL first and parses again with LL only when SLL fails. Queries that hit an ambiguous grammar decision parse 40–700x faster. The tree and every error message are unchanged.

How. src/query/frontend/opencypher/parser.hpp: the first pass runs PredictionMode::SLL with BailErrorStrategy. On ParseCancellationException the token stream is rewound and the query is parsed with PredictionMode::LL, DefaultErrorStrategy and the existing error listener, which words the syntax error as before. CypherParserTest.ValidQueryNeedsNoFullContextPrediction asserts that valid queries make no full-context prediction.

Why. At each decision the parser first predicts with SLL, which ignores the call stack, so its result is cached. In the default LL mode, a decision SLL finds ambiguous falls back to LL: the lookahead is simulated again with the real call stack, and that result is not cached, so every parse of the query pays for it again. For a true ambiguity the alternatives never separate, so the scan runs long and still ends at the lowest alternative. SLL can resolve the conflict itself by taking the lowest viable alternative at once. LL picks that same alternative unless the call stack rules it out, and then the SLL parse fails and the LL retry runs.

This is ANTLR's documented fast path: The Definitive ANTLR 4 Reference, section 13.7 "Maximizing Parser Speed", and Adaptive LL(*) Parsing (Parr, Harwell, Fisher, OOPSLA 2014), section 3.2. Theorem 6.5 there proves SLL either builds the tree LL would or fails; the grammar has no semantic predicates or actions.

Benchmarks and details are in the comment below.

@imilinovic

imilinovic commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Tracking

  • [Link to Epic/Issue]

Standard development

CI Testing Labels

  • Select the appropriate CI test labels (CI -build=build-name -test=test-suite)

Documentation checklist

  • Add the documentation label
  • Add the bug / feature label
  • Add the milestone for which this feature is intended
    • If not known, set for a later milestone
  • Write a release note, including added/changed clauses
    • perf: Cypher queries parse faster the first time they run, about 170x for a relationship with a property map and about 700x for an exists() pattern. #4963
  • [ Documentation PR link memgraph/documentation#XXXX ]
    • Is back linked to this development PR

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

This PR has potential conflicts with the following other open pull requests which modify the same files:

@imilinovic imilinovic added this to the mg-v3.14.0 milestone Sep 29, 2026
@imilinovic imilinovic added the CI -build=release -test=benchmark Run release build and benchmark on push label Sep 29, 2026
@imilinovic
imilinovic force-pushed the perf/two-stage-parsing branch from 413293a to 6b15c6f Compare September 29, 2026 10:05
@imilinovic imilinovic removed the CI -build=release -test=benchmark Run release build and benchmark on push label Sep 30, 2026
@imilinovic

imilinovic commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor Author

Benchmarks and details

Where LL pays. Parsing master's grammar in LL mode over the 2984 gql_behave statements, 17% make at least one full-context prediction:

rule statements typical shape parse LL → SLL
relationshipDetail 508 relationship with a property map, CREATE ()-[:T {p: 0}]->() ~500 µs → 3 µs
atom 441 a null literal, RETURN sum(null) ~120 µs → 3 µs
existsExpression 61 WHERE exists((n)-[:R]->()) ~7 ms → 10 µs
variableExpansion / relationshipLambda 54 / 29 weighted / all-shortest paths with lambdas ~10 ms → 35 µs

Whole corpus (warm DFA, best of 3): 958 ms LL vs 21 ms SLL. Statements without a full-context prediction are 1.4x faster too. The Movies-graph CREATE script parses in 1.3 ms instead of 147 ms. Three statements fail SLL and take the LL retry, paying one extra parse: two variable-length patterns where an unbounded * is followed directly by a property map ([r *{p: 1}]), and reduce(a=exists(()), …). A bounded pattern with a map (*1..2 {p: 1}) stays on SLL. Outside the corpus, a $param bound (*..$x) also takes the retry, for about 3% on that parse; the stripper keeps literal bounds, so it does not create either shape.

Uncached EXPLAIN of 100 list comprehensions, master → branch:

query master branch
[x IN xs WHERE x:A AND true | x] 76.4 ms 3.9 ms
reduce(… CASE WHEN x:A …) 56.7 ms 6.5 ms
comprehension without a conflict 4.8 ms 3.6 ms

gql_behave suite wall time (Release, 7 interleaved runs each, a fresh server per run, no concurrent build, identical pass/fail counts):

suite master median / best branch median / best change median / best
memgraph_V1 7.50 / 6.73 s 6.12 / 5.92 s −18% / −12% (faster in 6 of 7 runs)
openCypher_M09 3.59 / 3.39 s 3.64 / 3.09 s +1% / −9% (within the ±15% run-to-run noise)

Test time is not the point of the change; the numbers show the saving is visible end to end on the larger suite.

The parser ran in LL prediction mode. There, every decision SLL finds ambiguous
runs a full-context simulation, and a truly ambiguous input pays for one every
time: `[x IN xs WHERE x:A AND true | x]` parses 20x slower than the same
comprehension without the conjunct.

Parsing now starts in SLL mode with a bail-out error strategy, which resolves
such a decision without the simulation. ANTLR guarantees SLL either builds the
tree LL would or fails, so only a failure is parsed again in LL mode, which also
words the syntax error as before.

Uncached EXPLAIN of 100 comprehensions: `x:A AND true | x` 76.4 -> 3.9 ms,
`reduce(... CASE WHEN x:A ...)` 56.7 -> 6.5 ms, a comprehension without a
conflict 4.8 -> 3.6 ms. Error messages are unchanged.
The LL retry was only tested for a syntax error. Two valid shapes fail SLL, a
variable-length pattern with a property map and reduce() over exists(). They
must still parse, and the parse must pay for full-context prediction.
The SLL/LL tree guarantee holds because the grammar has no semantic
predicates or actions; the comment now says so. FullContextPredictions()
counts only LL-pass predictions, and the redundant tokens_.seek(0) is gone,
since Parser::reset() seeks the stream itself.

The fallback test swaps reduce(a = exists(()), ...), which the visitor
rejects, for a `$param` var-length bound that is valid and bails in SLL on a
no-viable-alternative error. The syntax-error test asserts the full message.
@imilinovic
imilinovic force-pushed the perf/two-stage-parsing branch from 832b901 to 7896e51 Compare September 30, 2026 10:45
@imilinovic
imilinovic marked this pull request as ready for review September 30, 2026 10:47
Copilot AI balanced review requested due to automatic review settings September 30, 2026 10:47

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@imilinovic
imilinovic requested a review from as51340 September 30, 2026 10:47
The constructor now tries ParseSLL() and falls back to ParseLL(), in place of
an early return from a try block and an empty catch. Each pass sets its own
prediction mode and error strategy.
@sonarqubecloud

Copy link
Copy Markdown

@imilinovic
imilinovic added this pull request to the merge queue Sep 30, 2026
Merged via the queue into master with commit 1735d12 Sep 30, 2026
25 checks passed
@imilinovic
imilinovic deleted the perf/two-stage-parsing branch September 30, 2026 12:13
@vpavicic vpavicic mentioned this pull request Oct 2, 2026
66 of 75 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants