Follow-up from the #403/#271 fix (#613 workstream A2).
What remains
On the query paths, TIMEOUT_EXCEEDED (159), TOO_SLOW (160), MEMORY_LIMIT_EXCEEDED (241) and WaveHouse's own clickhouse.query_timeout deadline are answered 503 clickhouse.unavailable, retryable, with Retry-After: 5 — except on /v1/query when the role set the matching cap (max_execution_time / max_memory_usage), where they are 400 clickhouse.limit_exceeded.
So on /v1/ops/query, on pipes, and on /v1/query for a role without caps, a query that is simply too heavy for the server's limits gets retried by the SDK (maxRetries 2, 5 s apart) and fails the same way each time. That's the doomed-retry shape #271 was about, narrowed to a class where the verdict really is ambiguous: the same codes also come from a busy server (Memory limit (total) exceeded, a timeout on an overloaded node), where a retry is correct.
Measured / inferred
- measured: the mapping above is what
internal/api/ch_errors.go does (unit table in ch_errors_test.go).
- inferred: ClickHouse's message separates
Memory limit (for query) from (total), and a timeout's elapsed time against the configured limit would separate "slow because heavy" from "slow because busy". Nothing reads either today.
Options
- Read the message for 241 (
for query / for user → limit_exceeded, total → unavailable). Brittle across ClickHouse versions.
- Treat
clickhouse.query_timeout expiry as limit_exceeded on every path (the query took the whole budget), keep 241 as unavailable.
- Leave it; the SDK's retries are bounded.
Trigger to act: a dogfooding report of retried timeouts, or the SDK retry counters (#94) showing retries on 503 clickhouse.unavailable that never succeed.
Follow-up from the #403/#271 fix (#613 workstream A2).
What remains
On the query paths,
TIMEOUT_EXCEEDED(159),TOO_SLOW(160),MEMORY_LIMIT_EXCEEDED(241) and WaveHouse's ownclickhouse.query_timeoutdeadline are answered503 clickhouse.unavailable, retryable, withRetry-After: 5— except on/v1/querywhen the role set the matching cap (max_execution_time/max_memory_usage), where they are400 clickhouse.limit_exceeded.So on
/v1/ops/query, on pipes, and on/v1/queryfor a role without caps, a query that is simply too heavy for the server's limits gets retried by the SDK (maxRetries2, 5 s apart) and fails the same way each time. That's the doomed-retry shape #271 was about, narrowed to a class where the verdict really is ambiguous: the same codes also come from a busy server (Memory limit (total) exceeded, a timeout on an overloaded node), where a retry is correct.Measured / inferred
internal/api/ch_errors.godoes (unit table inch_errors_test.go).Memory limit (for query)from(total), and a timeout's elapsed time against the configured limit would separate "slow because heavy" from "slow because busy". Nothing reads either today.Options
for query/for user→ limit_exceeded,total→ unavailable). Brittle across ClickHouse versions.clickhouse.query_timeoutexpiry as limit_exceeded on every path (the query took the whole budget), keep 241 as unavailable.Trigger to act: a dogfooding report of retried timeouts, or the SDK retry counters (#94) showing retries on 503
clickhouse.unavailablethat never succeed.