Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
41 changes: 41 additions & 0 deletions .changeset/20335-pg-aggregate-numbers.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
---
'@objectstack/driver-sql': patch
---

fix(driver-sql): `count` / `count_distinct` / `sum` / `avg` answer JS numbers on PostgreSQL and MySQL, as they do on SQLite and on the engine's rows path

Clause-②: no

`SqlDriver.aggregate` handed the SQL client's answer straight through. node-postgres parses
`bigint` (`count`, and `sum` over an integer column) and `numeric` (`sum` / `avg` over the
numeric family's exact-decimal column, `avg` over an integer column) to strings, and mysql2
does the same for `DECIMAL` (`SUM` / `AVG`). So one grouped query answered

{ "n": "2", "total": "500.000000000000000000000000000000" }

on PostgreSQL's native path and `{ "n": 2, "total": 500 }` on SQLite and on the rows path of
every dialect. The engine's `having` compares values as they
arrive, so `having { n: { $in: [2] } }` kept no group on PostgreSQL alone, and
`having { total: { $in: [500, 20] } }` kept no group on PostgreSQL or MySQL, while a string
comparand such as `{ total: { $lt: 'not-a-date' } }` kept every group there and none anywhere
else.

Those four functions now answer a JS number on every dialect, through `SqlDriver.aggregate`,
`engine.aggregate` and `POST /api/v1/data/:object/query`. The presentation is keyed on the
aggregate function the query asked for; it only rewrites a string, so SQLite's answers are
byte-identical to before. `min` / `max` are unchanged: they answer a value of the column and
keep that column's presentation (a declared numeric field was already a number).
Non-aggregate reads (`find()`, `distinct()`) are unchanged, and no connection-level type
parser is touched.

**Precision policy.** The answer is one JS number (an IEEE-754 double) on every dialect. A
`sum` / `avg` whose exact value needs more than a double's 15 to 17 significant digits, or an
integer total at or above 2^53, is rounded to the nearest double. That is the same bound
`find()` already puts on a read of the same exact-decimal column, and the bound the rows path
has always had. A total that fits keeps its exact value (`500`, `30.75`, `0.375`). The answer
is never a string, including for large totals: an answer whose type depended on its size would
break the same `having` or chart for exactly those totals.

A consumer that read these values through `Number(...)` gets the same number it computed
before. A consumer that compared them as strings, or checked `typeof value === 'string'`,
now receives a number.
Original file line number Diff line number Diff line change
@@ -0,0 +1,208 @@
// Copyright (c) 2026 ObjectStack. Licensed under the Apache-2.0 license.

/**
* [#20335] `count` / `count_distinct` / `sum` / `avg` answer a JS NUMBER from
* `SqlDriver.aggregate`, on every dialect — the value the engine's rows path
* and SQLite already answered.
*
* Measured on the base (`26daf0b036`) through this door, `engine.aggregate` and
* `POST /api/v1/data/:object/query`, `groupBy` customer, four groups:
*
* | dialect | `count` | `count_distinct` | `sum` / `avg` over number, currency, percent | `sum` / `avg` over rating (integer column) |
* |:--|:--|:--|:--|:--|
* | SQLite | number | number | number | number |
* | PostgreSQL 16.13 | `"2"` | `"2"` | `"500.000000000000000000000000000000"` | `"7"` / `"3.5000000000000000"` |
* | MySQL 8.0.46 | number | number | `"500.000000000000000000000000000000"` | `"7"` / `"3.5000"` |
*
* node-postgres parses `bigint` (OID 20) and `numeric` (OID 1700) to strings,
* and mysql2 does the same for `DECIMAL`; `min` / `max` over a declared numeric
* field were already numbers (the column's own `'number'` presentation, #16318)
* and are unchanged. The engine's `having` then compared `"2"` against `2`:
* `having { n: { $in: [2] } }` kept no group on PostgreSQL's native path and
* c1, c2 everywhere else (pinned at the engine and REST doors in
* `@objectstack/rest`'s `rest-aggregate-numeric-having.test.ts`).
*
* The fixture's values are dyadic fractions on purpose, so a JS double holds
* every sum and average EXACTLY: the expected values below are computed from the
* rows with JS arithmetic — the rows path's own arithmetic — and asserted with
* `toBe`, so a string, a boolean or a rounding difference each fail.
*
* The precision policy (one JS double, the loss beyond a double's precision
* declared — `AGGREGATE_ANSWER_KIND` in `sql-driver.ts`) is pinned by the last
* case of each cell: a total the exact-decimal column holds but a double cannot
* answers the nearest double, which is also what `find()` reads for the row.
*/

import { describe, it, expect, beforeAll, afterAll } from 'vitest';
import type { DriverQuery } from '@objectstack/spec/contracts';
import { SqlDriver } from './sql-driver.js';
import { DIALECT_CELLS, declareDialectCell, type DialectCell } from './live-dialect-matrix.testkit.js';

const TABLE = 'os20335_agg_numbers';

interface Row {
id: string;
customer_id: string;
amount: number;
price: number;
rate: number;
stars: number;
flag: boolean;
note: string;
}

const ROWS: readonly Row[] = [
{ id: 'o1', customer_id: 'c1', amount: 100, price: 10.25, rate: 0.25, stars: 3, flag: true, note: 'b' },
{ id: 'o2', customer_id: 'c1', amount: 400, price: 20.5, rate: 0.5, stars: 4, flag: false, note: 'a' },
{ id: 'o3', customer_id: 'c2', amount: 900, price: 30.75, rate: 0.75, stars: 5, flag: true, note: 'c' },
{ id: 'o4', customer_id: 'c2', amount: 300, price: 40, rate: 0.125, stars: 2, flag: true, note: 'c' },
{ id: 'o5', customer_id: 'c3', amount: 50, price: 5.5, rate: 0.5, stars: 1, flag: false, note: 'e' },
{ id: 'o6', customer_id: 'c4', amount: 20, price: 1.25, rate: 0.25, stars: 5, flag: false, note: 'f' },
];

/** number (integer-valued), currency, percent, and the one integer column of the family. */
const MEASURED = ['amount', 'price', 'rate', 'stars'] as const;
const GROUPS = ['c1', 'c2', 'c3', 'c4'] as const;

const byGroup = (g: string) => ROWS.filter((r) => r.customer_id === g);
const sumOf = (rows: readonly Row[], f: keyof Row) => rows.reduce((a, r) => a + Number(r[f]), 0);

function grouped(): DriverQuery {
const aggregations: Array<Record<string, unknown>> = [
{ function: 'count', alias: 'n' },
{ function: 'count_distinct', field: 'note', alias: 'nd' },
{ function: 'sum', field: 'flag', alias: 'sum_flag' },
{ function: 'avg', field: 'flag', alias: 'avg_flag' },
{ function: 'sum', field: 'spare', alias: 'sum_spare' },
{ function: 'avg', field: 'spare', alias: 'avg_spare' },
{ function: 'min', field: 'note', alias: 'min_note' },
];
for (const f of MEASURED) {
for (const fn of ['count', 'sum', 'avg', 'min', 'max']) aggregations.push({ function: fn, field: f, alias: `${fn}_${f}` });
}
return { groupBy: ['customer_id'], aggregations } as DriverQuery;
}

function declareCell(cell: DialectCell): void {
describe(`[#20335] driver-sql — aggregate counts and totals are numbers (${cell.label})`, () => {
let driver: SqlDriver;
let answers: Map<string, Record<string, unknown>>;

beforeAll(async () => {
driver = new SqlDriver(cell.config());
await driver.execute(`drop table if exists ${TABLE}`).catch(() => {});
await driver.initObjects([
{
name: TABLE,
fields: {
customer_id: { type: 'text' },
amount: { type: 'number' },
price: { type: 'currency' },
rate: { type: 'percent' },
stars: { type: 'rating' },
flag: { type: 'boolean' },
note: { type: 'text' },
// NULL in every row: `sum` folds to 0 (#15546) and `avg` stays null.
spare: { type: 'number' },
},
},
] as never);
for (const row of ROWS) await driver.create(TABLE, { ...row }, { bypassTenantAudit: true });
const rows = (await driver.aggregate(TABLE, grouped())) as Array<Record<string, unknown>>;
answers = new Map(rows.map((r) => [String(r.customer_id), r]));
});

afterAll(async () => {
await driver?.execute(`drop table if exists ${TABLE}`).catch(() => {});
await driver?.disconnect();
});

it('answers the four groups', () => {
expect([...answers.keys()].sort()).toEqual([...GROUPS]);
});

it('count and count_distinct are numbers, equal to the rows', () => {
for (const g of GROUPS) {
const a = answers.get(g)!;
const rows = byGroup(g);
expect(a.n, `${g} count(*)`).toBe(rows.length);
expect(a.nd, `${g} count_distinct(note)`).toBe(new Set(rows.map((r) => r.note)).size);
for (const f of MEASURED) expect(a[`count_${f}`], `${g} count(${f})`).toBe(rows.length);
}
});

it('sum and avg over number, currency, percent and rating are numbers, equal to the rows', () => {
for (const g of GROUPS) {
const a = answers.get(g)!;
const rows = byGroup(g);
for (const f of MEASURED) {
expect(a[`sum_${f}`], `${g} sum(${f})`).toBe(sumOf(rows, f));
expect(a[`avg_${f}`], `${g} avg(${f})`).toBe(sumOf(rows, f) / rows.length);
}
}
});

it('sum and avg over a boolean answer the #11152 numbers strictly', () => {
for (const g of GROUPS) {
const a = answers.get(g)!;
const rows = byGroup(g);
expect(a.sum_flag, `${g} sum(flag)`).toBe(sumOf(rows, 'flag'));
expect(a.avg_flag, `${g} avg(flag)`).toBe(sumOf(rows, 'flag') / rows.length);
}
});

it('an all-NULL aggregand: sum folds to the number 0, avg stays null', () => {
for (const g of GROUPS) {
expect(answers.get(g)!.sum_spare, `${g} sum(spare)`).toBe(0);
expect(answers.get(g)!.avg_spare, `${g} avg(spare)`).toBeNull();
}
});

it('min and max keep the column presentation — numbers for the numeric family, text for text', () => {
for (const g of GROUPS) {
const a = answers.get(g)!;
const rows = byGroup(g);
for (const f of MEASURED) {
expect(a[`min_${f}`], `${g} min(${f})`).toBe(Math.min(...rows.map((r) => Number(r[f]))));
expect(a[`max_${f}`], `${g} max(${f})`).toBe(Math.max(...rows.map((r) => Number(r[f]))));
}
expect(a.min_note, `${g} min(note)`).toBe([...rows.map((r) => r.note)].sort()[0]);
}
});

it('precision policy: a total a double cannot hold answers the nearest double, as find() does', async () => {
// Written by SQL, not by the driver: a JS number could not carry these
// values in the first place, which is the whole point of the case. The
// literals are numeric, so the exact-decimal column stores them exactly
// on PostgreSQL and MySQL (SQLite's REAL column rounds on write).
const EXACT = ['9007199254740993', '12345678901234567.123456789'];
for (const [i, literal] of EXACT.entries()) {
await driver.execute(`insert into ${TABLE} (id, customer_id, amount) values ('p${i}', 'p${i}', ${literal})`);
}
const rows = (await driver.aggregate(TABLE, {
where: { customer_id: { $in: ['p0', 'p1'] } },
groupBy: ['customer_id'],
aggregations: [
{ function: 'sum', field: 'amount', alias: 'total' },
{ function: 'avg', field: 'amount', alias: 'mean' },
],
} as DriverQuery)) as Array<Record<string, unknown>>;
const found = (await driver.find(TABLE, { where: { customer_id: { $in: ['p0', 'p1'] } } })) as Array<
Record<string, unknown>
>;
for (const [i, literal] of EXACT.entries()) {
const row = rows.find((r) => r.customer_id === `p${i}`)!;
expect(typeof row.total, `sum over ${literal} is a number, never a string`).toBe('number');
expect(row.total, `sum over ${literal}`).toBe(Number(literal));
expect(row.mean, `avg over ${literal}`).toBe(Number(literal));
expect(row.total, `the same bound find() reads for ${literal}`).toBe(
found.find((r) => r.customer_id === `p${i}`)!.amount,
);
}
});
});
}

for (const cell of DIALECT_CELLS) {
declareDialectCell(cell, 'aggregate numeric presentation (#20335)', declareCell);
}
85 changes: 80 additions & 5 deletions packages/drivers/driver-sql/src/sql-driver.ts
Original file line number Diff line number Diff line change
Expand Up @@ -1476,6 +1476,64 @@ const SQL_AGGREGATE_FUNCTIONS: ReadonlyMap<string, SqlAggregateLowering> = new M
['count_distinct', { sql: 'count', distinct: true }],
]);

/**
* [#20335] What each declared aggregate function ANSWERS — a derived `number`,
* or a value OF the aggregated column — and therefore which read presentation
* {@link SqlDriver.aggregate} gives its result column.
*
* - `'number'` — `count`, `count_distinct`, `sum`, `avg`. A count or a total is
* a number whatever the column held, and it is presented as one (`'number'`,
* the presenter `formatOutput` applies to a numeric field on a `find()` row).
* - `'column'` — `min`, `max`. The answer is one of the column's own values, so
* it takes that column's presentation ({@link SqlDriver.readPresentationKind}),
* exactly as before this table existed.
*
* Why the `'number'` half needs presenting at all: the SQL client hands a
* result back as the wire type of the SQL expression, not as the platform's
* value type. Measured on live PostgreSQL 16.13 and MySQL 8.0.46 through this
* driver's own connections: node-postgres parses `bigint` (OID 20 — `count`,
* and `sum` over an integer column) and `numeric` (OID 1700 — `sum` / `avg`
* over the exact-decimal numeric family, `avg` over an integer column) to
* STRINGS (`"2"`, `"500.000000000000000000000000000000"`), and mysql2 does the
* same for `DECIMAL` (`SUM` / `AVG`; its `COUNT` arrives as a number). The
* engine's rows path (`objectql`'s `in-memory-aggregation.ts`) and SQLite answer
* numbers for the same query, so `having { n: { $in: [2] } }` kept c1, c2 on
* those and no group on PostgreSQL's native path.
*
* Keyed on the function the query ASKED for, never on whether a value looks
* numeric, and deliberately not gated by dialect: the presenter only rewrites a
* STRING, so a client that already answers a number (better-sqlite3, mysql2's
* `COUNT`) passes through untouched — measured byte-identical on SQLite — and
* no list of "string-answering dialects" exists to drift.
*
* ## The precision policy — one JS number, the loss declared
*
* The answer is `Number(text)`: an IEEE-754 double, on every dialect. A `sum` /
* `avg` over the exact-decimal column (`numeric(65,30)` / `DECIMAL(65,30)`)
* whose value needs more than a double's ~15-17 significant digits, or an
* integer at or above 2^53, is ROUNDED to the nearest double — declared, not
* silent: it is the same bound `formatOutput` already puts on a `find()` read of
* that column (#16318, `valueSchemaFor`'s `z.number().finite()`, ADR-0104 D1),
* and the bound the rows path has always had (`toNumber` sums JS doubles). A
* value-dependent type — a number when it fits, a string when it does not — was
* rejected: it would reopen this defect for exactly the large totals, where a
* `having` `$in` or a chart silently stops matching. Only a string `Number()`
* reads as NaN (PostgreSQL's `numeric` `'NaN'`) is left as written, the
* presenter's existing rule.
*
* A `Record` over `AggregationFunction` on purpose: a function that joins the
* declared vocabulary without an answer here fails `tsc` rather than reaching a
* caller unpresented.
*/
const AGGREGATE_ANSWER_KIND: Readonly<Record<AggregationFunction, 'number' | 'column'>> = {
count: 'number',
count_distinct: 'number',
sum: 'number',
avg: 'number',
min: 'column',
max: 'column',
};

/**
* [#5907] The aggregate vocabulary the Query Protocol DECLARES, read from the
* spec rather than restated — `AggregationNodeSchema.function` is this enum, so
Expand Down Expand Up @@ -9858,8 +9916,10 @@ export class SqlDriver implements IDataDriver {
// GROUP BY bucket expression (#3773) and the result presentation (#3797).
const table = this.coercionKey(builder);

// Result columns that carry a raw column VALUE (rather than a count/total
// derived from one), keyed by the column name the caller will read.
// Result columns and the presentation each takes, keyed by the column name
// the caller will read: a raw column VALUE (a group key, `min`/`max`) takes
// its column's presentation, and [#20335] a count or total derived from one
// takes the `'number'` presentation (see `AGGREGATE_ANSWER_KIND`).
// Collected while the statement is built because that is the only point
// where a column name and its meaning are both known: a `min()` lands under
// its alias (never under the field name), and a date-BUCKETED column lands
Expand Down Expand Up @@ -10009,14 +10069,25 @@ export class SqlDriver implements IDataDriver {
// their NULL passes through. See {@link foldEmptyAggregateAnswers}.
const identity = emptyGroupValueFor(funcName);
if (identity !== undefined) foldedOutput.set(agg.alias, identity);
// [#20335] A count or a total is presented as the number it is, on
// every dialect: node-postgres hands `bigint` / `numeric` back as a
// string and mysql2 `DECIMAL`, so without this the native path
// answered `"2"` where the rows path and SQLite answer `2`. Keyed on
// the function asked for; the precision policy (one JS double, the
// loss above a double's precision declared) is stated on
// `AGGREGATE_ANSWER_KIND`. The fold above runs first, so a folded
// `0` is already a number and passes through.
if (AGGREGATE_ANSWER_KIND[funcName] === 'number') {
presentedOutput.set(agg.alias, 'number');
}
// `min`/`max` are the only supported functions that hand back a value
// OF the column rather than a count/total derived from it, so they are
// the only ones whose result still needs the column's presentation.
// `alias` is required by `AggregationNodeSchema`; the unaliased branch
// below lands under a dialect-dependent column name
// (`max("closed_at")` on SQLite, `max` on Postgres) and is defensive
// only, so it is deliberately not tracked.
if ((funcName === 'min' || funcName === 'max') && agg.field) {
if (AGGREGATE_ANSWER_KIND[funcName] === 'column' && agg.field) {
// [#11152] A BOOLEAN aggregand is the ruled exception to "the
// result still needs the column's presentation": the maintainer's
// 2026-08-28 ruling (superseding #11249's `false`/`true`, which
Expand Down Expand Up @@ -15128,7 +15199,9 @@ export class SqlDriver implements IDataDriver {
* the audit-stamp fold run everywhere (`datetime` and `audit_timestamp`
* since #13973, [ADR-0053 D-F1] — the former SQLite-only and the latter
* absent before, which handed the two live dialects' `Date` through), the
* numeric coercion is SQLite-only, and the boolean coercion runs on SQLite
* numeric coercion runs everywhere (#16318 for a declared numeric column;
* [#20335] for `aggregate()`'s counts and totals, `AGGREGATE_ANSWER_KIND`),
* and the boolean coercion runs on SQLite
* and MySQL (#11782 — the two dialects whose stored boolean is a number).
* {@link readPresentationKind} does the dialect gating for the scalar kinds,
* so by the time one arrives here the dialect is settled.
Expand Down Expand Up @@ -15196,7 +15269,9 @@ export class SqlDriver implements IDataDriver {

/**
* Apply {@link presentReadValue} to the result columns a caller of
* `aggregate()` will read as column VALUES — group keys, and `min`/`max`.
* `aggregate()` will read as column VALUES — group keys, and `min`/`max` —
* and [#20335] to its counts and totals (`count`, `count_distinct`, `sum`,
* `avg`), which take the `'number'` presenter (`AGGREGATE_ANSWER_KIND`).
*
* Which columns those are cannot be recovered from the rows — the driver has
* to be told, because the mapping from column name to meaning is only
Expand Down
Loading
Loading