diff --git a/.changeset/21042-analytics-native-aggregate-policies.md b/.changeset/21042-analytics-native-aggregate-policies.md new file mode 100644 index 00000000000..3b20d7cdd82 --- /dev/null +++ b/.changeset/21042-analytics-native-aggregate-policies.md @@ -0,0 +1,27 @@ +--- +'@objectstack/core': minor +'@objectstack/driver-sql': patch +'@objectstack/service-analytics': patch +--- + +fix: the analytics native-SQL path aggregates with the engine's own aggregate policies, so one query answers one number whichever strategy serves it: `sum` / `avg` accumulate in double, a PostgreSQL boolean aggregand is cast, and an all-NULL `sum` answers `0`. The operand policies move from `@objectstack/driver-sql` to `@objectstack/core` (#21042) + +Clause-②: yes (widening) + +**New exports.** `@objectstack/core` exports the aggregate operand policies, moved here from `@objectstack/driver-sql`, where they were module-private. The driver now imports them and emits byte-identical statements. + +- `AGGREGATE_ACCUMULATION`: what each declared aggregate function accumulates in on PostgreSQL and MySQL. `avg` accumulates in double; `sum` accumulates in double over a fractional column; the counts, `min` and `max` take the column as stored. +- `aggregandColumnClass(shape)`: the one column-class predicate those policies read, over a column's declared `{ type, multiple }`. It answers `'fractional'`, `'integral'`, `'boolean'`, or `undefined` for every other column, a multi-valued one included. The type `AggregandColumnClass` names the three classes. +- `POSTGRES_BOOLEAN_AGGREGAND_CAST`: the functions whose boolean aggregand is cast to `int` on PostgreSQL. These are `sum`, `avg`, `min` and `max`; the two counts are never cast. +- `doubleAccumulationOperand(operand, dialect)`: the column's text, parsed as a double, spelled for `'postgres'` or `'mysql'`. +- `aggregandOperandSql(func, columnClass, dialect, operand)`: the operand an aggregate wraps, with the cast inside the double operand. The type `AggregandSqlDialect` names its dialects (`'sqlite'`, `'postgres'`, `'mysql'`, `'unknown'`). + +**What changed.** `POST /api/v1/analytics/query` and `POST /api/v1/analytics/dataset/query` served by `NativeSQLStrategy` (the default on a SQL driver) skipped three policies `SqlDriver.aggregate()` applies. So the ObjectQL strategy and `engine.aggregate` answered differently for the same query. Measured on SQLite and PostgreSQL 16.13: + +- On PostgreSQL, `sum` / `avg` over an exact-decimal column, and `avg` over an integer one, added exact decimals. For example, `0.1 + 0.2` answered `0.3` and `11 / 9` answered `1.222222222222222`, where the engine answers `0.30000000000000004` and `1.2222222222222223`. The native statement now accumulates in double, as the driver does. +- On PostgreSQL, `sum` / `avg` / `min` / `max` over a boolean field answered `500` (`function sum(boolean) does not exist`). The native statement now casts the boolean aggregand to `int`, as the driver does, and answers the numbers the engine answers. +- On every dialect, a group whose aggregand is NULL in every row, and a measure-scoped `sum` that admits no row, answered `sum` `null` at the cube door. The strategy now folds a `null` answer to `emptyGroupValueFor` (`@objectstack/spec`) for every measure, so that `sum` answers `0`. `avg`, `min` and `max` over nothing stay `null`. The dataset door already answered `0`. + +This is no narrowing: each answer moves to the value the platform already declared for the same query. + +**What did not move.** `@objectstack/driver-sql`'s statements and answers are unchanged: a move-proof test compiles each aggregate function over each column class on SQLite, PostgreSQL and MySQL, and the statements equal the ones captured before the move. SQLite's native statement is unchanged, because neither operand policy applies there. A host that relays no field declarations to the analytics service, or names no SQL dialect, gets today's native arithmetic. diff --git a/packages/core/src/utils/aggregate-answer.test.ts b/packages/core/src/utils/aggregate-answer.test.ts index 687f9885480..f205144291c 100644 --- a/packages/core/src/utils/aggregate-answer.test.ts +++ b/packages/core/src/utils/aggregate-answer.test.ts @@ -11,11 +11,30 @@ * `numeric` to strings), and the large totals are the ones #20335 pinned at the * driver door. Each face pins the same presentation through its own door, in * its own package, on SQLite and PostgreSQL. + * + * [#21042] The operand policies, moved here from `driver-sql` so the + * analytics native-SQL face applies them too: `AGGREGATE_ACCUMULATION` + * (#20387), the column class it and the boolean cast read + * (`aggregandColumnClass`), the PostgreSQL boolean-aggregand cast (#11635) and + * the double operand. The SQL text below is the text `driver-sql` emitted + * before the move (its move proof, `sql-driver-21042-aggregate-policy-move + * .test.ts`, pins the whole statements); each face pins the answers through + * its own door. */ import { describe, it, expect } from 'vitest'; -import { AggregationFunction } from '@objectstack/spec/data'; -import { AGGREGATE_ANSWER_KIND, presentAsNumber } from './aggregate-answer'; +import { AggregationFunction, FieldType } from '@objectstack/spec/data'; +import { + AGGREGATE_ACCUMULATION, + AGGREGATE_ANSWER_KIND, + POSTGRES_BOOLEAN_AGGREGAND_CAST, + aggregandColumnClass, + aggregandOperandSql, + doubleAccumulationOperand, + presentAsNumber, + type AggregandColumnClass, + type AggregandSqlDialect, +} from './aggregate-answer'; describe('[#20889] AGGREGATE_ANSWER_KIND — what each declared aggregate function answers', () => { it('has exactly one row per declared aggregate function', () => { @@ -74,3 +93,148 @@ describe("[#20889] presentAsNumber — the 'number' presenter", () => { expect(presentAsNumber(date)).toBe(date); }); }); + +describe('[#20387, #21042] AGGREGATE_ACCUMULATION — what each declared aggregate function accumulates in', () => { + it('has exactly one row per declared aggregate function', () => { + expect(Object.keys(AGGREGATE_ACCUMULATION).sort()).toEqual([...AggregationFunction.options].sort()); + }); + + it('avg accumulates in double, sum in double over a fractional column, the rest as stored', () => { + expect(AGGREGATE_ACCUMULATION).toEqual({ + count: 'as-stored', + count_distinct: 'as-stored', + sum: 'double-over-fractional', + avg: 'double', + min: 'as-stored', + max: 'as-stored', + }); + }); +}); + +describe('[#11635, #21042] POSTGRES_BOOLEAN_AGGREGAND_CAST — the functions a PostgreSQL boolean aggregand is cast for', () => { + it('has exactly one row per declared aggregate function', () => { + expect(Object.keys(POSTGRES_BOOLEAN_AGGREGAND_CAST).sort()).toEqual([...AggregationFunction.options].sort()); + }); + + it('sum, avg, min and max cast; the two counts never do', () => { + expect(POSTGRES_BOOLEAN_AGGREGAND_CAST).toEqual({ + count: false, + count_distinct: false, + sum: true, + avg: true, + min: true, + max: true, + }); + }); +}); + +describe('[#20387, #11635, #21042] aggregandColumnClass — the one column-class predicate over a declared shape', () => { + const FRACTIONAL = ['number', 'currency', 'percent', 'slider', 'progress', 'summary']; + const INTEGRAL = ['rating']; + const BOOLEAN = ['boolean', 'toggle']; + + it('classes every declared field type: fractional, integral, boolean, or none', () => { + for (const type of FieldType.options) { + const want: AggregandColumnClass | undefined = FRACTIONAL.includes(type) + ? 'fractional' + : INTEGRAL.includes(type) + ? 'integral' + : BOOLEAN.includes(type) + ? 'boolean' + : undefined; + expect(aggregandColumnClass({ type }), type).toBe(want); + } + }); + + it('every member of the three classes is a declared field type, so none of them is a typo', () => { + for (const type of [...FRACTIONAL, ...INTEGRAL, ...BOOLEAN]) expect(FieldType.options, type).toContain(type); + }); + + it("classes driver-sql's internal column aliases: float is fractional, integer and int integral", () => { + expect(aggregandColumnClass({ type: 'float' })).toBe('fractional'); + expect(aggregandColumnClass({ type: 'integer' })).toBe('integral'); + expect(aggregandColumnClass({ type: 'int' })).toBe('integral'); + }); + + it('a multi-valued column (a JSON list) is in no class', () => { + expect(aggregandColumnClass({ type: 'select', multiple: true })).toBeUndefined(); + expect(aggregandColumnClass({ type: 'tags' })).toBeUndefined(); + expect(aggregandColumnClass({ type: 'lookup', multiple: true })).toBeUndefined(); + }); + + it('`multiple` on a type that cannot be multi-valued changes nothing, as storage ignores it', () => { + expect(aggregandColumnClass({ type: 'number', multiple: true })).toBe('fractional'); + expect(aggregandColumnClass({ type: 'rating', multiple: true })).toBe('integral'); + expect(aggregandColumnClass({ type: 'boolean', multiple: true })).toBe('boolean'); + }); + + it('no declaration is no class', () => { + expect(aggregandColumnClass(undefined)).toBeUndefined(); + expect(aggregandColumnClass(null)).toBeUndefined(); + expect(aggregandColumnClass({})).toBeUndefined(); + expect(aggregandColumnClass({ type: 'string' })).toBeUndefined(); + }); +}); + +describe('[#20387, #21042] doubleAccumulationOperand — the column text, parsed as a double', () => { + it('PostgreSQL and MySQL each spell it their way', () => { + expect(doubleAccumulationOperand('"amount"', 'postgres')).toBe('cast(cast("amount" as text) as double precision)'); + expect(doubleAccumulationOperand('`amount`', 'mysql')).toBe('cast(cast(`amount` as char) as double)'); + }); +}); + +describe('[#20387, #11635, #21042] aggregandOperandSql — the operand each face aggregates', () => { + const X = 'COL'; + const pg = (f: AggregationFunction, c: AggregandColumnClass | undefined) => aggregandOperandSql(f, c, 'postgres', X); + const my = (f: AggregationFunction, c: AggregandColumnClass | undefined) => aggregandOperandSql(f, c, 'mysql', X); + const PG_DOUBLE = (x: string) => `cast(cast(${x} as text) as double precision)`; + const MY_DOUBLE = (x: string) => `cast(cast(${x} as char) as double)`; + + it('PostgreSQL: sum over a fractional column and avg over every class accumulate in double', () => { + expect(pg('sum', 'fractional')).toBe(PG_DOUBLE(X)); + expect(pg('avg', 'fractional')).toBe(PG_DOUBLE(X)); + expect(pg('avg', 'integral')).toBe(PG_DOUBLE(X)); + expect(pg('sum', 'integral'), 'an integer total stays exact').toBe(X); + }); + + it('PostgreSQL: a boolean aggregand is cast to int for sum / avg / min / max, inside the double operand', () => { + expect(pg('sum', 'boolean')).toBe(`cast(${X} as int)`); + expect(pg('avg', 'boolean')).toBe(PG_DOUBLE(`cast(${X} as int)`)); + expect(pg('min', 'boolean')).toBe(`cast(${X} as int)`); + expect(pg('max', 'boolean')).toBe(`cast(${X} as int)`); + }); + + it('PostgreSQL: the counts and min / max over a numeric column take the column as stored', () => { + for (const c of ['fractional', 'integral', 'boolean'] as const) { + expect(pg('count', c), `count ${c}`).toBe(X); + expect(pg('count_distinct', c), `count_distinct ${c}`).toBe(X); + } + expect(pg('min', 'fractional')).toBe(X); + expect(pg('max', 'integral')).toBe(X); + }); + + it('MySQL: the same accumulation, and no boolean cast (a boolean is tinyint(1) there)', () => { + expect(my('sum', 'fractional')).toBe(MY_DOUBLE(X)); + expect(my('avg', 'integral')).toBe(MY_DOUBLE(X)); + expect(my('avg', 'boolean')).toBe(MY_DOUBLE(X)); + expect(my('sum', 'boolean')).toBe(X); + expect(my('min', 'boolean')).toBe(X); + expect(my('sum', 'integral')).toBe(X); + }); + + it('a column in no class is aggregated as stored on every dialect', () => { + for (const d of ['postgres', 'mysql', 'sqlite', 'unknown'] as const) { + for (const f of AggregationFunction.options) expect(aggregandOperandSql(f, undefined, d, X), `${d} ${f}`).toBe(X); + } + }); + + it('SQLite and an unnamed dialect apply neither policy: SQLite already adds doubles and stores booleans as 0 / 1', () => { + for (const d of ['sqlite', 'unknown'] as const satisfies readonly AggregandSqlDialect[]) { + for (const f of AggregationFunction.options) { + for (const c of ['fractional', 'integral', 'boolean'] as const) { + expect(aggregandOperandSql(f, c, d, X), `${d} ${f} ${c}`).toBe(X); + } + } + } + }); +}); diff --git a/packages/core/src/utils/aggregate-answer.ts b/packages/core/src/utils/aggregate-answer.ts index 755fe5fbcb2..99c876d2b97 100644 --- a/packages/core/src/utils/aggregate-answer.ts +++ b/packages/core/src/utils/aggregate-answer.ts @@ -30,9 +30,40 @@ * The table and its docblock below moved here from `driver-sql` unchanged, so * they speak of `SqlDriver.aggregate` and `formatOutput` in that driver's * voice. + * + * ## The operand policies (#21042) + * + * The same two faces must also AGGREGATE the same operand, or they answer + * different numbers for one query before any presenter runs: on PostgreSQL the + * analytics native-SQL face added exact decimals where `driver-sql` adds + * doubles (#20387), and answered `500` for a `sum` over a boolean that + * `driver-sql` casts first (#11635). So the operand half moved here too, from + * `driver-sql`, where it was module-private: + * + * - {@link AGGREGATE_ACCUMULATION} (moved with its docblock, in the driver's + * voice) — what each function accumulates in; + * - {@link aggregandColumnClass} — the ONE column-class predicate both + * policies read, over the declared shape `{ type, multiple }`; + * - {@link POSTGRES_BOOLEAN_AGGREGAND_CAST} — the functions a PostgreSQL + * boolean aggregand is cast for; + * - {@link doubleAccumulationOperand} — the double operand, the dialect a + * parameter; + * - {@link aggregandOperandSql} — the three composed, in the one order both + * faces emit: the cast inside, the double operand around it. + * + * `driver-sql` reads its own registries for a column's class (filled by the + * predicate), and the native-SQL face asks the predicate with the declaration + * the analytics host already relays (`declaredValueShape`). The stated cost: + * this package now carries a second piece of dialect SQL text, beside + * `json-membership-sql.ts`. */ -import type { AggregationFunction } from '@objectstack/spec/data'; +import { + isMultiValueField, + numericColumnFor, + NUMERIC_VALUE_TYPES, + type AggregationFunction, +} from '@objectstack/spec/data'; /** * [#20335] What each declared aggregate function ANSWERS — a derived `number`, @@ -118,3 +149,200 @@ export function presentAsNumber(value: unknown): unknown { } return value; } + +/** + * [#20387] What each declared aggregate function ACCUMULATES IN on PostgreSQL + * and MySQL — the arithmetic half of the one-double policy + * {@link AGGREGATE_ANSWER_KIND} states for the answer's type. + * + * - `'double'` — `avg`, over every declared numeric or boolean aggregand. + * - `'double-over-fractional'` — `sum`, over a declared column whose values are + * fractions (`fractionalNumericFields`: the exact-decimal family and the + * driver's `float` alias). A `sum` over an integer-valued column (`rating`, + * the `integer` / `int` aliases, a boolean) stays the database's exact + * integer total, rounded once to the double by the presenter. + * - `'as-stored'` — `count` / `count_distinct` (a count is an exact integer) + * and `min` / `max` (a value OF the column; no arithmetic happens). + * + * Why: the engine's rows path (`objectql`'s `in-memory-aggregation.ts`) and + * SQLite add JS doubles, while PostgreSQL's `numeric` and MySQL's `DECIMAL` + * add exact decimals. Measured on live PostgreSQL 16.13 and MySQL 8.0.46 over + * a `number` column holding `0.1` and `0.2`: `sum` answered `0.3` natively and + * `0.30000000000000004` on SQLite and every rows path, so + * `having { s: { $eq: 0.3 } }` kept the group on those two native faces only. + * `avg` over an INTEGER column diverged too, which is why `avg` is `'double'` + * whatever the column holds: MySQL rounds a `DECIMAL` average to + * `div_precision_increment` (4) places (`avg` of 1, 2, 2 answered `1.6667`), + * and PostgreSQL's `numeric` average rounds to 16 places before the presenter + * rounds again (`11 / 9` answered `1.2222222222222222`, JS + * `1.2222222222222223`; 10 of 27,962 integer pairs measured). + * + * The operand is the column's TEXT, parsed as a double — + * `cast(cast(x as text) as double precision)` on PostgreSQL, + * `cast(cast(x as char) as double)` on MySQL — because that is the value the + * SQL client hands `find()`, and so the value the rows path adds. For an + * exact-decimal column it is the plain cast (both servers convert a decimal to + * a double through its text); for a binary `real` / `FLOAT` column, which a + * table created before the exact-decimal columns still carries, the plain cast + * would widen the binary32 value (`0.1` → `0.10000000149011612`) where the + * client reads `0.1`. MySQL's `CAST(… AS DOUBLE)` needs 8.0.17 or later. + * + * ⚠️ Residual, stated: on PostgreSQL and MySQL the double sums are added in + * scan order, one after another, without compensation. SQLite (3.43+) adds with + * compensated (Kahan-Babuska-Neumaier) summation, and since #20489 so does the + * engine's rows path (`in-memory-aggregation.ts`, `compensatedSum`), so a group + * of three or more fractions can still differ in the last place between the + * PostgreSQL / MySQL native faces and those two (`0.1 + 0.2 + 0.3`: PostgreSQL / + * MySQL `0.6000000000000001`, SQLite and the rows path `0.6`). Two addends + * cannot differ, which is why the pin is `0.1 + 0.2`. + * + * A `Record` over `AggregationFunction` for the same reason as + * {@link AGGREGATE_ANSWER_KIND}: a function added to the vocabulary without an + * answer here fails `tsc`. + */ +export const AGGREGATE_ACCUMULATION: Readonly< + Record +> = { + count: 'as-stored', + count_distinct: 'as-stored', + sum: 'double-over-fractional', + avg: 'double', + min: 'as-stored', + max: 'as-stored', +}; + +/** + * [#21042] The class of an aggregated column, as the operand policies read it: + * + * - `'fractional'` — its values are fractions: the exact-decimal members of + * `NUMERIC_COLUMN_REPRESENTATION` (`numericColumnFor(type).kind === + * 'exact'`: `number`, `currency`, `percent`, `slider`, `progress`, + * `summary`) and `driver-sql`'s `float` alias. `sum` accumulates in double + * over these ({@link AGGREGATE_ACCUMULATION}). + * - `'integral'` — its values are integers: `rating`, and `driver-sql`'s + * `integer` / `int` aliases (how an introspected `bigint` reaches it). `sum` + * keeps the exact total; `avg` accumulates in double. + * - `'boolean'` — `boolean` / `toggle`. `avg` accumulates in double; on + * PostgreSQL every function but the counts casts it to `int` first + * ({@link POSTGRES_BOOLEAN_AGGREGAND_CAST}). + * + * Every other column — text, a date, a multi-valued (JSON) list, a column with + * no declaration — is in no class, and every policy leaves it as stored. + */ +export type AggregandColumnClass = 'fractional' | 'integral' | 'boolean'; + +/** + * [#20387, #11635, #21042] The ONE column-class predicate the aggregate operand + * policies read, over a column's declared shape `{ type, multiple }`. + * + * It replaces `driver-sql`'s module-private `isFractionalNumericType` and states + * the classes that driver's registries record for the same declaration: its + * fractional registry is filled by this predicate, and its numeric and boolean + * registries hold exactly the other two classes (`driver-sql`'s move proof + * pins the statements each class emits). SCALAR only: a column whose stored + * value is a list (`isMultiValueField`, the same question `driver-sql`'s + * storage asks) is in no class, and a `multiple` flag on a type that cannot be + * multi-valued changes nothing, as it changes nothing in storage. + * + * The three `driver-sql` aliases ride here because the class is the driver's + * as much as the declaration's: they are not `FieldType`s, and only an + * introspected (federated) column carries one. + */ +export function aggregandColumnClass( + shape: { readonly type?: unknown; readonly multiple?: unknown } | null | undefined, +): AggregandColumnClass | undefined { + const type = shape?.type; + if (typeof type !== 'string') return undefined; + if (isMultiValueField({ type, multiple: shape?.multiple === true })) return undefined; + if (type === 'boolean' || type === 'toggle') return 'boolean'; + if (type === 'float' || numericColumnFor(type)?.kind === 'exact') return 'fractional'; + if (type === 'integer' || type === 'int' || NUMERIC_VALUE_TYPES.has(type)) return 'integral'; + return undefined; +} + +/** + * [#11635] The functions whose BOOLEAN aggregand is cast to `int` on + * PostgreSQL — the one dialect that stores `Field.boolean` as a real `boolean` + * column and defines no `sum` / `avg` / `min` / `max` over it (SQLSTATE + * `42883`, measured on PG 16.13; SQLite stores 0/1 INTEGER and MySQL + * `tinyint(1)`, so both compute natively). The answers are ruled ([#11152]): + * all four answer numbers, `sum` / `avg` arithmetic over 1/0 and `min` / `max` + * the `0` / `1` the cast computes. `count` / `count_distinct` are deliberately + * NOT cast — `count` is defined over boolean everywhere, and their answers + * were correct before the cast existed and must not move. + * + * A `Record` over `AggregationFunction`, as {@link AGGREGATE_ACCUMULATION}: a + * function added to the vocabulary without an answer here fails `tsc`. + */ +export const POSTGRES_BOOLEAN_AGGREGAND_CAST: Readonly> = { + count: false, + count_distinct: false, + sum: true, + avg: true, + min: true, + max: true, +}; + +/** + * The SQL dialects the operand policies are stated for: the three a SQL face + * of the platform emits for, and `'unknown'` — the same four names + * `driver-sql`'s `SqlDialectName` and the analytics compilers' dialect carry. + * Neither policy applies on SQLite (it already adds doubles and stores a + * boolean as 0 / 1) or on a dialect nobody named. + */ +export type AggregandSqlDialect = 'sqlite' | 'postgres' | 'mysql' | 'unknown'; + +/** + * [#20387] The operand of a double-accumulated `sum` / `avg`: the column's + * text, parsed as a double — the value the SQL client hands `find()`, and so + * the value the rows path adds. See {@link AGGREGATE_ACCUMULATION} for why the + * text and not a plain cast. + */ +export function doubleAccumulationOperand(operand: string, dialect: 'postgres' | 'mysql'): string { + return dialect === 'postgres' + ? `cast(cast(${operand} as text) as double precision)` + : `cast(cast(${operand} as char) as double)`; +} + +/** [#20387] Whether `func` accumulates in double over a column of `columnClass` on `dialect`. */ +function accumulatesInDouble( + func: AggregationFunction, + columnClass: AggregandColumnClass | undefined, + dialect: AggregandSqlDialect, +): dialect is 'postgres' | 'mysql' { + if (dialect !== 'postgres' && dialect !== 'mysql') return false; + switch (AGGREGATE_ACCUMULATION[func]) { + case 'double': + return columnClass !== undefined; + case 'double-over-fractional': + return columnClass === 'fractional'; + case 'as-stored': + return false; + } +} + +/** + * [#20387, #11635, #21042] The operand an aggregate wraps, per the policies + * above: `operand` — the column as the caller already spells it (a knex + * identifier binding, a quoted column) — cast to `int` when it is a PostgreSQL + * boolean aggregand of a function {@link POSTGRES_BOOLEAN_AGGREGAND_CAST} + * names, and then wrapped in {@link doubleAccumulationOperand} when + * {@link AGGREGATE_ACCUMULATION} says the function accumulates in double over + * its class on PostgreSQL or MySQL. Otherwise `operand`, unchanged. + * + * One composition, so the faces cannot order the two differently: a PostgreSQL + * boolean `avg` is `cast(cast(cast(x as int) as text) as double precision)` on + * every face. A column in no class (`undefined`) is aggregated as stored — the + * answer a face that cannot read the column's declaration keeps. + */ +export function aggregandOperandSql( + func: AggregationFunction, + columnClass: AggregandColumnClass | undefined, + dialect: AggregandSqlDialect, + operand: string, +): string { + const cast = dialect === 'postgres' && columnClass === 'boolean' && POSTGRES_BOOLEAN_AGGREGAND_CAST[func] + ? `cast(${operand} as int)` + : operand; + return accumulatesInDouble(func, columnClass, dialect) ? doubleAccumulationOperand(cast, dialect) : cast; +} diff --git a/packages/drivers/driver-sql/src/sql-driver-21042-aggregate-policy-move.test.ts b/packages/drivers/driver-sql/src/sql-driver-21042-aggregate-policy-move.test.ts new file mode 100644 index 00000000000..47870fc9d35 --- /dev/null +++ b/packages/drivers/driver-sql/src/sql-driver-21042-aggregate-policy-move.test.ts @@ -0,0 +1,179 @@ +// Copyright (c) 2026 ObjectStack. Licensed under the Apache-2.0 license. + +/** + * [#21042] The move proof for `driver-sql`'s aggregate operand policies: the + * statement `SqlDriver.aggregate()` emits, per dialect, per aggregate function + * and per declared column class, is the one this driver emitted before the + * policies moved. + * + * Three policies were module-private here and moved to `@objectstack/core` + * (`utils/aggregate-answer.ts`), so the analytics native-SQL face applies the + * same rules from one implementation: + * + * - `AGGREGATE_ACCUMULATION` (#20387): `sum` over a fractional column and `avg` + * over every numeric or boolean column accumulate in double on PostgreSQL + * and MySQL, through the column's text; + * - the column class each policy reads, now one predicate over the declared + * shape `{ type, multiple }` (it used to be `isFractionalNumericType` here, + * beside the registry reads); + * - the PostgreSQL boolean-aggregand cast (#11635): `sum` / `avg` / `min` / + * `max` over a boolean column cast it to `int`, the counts never. + * + * Every expected expression below was captured from the driver BEFORE the move + * (base `d34aa58a2a`), compiled through the real `aggregate()` offline: the + * probe captures knex's `toSQL()` where the statement would execute, so no + * connection is opened. Both registration paths fill the registries the + * policies read — a managed object (`registerObjectMetadata`) and a federated + * one (`registerExternalObject`) — and each must emit the same expression. + * + * The columns are one of each class: an exact-decimal `number` and the + * driver's `float` alias (fractional), a `rating` and the `integer` alias + * (integer-valued), `boolean` and `toggle`, and three the policies leave + * alone — `text`, a multi-valued `select` (a JSON column) and a field that + * declares no type. SQLite applies none of the policies; its control row + * holds that. + */ + +import { describe, it, expect, afterAll } from 'vitest'; +import type { DriverOptions } from '@objectstack/spec/data'; +import { SqlDriver, type SqlDriverConfig } from './sql-driver.js'; + +const FIELDS: Record> = { + g: { type: 'text' }, + f_number: { type: 'number' }, + f_float: { type: 'float' }, + i_rating: { type: 'rating' }, + i_integer: { type: 'integer' }, + b_boolean: { type: 'boolean' }, + b_toggle: { type: 'toggle' }, + o_text: { type: 'text' }, + o_select_multi: { type: 'select', multiple: true }, + o_untyped: {}, +}; + +const FUNCTIONS = ['count', 'count_distinct', 'sum', 'avg', 'min', 'max'] as const; +type Fn = (typeof FUNCTIONS)[number]; + +/** The six expressions, in {@link FUNCTIONS} order, one column per row. */ +type Row = readonly [count: string, countDistinct: string, sum: string, avg: string, min: string, max: string]; + +interface Cell { + readonly dialect: 'sqlite' | 'postgres' | 'mysql'; + readonly config: SqlDriverConfig; + /** An identifier as the dialect quotes it. */ + readonly q: (name: string) => string; + /** The captured aggregate expression per column; columns absent here are not pinned on this dialect. */ + readonly expressions: Readonly>; +} + +const pgDouble = (x: string) => `cast(cast(${x} as text) as double precision)`; +const myDouble = (x: string) => `cast(cast(${x} as char) as double)`; + +const CELLS: readonly Cell[] = [ + { + dialect: 'postgres', + config: { client: 'pg', connection: {} } as SqlDriverConfig, + q: (n) => `"${n}"`, + expressions: { + f_number: ['count("f_number")', 'count(distinct "f_number")', `sum(${pgDouble('"f_number"')})`, `avg(${pgDouble('"f_number"')})`, 'min("f_number")', 'max("f_number")'], + f_float: ['count("f_float")', 'count(distinct "f_float")', `sum(${pgDouble('"f_float"')})`, `avg(${pgDouble('"f_float"')})`, 'min("f_float")', 'max("f_float")'], + i_rating: ['count("i_rating")', 'count(distinct "i_rating")', 'sum("i_rating")', `avg(${pgDouble('"i_rating"')})`, 'min("i_rating")', 'max("i_rating")'], + i_integer: ['count("i_integer")', 'count(distinct "i_integer")', 'sum("i_integer")', `avg(${pgDouble('"i_integer"')})`, 'min("i_integer")', 'max("i_integer")'], + b_boolean: ['count("b_boolean")', 'count(distinct "b_boolean")', 'sum(cast("b_boolean" as int))', `avg(${pgDouble('cast("b_boolean" as int)')})`, 'min(cast("b_boolean" as int))', 'max(cast("b_boolean" as int))'], + b_toggle: ['count("b_toggle")', 'count(distinct "b_toggle")', 'sum(cast("b_toggle" as int))', `avg(${pgDouble('cast("b_toggle" as int)')})`, 'min(cast("b_toggle" as int))', 'max(cast("b_toggle" as int))'], + o_text: ['count("o_text")', 'count(distinct "o_text")', 'sum("o_text")', 'avg("o_text")', 'min("o_text")', 'max("o_text")'], + o_select_multi: ['count("o_select_multi")', 'count(distinct "o_select_multi")', 'sum("o_select_multi")', 'avg("o_select_multi")', 'min("o_select_multi")', 'max("o_select_multi")'], + o_untyped: ['count("o_untyped")', 'count(distinct "o_untyped")', 'sum("o_untyped")', 'avg("o_untyped")', 'min("o_untyped")', 'max("o_untyped")'], + }, + }, + { + dialect: 'mysql', + config: { client: 'mysql2', connection: {} } as SqlDriverConfig, + q: (n) => `\`${n}\``, + expressions: { + f_number: ['count(`f_number`)', 'count(distinct `f_number`)', `sum(${myDouble('`f_number`')})`, `avg(${myDouble('`f_number`')})`, 'min(`f_number`)', 'max(`f_number`)'], + f_float: ['count(`f_float`)', 'count(distinct `f_float`)', `sum(${myDouble('`f_float`')})`, `avg(${myDouble('`f_float`')})`, 'min(`f_float`)', 'max(`f_float`)'], + i_rating: ['count(`i_rating`)', 'count(distinct `i_rating`)', 'sum(`i_rating`)', `avg(${myDouble('`i_rating`')})`, 'min(`i_rating`)', 'max(`i_rating`)'], + i_integer: ['count(`i_integer`)', 'count(distinct `i_integer`)', 'sum(`i_integer`)', `avg(${myDouble('`i_integer`')})`, 'min(`i_integer`)', 'max(`i_integer`)'], + b_boolean: ['count(`b_boolean`)', 'count(distinct `b_boolean`)', 'sum(`b_boolean`)', `avg(${myDouble('`b_boolean`')})`, 'min(`b_boolean`)', 'max(`b_boolean`)'], + b_toggle: ['count(`b_toggle`)', 'count(distinct `b_toggle`)', 'sum(`b_toggle`)', `avg(${myDouble('`b_toggle`')})`, 'min(`b_toggle`)', 'max(`b_toggle`)'], + o_text: ['count(`o_text`)', 'count(distinct `o_text`)', 'sum(`o_text`)', 'avg(`o_text`)', 'min(`o_text`)', 'max(`o_text`)'], + o_select_multi: ['count(`o_select_multi`)', 'count(distinct `o_select_multi`)', 'sum(`o_select_multi`)', 'avg(`o_select_multi`)', 'min(`o_select_multi`)', 'max(`o_select_multi`)'], + o_untyped: ['count(`o_untyped`)', 'count(distinct `o_untyped`)', 'sum(`o_untyped`)', 'avg(`o_untyped`)', 'min(`o_untyped`)', 'max(`o_untyped`)'], + }, + }, + { + // CONTROL: no policy applies on SQLite — every column is aggregated as stored. + dialect: 'sqlite', + config: { client: 'better-sqlite3', connection: { filename: ':memory:' }, useNullAsDefault: true }, + q: (n) => `\`${n}\``, + expressions: { + f_number: ['count(`f_number`)', 'count(distinct `f_number`)', 'sum(`f_number`)', 'avg(`f_number`)', 'min(`f_number`)', 'max(`f_number`)'], + i_rating: ['count(`i_rating`)', 'count(distinct `i_rating`)', 'sum(`i_rating`)', 'avg(`i_rating`)', 'min(`i_rating`)', 'max(`i_rating`)'], + b_boolean: ['count(`b_boolean`)', 'count(distinct `b_boolean`)', 'sum(`b_boolean`)', 'avg(`b_boolean`)', 'min(`b_boolean`)', 'max(`b_boolean`)'], + }, + }, +]; + +/** Captures the statement `aggregate()` would execute, and answers no rows. */ +class AggregateProbe extends SqlDriver { + captured: { sql: string; bindings: readonly unknown[] } | null = null; + + protected getBuilder(object: string, options?: DriverOptions) { + const builder = super.getBuilder(object, options); + // `aggregate()` awaits the builder exactly once; answer that await offline. + Object.defineProperty(builder, 'then', { + configurable: true, + value: (resolve: (rows: unknown[]) => unknown) => { + const { sql, bindings } = builder.toSQL(); + this.captured = { sql, bindings: [...bindings] }; + return Promise.resolve(resolve([])); + }, + }); + return builder; + } + + async compile(table: string, fn: Fn | 'count', field: string | undefined) { + this.captured = null; + await this.aggregate(table, { + groupBy: ['g'], + aggregations: [{ function: fn, ...(field ? { field } : {}), alias: 'a' }], + } as never); + return this.captured; + } +} + +const probes: AggregateProbe[] = []; +afterAll(async () => { + for (const probe of probes) await probe.disconnect().catch(() => {}); +}); + +for (const cell of CELLS) { + for (const fill of ['managed', 'external'] as const) { + describe(`[#21042] the aggregate statement is unchanged by the policy move (${cell.dialect}, ${fill} registration)`, () => { + const table = `t_${fill}`; + const probe = new AggregateProbe(cell.config); + probes.push(probe); + if (fill === 'managed') probe.registerObjectMetadata([{ name: table, fields: FIELDS }]); + else probe.registerExternalObject({ name: table, fields: FIELDS }); + const frame = (expr: string) => + `select ${cell.q('g')}, ${expr} as ${cell.q('a')} from ${cell.q(table)} group by ${cell.q('g')}`; + + it('the probe compiles for the dialect it names', () => { + expect(probe.dialectName).toBe(cell.dialect); + }); + + it('count(*) is untouched', async () => { + expect(await probe.compile(table, 'count', undefined)).toEqual({ sql: frame('count(*)'), bindings: [] }); + }); + + for (const [column, row] of Object.entries(cell.expressions)) { + it(`${column}: every function emits the pre-move expression`, async () => { + for (const [i, fn] of FUNCTIONS.entries()) { + expect(await probe.compile(table, fn, column), `${fn}(${column})`).toEqual({ sql: frame(row[i]!), bindings: [] }); + } + }); + } + }); + } +} diff --git a/packages/drivers/driver-sql/src/sql-driver.ts b/packages/drivers/driver-sql/src/sql-driver.ts index a7a0da5ef63..7651a8fd553 100644 --- a/packages/drivers/driver-sql/src/sql-driver.ts +++ b/packages/drivers/driver-sql/src/sql-driver.ts @@ -25,6 +25,10 @@ import { AggregationFunction, emptyGroupValueFor } from '@objectstack/spec/data' // presenter its counts and totals take — defined once in core, for this // driver's `aggregate()` and the analytics native-SQL face alike. import { AGGREGATE_ANSWER_KIND, presentAsNumber } from '@objectstack/core'; +// [#20387, #11635, #21042] What each aggregate function's operand ACCUMULATES +// IN, the boolean-aggregand cast, and the one column-class predicate both read +// — defined once in core too, for the same two faces. +import { aggregandColumnClass, aggregandOperandSql, type AggregandColumnClass } from '@objectstack/core'; import { STRUCTURED_JSON_TYPES, FILE_REFERENCE_TYPES, MULTI_OPTION_TYPES, NUMERIC_VALUE_TYPES, isMultiValueField } from '@objectstack/spec/data'; // [#16318] The per-field-type physical representation of the NUMERIC family. // `os generate migration` reads the SAME table, in both of its formats — that @@ -363,14 +367,11 @@ const NUMERIC_SCALAR_TYPES = new Set([ 'integer', 'int', 'float', ]); -/** - * [#20387] Whether a numeric field type's column holds FRACTIONS rather than - * integers: the exact-decimal members of `NUMERIC_COLUMN_REPRESENTATION` and - * the driver's `float` alias. Read into {@link SqlDriver.fractionalNumericFields}. - */ -function isFractionalNumericType(type: string): boolean { - return type === 'float' || numericColumnFor(type)?.kind === 'exact'; -} +// [#20387, #21042] Whether a column holds FRACTIONS — the exact-decimal members +// of `NUMERIC_COLUMN_REPRESENTATION` and this driver's `float` alias — is the +// `'fractional'` class of `aggregandColumnClass` (`@objectstack/core`), the one +// predicate the analytics native-SQL face asks too. Read into +// {@link SqlDriver.fractionalNumericFields}. /** * The builtin audit-timestamp columns every managed object carries. They are @@ -1539,66 +1540,13 @@ const SQL_AGGREGATE_FUNCTIONS: ReadonlyMap = new M // `'number'` presenter, so the analytics native-SQL face presents a count or a // total with this driver's own rule. -/** - * [#20387] What each declared aggregate function ACCUMULATES IN on PostgreSQL - * and MySQL — the arithmetic half of the one-double policy - * {@link AGGREGATE_ANSWER_KIND} states for the answer's type. - * - * - `'double'` — `avg`, over every declared numeric or boolean aggregand. - * - `'double-over-fractional'` — `sum`, over a declared column whose values are - * fractions (`fractionalNumericFields`: the exact-decimal family and the - * driver's `float` alias). A `sum` over an integer-valued column (`rating`, - * the `integer` / `int` aliases, a boolean) stays the database's exact - * integer total, rounded once to the double by the presenter. - * - `'as-stored'` — `count` / `count_distinct` (a count is an exact integer) - * and `min` / `max` (a value OF the column; no arithmetic happens). - * - * Why: the engine's rows path (`objectql`'s `in-memory-aggregation.ts`) and - * SQLite add JS doubles, while PostgreSQL's `numeric` and MySQL's `DECIMAL` - * add exact decimals. Measured on live PostgreSQL 16.13 and MySQL 8.0.46 over - * a `number` column holding `0.1` and `0.2`: `sum` answered `0.3` natively and - * `0.30000000000000004` on SQLite and every rows path, so - * `having { s: { $eq: 0.3 } }` kept the group on those two native faces only. - * `avg` over an INTEGER column diverged too, which is why `avg` is `'double'` - * whatever the column holds: MySQL rounds a `DECIMAL` average to - * `div_precision_increment` (4) places (`avg` of 1, 2, 2 answered `1.6667`), - * and PostgreSQL's `numeric` average rounds to 16 places before the presenter - * rounds again (`11 / 9` answered `1.2222222222222222`, JS - * `1.2222222222222223`; 10 of 27,962 integer pairs measured). - * - * The operand is the column's TEXT, parsed as a double — - * `cast(cast(x as text) as double precision)` on PostgreSQL, - * `cast(cast(x as char) as double)` on MySQL — because that is the value the - * SQL client hands `find()`, and so the value the rows path adds. For an - * exact-decimal column it is the plain cast (both servers convert a decimal to - * a double through its text); for a binary `real` / `FLOAT` column, which a - * table created before the exact-decimal columns still carries, the plain cast - * would widen the binary32 value (`0.1` → `0.10000000149011612`) where the - * client reads `0.1`. MySQL's `CAST(… AS DOUBLE)` needs 8.0.17 or later. - * - * ⚠️ Residual, stated: on PostgreSQL and MySQL the double sums are added in - * scan order, one after another, without compensation. SQLite (3.43+) adds with - * compensated (Kahan-Babuska-Neumaier) summation, and since #20489 so does the - * engine's rows path (`in-memory-aggregation.ts`, `compensatedSum`), so a group - * of three or more fractions can still differ in the last place between the - * PostgreSQL / MySQL native faces and those two (`0.1 + 0.2 + 0.3`: PostgreSQL / - * MySQL `0.6000000000000001`, SQLite and the rows path `0.6`). Two addends - * cannot differ, which is why the pin is `0.1 + 0.2`. - * - * A `Record` over `AggregationFunction` for the same reason as - * {@link AGGREGATE_ANSWER_KIND}: a function added to the vocabulary without an - * answer here fails `tsc`. - */ -const AGGREGATE_ACCUMULATION: Readonly< - Record -> = { - count: 'as-stored', - count_distinct: 'as-stored', - sum: 'double-over-fractional', - avg: 'double', - min: 'as-stored', - max: 'as-stored', -}; +// [#20387, #21042] `AGGREGATE_ACCUMULATION` — what each declared aggregate +// function ACCUMULATES IN on PostgreSQL and MySQL, the arithmetic half of the +// one-double policy — lives in `@objectstack/core` (`utils/aggregate-answer.ts`) +// with its docblock, beside the boolean-aggregand cast and the column-class +// predicate both read. {@link SqlDriver.aggregate} asks them through +// `aggregandOperandSql`, so the analytics native-SQL face accumulates the same +// operand with this driver's own rule. /** * [#5907] The aggregate vocabulary the Query Protocol DECLARES, read from the @@ -10041,25 +9989,19 @@ export class SqlDriver implements IDataDriver { // #11249's `false`/`true` for the order statistics) pins ALL FOUR as // numbers: `sum`/`avg` arithmetic over 1/0, `min`/`max` the `0`/`1` // the cast computes, presented as-is (see the presentation note - // below). `cast(?? as int)` keeps the column in a knex identifier - // binding exactly as the uncast form does. `count`/`count_distinct` - // are deliberately NOT cast (both lower to `count`, defined over - // boolean everywhere — their answers were correct before this and - // must not move). - const castBooleanAggregand = - this.isPostgres && - lowering.sql !== 'count' && - fieldExpr !== '*' && - table !== null && - (this.booleanFields[table]?.includes(fieldExpr) ?? false); - const columnExpr = castBooleanAggregand ? 'cast(?? as int)' : '??'; + // below). `count`/`count_distinct` are deliberately NOT cast (both + // lower to `count`, defined over boolean everywhere — their answers + // were correct before this and must not move). // [#20387] `sum` / `avg` accumulate in double on PostgreSQL and MySQL, // the arithmetic SQLite and the engine's rows path already use, so one // query answers one number on every face (`AGGREGATE_ACCUMULATION`). - // Still one `??` binding: the wrap is SQL text around it. - const argExpr = fieldExpr !== '*' && this.accumulatesInDouble(funcName, table, fieldExpr) - ? this.doubleAccumulationOperand(columnExpr) - : columnExpr; + // [#21042] Both policies, and the order they compose in, are + // `aggregandOperandSql` (`@objectstack/core`), read with this column's + // class from the registries below — the rule the analytics native-SQL + // face applies too. Still one `??` binding: the cast and the wrap are + // SQL text around it, so the column stays a knex identifier binding. + const aggregandClass = fieldExpr === '*' ? undefined : this.aggregandColumnClassOf(table, fieldExpr); + const argExpr = aggregandOperandSql(funcName, aggregandClass, this.dialectName, '??'); const rawFunc = lowering.distinct ? `${lowering.sql}(distinct ${argExpr})` : `${lowering.sql}(${argExpr})`; @@ -10181,37 +10123,26 @@ export class SqlDriver implements IDataDriver { } /** - * [#20387] Whether {@link aggregate} accumulates this aggregation's operand in - * double — `AGGREGATE_ACCUMULATION` read against the column's declaration. + * [#20387, #11635, #21042] The class {@link aggregate}'s operand policies read + * for one aggregated column (`aggregandColumnClass`, `@objectstack/core`), + * answered from this driver's registries of the column's declaration: + * `booleanFields` is the `'boolean'` class, `fractionalNumericFields` (filled + * by the predicate itself) the `'fractional'` one, and the rest of + * `numericFields` the `'integral'` one. Every other column — and an unknown + * table — is in no class, so it keeps the database's own arithmetic, as + * before. * - * PostgreSQL and MySQL only. SQLite stores the fractional family as REAL + * The policies apply on PostgreSQL and MySQL only (`aggregandOperandSql` + * reads the dialect). SQLite stores the fractional family as REAL * (`ColumnCompiler_SQLite3.prototype.decimal` is `'float'`), so its `sum` / - * `avg` already add doubles. A column this driver has no numeric or boolean - * declaration for keeps the database's own arithmetic, as before. - */ - protected accumulatesInDouble(func: AggregationFunction, table: string | null, field: string): boolean { - if (table === null || !(this.isPostgres || this.isMysql)) return false; - switch (AGGREGATE_ACCUMULATION[func]) { - case 'double': - return (this.numericFields[table]?.includes(field) ?? false) - || (this.booleanFields[table]?.includes(field) ?? false); - case 'double-over-fractional': - return this.fractionalNumericFields[table]?.includes(field) ?? false; - case 'as-stored': - return false; - } - } - - /** - * [#20387] The operand of a double-accumulated `sum` / `avg`: the column's - * text, parsed as a double — the value the SQL client hands `find()`, and so - * the value the rows path adds. See `AGGREGATE_ACCUMULATION` for why the text - * and not a plain cast. + * `avg` already add doubles. */ - protected doubleAccumulationOperand(operand: string): string { - return this.isPostgres - ? `cast(cast(${operand} as text) as double precision)` - : `cast(cast(${operand} as char) as double)`; + protected aggregandColumnClassOf(table: string | null, field: string): AggregandColumnClass | undefined { + if (table === null) return undefined; + if (this.booleanFields[table]?.includes(field)) return 'boolean'; + if (this.fractionalNumericFields[table]?.includes(field)) return 'fractional'; + if (this.numericFields[table]?.includes(field)) return 'integral'; + return undefined; } /** @@ -11478,8 +11409,9 @@ export class SqlDriver implements IDataDriver { if (NUMERIC_SCALAR_TYPES.has(type) && !isMultiValuedColumn(type, field)) numericCols.push(name); // [#16318] The authorable half only — see {@link numericValueFields}. if (NUMERIC_VALUE_TYPES.has(type) && !isMultiValuedColumn(type, field)) numericValueCols.push(name); - // [#20387] See {@link fractionalNumericFields}. - if (isFractionalNumericType(type) && !isMultiValuedColumn(type, field)) fractionalCols.push(name); + // [#20387, #21042] See {@link fractionalNumericFields}: the predicate's + // `'fractional'` class, scalar only by the predicate's own reading. + if (aggregandColumnClass({ type, multiple: field?.multiple }) === 'fractional') fractionalCols.push(name); if (type === 'date') dateCols.push(name); if (type === 'datetime') datetimeCols.push(name); if (type === 'time') timeCols.push(name); @@ -11583,8 +11515,9 @@ export class SqlDriver implements IDataDriver { if (NUMERIC_VALUE_TYPES.has(type) && !isMultiValuedColumn(type, field)) { numericValueCols.push(name); } - // [#20387] See {@link fractionalNumericFields}. - if (isFractionalNumericType(type) && !isMultiValuedColumn(type, field)) { + // [#20387, #21042] See {@link fractionalNumericFields}: the predicate's + // `'fractional'` class, scalar only by the predicate's own reading. + if (aggregandColumnClass({ type, multiple: field?.multiple }) === 'fractional') { fractionalCols.push(name); } if (type === 'date') { diff --git a/packages/rest/src/analytics-dataset-aggregate-policies-door.test.ts b/packages/rest/src/analytics-dataset-aggregate-policies-door.test.ts new file mode 100644 index 00000000000..db2346b126d --- /dev/null +++ b/packages/rest/src/analytics-dataset-aggregate-policies-door.test.ts @@ -0,0 +1,270 @@ +// Copyright (c) 2026 ObjectStack. Licensed under the Apache-2.0 license. + +/** + * [#21042] `POST /api/v1/analytics/dataset/query` answers each aggregate with + * the engine's own aggregate policies, so the route gives one number whichever + * strategy serves it — over a real `SqlDriver`, on SQLite and on PostgreSQL. + * + * ## Measured on the base, through this door + * + * An inline dataset over three groups, each measure asked alone, the route + * mounted twice: once over the analytics service `AnalyticsServicePlugin` + * composes by default (`NativeSQLStrategy` answers, one raw statement per + * query), once narrowed to the ObjectQL strategy (`engine.aggregate`). + * + * | measure | PostgreSQL 16, native | PostgreSQL 16, ObjectQL | SQLite, both | + * |:--|:--|:--|:--| + * | `sum` / `avg` over a `number` holding 0.1 and 0.2 | `0.3` / `0.15` | `0.30000000000000004` / `0.15000000000000002` | the ObjectQL answer | + * | `avg` over a `rating` of seven 1s and two 2s | `1.222222222222222` | `1.2222222222222223` | the ObjectQL answer | + * | `sum` / `avg` / `min` / `max` over a `boolean` | `500` | numbers | numbers | + * + * A group whose aggregand is NULL in every row answered `sum` `0` at this door + * on both faces already — `DatasetExecutor` fills it — and still does: the + * native face now folds it too, and the fill is idempotent on a folded row. + * + * The expected values are the engine's arithmetic over the fixture rows (JS + * doubles, in row order), with the card's literals pinned beside them. + * + * ## The composition, and the dialect axis of THIS file + * + * The analytics service is the one `AnalyticsServicePlugin` composes over a + * real `ObjectQL` engine. The SQLite cell always runs. The PostgreSQL cell runs + * where `OS_TEST_POSTGRES_URL` is set and is a named skip otherwise; no CI step + * provisions that variable for this package, so the live cell is red-capable + * and un-run in CI, and the PR that landed this file carries its local + * PostgreSQL 16 run. The live cell owns its table, dropped before and after. + */ + +import { describe, it, expect, beforeAll, afterAll } from 'vitest'; +import { ObjectQL } from '@objectstack/objectql'; +import { SqlDriver } from '@objectstack/driver-sql'; +import { AnalyticsServicePlugin, type AnalyticsService } from '@objectstack/service-analytics'; +import { RestServer } from './rest-server'; + +const OBJECT = 'rest_dataset_policy_ledger'; + +const LEDGER = { + name: OBJECT, + label: 'Dataset aggregate policy ledger', + fields: { + grp: { name: 'grp', type: 'text' as const }, + frac: { name: 'frac', type: 'number' as const }, + stars: { name: 'stars', type: 'rating' as const }, + flag: { name: 'flag', type: 'boolean' as const }, + }, +}; + +interface Row { + id: string; + grp: string; + frac: number | null; + stars: number | null; + flag: boolean | null; +} + +const ROWS: readonly Row[] = [ + { id: 'f1', grp: 'f', frac: 0.1, stars: 1, flag: true }, + { id: 'f2', grp: 'f', frac: 0.2, stars: 2, flag: false }, + ...[1, 1, 1, 1, 1, 1, 1, 2, 2].map((stars, k) => ({ id: `i${k}`, grp: 'i', frac: 1, stars, flag: k < 7 })), + ...[0, 1, 2].map((k) => ({ id: `n${k}`, grp: 'n', frac: null, stars: null, flag: null })), +]; + +const GROUPS = ['f', 'i', 'n'] as const; +type Group = (typeof GROUPS)[number]; + +/** The inline dataset the request carries — as a Studio preview or a widget posts it. */ +const DATASET = { + name: 'aggregate_policy_inline', + label: 'Aggregate policy inline', + object: OBJECT, + dimensions: [{ name: 'grp', field: 'grp', type: 'string' }], + measures: [ + { name: 'cnt', aggregate: 'count' }, + { name: 'sum_frac', aggregate: 'sum', field: 'frac' }, + { name: 'avg_frac', aggregate: 'avg', field: 'frac' }, + { name: 'sum_stars', aggregate: 'sum', field: 'stars' }, + { name: 'avg_stars', aggregate: 'avg', field: 'stars' }, + { name: 'sum_flag', aggregate: 'sum', field: 'flag' }, + { name: 'avg_flag', aggregate: 'avg', field: 'flag' }, + { name: 'min_flag', aggregate: 'min', field: 'flag' }, + { name: 'max_flag', aggregate: 'max', field: 'flag' }, + ], +}; +type Measure = (typeof DATASET.measures)[number]['name']; + +/** The rows path's arithmetic: JS doubles, added in row order. */ +function engineAnswer(measure: Measure, g: Group): number | null { + const rows = ROWS.filter((r) => r.grp === g); + if (measure === 'cnt') return rows.length; + const column = measure.endsWith('frac') ? 'frac' : measure.endsWith('stars') ? 'stars' : 'flag'; + const values = rows.map((r) => r[column]).filter((v) => v !== null).map(Number); + const sum = values.reduce((a, b) => a + b, 0); + if (measure.startsWith('sum')) return sum; + if (values.length === 0) return null; + if (measure.startsWith('avg')) return sum / values.length; + return measure.startsWith('min') ? Math.min(...values) : Math.max(...values); +} + +interface Cell { + id: 'sqlite' | 'pg'; + label: string; + env: string | null; + config: () => Record | null; +} + +const CELLS: readonly Cell[] = [ + { id: 'sqlite', label: 'sqlite', env: null, config: () => ({ client: 'better-sqlite3', connection: { filename: ':memory:' }, useNullAsDefault: true }) }, + { + id: 'pg', + label: 'live postgres', + env: 'OS_TEST_POSTGRES_URL', + config: () => (process.env.OS_TEST_POSTGRES_URL ? { client: 'pg', connection: process.env.OS_TEST_POSTGRES_URL } : null), + }, +]; + +type Face = 'native' | 'objectql'; + +const quiet = { debug() {}, info() {}, warn() {}, error() {}, child() { return quiet; } }; + +function createMockServer() { + const noop = () => {}; + return { get: noop, post: noop, put: noop, delete: noop, patch: noop, use: noop, listen: async () => {}, close: async () => {} }; +} + +function mockProtocol() { + return { + getDiscovery: async () => ({ version: 'v0', routes: { data: '', metadata: '' } }), + getMetaTypes: async () => [], + getMetaItems: async () => [], + }; +} + +function makeRes() { + const res: any = { + statusCode: 200, + body: undefined as any, + header: () => res, + status: (code: number) => { res.statusCode = code; return res; }, + json: (body: unknown) => { res.body = body; return res; }, + end: () => res, + }; + return res; +} + +describe('[#21042] the oracle reads the card', () => { + it('the engine arithmetic above answers the card literals', () => { + expect(engineAnswer('sum_frac', 'f')).toBe(0.30000000000000004); + expect(engineAnswer('avg_frac', 'f')).toBe(0.15000000000000002); + expect(engineAnswer('avg_stars', 'i')).toBe(1.2222222222222223); + expect(engineAnswer('sum_flag', 'i')).toBe(7); + expect(engineAnswer('avg_flag', 'i')).toBe(0.7777777777777778); + expect(engineAnswer('sum_frac', 'n')).toBe(0); + expect(engineAnswer('avg_frac', 'n')).toBeNull(); + }); +}); + +for (const cell of CELLS) { + const config = cell.config(); + describe.skipIf(!config)( + `[#21042] POST /api/v1/analytics/dataset/query — one number whichever strategy serves it — ${cell.label}${config ? '' : ` (skipped: set ${cell.env} to run this cell)`}`, + () => { + let driver: any; + let engine: ObjectQL; + /** Raw-SQL statements and engine aggregates that read THIS object. */ + const reads = { rawSql: 0, aggregate: 0 }; + const routes: Partial) => Promise<{ status: number; body: any }>>> = {}; + + const dropTable = async () => { + if (cell.id === 'pg') await driver?.execute(`drop table if exists ${OBJECT}`).catch(() => {}); + }; + + const ask = async (face: Face, measure: Measure) => { + const before = { ...reads }; + const res = await routes[face]!({ measures: [measure], dimensions: ['grp'] }); + return { ...res, rawSql: reads.rawSql - before.rawSql, aggregate: reads.aggregate - before.aggregate }; + }; + + beforeAll(async () => { + driver = new SqlDriver(config as any); + await dropTable(); + engine = new ObjectQL({ logger: quiet } as any); + engine.registerDriver(driver, true); + await engine.init(); + engine.registry.registerObject(LEDGER as any); + await engine.syncSchemas(); + for (const row of ROWS) await engine.insert(OBJECT, { ...row } as any); + + const realExecute = (engine as any).execute.bind(engine); + (engine as any).execute = (sql: unknown, opts?: { object?: string }) => { + if (opts?.object === OBJECT) reads.rawSql += 1; + return realExecute(sql, opts); + }; + const realAggregate = engine.aggregate.bind(engine); + (engine as any).aggregate = (object: string, ...rest: unknown[]) => { + if (object === OBJECT) reads.aggregate += 1; + return (realAggregate as any)(object, ...rest); + }; + + for (const [face, caps] of [ + ['native', undefined], + ['objectql', () => ({ nativeSql: false, objectqlAggregate: true, inMemory: false })], + ] as const) { + const registered: Record = {}; + await new AnalyticsServicePlugin({ ...(caps ? { queryCapabilities: caps } : {}) } as any).init({ + getService: (name: string) => (name === 'data' ? engine : registered[name]), + registerService: (name: string, svc: unknown) => { registered[name] = svc; }, + replaceService: (name: string, svc: unknown) => { registered[name] = svc; }, + hook: () => {}, + logger: quiet, + } as never); + const service = registered.analytics as AnalyticsService; + + const rest = new RestServer( + createMockServer() as any, mockProtocol() as any, { api: { requireAuth: false } } as any, + undefined, undefined, undefined, undefined, undefined, undefined, undefined, + undefined, undefined, undefined, undefined, + async () => service, + ); + (rest as any).resolveExecCtx = async () => ({ userId: 'test-user' }); + rest.registerRoutes(); + const route = rest.getRoutes().find((r: any) => r.method === 'POST' && r.path === '/api/v1/analytics/dataset/query'); + expect(route).toBeDefined(); + routes[face] = async (selection) => { + const res = makeRes(); + // What the wire carries: JSON, both ways. + const body = JSON.parse(JSON.stringify({ dataset: DATASET, selection })); + await route!.handler({ method: 'POST', params: {}, headers: {}, body, query: {} } as any, res); + return { status: res.statusCode, body: JSON.parse(JSON.stringify(res.body ?? null)) }; + }; + } + }); + + afterAll(async () => { + await dropTable(); + try { await engine?.destroy(); } catch { /* noop */ } + }); + + for (const { name: measure } of DATASET.measures) { + it(`${measure}: both strategies answer 200 and the engine's number in every group`, async () => { + const native = await ask('native', measure); + const objectql = await ask('objectql', measure); + expect(native.status, `native: ${JSON.stringify(native.body)}`).toBe(200); + expect(objectql.status, `objectql: ${JSON.stringify(objectql.body)}`).toBe(200); + expect(native.rawSql, 'NativeSQLStrategy served the native route').toBeGreaterThanOrEqual(1); + expect(native.aggregate, 'the native route asked no engine aggregate').toBe(0); + expect(objectql.rawSql, 'the ObjectQL route ran no raw statement').toBe(0); + expect(objectql.aggregate, 'the ObjectQL route asked the engine').toBeGreaterThanOrEqual(1); + const n = new Map((native.body.rows as Array>).map((r) => [String(r.grp), r])); + const o = new Map((objectql.body.rows as Array>).map((r) => [String(r.grp), r])); + expect([...n.keys()].sort()).toEqual([...GROUPS]); + expect([...o.keys()].sort()).toEqual([...GROUPS]); + for (const g of GROUPS) { + const want = engineAnswer(measure, g); + expect(n.get(g)![measure], `${g}: native answers ${want}, never ${JSON.stringify(n.get(g)![measure])}`).toBe(want); + expect(o.get(g)![measure], `${g}: objectql answers ${want}`).toBe(want); + } + }); + } + }, + ); +} diff --git a/packages/services/service-analytics/src/__tests__/cube-measure-field-type-door.test.ts b/packages/services/service-analytics/src/__tests__/cube-measure-field-type-door.test.ts index 5298fb00f00..627421ab5c4 100644 --- a/packages/services/service-analytics/src/__tests__/cube-measure-field-type-door.test.ts +++ b/packages/services/service-analytics/src/__tests__/cube-measure-field-type-door.test.ts @@ -249,19 +249,17 @@ for (const cell of CELLS) { // [#21044] The one rule declines the boolean class (three readings // disagree about what the aggregate returns), so the producer's `number` - // stands — the word SQLite's 0 / 1 answer fits. Not run on PostgreSQL's - // native face, where `max(boolean)` does not exist and the statement is - // a 500 for an ACCEPTED pair: a defect of the native lowering, reported - // on its own and not this door's. - it.skipIf(cell.id === 'pg' && face === 'native')( - `${face}: an accepted boolean pair is served and keeps number — the one rule declines it`, - async () => { - const { res, err } = await read(face, CUBE.name, ['max_flag']); - expect(err, err?.message).toBeUndefined(); - expect(res!.rows[0]!.max_flag).toBe(1); - expect(res!.fields.find((f) => f.name === 'max_flag')?.type).toBe('number'); - }, - ); + // stands — the word SQLite's 0 / 1 answer fits. [#21042] Runs on every + // cell and face: PostgreSQL defines no `max(boolean)`, and the native + // face now casts the boolean aggregand to `int` as `driver-sql` does, + // through the one policy both read (`@objectstack/core`), so the + // accepted pair answers `1` there too instead of a 500. + it(`${face}: an accepted boolean pair is served and keeps number — the one rule declines it`, async () => { + const { res, err } = await read(face, CUBE.name, ['max_flag']); + expect(err, err?.message).toBeUndefined(); + expect(res!.rows[0]!.max_flag).toBe(1); + expect(res!.fields.find((f) => f.name === 'max_flag')?.type).toBe('number'); + }); it(`${face}: a suffix-inferred measure is judged as an authored one — on the configured cube and on an inferred cube`, async () => { for (const cube of [CUBE.name, OBJECT]) { diff --git a/packages/services/service-analytics/src/__tests__/native-sql-aggregate-policies.test.ts b/packages/services/service-analytics/src/__tests__/native-sql-aggregate-policies.test.ts new file mode 100644 index 00000000000..bb6f624a16b --- /dev/null +++ b/packages/services/service-analytics/src/__tests__/native-sql-aggregate-policies.test.ts @@ -0,0 +1,303 @@ +// Copyright (c) 2026 ObjectStack. Licensed under the Apache-2.0 license. + +/** + * [#21042] The native-SQL strategy answers an aggregate with the engine's own + * aggregate policies, so one query gives one number whichever strategy serves + * it — at the cube door (`AnalyticsService.query`, what + * `POST /api/v1/analytics/query` relays verbatim) and at the dataset door + * (`AnalyticsService.queryDataset`). + * + * ## Measured on the base, through these doors + * + * The native face skipped three policies `driver-sql`'s own `aggregate()` + * applies, and the engine and the ObjectQL face therefore answered otherwise: + * + * | policy | native face | ObjectQL face / `engine.aggregate` | + * |:--|:--|:--| + * | #20387, double accumulation (PostgreSQL) | `sum` of 0.1 and 0.2 `0.3`, their `avg` `0.15`; `avg` of seven 1s and two 2s `1.222222222222222` | `0.30000000000000004`, `0.15000000000000002`, `1.2222222222222223` | + * | #11635, the boolean-aggregand cast (PostgreSQL) | `sum` / `avg` / `min` / `max` over a boolean: `500` (`function sum(boolean) does not exist`) | numbers, per #11152's ruling | + * | #15546, the empty-sum fold (every dialect) | a group whose aggregand is NULL in every row, and a measure-scoped `sum` with no admitted row, answer `sum` `null` at the cube door | `0` (the dataset door's `DatasetExecutor` fill also answered `0`) | + * + * `avg` / `min` / `max` over nothing stay `null` on every face, as ruled. + * + * ## What the fix does, and what these pins hold + * + * The policies live once, in `@objectstack/core` (`utils/aggregate-answer.ts`), + * and both faces read them: the native compile wraps each measure's column with + * the operand the policy names for its aggregate, its column's declared class + * and the dialect, and the native shaping point folds a `null` answer to + * `emptyGroupValueFor` (`@objectstack/spec`) for every measure, measure-scoped + * ones included, before the number presenter. + * + * Each measure is asked on its own, on each face and at each door, and every + * group's answer is held to the value the engine computes for those rows — so + * a policy one face skips fails exactly that measure's case, and a face that + * errors fails it too. At the dataset door a measure rides with the base + * `cnt`, so the executor's main statement reports every group: a selection of + * measure-scoped measures alone reports only the groups their filter admits. The two faces are also held equal to each other, cell for cell. + * One cell is NOT held to a value: at the dataset door a measure-scoped `avg` + * is absent from a group its supplementary query reported no row for, on both + * faces (no divergence, and no ruled answer); that cell is held only to "both + * faces agree". + * + * ## The dialect axis of THIS file + * + * The SQLite cell always runs; there, the fold is the policy that diverged. The + * PostgreSQL cell runs where `OS_TEST_POSTGRES_URL` is set and is a named skip + * otherwise; no CI step provisions that variable for this package, so the live + * cell is red-capable and un-run in CI, and the PR that landed this file carries + * its local PostgreSQL 16 run. The operand the PostgreSQL / MySQL cells rely on + * is pinned per dialect in CI by `@objectstack/core`'s `aggregate-answer.test.ts` + * and by `driver-sql`'s move proof. MySQL is not a cell here. The live cell owns + * its table, dropped before and after. + */ + +import { describe, it, expect, beforeAll, afterAll } from 'vitest'; +import { ObjectQL } from '@objectstack/objectql'; +import { SqlDriver } from '@objectstack/driver-sql'; +import type { AnalyticsService } from '../analytics-service.js'; +import { AnalyticsServicePlugin } from '../plugin.js'; + +const OBJECT = 'os21042_policy_ledger'; + +const LEDGER = { + name: OBJECT, + label: 'Aggregate policy ledger', + fields: { + grp: { name: 'grp', type: 'text' as const }, + tag: { name: 'tag', type: 'text' as const }, + // An exact-decimal column (`numeric(65,30)` on PostgreSQL): a FRACTIONAL class. + frac: { name: 'frac', type: 'number' as const }, + // An integer column: `sum` keeps the exact total, `avg` accumulates in double. + stars: { name: 'stars', type: 'rating' as const }, + flag: { name: 'flag', type: 'boolean' as const }, + }, +}; + +interface Row { + id: string; + grp: string; + tag: string; + frac: number | null; + stars: number | null; + flag: boolean | null; +} + +const NINE_STARS = [1, 1, 1, 1, 1, 1, 1, 2, 2]; + +const ROWS: readonly Row[] = [ + // `f`: `0.1 + 0.2` — two addends, so no summation order can move the last place. + { id: 'f1', grp: 'f', tag: 'x', frac: 0.1, stars: 1, flag: true }, + { id: 'f2', grp: 'f', tag: 'x', frac: 0.2, stars: 2, flag: false }, + // `i`: 11 / 9 over an integer column; 7 / 9 over a boolean; no `x` tag. + ...NINE_STARS.map((stars, k) => ({ id: `i${k}`, grp: 'i', tag: 'y', frac: 1, stars, flag: k < 7 })), + // `n`: every aggregand NULL in every row. + ...[0, 1, 2].map((k) => ({ id: `n${k}`, grp: 'n', tag: 'z', frac: null, stars: null, flag: null })), +]; + +const GROUPS = ['f', 'i', 'n'] as const; +type Group = (typeof GROUPS)[number]; + +/** + * The dataset both doors read. The cube door reaches its measure-scoped + * filters through the compiled dataset the service registered under its name. + */ +const DATASET = { + name: 'os21042_policy_ds', + label: 'Aggregate policy dataset', + object: OBJECT, + dimensions: [{ name: 'grp', field: 'grp', type: 'string' }], + measures: [ + { name: 'cnt', aggregate: 'count' }, + { name: 'sum_frac', aggregate: 'sum', field: 'frac' }, + { name: 'avg_frac', aggregate: 'avg', field: 'frac' }, + { name: 'sum_stars', aggregate: 'sum', field: 'stars' }, + { name: 'avg_stars', aggregate: 'avg', field: 'stars' }, + { name: 'sum_flag', aggregate: 'sum', field: 'flag' }, + { name: 'avg_flag', aggregate: 'avg', field: 'flag' }, + { name: 'min_flag', aggregate: 'min', field: 'flag' }, + { name: 'max_flag', aggregate: 'max', field: 'flag' }, + // Measure-scoped: only `f` holds an `x` tag. + { name: 'x_cnt', aggregate: 'count', filter: { tag: 'x' } }, + { name: 'x_sum_frac', aggregate: 'sum', field: 'frac', filter: { tag: 'x' } }, + { name: 'x_avg_frac', aggregate: 'avg', field: 'frac', filter: { tag: 'x' } }, + ], +}; +const MEASURES = DATASET.measures.map((m) => m.name); +type Measure = (typeof MEASURES)[number]; + +/** The rows path's own arithmetic: JS doubles, added in row order (two addends at most differ nowhere). */ +function engineAnswer(measure: Measure, g: Group): number | null { + const all = ROWS.filter((r) => r.grp === g); + const admitted = measure.startsWith('x_') ? all.filter((r) => r.tag === 'x') : all; + const column = measure.endsWith('frac') ? 'frac' : measure.endsWith('stars') ? 'stars' : 'flag'; + const values = admitted.map((r) => r[column]).filter((v) => v !== null).map(Number); + const sum = values.reduce((a, b) => a + b, 0); + if (measure === 'cnt' || measure === 'x_cnt') return admitted.length; + if (measure.startsWith('sum') || measure === 'x_sum_frac') return sum; + if (values.length === 0) return null; + if (measure.startsWith('avg') || measure === 'x_avg_frac') return sum / values.length; + if (measure === 'min_flag') return Math.min(...values); + return Math.max(...values); +} + +/** The card's literals, so the oracle above cannot drift with them. */ +const LITERALS: ReadonlyArray = [ + ['sum_frac', 'f', 0.30000000000000004], + ['avg_frac', 'f', 0.15000000000000002], + ['avg_stars', 'i', 1.2222222222222223], + ['avg_flag', 'i', 0.7777777777777778], + ['sum_flag', 'i', 7], + ['min_flag', 'i', 0], + ['max_flag', 'i', 1], + ['sum_frac', 'n', 0], + ['sum_stars', 'n', 0], + ['sum_flag', 'n', 0], + ['avg_frac', 'n', null], + ['max_flag', 'n', null], + ['cnt', 'n', 3], + ['x_sum_frac', 'i', 0], + ['x_cnt', 'i', 0], +]; + +/** The one cell a value is not held for: the dataset door omits it (see the header). */ +const datasetDoorOmits = (measure: Measure, g: Group) => measure === 'x_avg_frac' && g !== 'f'; + +interface Cell { + id: 'sqlite' | 'pg'; + label: string; + env: string | null; + config: () => Record | null; +} + +const CELLS: readonly Cell[] = [ + { id: 'sqlite', label: 'sqlite', env: null, config: () => ({ client: 'better-sqlite3', connection: { filename: ':memory:' }, useNullAsDefault: true }) }, + { + id: 'pg', + label: 'live postgres', + env: 'OS_TEST_POSTGRES_URL', + config: () => (process.env.OS_TEST_POSTGRES_URL ? { client: 'pg', connection: process.env.OS_TEST_POSTGRES_URL } : null), + }, +]; + +type Face = 'native' | 'objectql'; +type Door = 'cube' | 'dataset'; + +const quiet = { debug() {}, info() {}, warn() {}, error() {}, child() { return quiet; } }; + +describe('[#21042] the oracle reads the card', () => { + it('the engine arithmetic above answers the card literals', () => { + for (const [m, g, want] of LITERALS) expect(engineAnswer(m, g), `${m} ${g}`).toBe(want); + }); +}); + +for (const cell of CELLS) { + const config = cell.config(); + describe.skipIf(!config)( + `[#21042] analytics native SQL applies the engine's aggregate policies (${cell.label})${config ? '' : ` (skipped: set ${cell.env} to run this cell)`}`, + () => { + let driver: any; + let engine: ObjectQL; + /** Raw-SQL statements and engine aggregates that read THIS object. */ + const reads = { rawSql: 0, aggregate: 0 }; + const services: Partial> = {}; + + const dropTable = async () => { + if (cell.id === 'pg') await driver?.execute(`drop table if exists ${OBJECT}`).catch(() => {}); + }; + + /** One measure, on one face, at one door: the answer per group, and which strategy served it. */ + const read = async (face: Face, door: Door, measure: Measure) => { + const before = { ...reads }; + const measures = door === 'dataset' && measure !== 'cnt' ? ['cnt', measure] : [measure]; + const selection = { measures, dimensions: ['grp'] }; + const outcome = await (door === 'cube' + ? services[face]!.query({ cube: DATASET.name, ...selection } as any) + : services[face]!.queryDataset(DATASET as any, selection as any) + ).then( + (res) => ({ rows: res.rows as Array>, err: undefined as (Error & { code?: string }) | undefined }), + (err) => ({ rows: [] as Array>, err: err as Error & { code?: string } }), + ); + return { + ...outcome, + byGroup: new Map(outcome.rows.map((r) => [String(r.grp), r])), + rawSql: reads.rawSql - before.rawSql, + aggregate: reads.aggregate - before.aggregate, + }; + }; + + beforeAll(async () => { + driver = new SqlDriver(config as any); + await dropTable(); + engine = new ObjectQL({ logger: quiet } as any); + engine.registerDriver(driver, true); + await engine.init(); + engine.registry.registerObject(LEDGER as any); + await engine.syncSchemas(); + for (const row of ROWS) await engine.insert(OBJECT, { ...row } as any); + + const realExecute = (engine as any).execute.bind(engine); + (engine as any).execute = (sql: unknown, opts?: { object?: string }) => { + if (opts?.object === OBJECT) reads.rawSql += 1; + return realExecute(sql, opts); + }; + const realAggregate = engine.aggregate.bind(engine); + (engine as any).aggregate = (object: string, ...rest: unknown[]) => { + if (object === OBJECT) reads.aggregate += 1; + return (realAggregate as any)(object, ...rest); + }; + + // The plugin's own composition over the real engine. `native`: both + // auto-bridges live, so NativeSQLStrategy answers. `objectql`: narrowed + // to the engine-aggregate path. + for (const [face, caps] of [ + ['native', undefined], + ['objectql', () => ({ nativeSql: false, objectqlAggregate: true, inMemory: false })], + ] as const) { + const registered: Record = {}; + await new AnalyticsServicePlugin({ ...(caps ? { queryCapabilities: caps } : {}) } as any).init({ + getService: (name: string) => (name === 'data' ? engine : registered[name]), + registerService: (name: string, svc: unknown) => { registered[name] = svc; }, + replaceService: (name: string, svc: unknown) => { registered[name] = svc; }, + hook: () => {}, + logger: quiet, + } as never); + const service = registered.analytics as AnalyticsService; + // The configuration door: the cube door reads the dataset's measure filters by name. + service.registerDataset(DATASET as any); + services[face] = service; + } + }); + + afterAll(async () => { + await dropTable(); + try { await engine?.destroy(); } catch { /* noop */ } + }); + + for (const door of ['cube', 'dataset'] as const) { + for (const measure of MEASURES) { + it(`${door} door, ${measure}: both faces answer the engine's number in every group`, async () => { + const native = await read('native', door, measure); + const objectql = await read('objectql', door, measure); + expect(native.err, `native: ${native.err?.code} ${native.err?.message}`).toBeUndefined(); + expect(objectql.err, `objectql: ${objectql.err?.code} ${objectql.err?.message}`).toBeUndefined(); + expect(native.rawSql, 'NativeSQLStrategy served the native face').toBeGreaterThanOrEqual(1); + expect(native.aggregate, 'the native face asked no engine aggregate').toBe(0); + expect(objectql.rawSql, 'the ObjectQL face ran no raw statement').toBe(0); + expect(objectql.aggregate, 'the ObjectQL face asked the engine').toBeGreaterThanOrEqual(1); + expect([...native.byGroup.keys()].sort(), 'native groups').toEqual([...GROUPS]); + expect([...objectql.byGroup.keys()].sort(), 'objectql groups').toEqual([...GROUPS]); + for (const g of GROUPS) { + const n = native.byGroup.get(g)![measure]; + const o = objectql.byGroup.get(g)![measure]; + expect(n, `${g}: the native face equals the ObjectQL face`).toBe(o); + if (door === 'dataset' && datasetDoorOmits(measure, g)) continue; + expect(n, `${g}: native answers the engine's number, never ${JSON.stringify(n)}`).toBe(engineAnswer(measure, g)); + expect(o, `${g}: objectql answers the engine's number`).toBe(engineAnswer(measure, g)); + } + }); + } + } + }, + ); +} diff --git a/packages/services/service-analytics/src/__tests__/native-sql-measure-number-presentation.test.ts b/packages/services/service-analytics/src/__tests__/native-sql-measure-number-presentation.test.ts index 4d7249dd9d4..033dd212dce 100644 --- a/packages/services/service-analytics/src/__tests__/native-sql-measure-number-presentation.test.ts +++ b/packages/services/service-analytics/src/__tests__/native-sql-measure-number-presentation.test.ts @@ -48,10 +48,10 @@ * are computed from the rows with JS arithmetic and asserted with `toBe`: a * string, a boolean or a rounding difference each fail. The precision case * holds #20335's one-double policy — a total beyond a double answers the - * nearest double, equal to SQLite and to the ObjectQL face. ⛔ No case pins the - * SUM / AVG accumulation of non-dyadic fractions (`0.1 + 0.2`, `11 / 9`): the - * native statement does not carry #20387's double accumulation, a different - * defect shape that is tracked on its own. + * nearest double, equal to SQLite and to the ObjectQL face. The SUM / AVG + * accumulation of non-dyadic fractions (`0.1 + 0.2`, `11 / 9`) is a different + * question — what the statement adds, not how its answer is presented — and is + * pinned beside this file in `native-sql-aggregate-policies.test.ts` (#21042). * * ## The dialect axis of THIS file * diff --git a/packages/services/service-analytics/src/strategies/native-sql-strategy.ts b/packages/services/service-analytics/src/strategies/native-sql-strategy.ts index 9e382c7811d..3c20edb4734 100644 --- a/packages/services/service-analytics/src/strategies/native-sql-strategy.ts +++ b/packages/services/service-analytics/src/strategies/native-sql-strategy.ts @@ -30,6 +30,11 @@ import { nextUtcCalendarDay, resolveAnalyticsDateRangeString, isUnboundedAbove } // [#20889] What each aggregate function ANSWERS, and the `'number'` presenter — // the rule `driver-sql`'s own `aggregate()` applies, defined once in core. import { AGGREGATE_ANSWER_KIND, presentAsNumber } from '@objectstack/core'; +// [#21042] What each aggregate's operand accumulates in, the PostgreSQL +// boolean-aggregand cast, and the one column-class predicate both read — the +// operand rule `driver-sql`'s own `aggregate()` applies, defined once in core. +import { aggregandColumnClass, aggregandOperandSql } from '@objectstack/core'; +import { emptyGroupValueFor } from '@objectstack/spec/data'; import { explicitDateRangeWindow } from '../date-range-array-arm.js'; /** @@ -660,6 +665,35 @@ export class NativeSQLStrategy implements AnalyticsStrategy { const rows = await ctx.executeRawSql!(objectName, sql, params); + // [#21042, #15546] A measure whose aggregate answers NULL over nothing — SQL + // `SUM` over a group whose aggregand is NULL in every row, or over no row a + // measure-scoped filter admits — answers the identity the platform declares + // for that aggregate over NOTHING (`emptyGroupValueFor`, spec + // `data/aggregation-policy.ts`): summing nothing is `0`, a measured fact. + // The engine and the ObjectQL face fold it (`driver-sql`'s + // `foldEmptyAggregateAnswers`, the rows path); this face answered `null` for + // the same group. Read from the policy, never restated, for EVERY measure — + // a measure-scoped one carries its aggregate in the same `type` — so + // `avg` / `min` / `max` (no identity) and the expression metric types + // (`undefined` too) keep their NULL. Only `null` folds, before the + // presenter, in `driver-sql`'s order: an `undefined` would be a column + // never projected, a different defect that must stay visible. The dataset + // door's `DatasetExecutor` fill still runs after this and is idempotent on + // a folded row. + const folds: Array = []; + for (const member of query.measures ?? []) { + const identity = emptyGroupValueFor(this.lookupMember(cube, member, 'measure')?.type); + if (identity !== undefined) folds.push([member, identity]); + } + if (folds.length > 0 && Array.isArray(rows)) { + for (const row of rows) { + if (!row || typeof row !== 'object') continue; + for (const [member, identity] of folds) { + if (row[member] === null) row[member] = identity; + } + } + } + // [#20889] A measure column `fields[]` declares `number` answers a number, // on every dialect. The SQL client hands an aggregate back as the wire type // of its expression: node-postgres parses `bigint` (`count`, `sum` over an @@ -771,7 +805,7 @@ export class NativeSQLStrategy implements AnalyticsStrategy { ctx, ) : null; - const aggExpr = this.resolveMeasureSql(cube, measure, tableName, joins, predicate); + const aggExpr = this.resolveMeasureSql(cube, measure, tableName, joins, predicate, ctx); selectClauses.push(`${aggExpr} AS "${measure}"`); } } @@ -1142,13 +1176,17 @@ export class NativeSQLStrategy implements AnalyticsStrategy { * @param predicate - The measure's own scoped filter, already compiled to a * SQL boolean (`null` = the measure declares none, or declares one that * constrains nothing — `compileFilterNode`'s TRUE). #10298. + * @param ctx - [#21042] The host's answers this compile reads for the + * aggregand: the column's declared shape (`declaredValueShape`) and the + * dialect the statement runs on (`sqlDialect`). */ private resolveMeasureSql( cube: Cube, member: string, parentTable: string, joins: StatementJoins, - predicate: string | null = null, + predicate: string | null, + ctx: StrategyContext, ): string { const measure = this.lookupMember(cube, member, 'measure') as | { sql: string; type: string } @@ -1173,10 +1211,40 @@ export class NativeSQLStrategy implements AnalyticsStrategy { ); } - const col = measure.sql === '*' + const column = measure.sql === '*' ? '*' : this.qualifyAndRegisterJoin(measure.sql, parentTable, joins, cube); + // [#21042] The OPERAND each aggregate wraps, per the engine's own policies + // (`aggregandOperandSql`, `@objectstack/core`, the rule `driver-sql`'s + // `aggregate()` applies): on PostgreSQL and MySQL `sum` over a fractional + // column and `avg` over every numeric or boolean one accumulate in double + // (#20387), and on PostgreSQL a boolean aggregand is cast to `int` for + // `sum` / `avg` / `min` / `max` (#11635), never for the counts. Without it + // this face added exact decimals where the engine adds doubles, and + // answered `500` for a boolean `sum` the engine answers. The column's + // class is the one predicate's, over the declaration the host relays for + // the object the column lives on — the base object, or the object a + // relationship path's last hop reads ({@link columnObjectOf}, the one hop + // resolver). An expression, a column the host cannot describe, or a host + // that names no dialect gets no class or no policy, and is aggregated as + // stored. The expression metric types are not aggregates and are never + // wrapped. + const col = column === '*' || !Object.prototype.hasOwnProperty.call(AGGREGATE_ANSWER_KIND, measure.type) + ? column + : aggregandOperandSql( + measure.type as AggregationFunction, + IDENTIFIER_PATH.test(measure.sql) + ? aggregandColumnClass( + declaredValueShapeResolver(ctx, columnObjectOf(cube, parentTable, measure.sql, joins.referenceOf))?.( + measure.sql.slice(measure.sql.lastIndexOf('.') + 1), + ), + ) + : undefined, + sqlDialectFor(ctx, parentTable), + column, + ); + if (predicate !== null) { const wrapConditional = CONDITIONAL_AGGREGATE_SQL[measure.type]; if (wrapConditional) return wrapConditional(col, predicate);