Skip to content

Zero-cost table(view) - #3246

Merged
texodus merged 5 commits into
masterfrom
zero-cost-table-view
Oct 3, 2026
Merged

texodus merged 5 commits into
masterfrom
zero-cost-table-view

Conversation

@texodus

@texodus texodus commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

This PR makes client.table(view) an engine-supported API and optimizes it to be effectively zero-cost, sharing memory with the underlying View it wraps. In addition, two related breaking API changes have been introduced which cover functional gaps in making table(view) symmetric (in preparation for adding it to the UI).

client.table(view) optimization

It has always been possible to flatten a View into a Table by passing the View to the client.table() function. Internally however, this mechanism had no true special treatment - the View was serialized and de-serialized via Arrow to populate another copy. This was slow (relatively) and memory wasteful.

Now, client.table(view) (when called with the same Client that owned the View) will create a special derived Table, allocating only metadata but sharing its columns with the input View, with a few changes:

  • Read-only: update(), remove(), clear() and replace() raise, and so do the index and limit options. These could technically fall back to the cost of a copy, but I'm not sure its work it.
  • view.delete() raises while a derived table still reads the view; delete the derived table first. Previously the copy became a frozen snapshot.
  • Parent sort is ignored. Rows follow source insertion order. This is similarly a consequence of sharing the underlying (which is not sorted in-place). It is easy enough to manually inherit the sort but this is up for debate.
  • Rows ejected from the parent's filter are removed, where the copy kept them with stale values.
  • Updates in the engien are processed in the same engine cycle as the source, not eventually.
  • A view that omits the index column is keyed by row identity, so updates apply in place instead of appending.
  • Pivoted views mirror the view's rollup modes. Subtotal and total rows are included under rollup; set group_rollup_mode: "flat" for leaves only. The old copy appended its changed rows on every tick.
  • The schema is frozen at creation. A split_by value that appears later is dropped, and one that vanishes goes all-null. The new schema option on client.table(view, { schema }) pre-declares columns.
  • Cross-client views keep the legacy copy implementation.
  • Virtual servers: table(view) now errors unless the handler reports the new view_derivations feature and implements view_make_table. DuckDB and Postgres implement this (View are already TEMPORARY TABLE on these engines so this is free but static).

Flip Arrow/columns readable output structure

Formerly, a config like this:

{
    "group_by": ["State"],
    "columns": ["Sales"]
}

... would have produced a View with columns named "State (Group by 1)", "Sales" in CSV and human-readable Arrow output, which is awkward for recomposing when used as input to another Table. Now, the columns will just be "State" and "Sales". If the config has a conflict, the columns column will get qualified instead:

{
    "group_by": ["State"],
    "columns": ["State"],
    "aggregates": {"State": "count"}
}

... will return columns "State" and "State (count)".

view.column_paths() returns structured paths

Previously a string[] of "a|b|Sales"; it is now area[level][column] of typed scalars, in JS, Python and Rust (Vec<Vec<Scalar>>). Because column names in flattened Tables may now have separator values routinely, splicing by the separator character no longer works for separating group levels from column names.

  • This is a column-oriented list.
  • Values are no longer stringified. A datetime split value is epoch milliseconds, floats are full precision (not %g with 6 significant digits), and null is null, not "null" (in headers).
  • Virtual servers still produce string values at every level, since handlers return joined paths (TODO).
  • Still joined with |: to_columns/to_json keys, CSV headers, and derived-table column names.
  • Datagrid split_by and group_by headers use the column's number_format/date_format (split_by doesn't respect the date formatting settings #2922), so header text changes. The plugin's column_config_schema receives a trailing role argument ("column" | "group_by" | "split_by"), which is additive.

sum family null semantics

Perspective has a family of sum aggregates that exposed an asymmetry when used with table(view), namely that a sum of a column of exclusively null values is 0, which caused empty groups from the parent View (caused by a split_by eliminating a cell) to show as 0 in the derived Table. This could be cured by using sum not null as an aggregate instead, but sum being the default makes this behavior surprising.

As a fix, sum now works like sum not null was intended to work - returning null for groups with no valid values instead of 0

  • sum returns null for a group with no non-null values; it used to return 0. This matches SQL, the SQL virtual servers, and the engine's own mean/min/max behavior.
  • sum or zero is new and keeps the old 0 behaviour.
  • sum not null is a parse-only legacy alias for sum.
  • sum abs, abs sum, pct sum parent and pct sum total follow the same
    null behavior change. A null sum used to render as 0%, eg.

All of these sum aggregates used the incremental aggregate path now as well, making them all equally fast.

Internal change, refactoring a a bunch of tests

  • viewer.restore() has been refactored to guarantee atomic-update-or-error semantics going forward (and has better error handling as a result for broken configs).
  • Engine expression calculation has been updated to fully integrate with the incremental aggregate path instead of force-recalculate on update, making it quite a bit faster.
  • Upgrade emscripten to 5.0.3 to match pyodide bump in Upgrade Pyodide build to Python 3.14 and Emscripten 5 #3245

Signed-off-by: Andrew Stein <steinlink@gmail.com>
Signed-off-by: Andrew Stein <steinlink@gmail.com>

# Conflicts:
#	rust/perspective-client/src/rust/virtual_server/generic_sql_model/tests.rs
Signed-off-by: Andrew Stein <steinlink@gmail.com>

# Conflicts:
#	packages/viewer-datagrid/src/ts/custom_elements/datagrid.ts
@texodus texodus added enhancement Feature requests or improvements breaking labels Sep 30, 2026
@texodus
texodus force-pushed the zero-cost-table-view branch 5 times, most recently from a5c1c38 to 8bdba05 Compare October 1, 2026 04:40
Signed-off-by: Andrew Stein <steinlink@gmail.com>
@texodus
texodus force-pushed the zero-cost-table-view branch 2 times, most recently from 77632a3 to 63e0fa0 Compare October 2, 2026 19:12
Signed-off-by: Andrew Stein <steinlink@gmail.com>
@texodus
texodus force-pushed the zero-cost-table-view branch from 63e0fa0 to 966a8cd Compare October 2, 2026 20:05
@texodus
texodus merged commit 6dbc6b8 into master Oct 3, 2026
40 checks passed
@texodus
texodus deleted the zero-cost-table-view branch October 3, 2026 00:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

breaking enhancement Feature requests or improvements

Development

Successfully merging this pull request may close these issues.

1 participant