Skip to content

fix(table_diff): resolve key and skip column names against engine-reported casing - #6110

Open
akshayram1 wants to merge 1 commit into
SQLMesh:mainfrom
akshayram1:fix/table-diff-key-column-case
Open

akshayram1 wants to merge 1 commit into
SQLMesh:mainfrom
akshayram1:fix/table-diff-key-column-case

Conversation

@akshayram1

@akshayram1 akshayram1 commented Oct 1, 2026 •

Copy link
Copy Markdown

Description

Fixes #6067.

TableDiff normalizes user-supplied on and skip_columns names with the connection dialect. On BigQuery and DuckDB that lowercases them, but adapter.columns() returns the stored casing (for example KEY1). The exact dict lookups then fail:

  • on keys raise KeyError: 'key1' in _key_expression.
  • skip_columns are silently not skipped.

This PR follows option 2 from the issue: resolve each user-supplied name against the engine-reported schema.

  • An exact match always wins, so case-sensitive engines behave as before.
  • Otherwise a single case-insensitive match is used.
  • If several columns match and none matches exactly, it raises a SQLMeshError saying the name is ambiguous.
  • A key column that doesn't exist raises a SQLMeshError that lists the available columns, instead of a KeyError.
  • A skip column that doesn't exist in a table is still ignored for that table, as before.

Names are resolved separately for the source and target tables, so their casing can differ (for example KEY1 vs key1). This covers the list form of on, the expression form (s.KEY1 = t.KEY1) and skip_columns.

key_columns now works on a copy of the on expression, so it no longer changes the expression the caller passed in.

Test Plan

New DuckDB tests in tests/core/test_table_diff.py:

  • Uppercase, lowercase and mixed-case keys, with one or several columns, in both the list and expression forms of on.
  • Different key casing in the source and target tables.
  • Skip columns with uppercase names, and skip columns present in only one table.
  • Errors for a missing column and for an ambiguous name.
  • Exact match winning over a case-insensitive match.
  • The on expression not being changed.
  • The generated SQL referencing the resolved column names.

The first new tests failed before the fix (8 of 8). After the fix, tests/core/test_table_diff.py passes: 28 tests.

make fast-test: 2621 passed. 6 failed, all in tests/core/test_connection_config.py::*pyodbc*. They fail the same way on unmodified main in my local environment, because unixodbc isn't installed (libodbc.2.dylib not loaded), so they are unrelated to this change.

Known gap: the tests only run on DuckDB. I didn't run anything against a case-sensitive engine such as Postgres or Snowflake. Resolved names go through the existing quote_identifiers pass, so uppercase columns should be emitted quoted there, but no test proves it.

Checklist

  • I have run make style and fixed any issues
  • I have added tests for my changes (if applicable)
  • All existing tests pass (make fast-test): everything passes except the 6 local pyodbc failures described above
  • My commits are signed off (git commit -s) per the DCO

…orted casing

User-supplied `on` and `skip_columns` names are normalized with the
connection dialect, which lowercases them on engines like BigQuery and
DuckDB, while `adapter.columns()` reports the stored casing. Exact
lookups then failed with a KeyError for keys and silently ignored skip
columns.

Resolve each name against the source and target schemas: an exact match
wins, otherwise a unique case-insensitive match is used. Ambiguous names
and missing key columns raise a clear SQLMeshError. Also stop mutating
the caller's `on` expression.

Fixes SQLMesh#6067

Signed-off-by: Akshay Chame <akshaychame2@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

table_diff: KeyError when -o key columns are not lower case on Bigquery

1 participant