Skip to content

OpenLineage: record sql_metadata for Polars database I/O - #1743

Open
Dev-iL wants to merge 4 commits into
apache:mainfrom
SummitSG-LLC:2609/polars-sql-metadata
Open

Dev-iL wants to merge 4 commits into
apache:mainfrom
SummitSG-LLC:2609/polars-sql-metadata

Conversation

@Dev-iL

@Dev-iL Dev-iL commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

PolarsDatabaseReader and PolarsDatabaseWriter described database I/O with file metadata only, so OpenLineageAdapter(sql_dataset_identity="datasource") could not identify the database and emitted fake FileSystem datasets .

This PR adds sql_metadata next to file_metadata: legacy mode emits exactly what it did before, and datasource mode names the real table (for example sqlite://<file> + orders).

This does not fully fix #1740, since recognizing more connection types in get_sql_source (ADBC, connectorx URIs, cursors, Sessions) is not handled.

Changes

Detailed changes

Post-1.0 classes (hamilton/plugins/polars_post_1_0_0_extensions.py)

  • load_data and save_data return {**get_file_and_dataframe_metadata(...), **get_sql_metadata(..., db_connection=self.connection, operation="read"|"write")}. The adapter checks file_metadata before sql_metadata, so legacy output is unchanged.
  • Read row count is len(df) only for a pl.DataFrame; iter_batches=True returns a generator, so rows is None. Write row count is the int write_database returns. get_sql_metadata's row counting is not broadened.
  • The writer unpacks table_name the way Polars does (csv.reader with . as the delimiter): the last component is the table, the one before it the schema, and a leading catalog is dropped because the connection identifies it. main.orders, "my-tbl" and "a.b" therefore resolve to the same dataset as a query reading that table. file_metadata["path"] keeps the name as given.
  • Nothing added after a successful read or write can raise: connection inspection already goes through get_sql_source, which catches errors, and where the new parsing could fail Polars raises first.

Pre-1.0 classes (hamilton/plugins/polars_pre_1_0_0_extension.py)

  • Keep file_metadata and add a connection-less sql_metadata (source is None) whose notes reads "Datasource lineage for Polars database I/O requires polars>=1.0; upgrade polars". In datasource mode the adapter emits no dataset and logs the note as a warning.

Docs (docs/reference/lifecycle-hooks/OpenLineageAdapter.rst)

  • The SQL datasets overview lists the Polars database materializers as recording the datasource.
  • Legacy mode keeps their earlier names: the read is named after the loader node with the query as its dataSource URI and no sql job facet, and the write is named after the table.
  • Datasource mode for them needs polars>=1.0; older Polars logs an upgrade note and emits no dataset.
  • Writes to a table whose name is not a plain identifier are parsed, so a name that parses as a statement naming tables is reported as those tables; the page says to give such tables plain names.

Tests

  • tests/plugins/test_h_openlineage.py: a driver run with OpenLineageAdapter in datasource mode asserts the SQLite namespace and name for the write and the read, the sql job facet, and no storage facet. A default-mode run pins the exact legacy namespace, names and facet keys, and fails if the adapter's file_metadata/sql_metadata branches are reordered. The pre-1.0 classes are checked for source is None, no dataset and the upgrade warning. Qualified and quoted table names are checked for write/read identity agreement. A fixture supplies the names Polars only imports for type checking, because @load_from.database and @save_to.database cannot resolve the class type hints otherwise (this fails on main too).
  • tests/plugins/test_polars_extensions.py: row counts for the plain read, the write and the iter_batches=True read, for both the post-1.0 and pre-1.0 classes; file_metadata is still returned; the table and schema the writer records for unqualified, qualified and quoted names.

Spec

  • writeups/specs/2609-01-polars-sql-metadata.md is the manifest this change was built and verified against.

How I tested this

  • pre-commit is clean. pytest tests/plugins/test_polars_extensions.py tests/plugins/test_polars_lazyframe_extensions.py tests/plugins/test_h_openlineage.py tests/io/test_utils.py gives 122 passed, 1 skipped (the Postgres test that needs HAMILTON_TEST_POSTGRES_URL, skipped on main too).
  • Legacy datasets and file_metadata for every Polars database reader and writer were compared against main and are identical.

Notes

  • Polars database I/O now also records sql_metadata. Default-mode OpenLineage users may see the identity FutureWarning, silenced by sql_dataset_identity="legacy".

Follow-up work

Neither item below is needed for this change, and neither is a regression: legacy output is unchanged.

1. Tell the converter when a write target is a literal table name. sql_datasets cannot tell a table called INSERT INTO orders SELECT * FROM customers from a statement, because get_sql_metadata also accepts statements from custom savers and files them under either key. A write to such a quoted name in datasource mode emits orders (or nothing for SELECT * FROM orders) instead of the table written, so the write and read identities disagree. The pandas to_sql writer behaves the same on main. The fix belongs at the shared boundary: an explicit marker from the writers that the target is an identifier, honored by sql_datasets, with a metadata version bump and write/read identity regressions for these names. It needs a maintainer's view on the metadata schema, since orchestrator providers reuse sql_datasets. Until then the OpenLineage reference page says to give such tables plain names.

2. Recognize more connection types in get_sql_source. ADBC connections, connectorx URIs, Polars cursors and SQLAlchemy sessions are not recognized, so Polars database I/O through them records no datasource and emits no dataset in datasource mode. This is the part of #1740 this PR leaves open.

Checklist

  • PR has an informative and human-readable title (this will be pulled into the release notes)
  • Changes are limited to a single goal (no scope creep)
  • Code passed the pre-commit check & code is left cleaner/nicer than when first encountered.
  • Any change in functionality is tested
  • New functions are documented (with a description, list of inputs, and expected output)
  • Placeholder code is flagged / future TODOs are captured in comments
  • Project documentation has been updated if adding/changing functionality.

Dev-iL and others added 4 commits September 29, 2026 12:56
PolarsDatabaseReader and PolarsDatabaseWriter described database I/O with
file metadata only (get_file_and_dataframe_metadata on the query or table
name), so OpenLineageAdapter(sql_dataset_identity="datasource") could not
identify the database and emitted fake FileSystem datasets.

Post-1.0 classes now add get_sql_metadata alongside file_metadata, so legacy
output is unchanged and datasource mode names the SQLite/Postgres table. The
writer unpacks table_name the way Polars does (quote-aware, [[catalog.]schema.]table)
so a write resolves to the same dataset as a query reading that table. The
pre-1.0 classes keep file metadata and add a connection-less sql_metadata whose
note asks datasource users to upgrade to polars>=1.0.

Adds the spec under writeups/specs/ and documents the behaviour on the
OpenLineageAdapter reference page.

Release note: Polars database I/O now also records `sql_metadata`;
default-mode OpenLineage users may see the identity `FutureWarning`, silenced
by `sql_dataset_identity="legacy"`.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
polars_pre_1_0_0_extension.py falls back to `CsvEncoding = type` when
polars.type_aliases lacks it, but never defined `SchemaDefinition`, which
PolarsJSONReader annotates its fields with. Python 3.10 to 3.13 evaluate
that annotation at import, so importing the module under Polars 1.0 or later
raised NameError; Python 3.14's lazy annotations hid it. The new pre-1.0
row-count test imports the module, so CI failed on 3.10 to 3.13 only.

Define the same `type` fallback for SchemaDefinition. Make the test fixture
that resolves polars' TYPE_CHECKING-only connection names version-independent,
since polars 1.44 names them differently from 1.41.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The Python 3.14 job resolved fonttools 4.66.1 in the minute it was being
uploaded, when only its cp311 and cp312 wheels existed, and failed at
`uv sync` before running any test. The wheels are complete now.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…n OpenLineage

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@Dev-iL
Dev-iL requested a review from skrawcz October 4, 2026 08:30

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Polars database reader/writer record file metadata, so SQL lineage is missing

1 participant