Skip to content

NaN in FLOAT/DOUBLE columns is returned as None (same as NULL) when pandas is enabled #978

Description

@maharanay22

With pandas enabled (the default), a floating-point NaN in a FLOAT or DOUBLE column is returned as None, so it can't be told apart from SQL NULL. With _disable_pandas=True the same data comes back correctly as float('nan').

Example: a DOUBLE column containing [NaN, NULL, inf, 1.5]

  • default (pandas enabled): [None, None, inf, 1.5]
  • _disable_pandas=True: [nan, None, inf, 1.5]

Reproduced on pandas 2.3.3 and 3.0.6 (pyarrow 25, connector main at 61b9a7f).

Cause: ResultSet._convert_arrow_table (src/databricks/sql/result_set.py) maps float32/float64 to pandas' nullable Float32Dtype/Float64Dtype. That conversion turns IEEE NaN into pd.NA, and df.to_numpy(na_value=None, dtype="object") then turns it into None.

This changes user-visible results, so I'm opening this issue alongside the fix: I have a fix with regression tests ready and am opening the PR now. It restores NaN from the Arrow column after the pandas conversion, keeps NULL as None, and matches the _disable_pandas=True output.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions