With pandas enabled (the default), a floating-point NaN in a FLOAT or DOUBLE column is returned as None, so it can't be told apart from SQL NULL. With _disable_pandas=True the same data comes back correctly as float('nan').
Example: a DOUBLE column containing [NaN, NULL, inf, 1.5]
- default (pandas enabled):
[None, None, inf, 1.5]
_disable_pandas=True: [nan, None, inf, 1.5]
Reproduced on pandas 2.3.3 and 3.0.6 (pyarrow 25, connector main at 61b9a7f).
Cause: ResultSet._convert_arrow_table (src/databricks/sql/result_set.py) maps float32/float64 to pandas' nullable Float32Dtype/Float64Dtype. That conversion turns IEEE NaN into pd.NA, and df.to_numpy(na_value=None, dtype="object") then turns it into None.
This changes user-visible results, so I'm opening this issue alongside the fix: I have a fix with regression tests ready and am opening the PR now. It restores NaN from the Arrow column after the pandas conversion, keeps NULL as None, and matches the _disable_pandas=True output.
With pandas enabled (the default), a floating-point
NaNin a FLOAT or DOUBLE column is returned asNone, so it can't be told apart from SQLNULL. With_disable_pandas=Truethe same data comes back correctly asfloat('nan').Example: a DOUBLE column containing
[NaN, NULL, inf, 1.5][None, None, inf, 1.5]_disable_pandas=True:[nan, None, inf, 1.5]Reproduced on pandas 2.3.3 and 3.0.6 (pyarrow 25, connector
mainat 61b9a7f).Cause:
ResultSet._convert_arrow_table(src/databricks/sql/result_set.py) maps float32/float64 to pandas' nullableFloat32Dtype/Float64Dtype. That conversion turns IEEE NaN intopd.NA, anddf.to_numpy(na_value=None, dtype="object")then turns it intoNone.This changes user-visible results, so I'm opening this issue alongside the fix: I have a fix with regression tests ready and am opening the PR now. It restores NaN from the Arrow column after the pandas conversion, keeps NULL as
None, and matches the_disable_pandas=Trueoutput.