The GitHub Actions job "Required Checks" on texera.git/feat/parquet-source has 
failed.
Run started by GitHub user kz930 (triggered by kz930).

Head commit for run:
e2181309caee911205852785bfb815c5121a34ed / kary zheng <[email protected]>
fix(operator): read a Parquet file into the dtypes that keep its nulls

A nullable long lost precision in the exported script: pandas widens a holed
integer column through a float, where every value past 2^53 is rounded, so
9007199254740993 came back as ...992. The executor reads the exact long off the
same file. A holed 32-bit integer went the same way, its column arriving as a
float where the engine keeps INTEGER.

The read asks for the nullable dtypes, as the Arrow source already does, and the
normalization beside it is named in the nullable spelling and widens into the
nullable dtypes in turn, so a hole stays a hole rather than the NaN a numpy
column would have to write it as.

The parity test grows a row with a hole in every column, which is what costs a
numpy column its type. That test could not have caught this before: it never
bound the `sourceFile` its own body names, so the script stopped there and the
comparison never ran. It binds it now, and its driver reads pd.NA and NaT as the
null the executor hands over.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

Report URL: https://github.com/apache/texera/actions/runs/35431607383

With regards,
GitHub Actions via GitBox

Reply via email to