kz930 opened a new issue, #7550:
URL: https://github.com/apache/texera/issues/7550

   ### What happened?
   
   One blank cell in a numeric CSV column makes CSV File Scan type the whole 
column as STRING, so every downstream operator that does arithmetic on it 
fails, including on the rows that hold perfectly good numbers.
   
   The cause is that the schema inference and the execution disagree about what 
a blank cell is. `CSVScanSourceOpDesc.scala:116` sets 
`csvSetting.setNullValue("")` before inferring, so a blank reads as an empty 
string: `tryParseDouble("")` fails, `tryParseBoolean("")` fails, and 
`inferField` lands on `tryParseString()`. `CSVScanSourceOpExec.scala:95-100` 
builds its parser without `setNullValue`, so at execution the same blank reads 
as null, which is what `AttributeTypeUtils.parseField` is written to pass 
through.
   
   Either convention is defensible on its own. What breaks is having one for 
the schema and another for the data, and the schema side is the one that 
silently downgrades a numeric column.
   
   Two files that differ only in one blank cell take the same workflow down two 
different paths:
   
   ```
   id,petal_length,petal_width      id,petal_length,petal_width
   1,2.6,0.75                       1,2.6,0.75
   2,,1.3                           2,3.9,1.3
   3,5.2,1.85                       3,5.2,1.85
   ```
   
   With the blank, Iris Logistic Regression takes 2 rows, emits 1 and pauses. 
Without it, the same operator with the same configuration takes 3 and emits 3. 
The failure is on row 1, whose values are `2.6` and `0.75`, because they arrive 
as `array([['2.6', '0.75']], dtype='<U32')` rather than as numbers.
   
   Expected: a blank cell should not change the inferred type of a column. 
Inference and execution should agree on what a blank is.
   
   ### How to reproduce?
   
   Upload the two CSVs above. Build `CSV File Scan` to `Hugging Face Iris 
Logistic Regression` with Petal Length Cm Attribute = `petal_length` and Petal 
Width Cm Attribute = `petal_width`. Run it once against each file. The one with 
the blank cell fails on its first row; the one without it completes.
   
   Any operator that does arithmetic on a scanned numeric column shows the same 
thing; Iris is just the shortest workflow that reaches it.
   
   ### Version/Branch
   
   1.3.0-incubating-SNAPSHOT (main)
   
   ### Relevant log output
   
   ```shell
   2026-08-11 00:51:41.181 | ERROR | 
core.architecture.managers.context:report_exception:104 - ufunc 'subtract' did 
not contain a loop with signature matching types (dtype('<U32'), 
dtype('float64')) -> None
                    │          └ [3.72666667, 1.17619048]
                    └ array([['2.6', '0.75']], dtype='<U32')
   
   numpy._core._exceptions._UFuncNoLoopError: ufunc 'subtract' did not contain 
a loop with signature matching types (dtype('<U32'), dtype('float64')) -> None
   ```
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to