The GitHub Actions job "Required Checks" on texera.git/main has succeeded.
Run started by GitHub user github-merge-queue[bot] (triggered by 
github-merge-queue[bot]).

Head commit for run:
566efb81eb33cdcaab1a759ee26951c5514f1b77 / Eugene Gu <[email protected]>
feat(workflow-operator): report scan rows that fail schema parsing (#6947)

### What changes were proposed in this PR?

CSV/JSONL scan sources infer each column's type from only the first
`INFER_READ_LIMIT` (=100) rows. When a later row held a value that did
not parse as the inferred type, the execs caught the error, mapped the
row to `null`/`None`, and filtered it out — silently dropping the entire
row with no error, warning, count, or log. Every downstream count,
aggregate, and join then ran on a silently truncated dataset. The bug is
the silence, not the rejection.

Following the consensus in discussion #6324 (skip + surface, rather than
the fail-fast approach originally proposed in #6323), a scan now still
skips the unparsable row — the run completes and existing workflows keep
working — but each skipped row is surfaced to the user as a console
warning naming the row number, offending value, column, and expected
type, for example:

> WARNING: skipped row 150 — value '55.5' in column 'age' cannot be read
as INTEGER. Column types were inferred from an initial sample of 100
rows, and this value does not match.

<img width="1137" height="788" alt="Screenshot 2026-08-05 at 2 10 49 PM"
src="https://github.com/user-attachments/assets/171f741e-b11b-4567-8c62-d6939b9faa58";
/>

Changes:

- Add `ScanRowParseError`, which builds the per-row warning by
re-parsing the failing row's fields to identify the offending column,
with a generic fallback message when no single column can be identified
(e.g. a malformed JSON line). The message names the inference sample
size actually used — `INFER_READ_LIMIT` capped by the user-set row limit
(`ScanSourceOpDesc.inferSampleSize`) — rather than a hard-coded 100.
- Add `SkippedRowReporter`, shared by all scan variants: it reports the
first 100 skipped rows individually and then appends a single summary
line carrying the true total, so a heavily malformed file cannot flood
the console or exhaust memory.
- Record-and-skip in `CSVScanSourceOpExec`, `JSONLScanSourceOpExec`,
`ParallelCSVScanSourceOpExec`, and `CSVOldScanSourceOpExec`. The
legitimate null paths (an empty cell parsed to null, a blank/all-null
line, an exhausted block) are untouched and are not reported. Row
numbers: CSV reuses its existing counter, JSONL and csvOld gain absolute
line/data-row numbers (a configured `offset` is added back in, so the
reported number always matches the file), and ParallelCSV reports none
because it is byte-partitioned across workers.
- Add `OperatorExecutor.getWarnings` (defaults to empty), a generic
non-fatal warning channel that any executor can opt into.
- Emit the warnings at `FinalizeExecutor` in `DataProcessor` as PRINT
console messages via the new `ErrorUtils.mkPrintConsoleMessage`. The
titles keep the `WARNING: ` prefix the UI keys on, and, unlike the
executor exception path, this does not pause the run.

Inference logic is unchanged (the 100-row sampling is untouched). Scala
only — no UI change.

Behavior-change note for reviewers: workflows that previously ran to
completion while silently discarding malformed rows now still complete,
but they surface a console warning for each skipped row (capped at 100
detailed entries plus a summary). One case worth calling out: a csvOld
scan over a file with an empty cell in a numeric column previously
dropped that row silently and will now report it — this is exactly the
silent loss this PR exists to surface, not a regression.

Known limitations: warnings are collected in the executor and emitted
only when it finalizes — a run that fails or is killed mid-scan does not
surface them, and nothing is shown while the scan is still running.
Streaming emission as rows are skipped is possible follow-up work; the
finalize-only design follows the consensus in #6324.

### Any related issues, documentation, discussions?

Closes #6279. Design discussion: #6324. This supersedes the fail-fast
approach in #6323, which will be closed in favor of this PR.

### How was this PR tested?

Scan-operator unit tests (skip-and-report naming row/value/column/type,
the 100-detail cap with its summary line, empty-cell and blank-line
non-reporting, the malformed-JSON fallback, absolute line numbers when
an `offset` is set, and the warning naming the actual inference sample
size when a row `limit` caps it):

```
sbt "WorkflowOperator/testOnly \
  org.apache.texera.amber.operator.source.scan.csv.CSVScanSourceOpExecSpec \
  org.apache.texera.amber.operator.source.scan.json.JSONLScanSourceOpExecSpec \
  
org.apache.texera.amber.operator.source.scan.csv.ParallelCSVScanSourceOpExecSpec
 \
  
org.apache.texera.amber.operator.source.scan.csvOld.CSVOldScanSourceOpExecSpec"
```

Engine-side emission path (`DataProcessorSpec`): a new test asserts the
warnings reach the coordinator as PRINT console messages at finalize and
that the run is not paused.

Full local gate (lint + format + backend unit tests):

```
sbt "scalafixAll --check"   # pass
sbt scalafmtCheckAll        # pass
AMBER_TEST_FILTER=skip-integration sbt WorkflowExecutionService/test   # pass
sbt WorkflowOperator/test WorkflowCore/test WorkflowCompilingService/test   # 
pass
```

Manual, in the Texera UI, over CSV and JSONL datasets derived from a
real 403-row file whose `age` column is inferred as INTEGER: verified a
single bad row (skipped, one warning naming row/value/column/type), many
bad rows (100 detailed warnings plus a summary line), the clean file
(all rows, no warning), an empty cell (kept as a null row, no warning),
and a malformed JSONL line (skipped, generic fallback warning). In every
case the scan completed instead of failing.

Rebased onto main after #7247 (the JSONL worker-slice offset fix) landed

### Was this PR authored or co-authored using generative AI tooling?

Co-authored by: Claude Code (Fable 5)

---------

Signed-off-by: Eugene Gu <[email protected]>
Co-authored-by: Copilot Autofix powered by AI 
<[email protected]>

Report URL: https://github.com/apache/texera/actions/runs/36946209365

With regards,
GitHub Actions via GitBox

Reply via email to