JingsongLi commented on code in PR #9705:
URL: https://github.com/apache/paimon/pull/9705#discussion_r4001783586
##########
paimon-core/src/main/java/org/apache/paimon/operation/RawFileSplitRead.java:
##########
@@ -212,26 +221,85 @@ public RecordReader<InternalRow> createReader(
Builder formatReaderMappingBuilder =
createFormatReaderMappingBuilder(outputRowType, topN, limit);
+ boolean hasFilter = filters != null && !filters.isEmpty();
+ boolean hasDv = hasDeletionVector(files, dvFactories);
+ boolean fullScanRange =
+ rowRange != null && !hasFilter && topN == null && limit ==
null && !hasDv;
Review Comment:
[P2] Keep effective-row fallback when files may be ignored
With scan.ignore-lost-files=true, I wrote two Parquet files containing 0–4
and 5–9 and removed the first after planning. The ordinary reader returns
[5,6,7,8,9], but createReader(split, RowRange.of(1,2)) returns [] rather than
[6,7]. This eligibility condition lets the pushdown count the missing file's
five manifest rows and skip the surviving file before the missing-file reader
can emit zero rows; the outer effective-row skip/limit is disabled.
Disabling this pushdown when lost/corrupt files may be ignored makes the
same real write/read regression pass. Please preserve the actual-output
fallback for these settings, guard the equivalent data-evolution range
calculation, and test ranged output against the corresponding slice of an
ordinary read.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]