danny0405 commented on code in PR #19948:
URL: https://github.com/apache/hudi/pull/19948#discussion_r4068361527
##########
hudi-spark-datasource/hudi-spark-common/src/main/scala/org/apache/hudi/MergeOnReadIncrementalRelationV2.scala:
##########
@@ -150,6 +153,13 @@ case class MergeOnReadIncrementalRelationV2(override val
sqlContext: SQLContext,
}
}
+ /**
+ * Returns the requested-time range selected by the completion-time query
analysis. The file-group
+ * reader needs this in addition to Spark's required filters so that
out-of-range log records do
+ * not participate in record merging and mask an earlier in-range version of
the same key.
+ */
+ override def getInstantRange: HOption[InstantRange] =
queryContext.getInstantRange
Review Comment:
Could we scope the new instant-range propagation to full-table-scan
fallback? The reported bug occurs when fallback selects the latest file slices,
allowing out-of-range log blocks to participate in merging. This would keep the
fix focused on that case and preserve the existing behavior of normal
incremental reads.
In `MergeOnReadIncrementalRelationV2`:
```scala
override def getInstantRange: HOption[InstantRange] =
if (fullTableScan) queryContext.getInstantRange
else HOption.empty()
```
Then use the same accessor in `composeRDD`:
```scala
instantRangeOpt = getInstantRange,
```
The HadoopFsRelation factory already calls this accessor, so both reader
paths would follow the same condition. We can keep
`incrementalSpanRecordFilters` unchanged and retain the fallback regression
coverage for both paths and table versions 6/8, alongside the normal
incremental-read tests.
This is a scope suggestion for the fallback fix, rather than a claim that
file selection guarantees every log block is in range on all normal-read paths.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]