weijietong commented on PR #9705: URL: https://github.com/apache/paimon/pull/9705#issuecomment-5657843715
In our use case, the training framework calculates the total number of rows within a partition based on a snapshot and determines the number of row ranges to be read for each data split. Without this new API, the initial implementation would have relied on the framework obtaining row-by-row iteration via `RecordReader<InternalRow> createReader(Split split)` and then manually implementing functionality similar to a `RangeSkipReader`. Given our Parquet-based storage format—specifically when dealing with large Parquet files but reading only a small number of rows defined by the row ranges—leveraging the new API and row-range pushdown optimization allows us to achieve a significant reduction in training time compared to the previous 40-minute duration. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
