weijietong commented on PR #9705:
URL: https://github.com/apache/paimon/pull/9705#issuecomment-5657843715

   In our use case, the training framework calculates the total number of rows 
within a partition based on a snapshot and determines the number of row ranges 
to be read for each data split. Without this new API, the initial 
implementation would have relied on the framework obtaining row-by-row 
iteration via `RecordReader<InternalRow> createReader(Split split)` and then 
manually implementing functionality similar to a `RangeSkipReader`. Given our 
Parquet-based storage format—specifically when dealing with large Parquet files 
but reading only a small number of rows defined by the row ranges—leveraging 
the new API and row-range pushdown optimization allows us to achieve a 
significant reduction in training time compared to the previous 40-minute 
duration.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to