LuciferYang opened a new pull request, #10039: URL: https://github.com/apache/paimon/pull/10039
### Purpose close #10038 `TopNDataSplitEvaluator` supports Spark single-key TopN pushdown (`ORDER BY col LIMIT n` on an append-only table, a min/max-comparable sort type, limit ≤ 100). It orders the splits by the sort column's aggregate min/max/null-count and keeps the best `limit`. It reads those aggregates from `DataSplit.minValue` / `maxValue` / `nullCount`, which aggregate across a split's files by silently skipping files without statistics. A split containing a stats-less file (a `stats.mode=counts` file, or one written before the column was added) therefore gets an aggregate that looks like a true bound but is not: the file's real values are ignored. Ordering by that fabricated bound can prune a split that may hold a row in the true top-N. On a pure counts-mode table every bound is fabricated, all splits tie, and an arbitrary subset survives, so the split with the true top row can be dropped and `ORDER BY ... LIMIT n` silently misses rows. This aggregates the statistics per file inside the evaluator and tracks completeness. If any file of a split lacks min/max or null-count for the sort column, the split has no trustworthy bound and is read unconditionally (never pruned), the same way splits that fail the min/max gate are already handled. A file that is provably all-null for the column (its null count equals its row count) legitimately contributes nothing to min/max and does not make the split incomplete. For a split whose files all have statistics, the aggregated min/max/null-count is the same as before, so fully-statted tables are unaffected. ### Tests `TableScanTest.testPushDownTopNMultiFileSplitWithMixedStatsIsAlwaysRead`: one split holds a full-statistics file (min 50) plus a counts-mode file with no min/max. Before the fix `DataSplit.minValue` skipped the counts file and reported 50, so at `LIMIT 1` the split ranked last and was pruned; the test asserts it is now always read (the counts file may hold a value below every other split's min). It fails against the pre-fix code, which returns only the other split. `TableScanTest.testPushDownTopNCountsModeStatsDisablesPruning`: every split is counts-mode (min/max unknown), so no bound is trustworthy and all splits must be read; the pre-fix code keeps an arbitrary subset at `LIMIT 1`. `TableScanTest.testPushDownTopNNullsLastKeepsStatsUnknownSplitFirst` (updated): an unknown-bound split is always read, and the best known-bound split still wins the remaining limit slot. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
