LuciferYang opened a new issue, #10038: URL: https://github.com/apache/paimon/issues/10038
### Search before asking - [X] I searched in the [issues](https://github.com/apache/paimon/issues) and found nothing similar. ### Paimon version master ### Compute Engine Spark (TopN / ORDER BY ... LIMIT pushdown) ### Minimal reproduce step On an append-only table, run a query that pushes a single-key TopN into Paimon: `SELECT * FROM t ORDER BY col LIMIT n` where the number of splits exceeds `n`. Arrange for a split whose sort-column statistics are incomplete: either a pure `stats.mode=counts` layout (no file has min/max for the column), or a split holding a mix of files where at least one file lacks min/max (a counts-mode file, or a file written before the column was added) alongside files that do have stats. `TopNDataSplitEvaluator` gates: append-only table, one sort key, a min/max-comparable type, limit ≤ 100. ### What doesn't meet your expectations? `ORDER BY ... LIMIT n` silently returns fewer / wrong rows — the split holding the true top row is pruned. `TopNDataSplitEvaluator` orders the splits by the sort column's aggregate min/max/null-count and keeps the best `limit`. It reads those aggregates from `DataSplit.minValue` / `maxValue` / `nullCount`, which aggregate across a split's files by silently skipping files that have no statistics. So a split containing a stats-less file yields a value that looks like a true bound but is not one: the file's real values are ignored. Ordering by that fabricated bound can rank a split below the `limit` cutoff and prune it, even though its stats-less file may hold a row in the true top-N. On a pure counts-mode table every split's bound is fabricated, they all tie, and an arbitrary subset survives — the split with the true top row can be the one dropped. Expected: a split whose sort-column statistics are incomplete (any file lacks min/max or null-count for the column) has no trustworthy bound, so it must be read rather than pruned. ### Anything else? Fix direction: aggregate the statistics per file inside the evaluator instead of relying on `DataSplit.minValue` and friends. If any file of a split lacks min/max or null-count for the sort column, treat the split like a gate failure and always read it. A file that is provably all-null for the column (null count equals its row count) legitimately contributes nothing to min/max and does not make the split incomplete. ### Are you willing to submit a PR? - [X] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
