LuciferYang opened a new issue, #10038:
URL: https://github.com/apache/paimon/issues/10038

   ### Search before asking
   
   - [X] I searched in the [issues](https://github.com/apache/paimon/issues) 
and found nothing similar.
   
   ### Paimon version
   
   master
   
   ### Compute Engine
   
   Spark (TopN / ORDER BY ... LIMIT pushdown)
   
   ### Minimal reproduce step
   
   On an append-only table, run a query that pushes a single-key TopN into 
Paimon: `SELECT * FROM t ORDER BY col LIMIT n` where the number of splits 
exceeds `n`. Arrange for a split whose sort-column statistics are incomplete: 
either a pure `stats.mode=counts` layout (no file has min/max for the column), 
or a split holding a mix of files where at least one file lacks min/max (a 
counts-mode file, or a file written before the column was added) alongside 
files that do have stats.
   
   `TopNDataSplitEvaluator` gates: append-only table, one sort key, a 
min/max-comparable type, limit ≤ 100.
   
   ### What doesn't meet your expectations?
   
   `ORDER BY ... LIMIT n` silently returns fewer / wrong rows — the split 
holding the true top row is pruned.
   
   `TopNDataSplitEvaluator` orders the splits by the sort column's aggregate 
min/max/null-count and keeps the best `limit`. It reads those aggregates from 
`DataSplit.minValue` / `maxValue` / `nullCount`, which aggregate across a 
split's files by silently skipping files that have no statistics. So a split 
containing a stats-less file yields a value that looks like a true bound but is 
not one: the file's real values are ignored. Ordering by that fabricated bound 
can rank a split below the `limit` cutoff and prune it, even though its 
stats-less file may hold a row in the true top-N. On a pure counts-mode table 
every split's bound is fabricated, they all tie, and an arbitrary subset 
survives — the split with the true top row can be the one dropped.
   
   Expected: a split whose sort-column statistics are incomplete (any file 
lacks min/max or null-count for the column) has no trustworthy bound, so it 
must be read rather than pruned.
   
   ### Anything else?
   
   Fix direction: aggregate the statistics per file inside the evaluator 
instead of relying on `DataSplit.minValue` and friends. If any file of a split 
lacks min/max or null-count for the sort column, treat the split like a gate 
failure and always read it. A file that is provably all-null for the column 
(null count equals its row count) legitimately contributes nothing to min/max 
and does not make the split incomplete.
   
   ### Are you willing to submit a PR?
   
   - [X] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to