cdelmonte-zg commented on issue #16919:
URL: https://github.com/apache/datafusion/issues/16919#issuecomment-6011469827

   Hi @alamb , I spent some more time on this and put the experiments in a 
small lab: https://github.com/cdelmonte-zg/ordered-file-groups. All the numbers 
can be regenerated. It uses DataFusion 55.1.0 at `e1aa7d956`, the fair pool, 
synthetic data and one machine. The details are in `NARRATIVE.md` and 
`RESULTS.md`.
   
   In this case, the ordering gets dropped in `ListingTable`. It builds ordered 
file groups, but rejects them when there are more than `target_partitions`. The 
scan then loses its advertised ordering.
   
   Removing that check helped with the deduplication query. The base case has 
12 files, 600,000 rows and needs four ordered groups, with a target of 2. At 
256 MB it went from 0.603 s to 0.351 s. The original's final aggregate hits its 
fair-pool quota and spills even though the pool still has room. At 2 GB and 
above, that aggregate stops spilling and the timings overlap. With more rows 
and a large enough pool, the original is faster.
   
   There is a downside with many groups. With about 1,200, accepting them is 
still faster at 256 MB (0.587 against 0.722 s), but slower at 128 MB. It also 
fails with an open-file limit of 4,096. So I don't think we can just remove the 
check.
   
   Raising `target_partitions` works with a few groups, but runs into trouble 
too. With 120 groups, it completed 0 of 10 runs at 256 MB and 2 of 10 at 2 GB. 
With 1,196 groups at 2 GB, none completed. Turning 
`split_file_groups_by_statistics` off, with the same target, gave 10 of 10 
completions. At that budget, changing to the unordered plan was enough to 
complete while keeping the high target.
   
   **Two things I'd like to ask**:
   
   1. Could `EXPLAIN` show why the ordering was dropped? At the moment I only 
found the reason in a debug log line.
   2. The grouping requires `min > previous_max`, but 
`MinMaxStatistics::is_sorted` accepts touching files (`max <= next_min`). Is 
that difference intentional? In one of the lab's layouts it means five groups 
instead of four.
   
   I can send a PR for either if that would be useful.
   
   Also, the numbers in my earlier comment came from a different setup. In this 
base case the original completes at every pool tried, from 128 MB to 4 GB. It 
spills at each budget, though at the larger ones the remaining spills are in 
the sort.
   
   @zheniasigayev, are your production files sorted by both columns? I'd still 
be interested to know.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to