cdelmonte-zg commented on issue #16919: URL: https://github.com/apache/datafusion/issues/16919#issuecomment-6011469827
Hi @alamb , I spent some more time on this and put the experiments in a small lab: https://github.com/cdelmonte-zg/ordered-file-groups. All the numbers can be regenerated. It uses DataFusion 55.1.0 at `e1aa7d956`, the fair pool, synthetic data and one machine. The details are in `NARRATIVE.md` and `RESULTS.md`. In this case, the ordering gets dropped in `ListingTable`. It builds ordered file groups, but rejects them when there are more than `target_partitions`. The scan then loses its advertised ordering. Removing that check helped with the deduplication query. The base case has 12 files, 600,000 rows and needs four ordered groups, with a target of 2. At 256 MB it went from 0.603 s to 0.351 s. The original's final aggregate hits its fair-pool quota and spills even though the pool still has room. At 2 GB and above, that aggregate stops spilling and the timings overlap. With more rows and a large enough pool, the original is faster. There is a downside with many groups. With about 1,200, accepting them is still faster at 256 MB (0.587 against 0.722 s), but slower at 128 MB. It also fails with an open-file limit of 4,096. So I don't think we can just remove the check. Raising `target_partitions` works with a few groups, but runs into trouble too. With 120 groups, it completed 0 of 10 runs at 256 MB and 2 of 10 at 2 GB. With 1,196 groups at 2 GB, none completed. Turning `split_file_groups_by_statistics` off, with the same target, gave 10 of 10 completions. At that budget, changing to the unordered plan was enough to complete while keeping the high target. **Two things I'd like to ask**: 1. Could `EXPLAIN` show why the ordering was dropped? At the moment I only found the reason in a debug log line. 2. The grouping requires `min > previous_max`, but `MinMaxStatistics::is_sorted` accepts touching files (`max <= next_min`). Is that difference intentional? In one of the lab's layouts it means five groups instead of four. I can send a PR for either if that would be useful. Also, the numbers in my earlier comment came from a different setup. In this base case the original completes at every pool tried, from 128 MB to 4 GB. It spills at each budget, though at the larger ones the remaining spills are in the sort. @zheniasigayev, are your production files sorted by both columns? I'd still be interested to know. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
