dwsmith1983 commented on code in PR #6753:
URL: https://github.com/apache/datafusion-comet/pull/6753#discussion_r4214944685
##########
spark/src/main/scala/org/apache/spark/sql/comet/CometScanUtils.scala:
##########
@@ -45,4 +51,94 @@ object CometScanUtils {
case _ => false
}
}
+
+ /**
+ * Bin packs `files` with `pack` separately for each object store `storeKey`
names, so no
+ * partition holds files from two stores, and numbers the partitions from 0.
Files that share
+ * one store are packed exactly as `pack` packs them.
+ */
+ def packFilesPerStore[K](files: Seq[PartitionedFile], storeKey:
PartitionedFile => K)(
+ pack: Seq[PartitionedFile] => Seq[FilePartition]): Seq[FilePartition] = {
+ val keyed = files.map(file => (storeKey(file), file))
+ val stores = keyed.map(_._1).distinct
+ if (stores.size <= 1) {
+ pack(files)
+ } else {
+ stores
+ .flatMap(store => pack(keyed.collect { case (key, file) if key ==
store => file }))
Review Comment:
> Please collect files into groups once, preserving first-seen store order
and file order, then invoke `pack` for each group.
Done in e909c077b. `CometScanUtils` groups the files by store in one pass,
keeping first-seen store order and file order, and both `packFilesPerStore` and
`splitPartitionsByStore` use it. The layouts are unchanged.
`CometScanSchemeFallbackSuite` now bounds the store-key comparisons for 2000
files over 200 stores, which fails with the per-store pass, and the seeded
property test checks both orders.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]