naveenp2708 opened a new pull request, #18383: URL: https://github.com/apache/iceberg/pull/18383
Adds a test for a storage-partitioned join where a `DISTINCT` or `GROUP BY` runs directly on a partially clustered join output. The existing `testAggregates` only covers an aggregate over a single scan, not over a join, so this fills that gap. Before SPARK-55848 this case silently returned wrong counts (1200 rows instead of 200, multiplied by the partition split factor). It fails on Spark 4.0.2 and passes on 4.0.3, which is the fix version. Test only, no production change. One question: the SPJ tests here run with partial clustering off right now, so I wasn't sure if this coverage belongs here or if you consider it Spark side. Happy to move or drop it if you'd rather. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
