peter-toth commented on code in PR #58969:
URL: https://github.com/apache/spark/pull/58969#discussion_r4074256256


##########
sql/core/src/test/scala/org/apache/spark/sql/execution/benchmark/AsOfJoinBenchmark.scala:
##########
@@ -125,6 +125,9 @@ object AsOfJoinBenchmark extends SqlBasedBenchmark {
     runBenchmark("AS-OF Join Benchmark") {
       // 10K left x 10K right, 100 groups - both paths feasible
       asOfJoinBenchmark(leftRows = 10000, rightRows = 10000, numGroups = 100)
+      // Few large groups: each left row scans ~1K buffered right rows, which 
is where the
+      // per-left-row cost of the right-side group scan dominates.

Review Comment:
   **Finding 1.** `findBestBackwardForward` (`SortMergeAsOfJoinExec.scala:385`) 
stops at the first right row that fails the as-of condition once it already has 
a match, so a left row scans a prefix of the buffer rather than all of it. The 
`~1K` figure is the buffer size — which is what the PR description says, 
"buffers ~1000 right rows per group" — not the scan length.
   
   Simulating the fixture exactly (left `ts = id`, right `ts = id * 3 / 2`, 
group `id % G`, 10000 rows per side, applying the loop's own termination rule):
   
   ```
   groups=100  group size=100   total comparisons=353,070    avg per left 
row=35.3   max=100   full-buffer scans=98
   groups=10   group size=1000  total comparisons=3,351,990  avg per left 
row=335.2  max=1000  full-buffer scans=8
   ```
   
   So 9.5x the scan work, ~335 rows per left row on average, and the full 1000 
only for the 8 left rows that have no match in their group at all — those are 
the ones where `bestMatch` stays null, so the `else if (bestMatch != null)` 
early return never fires and the loop runs the buffer out. If you ever want the 
every-row-scans-the-whole-buffer shape, that is the one to build: a group whose 
left keys all sort below every right key.
   
   ```suggestion
         // Few large groups: ~1K right rows buffered per group, and each left 
row rescans a
         // prefix of that buffer, which is where the per-left-row group-scan 
cost dominates.
   ```
   



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to