PointKernel commented on PR #15: URL: https://github.com/apache/datasketches-cuda/pull/15#issuecomment-5670353548
I suggest proceeding with #12 for the initial Theta migration and following up separately on this prototype's device-resident state and `update_async` work. Since this prototype was written, #12 and #14 have both gained bounded scratch memory: screening uses a survivor budget of `max(2^22, 16*k)` hashes, with capacity-guarded writes and smaller-chunk retries on overflow. For the measured `2^31`-key updates at `lg_k=12`, peak additional memory is now roughly **56–119 MiB**, down from about **16 GiB**, excluding the input. The input-sized scratch comparison in this PR's description therefore predates those changes. The latest recorded comparison after that change is below: RTX PRO 6000 Blackwell, CUDA 13.3, U64, `lg_k=12`, three interleaved rounds on an idle GPU. These are the September 8–9 measurements, not a new benchmark run today. The measured Theta code corresponds to the current heads (#12 `771e142`, #14 `f76c76c`, #15 `02f05fb`; #14 has subsequent formatting-only changes). | Workload | #12 | #14 | #15 | |---|---:|---:|---:| | Cold, `2^31` unique keys | 11.138 ms | 11.198 ms | 11.056 ms | | Warm, `2^31` unique keys | 10.574 ms | 10.571 ms | 10.739 ms | | Cold, 100M unique keys | 0.842 ms | 0.853 ms | 0.951 ms | | Cold, `2^24` keys in runs of 2048 | 0.583 ms | 0.583 ms | 0.215 ms | The grouped-input win here is substantial, while the large unique-input results are close. The grouped row contains 8192 distinct keys; these rows compare implementations on identical inputs, rather than isolating locality at a fixed cardinality. The warm row repeatedly updates the same sketch with the same keys. All rows use synchronous `update`, so they do not measure the benefit of `update_async`. #12 gives us the broader starting point: `lg_k=5..26`, passing CUDA 12.0/SM70 and CUDA 13.1/SM100 builds, and competitive performance with bounded scratch. This prototype currently supports only `lg_k=12`, requires SM90+ intrinsics, and reserves about 17.7 MiB of candidate arrays per sketch on a 188-SM GPU, including empty sketches. Its asynchronous state model is worth carrying forward with reusable/right-sized workspace and broader precision/toolchain support. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
