On Thu, 3 Sep 2026 06:56:05 GMT, Tushar saini <[email protected]> wrote:
> ## Summary > > [`JDK-8223933`](https://bugs.openjdk.org/browse/JDK-8223933) > > `Stream.distinct()` must remove duplicates according to `Object.equals`. On > sorted streams it only compared each element to the previous one, which fails > when `compareTo` is inconsistent with `equals`, so duplicates can still be > emitted. > > ## Fix > Keep the sorted-path consecutive check, and also track emitted elements in a > `HashSet` so uniqueness always follows `equals`. Add a regression test for > the reported case and for equals-duplicates in different sort groups. > > --------- > - [x] I confirm that I make this contribution in accordance with the [OpenJDK > Interim AI Policy](https://openjdk.org/legal/ai). > If possible, it would be good to pre-size this hashset when an estimated size > is available for the current stream If the stream is not `SIZED`, the estimated size could be completely off, orders of magnitude different. And even with `SIZED` stream, the `distinct()` call assumes that there are some duplicates possible. If the number of duplicates is big (which is not known in advance), you'll end up allocating and zeroing unnecessary amount of memory. I would stick to the default behavior here. ------------- PR Comment: https://git.openjdk.org/jdk/pull/32670#issuecomment-5525576785
