On Thu, 3 Sep 2026 06:56:05 GMT, Tushar saini <[email protected]> wrote:
> ## Summary > > [`JDK-8223933`](https://bugs.openjdk.org/browse/JDK-8223933) > > `Stream.distinct()` must remove duplicates according to `Object.equals`. On > sorted streams it only compared each element to the previous one, which fails > when `compareTo` is inconsistent with `equals`, so duplicates can still be > emitted. > > ## Fix > Keep the sorted-path consecutive check, and also track emitted elements in a > `HashSet` so uniqueness always follows `equals`. Add a regression test for > the reported case and for equals-duplicates in different sort groups. > > --------- > - [x] I confirm that I make this contribution in accordance with the [OpenJDK > Interim AI Policy](https://openjdk.org/legal/ai). > > If possible, it would be good to pre-size this hashset when an estimated > > size is available for the current stream > > If the number of duplicates is big (which is not known in advance), you'll > end up allocating and zeroing unnecessary amount of memory. I would stick to > the default behavior here. Fair enough, it's true we cannot know the cardinality of the items in the stream. We have to choose a default that is not too bad for these two opposite use cases. ------------- PR Comment: https://git.openjdk.org/jdk/pull/32670#issuecomment-5525796865
