On Thu, 3 Sep 2026 06:56:05 GMT, Tushar saini <[email protected]> wrote:

> ## Summary
> 
> [`JDK-8223933`](https://bugs.openjdk.org/browse/JDK-8223933)
> 
> `Stream.distinct()` must remove duplicates according to `Object.equals`. On 
> sorted streams it only compared each element to the previous one, which fails 
> when `compareTo` is inconsistent with `equals`, so duplicates can still be 
> emitted.
> 
> ## Fix
> Keep the sorted-path consecutive check, and also track emitted elements in a 
> `HashSet` so uniqueness always follows `equals`. Add a regression test for 
> the reported case and for equals-duplicates in different sort groups.
> 
> ---------
> - [x] I confirm that I make this contribution in accordance with the [OpenJDK 
> Interim AI Policy](https://openjdk.org/legal/ai).

> > If possible, it would be good to pre-size this hashset when an estimated 
> > size is available for the current stream
> 
> If the number of duplicates is big (which is not known in advance), you'll 
> end up allocating and zeroing unnecessary amount of memory. I would stick to 
> the default behavior here.

Fair enough, it's true we cannot know the cardinality of the items in the 
stream. We have to choose a default that is not too bad for these two opposite 
use cases.

-------------

PR Comment: https://git.openjdk.org/jdk/pull/32670#issuecomment-5525796865

Reply via email to