While testing cross-language binary compatibility, we found that the array-of-strings (AoS) tuple sketch hashes keys differently across languages (apache/datasketches-cpp#533, apache/datasketches-go#191):
- Java joins the key strings with "," and hashes the result as UTF-16LE code units (XxHash.hashCharArr). - C++ and Go join the same way but hash the UTF-8 bytes. The same keys therefore produce different hashes, so Java and C++/Go AoS sketches can't be combined correctly. To unblock 5.3.0, we're reverting the unreleased C++ AoS sketch (#476), and the rewrite will follow once we agree on a plan. Go released AoS in v0.2.0. This is related to, but separate from, our February discussion on UTF-8 validation of stored strings: https://lists.apache.org/thread/8p36zbmjp57fcjq7tlss69zy67459jv1 (continued at https://lists.apache.org/thread/yj7mfdg2rttbfosss736t7qtp1wt2fwk) What the February discussion missed was that Java's AoS chose *UTF-16LE, no BOM *encoding years earlier. Nonetheless, here we are: We need to choose one encoding for Tuple AoS for all languages: Option A: UTF-16LE[1] for Tuple AoS for all languages (Java's current behavior). - Existing Java AoS sketches stay valid. Java has hashed AoS keys this way since the sketch was first released (sketches-core 0.13.2, April 2019), so there may be years of stored sketches. - C++ and Go convert UTF-8 to UTF-16 before hashing. That's using a small self-contained routine with no dependency, verified to reproduce Java's hashes. Go's v0.2.0 sketches become incompatible (pre-1.0). - AoS stays the one place in the library that hashes strings as UTF-16. Every other string hash (theta, HLL, CPC, Count-Min, Bloom, tuple) uses UTF-8. Option B: UTF-8 everywhere. - Consistent with every other string hash in the library, and with how C++, Go, Rust and Python store strings. - Java's existing AoS sketches become incompatible with new ones. The serialized format doesn't record the encoding, so old and new Java sketches would combine silently and incorrectly, with no error. This would need a major version and clear migration guidance. My preference is Option A, because it protects existing Java users' data. But I'd like to hear other views, particularly on whether consistency is worth a break in compatibility. I normally would suggest that this discussion stay open for 2 weeks, but I will be traveling and not be back until Monday, October 26th. So I suggest we keep this open until then. Lee [1] The C++ committee deprecated the standard library's converter header <codecvt> in C++17 and removed it in C++26. C++ isn't dropping UTF-16LE itself, it is a fixed Unicode standard. What's going away is the standard library's *converter* header, <codecvt>. C++ is just removing one poorly designed tool for converting to it. A hand-written routine (~30 lines) does the same conversion and doesn't depend on that header.
