While testing cross-language binary compatibility, we found that the
array-of-strings (AoS) tuple sketch hashes keys differently across
languages (apache/datasketches-cpp#533, apache/datasketches-go#191):

- Java joins the key strings with "," and hashes the result as
UTF-16LE code units (XxHash.hashCharArr).
- C++ and Go join the same way but hash the UTF-8 bytes.

The same keys therefore produce different hashes, so Java and C++/Go AoS
sketches can't be combined correctly. To unblock 5.3.0, we're reverting
the unreleased C++ AoS sketch (#476), and the rewrite will follow once we
agree on a plan. Go released AoS in v0.2.0.

This is related to, but separate from, our February discussion on
UTF-8 validation of stored strings:
https://lists.apache.org/thread/8p36zbmjp57fcjq7tlss69zy67459jv1
(continued at
https://lists.apache.org/thread/yj7mfdg2rttbfosss736t7qtp1wt2fwk)

What the February discussion missed was that Java's AoS chose *UTF-16LE, no
BOM *encoding years earlier.

Nonetheless, here we are: We need to choose one encoding for Tuple AoS for
all languages:

Option A: UTF-16LE[1] for Tuple AoS for all languages (Java's current
behavior).

- Existing Java AoS sketches stay valid. Java has hashed AoS keys this way
  since the sketch was first released (sketches-core 0.13.2, April 2019),
  so there may be years of stored sketches.

- C++ and Go convert UTF-8 to UTF-16 before hashing. That's using a small
  self-contained routine with no dependency, verified to reproduce Java's
  hashes. Go's v0.2.0 sketches become incompatible (pre-1.0).

- AoS stays the one place in the library that hashes strings as UTF-16.
  Every other string hash (theta, HLL, CPC, Count-Min, Bloom, tuple) uses
UTF-8.

Option B: UTF-8 everywhere.

- Consistent with every other string hash in the library, and with how
  C++, Go, Rust and Python store strings.

- Java's existing AoS sketches become incompatible with new ones. The
  serialized format doesn't record the encoding, so old and new Java
  sketches would combine silently and incorrectly, with no error. This
  would need a major version and clear migration guidance.

My preference is Option A, because it protects existing Java users' data.
But I'd like to hear other views, particularly on whether consistency is
worth a break in compatibility.

I normally would suggest that this discussion stay open for 2 weeks, but I
will be traveling and not be
back until Monday, October 26th. So I suggest we keep this open until
then.

Lee

[1] The C++ committee deprecated the standard library's converter header
<codecvt> in C++17 and removed it in C++26.
C++ isn't dropping UTF-16LE itself, it is a fixed Unicode standard.
What's going away is the standard library's *converter* header, <codecvt>.
C++ is just removing one poorly designed tool for converting to it.
A hand-written routine (~30 lines) does the same conversion and doesn't
depend on that header.

Reply via email to