richardcocks opened a new issue, #4056:
URL: https://github.com/apache/iggy/issues/4056

   ### Bug description
   
   The C# SDK and Java SDK build the length prefix for a string identifier from 
the string's `.Length` / `.length()`, which in each case count UTF-16 code 
units. They then write the value as UTF-8 encoded bytes in C# or platform 
default encoded in Java 17. For any non-ASCII name these diverge and the 
identifier is encoded wrong.
   
   There is also a size guard in the C# SDK which checks the code unit count 
rather than the byte count.
   
   The server expects a 1 to 255 **byte** UTF-8 name (`WireName`, 
`core/binary_protocol/src/primitives/identifier.rs`), so this is a client-side 
encoding bug; nothing on the server changes.
   
   ## C# `Iggy_SDK/Identifier.cs`
   
   ```csharp
   Length = value.Length,                  // UTF-16 code units
   Value  = Encoding.UTF8.GetBytes(value)  // UTF-8 bytes
   ```
   
   Callers in `Contracts/Tcp/TcpContracts.cs` size their buffers from `Length`, 
so a non-ASCII id is either truncated (`GetBytesFromIdentifier`), throws 
`ArgumentException: Destination is too short` (tight single-field buffers like 
`GetUser`), or overruns into the next field (multi-field buffers like 
`UpdateStream`). In the last case the server reads a shorter, different 
identifier.
   
   Reflecting the internal serializers off the shipping DLL:
   
   ```
   "café"       len 4, 5 bytes   -> truncated to invalid UTF-8, server rejects
   "naïve-café" len 10, 12 bytes -> UpdateStream sends streamId "naïve-caf"
   "日本語"      len 3, 9 bytes    -> UpdateStream sends streamId "日"
   ```
   
   ## Java `BytesSerializer.java`
   
   In Java, the buffer grows, so it doesn't throw, instead it writes all the 
bytes but declares the code-unit count, so every non-ASCII id produces a 
desynced frame that the server misreads.
   
   ```java
   ByteBuf buffer = Unpooled.buffer(2 + identifier.getName().length()); // 
UTF-16 code units
   buffer.writeByte(identifier.getName().length());                     // 
length prefix
   buffer.writeBytes(identifier.getName().getBytes());                  // 
UTF-8 bytes ( or platform-default on Java 17 ) 
   ```
   Produces:
   ```
   café       len 4, 5 bytes  -> 02 04 63 61 66 C3 A9      (declares 4, 5 
follow; A9 spills)
   naïve-café len 10, 12 bytes -> server reads id "naïve-caf", C3 A9 spill into 
next field
   日本語      len 3, 9 bytes   -> server reads id "日", 6 bytes spill into next 
field
   ```
   
   Additionally, `identifier.getName().getBytes()` is platform dependent in 
Java 17, ( UTF-8 in Java 18 after JEP 400 ), so 
`buffer.writeBytes(identifier.getName().getBytes());` will differ between 
platforms.
   
   ### Affected area / component
   
   C# SDK
   
   ### Deployment
   
   Not applicable
   
   ### Versions
   
   C# SDK and Java SDK at `master` commit `be2265c9e` 
   
   ### Hardware / environment
   
   _No response_
   
   ### Sample code
   
   _No response_
   
   ### Logs
   
   _No response_
   
   ### Iggy server config
   
   _No response_
   
   ### Reproduction
   
   ```
   // Apache.Iggy .NET SDK — Identifier.String UTF-16-code-unit / UTF-8-byte 
length mismatch.
   //
   // Self-contained: no SDK reference, no project needed. Run with the .NET 10 
SDK:
   //     dotnet run IdentifierRepro.cs
   // or drop this file into any `dotnet new console` project as Program.cs.
   //
   // It replicates the SDK's serialization logic line-for-line (source of 
truth in
   // the current tree, cited per method):
   //   Iggy_SDK/Identifier.cs:79                 Identifier.String
   //   Iggy_SDK/Utils/TcpMessageStreamHelpers.cs:54  GetBytesFromIdentifier
   //   Iggy_SDK/Extensions/Extensions.cs:87      WriteBytesFromIdentifier
   //   Iggy_SDK/Contracts/Tcp/TcpContracts.cs    GetUser (tight buffer), 
UpdateStream (multi-field)
   //
   // `Length` is taken from string.Length (UTF-16 code units) while the value 
is
   // UTF-8 bytes, so for any non-ASCII name the wire length disagrees with the
   // payload. Callers size their buffers from `Length`, so a non-ASCII id is
   // truncated, throws, or overruns into the next field.
   
   using System.Text;
   
   try { Console.OutputEncoding = Encoding.UTF8; } catch { /* console may not 
allow it */ }
   
   foreach (var name in new[] { "orders", "café", "naïve-café", "日本語" })
       Report(name);
   
   CapCheck();
   
   // --- SDK logic, replicated 
---------------------------------------------------
   
   // Identifier.String: guard checks value.Length (UTF-16 code units); Length 
is set
   // from it while Value is the UTF-8 bytes. (Identifier.cs:79)
   static (int Length, byte[] Value) MakeStringIdentifier(string value)
   {
       if (value.Length is 0 or > 255)
           throw new ArgumentException("Value has incorrect size, must be 
between 1 and 255", nameof(value));
       return (value.Length, Encoding.UTF8.GetBytes(value));
   }
   
   // GetBytesFromIdentifier: stackalloc[2 + Length], copies only Length bytes 
of the
   // value -> silent truncation. (TcpMessageStreamHelpers.cs:54)
   static byte[] GetBytesFromIdentifier(int length, byte[] value)
   {
       Span<byte> bytes = stackalloc byte[2 + length];
       bytes[0] = 2; // IdKind.String
       bytes[1] = (byte)length;
       for (var i = 0; i < length; i++)
           bytes[i + 2] = value[i];
       return bytes.ToArray();
   }
   
   // GetUser: tight single-field buffer stackalloc[Length + 2], then
   // WriteBytesFromIdentifier does Value.CopyTo(bytes[2..]) -> throws when the 
UTF-8
   // value is longer than Length. (TcpContracts.GetUser + Extensions.cs:87)
   static byte[] GetUser(int length, byte[] value)
   {
       var bytes = new byte[length + 2];
       bytes[0] = 2;
       bytes[1] = (byte)length;
       value.CopyTo(bytes.AsSpan(2)); // destination room == Length; throws if 
value longer
       return bytes;
   }
   
   // UpdateStream: multi-field buffer sized from char counts. 
WriteBytesFromIdentifier
   // over-copies the UTF-8 value into the name region; the name is then 
written at
   // position 2 + Length. The length prefix still says Length, so the server 
reads a
   // shorter, different identifier. (TcpContracts.UpdateStream)
   static byte[] UpdateStream(int idLength, byte[] idValue, string name)
   {
       var nameUtf8 = Encoding.UTF8.GetBytes(name);
       var bytes = new byte[idLength + nameUtf8.Length + 3];
       bytes[0] = 2;
       bytes[1] = (byte)idLength;
       idValue.CopyTo(bytes.AsSpan(2));               // may overrun into the 
name region
       var position = 2 + idLength;
       bytes[position] = (byte)nameUtf8.Length;
       nameUtf8.CopyTo(bytes.AsSpan(position + 1));
       return bytes;
   }
   
   // --- reporting 
---------------------------------------------------------------
   
   static void Report(string name)
   {
       var (length, value) = MakeStringIdentifier(name);
       Console.WriteLine($"==== \"{name}\" ====");
       Console.WriteLine($"  Length (UTF-16 code units = wire length byte) : 
{length}");
       Console.WriteLine($"  Value.Length (actual UTF-8 bytes)             : 
{value.Length}");
       Console.WriteLine($"  mismatch                                      : 
{(length != value.Length ? "YES" : "no")}");
   
       var wire = GetBytesFromIdentifier(length, value);
       var payload = wire[2..];
       Console.WriteLine($"  GetBytesFromIdentifier wire : {Hex(wire)}");
       Console.WriteLine($"    payload {Hex(payload)}  valid UTF-8? 
{IsUtf8(payload, out var dec)}"
                       + (dec is null ? "" : $" -> \"{dec}\"")
                       + (value.Length != payload.Length ? $"   [truncated, 
dropped {value.Length - payload.Length} byte(s)]" : ""));
       Console.WriteLine($"    correct encoding would be : 
{Hex(Expected(name))}");
   
       Console.WriteLine($"  GetUser                     : {Try(() => 
GetUser(length, value), out var u)}"
                       + (u is null ? "" : $"  {Hex(u)}"));
   
       if (Try(() => UpdateStream(length, value, "topic"), out var upd) == "ok" 
&& upd is not null)
       {
           var (serverId, serverName) = DecodeUpdateStreamAsServer(upd);
           Console.WriteLine($"  UpdateStream                : {Hex(upd)}");
           Console.WriteLine($"    server reads streamId=\"{serverId}\" 
name=\"{serverName}\""
                           + (serverId != name ? $"   [WRONG — asked for 
\"{name}\"]" : ""));
       }
   
       Console.WriteLine();
   }
   
   static void CapCheck()
   {
       var big = new string('あ', 200); // 200 code units, 600 UTF-8 bytes
       Console.WriteLine("==== cap check: 200x 'あ' ====");
       Console.WriteLine($"  codeUnits={big.Length}  
utf8Bytes={Encoding.UTF8.GetByteCount(big)}");
       Console.WriteLine($"  passes `value.Length is 0 or > 255` (code-unit) 
guard? {big.Length is > 0 and <= 255}");
       Console.WriteLine("  ...but 600 bytes exceeds the server's 255-BYTE 
WireName cap.");
   }
   
   // --- helpers 
-----------------------------------------------------------------
   
   static string Hex(byte[] b) => b.Length == 0 ? "(empty)" : 
Convert.ToHexString(b).Chunk(2).Select(c => new string(c)).Aggregate((a, c) => 
a + " " + c);
   
   static string IsUtf8(byte[] b, out string? decoded)
   {
       try { decoded = new UTF8Encoding(false, true).GetString(b); return 
"yes"; }
       catch (DecoderFallbackException) { decoded = null; return "NO (server 
rejects: WireError::InvalidUtf8)"; }
   }
   
   // What Identifier.String should have produced: length from the UTF-8 byte 
count.
   static byte[] Expected(string name)
   {
       var v = Encoding.UTF8.GetBytes(name);
       var b = new byte[2 + v.Length];
       b[0] = 2;
       b[1] = (byte)v.Length;
       v.CopyTo(b, 2);
       return b;
   }
   
   // Parse an UpdateStream frame the way the server's wire decoder does.
   static (string id, string name) DecodeUpdateStreamAsServer(byte[] f)
   {
       int len = f[1];
       var idBytes = f[2..(2 + len)];
       int namePos = 2 + len;
       int nameLen = f[namePos];
       var nameBytes = f[(namePos + 1)..(namePos + 1 + nameLen)];
       string Dec(byte[] x) { try { return new UTF8Encoding(false, 
true).GetString(x); } catch { return "<invalid-utf8>"; } }
       return (Dec(idBytes), Dec(nameBytes));
   }
   
   static string Try(Func<byte[]> f, out byte[]? result)
   {
       try { result = f(); return "ok"; }
       catch (Exception ex) { result = null; return $"THROWS 
{ex.GetType().Name}: {ex.Message}"; }
   }
   ```
   
   
   ### Contribution
   
   - [ ] I'm willing to submit a pull request to fix this bug
   
   ### Good first issue
   
   - [ ] I think this could be a good first issue for a new contributor


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to