richardcocks opened a new issue, #4056:
URL: https://github.com/apache/iggy/issues/4056
### Bug description
The C# SDK and Java SDK build the length prefix for a string identifier from
the string's `.Length` / `.length()`, which in each case count UTF-16 code
units. They then write the value as UTF-8 encoded bytes in C# or platform
default encoded in Java 17. For any non-ASCII name these diverge and the
identifier is encoded wrong.
There is also a size guard in the C# SDK which checks the code unit count
rather than the byte count.
The server expects a 1 to 255 **byte** UTF-8 name (`WireName`,
`core/binary_protocol/src/primitives/identifier.rs`), so this is a client-side
encoding bug; nothing on the server changes.
## C# `Iggy_SDK/Identifier.cs`
```csharp
Length = value.Length, // UTF-16 code units
Value = Encoding.UTF8.GetBytes(value) // UTF-8 bytes
```
Callers in `Contracts/Tcp/TcpContracts.cs` size their buffers from `Length`,
so a non-ASCII id is either truncated (`GetBytesFromIdentifier`), throws
`ArgumentException: Destination is too short` (tight single-field buffers like
`GetUser`), or overruns into the next field (multi-field buffers like
`UpdateStream`). In the last case the server reads a shorter, different
identifier.
Reflecting the internal serializers off the shipping DLL:
```
"café" len 4, 5 bytes -> truncated to invalid UTF-8, server rejects
"naïve-café" len 10, 12 bytes -> UpdateStream sends streamId "naïve-caf"
"日本語" len 3, 9 bytes -> UpdateStream sends streamId "日"
```
## Java `BytesSerializer.java`
In Java, the buffer grows, so it doesn't throw, instead it writes all the
bytes but declares the code-unit count, so every non-ASCII id produces a
desynced frame that the server misreads.
```java
ByteBuf buffer = Unpooled.buffer(2 + identifier.getName().length()); //
UTF-16 code units
buffer.writeByte(identifier.getName().length()); //
length prefix
buffer.writeBytes(identifier.getName().getBytes()); //
UTF-8 bytes ( or platform-default on Java 17 )
```
Produces:
```
café len 4, 5 bytes -> 02 04 63 61 66 C3 A9 (declares 4, 5
follow; A9 spills)
naïve-café len 10, 12 bytes -> server reads id "naïve-caf", C3 A9 spill into
next field
日本語 len 3, 9 bytes -> server reads id "日", 6 bytes spill into next
field
```
Additionally, `identifier.getName().getBytes()` is platform dependent in
Java 17, ( UTF-8 in Java 18 after JEP 400 ), so
`buffer.writeBytes(identifier.getName().getBytes());` will differ between
platforms.
### Affected area / component
C# SDK
### Deployment
Not applicable
### Versions
C# SDK and Java SDK at `master` commit `be2265c9e`
### Hardware / environment
_No response_
### Sample code
_No response_
### Logs
_No response_
### Iggy server config
_No response_
### Reproduction
```
// Apache.Iggy .NET SDK — Identifier.String UTF-16-code-unit / UTF-8-byte
length mismatch.
//
// Self-contained: no SDK reference, no project needed. Run with the .NET 10
SDK:
// dotnet run IdentifierRepro.cs
// or drop this file into any `dotnet new console` project as Program.cs.
//
// It replicates the SDK's serialization logic line-for-line (source of
truth in
// the current tree, cited per method):
// Iggy_SDK/Identifier.cs:79 Identifier.String
// Iggy_SDK/Utils/TcpMessageStreamHelpers.cs:54 GetBytesFromIdentifier
// Iggy_SDK/Extensions/Extensions.cs:87 WriteBytesFromIdentifier
// Iggy_SDK/Contracts/Tcp/TcpContracts.cs GetUser (tight buffer),
UpdateStream (multi-field)
//
// `Length` is taken from string.Length (UTF-16 code units) while the value
is
// UTF-8 bytes, so for any non-ASCII name the wire length disagrees with the
// payload. Callers size their buffers from `Length`, so a non-ASCII id is
// truncated, throws, or overruns into the next field.
using System.Text;
try { Console.OutputEncoding = Encoding.UTF8; } catch { /* console may not
allow it */ }
foreach (var name in new[] { "orders", "café", "naïve-café", "日本語" })
Report(name);
CapCheck();
// --- SDK logic, replicated
---------------------------------------------------
// Identifier.String: guard checks value.Length (UTF-16 code units); Length
is set
// from it while Value is the UTF-8 bytes. (Identifier.cs:79)
static (int Length, byte[] Value) MakeStringIdentifier(string value)
{
if (value.Length is 0 or > 255)
throw new ArgumentException("Value has incorrect size, must be
between 1 and 255", nameof(value));
return (value.Length, Encoding.UTF8.GetBytes(value));
}
// GetBytesFromIdentifier: stackalloc[2 + Length], copies only Length bytes
of the
// value -> silent truncation. (TcpMessageStreamHelpers.cs:54)
static byte[] GetBytesFromIdentifier(int length, byte[] value)
{
Span<byte> bytes = stackalloc byte[2 + length];
bytes[0] = 2; // IdKind.String
bytes[1] = (byte)length;
for (var i = 0; i < length; i++)
bytes[i + 2] = value[i];
return bytes.ToArray();
}
// GetUser: tight single-field buffer stackalloc[Length + 2], then
// WriteBytesFromIdentifier does Value.CopyTo(bytes[2..]) -> throws when the
UTF-8
// value is longer than Length. (TcpContracts.GetUser + Extensions.cs:87)
static byte[] GetUser(int length, byte[] value)
{
var bytes = new byte[length + 2];
bytes[0] = 2;
bytes[1] = (byte)length;
value.CopyTo(bytes.AsSpan(2)); // destination room == Length; throws if
value longer
return bytes;
}
// UpdateStream: multi-field buffer sized from char counts.
WriteBytesFromIdentifier
// over-copies the UTF-8 value into the name region; the name is then
written at
// position 2 + Length. The length prefix still says Length, so the server
reads a
// shorter, different identifier. (TcpContracts.UpdateStream)
static byte[] UpdateStream(int idLength, byte[] idValue, string name)
{
var nameUtf8 = Encoding.UTF8.GetBytes(name);
var bytes = new byte[idLength + nameUtf8.Length + 3];
bytes[0] = 2;
bytes[1] = (byte)idLength;
idValue.CopyTo(bytes.AsSpan(2)); // may overrun into the
name region
var position = 2 + idLength;
bytes[position] = (byte)nameUtf8.Length;
nameUtf8.CopyTo(bytes.AsSpan(position + 1));
return bytes;
}
// --- reporting
---------------------------------------------------------------
static void Report(string name)
{
var (length, value) = MakeStringIdentifier(name);
Console.WriteLine($"==== \"{name}\" ====");
Console.WriteLine($" Length (UTF-16 code units = wire length byte) :
{length}");
Console.WriteLine($" Value.Length (actual UTF-8 bytes) :
{value.Length}");
Console.WriteLine($" mismatch :
{(length != value.Length ? "YES" : "no")}");
var wire = GetBytesFromIdentifier(length, value);
var payload = wire[2..];
Console.WriteLine($" GetBytesFromIdentifier wire : {Hex(wire)}");
Console.WriteLine($" payload {Hex(payload)} valid UTF-8?
{IsUtf8(payload, out var dec)}"
+ (dec is null ? "" : $" -> \"{dec}\"")
+ (value.Length != payload.Length ? $" [truncated,
dropped {value.Length - payload.Length} byte(s)]" : ""));
Console.WriteLine($" correct encoding would be :
{Hex(Expected(name))}");
Console.WriteLine($" GetUser : {Try(() =>
GetUser(length, value), out var u)}"
+ (u is null ? "" : $" {Hex(u)}"));
if (Try(() => UpdateStream(length, value, "topic"), out var upd) == "ok"
&& upd is not null)
{
var (serverId, serverName) = DecodeUpdateStreamAsServer(upd);
Console.WriteLine($" UpdateStream : {Hex(upd)}");
Console.WriteLine($" server reads streamId=\"{serverId}\"
name=\"{serverName}\""
+ (serverId != name ? $" [WRONG — asked for
\"{name}\"]" : ""));
}
Console.WriteLine();
}
static void CapCheck()
{
var big = new string('あ', 200); // 200 code units, 600 UTF-8 bytes
Console.WriteLine("==== cap check: 200x 'あ' ====");
Console.WriteLine($" codeUnits={big.Length}
utf8Bytes={Encoding.UTF8.GetByteCount(big)}");
Console.WriteLine($" passes `value.Length is 0 or > 255` (code-unit)
guard? {big.Length is > 0 and <= 255}");
Console.WriteLine(" ...but 600 bytes exceeds the server's 255-BYTE
WireName cap.");
}
// --- helpers
-----------------------------------------------------------------
static string Hex(byte[] b) => b.Length == 0 ? "(empty)" :
Convert.ToHexString(b).Chunk(2).Select(c => new string(c)).Aggregate((a, c) =>
a + " " + c);
static string IsUtf8(byte[] b, out string? decoded)
{
try { decoded = new UTF8Encoding(false, true).GetString(b); return
"yes"; }
catch (DecoderFallbackException) { decoded = null; return "NO (server
rejects: WireError::InvalidUtf8)"; }
}
// What Identifier.String should have produced: length from the UTF-8 byte
count.
static byte[] Expected(string name)
{
var v = Encoding.UTF8.GetBytes(name);
var b = new byte[2 + v.Length];
b[0] = 2;
b[1] = (byte)v.Length;
v.CopyTo(b, 2);
return b;
}
// Parse an UpdateStream frame the way the server's wire decoder does.
static (string id, string name) DecodeUpdateStreamAsServer(byte[] f)
{
int len = f[1];
var idBytes = f[2..(2 + len)];
int namePos = 2 + len;
int nameLen = f[namePos];
var nameBytes = f[(namePos + 1)..(namePos + 1 + nameLen)];
string Dec(byte[] x) { try { return new UTF8Encoding(false,
true).GetString(x); } catch { return "<invalid-utf8>"; } }
return (Dec(idBytes), Dec(nameBytes));
}
static string Try(Func<byte[]> f, out byte[]? result)
{
try { result = f(); return "ok"; }
catch (Exception ex) { result = null; return $"THROWS
{ex.GetType().Name}: {ex.Message}"; }
}
```
### Contribution
- [ ] I'm willing to submit a pull request to fix this bug
### Good first issue
- [ ] I think this could be a good first issue for a new contributor
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]