eskabetxe commented on PR #29422:
URL: https://github.com/apache/flink/pull/29422#issuecomment-6045665168

   Hi @gaborgsomogyi, thanks for looking!
   
   **Checkpoint vs savepoint:** yes, identical behavior. Both store the same 
`FileSinkCommittable` bytes through the same serializer chain 
(`FileSinkCommittableSerializer` -> `OutputStreamBasedPartFileWriter` 
recoverable serializers -> the filesystem's `RecoverableWriter` serializer); a 
savepoint is just a portable encoding of the same bytes. So restoring either 
with `flink-s3-fs-native` hits the same legacy-decode branch in 
`CommitterOperator#initializeState` — that's also where the original 
`EOFException` came from.
   
   **Endianness:** nothing is hardcoded ad-hoc — each format's endianness comes 
from the API that writes it. The legacy `S3RecoverableSerializer` 
(flink-s3-fs-base) writes via `ByteBuffer.order(LITTLE_ENDIAN)` (see its 
`serialize`), so the new decoder reads it back with 
`ByteBuffer.wrap(serialized).order(LITTLE_ENDIAN)` in 
`isLegacyFormat`/`decodeLegacy`. The native format uses 
`DataOutputStream`/`DataInputStream`, which the Java spec defines as 
big-endian. Detection is unambiguous in practice: mistaking native data for the 
legacy magic would require the first `writeUTF` length (big-endian, i.e. the 
object name's length) to be `0x3214` = 12820 bytes, far beyond the 1024-byte S3 
key limit.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to