eskabetxe commented on PR #29422: URL: https://github.com/apache/flink/pull/29422#issuecomment-6045665168
Hi @gaborgsomogyi, thanks for looking! **Checkpoint vs savepoint:** yes, identical behavior. Both store the same `FileSinkCommittable` bytes through the same serializer chain (`FileSinkCommittableSerializer` -> `OutputStreamBasedPartFileWriter` recoverable serializers -> the filesystem's `RecoverableWriter` serializer); a savepoint is just a portable encoding of the same bytes. So restoring either with `flink-s3-fs-native` hits the same legacy-decode branch in `CommitterOperator#initializeState` — that's also where the original `EOFException` came from. **Endianness:** nothing is hardcoded ad-hoc — each format's endianness comes from the API that writes it. The legacy `S3RecoverableSerializer` (flink-s3-fs-base) writes via `ByteBuffer.order(LITTLE_ENDIAN)` (see its `serialize`), so the new decoder reads it back with `ByteBuffer.wrap(serialized).order(LITTLE_ENDIAN)` in `isLegacyFormat`/`decodeLegacy`. The native format uses `DataOutputStream`/`DataInputStream`, which the Java spec defines as big-endian. Detection is unambiguous in practice: mistaking native data for the legacy magic would require the first `writeUTF` length (big-endian, i.e. the object name's length) to be `0x3214` = 12820 bytes, far beyond the 1024-byte S3 key limit. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
