jackylee-ch opened a new pull request, #10198: URL: https://github.com/apache/paimon/pull/10198
## Purpose Flink has a `create_tag_from_watermark` procedure but Spark has none, so a Spark user cannot tag the first snapshot whose watermark reaches a target — even though Spark already exposes `rollback_to_watermark` and `create_tag_from_timestamp`. It is useful when a Flink writer commits event-time watermarks and a Spark job needs a consistent tag at a watermark boundary. ## Change Add `CreateTagFromWatermarkProcedure`, registered as `create_tag_from_watermark`. It mirrors the existing `CreateTagFromTimestampProcedure` structure and the Flink procedure's logic: resolve `SnapshotManager.laterOrEqualWatermark`, fall back to the earliest tag whose watermark already covers the target, else raise `SnapshotNotExistException`. ## Tests `CreateTagFromWatermarkProcedureTest` commits three snapshots with watermarks 1000/2000/3000 (through the stream write builder, since a Spark batch write carries no watermark) and asserts the resolved snapshot for a below/equal watermark and a `SnapshotNotExistException` past the last one. Written with Claude Code; verification is mine. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
