jackylee-ch opened a new pull request, #10198:
URL: https://github.com/apache/paimon/pull/10198

   ## Purpose
   
   Flink has a `create_tag_from_watermark` procedure but Spark has none, so a
   Spark user cannot tag the first snapshot whose watermark reaches a target —
   even though Spark already exposes `rollback_to_watermark` and
   `create_tag_from_timestamp`. It is useful when a Flink writer commits
   event-time watermarks and a Spark job needs a consistent tag at a watermark
   boundary.
   
   ## Change
   
   Add `CreateTagFromWatermarkProcedure`, registered as
   `create_tag_from_watermark`. It mirrors the existing
   `CreateTagFromTimestampProcedure` structure and the Flink procedure's logic:
   resolve `SnapshotManager.laterOrEqualWatermark`, fall back to the earliest 
tag
   whose watermark already covers the target, else raise
   `SnapshotNotExistException`.
   
   ## Tests
   
   `CreateTagFromWatermarkProcedureTest` commits three snapshots with watermarks
   1000/2000/3000 (through the stream write builder, since a Spark batch write
   carries no watermark) and asserts the resolved snapshot for a below/equal
   watermark and a `SnapshotNotExistException` past the last one.
   
   Written with Claude Code; verification is mine.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to