[
https://issues.apache.org/jira/browse/FLINK-40697?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Yanquan Lv resolved FLINK-40697.
--------------------------------
Fix Version/s: cdc-3.7.0
Resolution: Fixed
Merged in master via e5c58adb4ec1c7a37331da8c25234237afd9a683.
> Release finished snapshot split metadata in the incremental source framework
> (flink-cdc-base)
> ---------------------------------------------------------------------------------------------
>
> Key: FLINK-40697
> URL: https://issues.apache.org/jira/browse/FLINK-40697
> Project: Flink
> Issue Type: Improvement
> Components: Flink CDC
> Reporter: Spoorthi Basu
> Assignee: Spoorthi Basu
> Priority: Major
> Labels: Flink-CDC, pull-request-available
> Fix For: cdc-3.7.0
>
>
> After the snapshot phase finishes and the stream split is assigned, the
> incremental source framework ({{flink-cdc-base}}) keeps the full
> snapshot-split metadata (assigned splits, finished offsets, table schemas)
> for the whole lifetime of the job and rebuilds it on every checkpoint. On a
> large multi-table job this keeps JobManager memory high and inflates
> checkpoint size.
> FLINK-39775 fixed this for the MySQL connector. This ticket generalizes the
> same mechanism into {{flink-cdc-base}} so the JDBC incremental dialects
> (Postgres, SqlServer, Oracle, Db2) benefit from it rather than each connector
> solving it separately.
> The release is opt-in behind a new option
> {{scan.incremental.snapshot.metadata.release.enabled}} (default {{false}}),
> so existing jobs are unaffected. It is tied to a completed checkpoint and
> guarded against a stale report from a failed-over reader. The enumerator
> checkpoint serializer is bumped only for jobs that enable the option, so
> default-off jobs keep the current format and stay restorable by an older
> connector build. The option is exposed through both the DataStream builders
> and Flink SQL.
> Relates to FLINK-39775.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)