[
https://issues.apache.org/jira/browse/FLINK-40432?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated FLINK-40432:
-----------------------------------
Labels: pull-request-available (was: )
> Parallel snapshot/CDC split by long handle space is skewed for AUTO_INCREMENT
> PKs and can OOM
> ---------------------------------------------------------------------------------------------
>
> Key: FLINK-40432
> URL: https://issues.apache.org/jira/browse/FLINK-40432
> Project: Flink
> Issue Type: Bug
> Components: Flink CDC
> Reporter: lifulong
> Priority: Major
> Labels: pull-request-available
>
> h2. Problem
>
> TiDB CDC splits a table by evenly partitioning the entire \{{long}} handle
> space
> (\{{Long.MIN_VALUE}} ~ \{{Long.MAX_VALUE}}) across parallel subtasks.
>
> For clustered \{{AUTO_INCREMENT}} primary keys, real handles only occupy a
> tiny slice of that
> space (typically near 0). With parallelism 8, almost all rows land on a
> single subtask; the
> other subtasks get empty key ranges and stay idle. The busy subtask then OOMs
> during
> {\{KVClient.scan}}.
>
> h2. Expected Behavior
>
> Snapshot and incremental CDC should scale with the table's real data
> distribution (TiKV
> regions), not the theoretical \{{long}} domain.
>
> h2. Proposed Solution
>
> Split the table by actual TiKV regions and assign a contiguous region chunk
> to each
> subtask.
>
> h2. Related
>
> * PR: https://github.com/apache/flink-cdc/pull/4508
--
This message was sent by Atlassian Jira
(v8.20.10#820010)