+1 (non-binding)

Hi Chanhae,

The proposed approaches look very promising for handling skewed data
distributions in practice.

Looking forward to the FLIP progress.

Best,
Sumin Joo

2026년 7월 7일 (화) 오후 7:09, Chanhae Oh <[email protected]>님이 작성:

> Hi all,
>
> Following up on FLINK-36812 [1], I've drafted a FLIP to address the
> limitations of the current scan.partition.column implementation.
> The existing approach splits data by equal numeric intervals, which breaks
> down under two common scenarios: skewed data distribution and the absence
> of a suitable numeric column (e.g., string-typed partition keys).
>
> Two extensions are proposed:
>
> 1. Physical ID Partitioning (scan.partition.use-physical-id)
> For databases that expose a physical row identifier (Oracle ROWID,
> PostgreSQL ctid), the connector can derive partition bounds automatically
> from DB-native metadata.
> Since rows are physically distributed across storage blocks, this produces
> naturally balanced splits with no user-defined partition column required.
>
> 2. Boundary Query (scan.partition.boundary-query)
> A SQL query returning ordered split points for the column specified in
> scan.partition.column. The connector uses these values to generate range
> predicates, enabling distribution-aware parallel reads for any column type
> — including strings.
> scan.partition.num is required and determines how many partitions are
> created.
>  This is conceptually close to Apache Sqoop's --boundary-query, extended
> to support string-type columns natively.
>
> Both extensions leave existing behavior unchanged.
>
> Full design: [2]
>
> Feedback and questions are welcome.
>
> Thanks,
> Chanhae Oh
>
> [1] https://issues.apache.org/jira/browse/FLINK-36812
> [2]
> https://docs.google.com/document/d/1fgnYuOo9U6PPuwJgvnmdfGfV70hhO2xBTSo0OR46r1g/edit?usp=sharing
>
>

Reply via email to