This is an automated email from the ASF dual-hosted git repository.
davidzollo pushed a commit to branch dev
in repository https://gitbox.apache.org/repos/asf/seatunnel.git
The following commit(s) were added to refs/heads/dev by this push:
new 5eb905b2d4 [Docs][Connector-V2] Document PostgreSQL/Opengauss CDC
stop.mode; bring Opengauss startup.mode in line with PostgreSQL (#11707)
5eb905b2d4 is described below
commit 5eb905b2d4ced290fbd6b9f292d63b05b1d4f70e
Author: Daniel <[email protected]>
AuthorDate: Fri Aug 14 15:19:55 2026 +0800
[Docs][Connector-V2] Document PostgreSQL/Opengauss CDC stop.mode; bring
Opengauss startup.mode in line with PostgreSQL (#11707)
Co-authored-by: danielnadean
<[email protected]>
Co-authored-by: Claude Sonnet 5 <[email protected]>
---
docs/en/connectors/source/Opengauss-CDC.md | 3 ++-
docs/en/connectors/source/PostgreSQL-CDC.md | 1 +
docs/zh/connectors/source/Opengauss-CDC.md | 3 ++-
docs/zh/connectors/source/PostgreSQL-CDC.md | 1 +
4 files changed, 6 insertions(+), 2 deletions(-)
diff --git a/docs/en/connectors/source/Opengauss-CDC.md
b/docs/en/connectors/source/Opengauss-CDC.md
index 2d166a4647..e7f258d861 100644
--- a/docs/en/connectors/source/Opengauss-CDC.md
+++ b/docs/en/connectors/source/Opengauss-CDC.md
@@ -75,7 +75,8 @@ select 'ALTER TABLE ' || schemaname || '.' || tablename || '
REPLICA IDENTITY FU
| table-names | List | Yes, if
`table-pattern` is not used | - | Tables to monitor. Use the fully
qualified `database.schema.table` format, for example:
`opengauss_cdc.inventory.orders`.
[...]
| table-pattern | String | Yes, if `table-names`
is not used | - | Regular expression for tables to monitor. Use the
fully qualified table name in the pattern, for example:
`opengauss_cdc\\.inventory\\..*`. `table-names` and `table-pattern` are
mutually exclusive.
[...]
| table-names-config | List | No | - |
Per-table config list. Example: `[{"table": "db1.schema1.table1","primaryKeys":
["key1"],"snapshotSplitColumn": "key2"}]`. Use `primaryKeys` for tables without
a physical primary key. `snapshotSplitColumn` must be a unique key; otherwise
SeaTunnel ignores it and selects a split column internally.
[...]
-| startup.mode | Enum | No | INITIAL |
Optional startup mode for Opengauss CDC consumer, valid enumerations are
`initial`, `earliest`, `latest`. <br/> `initial`: Synchronize historical data
at startup, and then synchronize incremental data.<br/> `earliest`: Startup
from the earliest offset possible.<br/> `latest`: Startup from the latest
offset.
[...]
+| startup.mode | Enum | No | INITIAL |
Optional startup mode for Opengauss CDC consumer, valid enumerations are
`initial`, `snapshot-only`, `committed-offset`, `earliest` and `latest`. <br/>
`initial`: Synchronize historical data at startup, and then synchronize
incremental data.<br/> `snapshot-only`: Synchronize historical data at startup
and finish as a bounded job without entering WAL streaming.<br/>
`committed-offset`: Skip snapshot data and st [...]
+| stop.mode | Enum | No | NEVER |
Optional stop mode for Opengauss CDC consumer. The only valid enumeration is
`never`: the source keeps streaming WAL changes and never stops on its own once
it reaches the incremental phase.
[...]
| snapshot.split.size | Integer | No | 8096 |
The split size (number of rows) of table snapshot, captured tables are split
into multiple splits when read the snapshot of table.
[...]
| snapshot.fetch.size | Integer | No | 1024 |
The maximum fetch size for per poll when read table snapshot.
[...]
| slot.name | String | No | seatunnel
| The Opengauss logical decoding slot name. Use a different slot name for each
CDC job that reads from the same Opengauss instance.
[...]
diff --git a/docs/en/connectors/source/PostgreSQL-CDC.md
b/docs/en/connectors/source/PostgreSQL-CDC.md
index 5cd1fa8757..8ea2ca1850 100644
--- a/docs/en/connectors/source/PostgreSQL-CDC.md
+++ b/docs/en/connectors/source/PostgreSQL-CDC.md
@@ -98,6 +98,7 @@ ALTER TABLE your_table_name REPLICA IDENTITY FULL;
| table-pattern | String | Yes, if `table-names`
is not used | - | Regular expression for tables to monitor. Use the
fully qualified table name in the pattern, for example:
`postgres_cdc\\.inventory\\..*`. `table-names` and `table-pattern` are mutually
exclusive.
[...]
| table-names-config | List | No | - |
Per-table config list. Example: `[{"table": "db1.schema1.table1","primaryKeys":
["key1"],"snapshotSplitColumn": "key2"}]`. Use `primaryKeys` for tables without
a physical primary key. `snapshotSplitColumn` must be a unique key; otherwise
SeaTunnel ignores it and selects a split column internally.
[...]
| startup.mode | Enum | No | INITIAL |
Optional startup mode for PostgreSQL CDC consumer, valid enumerations are
`initial`, `snapshot-only`, `committed-offset`, `earliest` and `latest`. <br/>
`initial`: Synchronize historical data at startup, and then synchronize
incremental data.<br/> `snapshot-only`: Synchronize historical data at startup
and finish as a bounded job without entering WAL streaming.<br/>
`committed-offset`: Skip snapshot data and s [...]
+| stop.mode | Enum | No | NEVER |
Optional stop mode for PostgreSQL CDC consumer. The only valid enumeration is
`never`: the source keeps streaming WAL changes and never stops on its own once
it reaches the incremental phase.
[...]
| snapshot.split.size | Integer | No | 8096 |
The split size (number of rows) of table snapshot, captured tables are split
into multiple splits when read the snapshot of table.
[...]
| snapshot.fetch.size | Integer | No | 1024 |
The maximum fetch size for per poll when read table snapshot.
[...]
| slot.name | String | No | seatunnel
| The PostgreSQL logical decoding slot name. Use a different slot name for each
CDC job that reads from the same PostgreSQL instance.
[...]
diff --git a/docs/zh/connectors/source/Opengauss-CDC.md
b/docs/zh/connectors/source/Opengauss-CDC.md
index c3c0cfe874..6ac3046edd 100644
--- a/docs/zh/connectors/source/Opengauss-CDC.md
+++ b/docs/zh/connectors/source/Opengauss-CDC.md
@@ -74,7 +74,8 @@ select 'ALTER TABLE ' || schemaname || '.' || tablename || '
REPLICA IDENTITY FU
| table-names | 列表 | 二选一 | - |
需要监控的表。请使用完整的 `database.schema.table` 格式,例如:`opengauss_cdc.inventory.orders`。
|
| table-pattern | 字符串 | 二选一 | - |
需要监控的表名正则表达式。正则需要匹配完整表名,例如:`opengauss_cdc\\.inventory\\..*`。`table-names` 和
`table-pattern` 互斥。
|
| table-names-config | 列表 | 否 | - |
表级配置列表。例如:`[{"table": "db1.schema1.table1","primaryKeys":
["key1"],"snapshotSplitColumn": "key2"}]`。无物理主键表可通过 `primaryKeys`
指定唯一键。`snapshotSplitColumn` 必须是唯一键,否则 SeaTunnel 会忽略该配置并自动选择拆分列。
|
-| startup.mode | 枚举 | 否 | INITIAL |
Opengauss CDC消费者的可选启动模式, 有效的枚举是`initial`, `earliest`, `latest`. <br/>
`initial`: 启动时同步历史数据,然后同步增量数据 <br/> `earliest`: 从可能的最早偏移量启动 <br/> `latest`:
从最近的偏移量启动 |
+| startup.mode | 枚举 | 否 | INITIAL |
Opengauss CDC消费者的可选启动模式,有效枚举为
`initial`、`snapshot-only`、`committed-offset`、`earliest` 和 `latest`。<br/>
`initial`: 启动时同步历史数据,然后同步增量数据。<br/> `snapshot-only`: 仅同步启动时的历史数据,然后以有界任务结束,不进入
WAL 流读取。<br/> `committed-offset`: 跳过快照数据,从配置的复制槽已提交 LSN 开始读取 WAL。该模式要求显式配置
`slot.name`,如果复制槽不存在或没有可用的已提交 LSN,则启动失败。<br/> `earliest`: 从可能的最早偏移量启动。<br/>
`latest`: 从最新偏移量启动。 |
+| stop.mode | 枚举 | 否 | NEVER |
Opengauss CDC 消费者的可选停止模式。唯一有效的枚举值是 `never`:一旦进入增量阶段,数据源会持续读取 WAL 变更,不会自行停止。
|
| snapshot.split.size | 整型 | 否 | 8096 |
表快照的分割大小(行数),在读取表的快照时,捕获的表被分割成多个split
|
| snapshot.fetch.size | 整型 | 否 | 1024 |
读取表快照时,每次轮询的最大读取大小
|
| slot.name | 字符串 | 否 | seatunnel |
Opengauss 逻辑解码槽名称。同一个 Opengauss 实例上如果有多个 CDC 任务,请为每个任务配置不同的 `slot.name`。
|
diff --git a/docs/zh/connectors/source/PostgreSQL-CDC.md
b/docs/zh/connectors/source/PostgreSQL-CDC.md
index 602e0e060c..f055672a39 100644
--- a/docs/zh/connectors/source/PostgreSQL-CDC.md
+++ b/docs/zh/connectors/source/PostgreSQL-CDC.md
@@ -96,6 +96,7 @@ ALTER TABLE your_table_name REPLICA IDENTITY FULL;
| table-pattern | String | 二选一 | - |
需要监控的表名正则表达式。正则需要匹配完整表名,例如:`postgres_cdc\\.inventory\\..*`。`table-names` 和
`table-pattern` 互斥。
[...]
| table-names-config | List | 否 | - |
表级配置列表。例如:`[{"table": "db1.schema1.table1","primaryKeys":
["key1"],"snapshotSplitColumn": "key2"}]`。无物理主键表可通过 `primaryKeys`
指定唯一键。`snapshotSplitColumn` 必须是唯一键,否则 SeaTunnel 会忽略该配置并自动选择拆分列。
[...]
| startup.mode | Enum | 否 | INITIAL |
PostgreSQL CDC 消费者的可选启动模式,有效枚举为
`initial`、`snapshot-only`、`committed-offset`、`earliest` 和 `latest`。<br/>
`initial`: 启动时同步历史数据,然后同步增量数据。<br/> `snapshot-only`: 仅同步启动时的历史数据,然后以有界任务结束,不进入
WAL 流读取。<br/> `committed-offset`: 跳过快照数据,从配置的复制槽已提交 LSN 开始读取 WAL。该模式要求显式配置
`slot.name`,如果复制槽不存在或没有可用的已提交 LSN,则启动失败。<br/> `earliest`: 从可能的最早偏移量启动。<br/>
`latest`: 从最新偏移量启动。 |
+| stop.mode | Enum | 否 | NEVER |
PostgreSQL CDC 消费者的可选停止模式。唯一有效的枚举值是 `never`:一旦进入增量阶段,数据源会持续读取 WAL 变更,不会自行停止。
[...]
| snapshot.split.size | Integer | 否 | 8096 |
表快照的拆分大小(行数),捕获的表在读取表快照时被拆分成多个拆分。
[...]
| snapshot.fetch.size | Integer | 否 | 1024 |
读取表快照时每次轮询的最大获取大小。
[...]
| slot.name | String | 否 | seatunnel |
PostgreSQL 逻辑解码槽名称。同一个 PostgreSQL 实例上如果有多个 CDC 任务,请为每个任务配置不同的 `slot.name`。
|