This is an automated email from the ASF dual-hosted git repository.
corgy-w pushed a commit to branch dev
in repository https://gitbox.apache.org/repos/asf/seatunnel.git
The following commit(s) were added to refs/heads/dev by this push:
new 5d641fcfda [Docs][Connector-V2] Improve Milvus Prometheus GraphQL
Feishu and Fluss connector docs (#11728)
5d641fcfda is described below
commit 5d641fcfda390f0ecb7ba2479126ef70aeb2f957
Author: Daniel Carter <[email protected]>
AuthorDate: Sun Aug 16 13:08:05 2026 +0800
[Docs][Connector-V2] Improve Milvus Prometheus GraphQL Feishu and Fluss
connector docs (#11728)
---
docs/en/connectors/sink/Feishu.md | 2 +-
docs/en/connectors/sink/Fluss.md | 10 ++++++----
docs/en/connectors/source/Milvus.md | 6 +++---
docs/en/connectors/source/MySQL-CDC.md | 17 +++++++++++++++--
docs/en/connectors/source/Oracle-CDC.md | 10 +++++++++-
docs/en/connectors/source/PostgreSQL-CDC.md | 19 ++++++++++++++++---
docs/en/connectors/source/Prometheus.md | 2 +-
docs/zh/connectors/source/MySQL-CDC.md | 15 ++++++++++++---
docs/zh/connectors/source/Oracle-CDC.md | 6 +++++-
docs/zh/connectors/source/PostgreSQL-CDC.md | 15 ++++++++++++---
10 files changed, 80 insertions(+), 22 deletions(-)
diff --git a/docs/en/connectors/sink/Feishu.md
b/docs/en/connectors/sink/Feishu.md
index 8c5ca6491a..155e9dc753 100644
--- a/docs/en/connectors/sink/Feishu.md
+++ b/docs/en/connectors/sink/Feishu.md
@@ -222,7 +222,7 @@ source {
transform {
Sql {
plugin_input = "raw"
- query = "SELECT 'text' AS msg_type, named_struct('text', concat('User ',
name, ' is ', cast(age as string), ' years old')) AS content FROM raw"
+ query = "SELECT 'text' AS msg_type, MAP('text', concat('User ', name, ' is
', cast(age as string), ' years old')) AS content FROM raw"
}
}
diff --git a/docs/en/connectors/sink/Fluss.md b/docs/en/connectors/sink/Fluss.md
index fd977b3fc4..186269b8cf 100644
--- a/docs/en/connectors/sink/Fluss.md
+++ b/docs/en/connectors/sink/Fluss.md
@@ -239,9 +239,11 @@ sink {
### Replicate CDC changes across databases
Use placeholders to replicate CDC changes from one Fluss namespace to another.
-The `schema_name` placeholder resolves to the upstream database name, and
-`table_name` to the upstream table name, so a single sink configuration can fan
-out many tables.
+Fluss only has a database + table identifier (no schema), so the upstream
+`schemaName` is always null for a Fluss-to-Fluss topology. Use
`${database_name}`
+to capture the upstream Fluss database and `${table_name}` to capture the
+upstream Fluss table; the placeholder values are rewritten by the engine before
+the sink connects.
```hocon
env {
@@ -264,7 +266,7 @@ sink {
Fluss {
plugin_input = "fluss_cdc"
bootstrap.servers = "fluss-coordinator:9123"
- database = "sink_db_${schema_name}"
+ database = "sink_db_${database_name}"
table = "sink_${table_name}"
multi_table_sink_replica = 2
}
diff --git a/docs/en/connectors/source/Milvus.md
b/docs/en/connectors/source/Milvus.md
index 85bb6e87df..4c9825e6fa 100644
--- a/docs/en/connectors/source/Milvus.md
+++ b/docs/en/connectors/source/Milvus.md
@@ -55,14 +55,14 @@ Common use cases:
| database | String | No | `default` | Source database.
|
| collection | String | No | - | Source collection. If it is
set, only this collection is read. If it is not set, all collections under
`database` are read. The legacy alias `collection_name` is also accepted.
|
| batch_size | Integer | No | 1000 | Number of records to fetch
from Milvus in one batch. A larger value improves throughput but uses more
memory; set it to a smaller value when records contain large vector payloads.
|
-| rate_limit | Integer | No | 1000000 | Maximum number of records
the reader requests from Milvus per second. Use this to throttle a streaming
job against the Milvus quota (QPS) or gRPC message-size limit. Set to `-1` to
disable throttling. |
+| rate_limit | Integer | No | 1000000 | Server-side query rate limit
(QPS) applied to the source collection via the Milvus
`collection.queryRate.max.qps` property. The reader mutates this
collection-wide setting while the job is running, so it affects every client of
the collection, not just this SeaTunnel job. Set to `-1` to disable.
|
## Notes
- `database` defaults to `default`, so simple local Milvus jobs do not need to
set it.
- `collection` is optional. Set it when the job should read exactly one
collection.
- `batch_size` controls the per-fetch page size, not the parallelism of
readers. Tune it together with `parallelism` to balance throughput and memory.
-- `rate_limit` is a server-side hint that protects against Milvus `GRPC limit`
errors when reading large volumes of vector data. Leave it at the default
unless you see rate-limit or gRPC errors in the logs.
+- `rate_limit` mutates the server-side `collection.queryRate.max.qps` property
on every collection the job reads, so the new limit applies to all clients of
that collection while the job is running. The reader resets the property to
`-1` on close, but a job crash before `close()` will leave the collection
throttled until it is restored manually. Leave it at the default unless you
observe throttling errors in the logs.
- When `collection` is not set, the source discovers all collections in
`database` and exposes each collection as a separate SeaTunnel table.
- The source splits work by Milvus partition. Collections with a partition key
are read with one split; collections without a partition key are split by
partition name and assigned across readers.
- When the source reads a collection with partitions, downstream Milvus sink
can use that metadata to create the same partition names on the target
collection.
@@ -205,7 +205,7 @@ sink {
### Throttle Reads Against a Shared Cluster
When the Milvus cluster is shared with other jobs, lower `rate_limit` and
`batch_size`
-so the source does not exceed the cluster's gRPC message-size limit.
+so the source does not exceed the cluster's per-collection query quota.
```bash
env {
diff --git a/docs/en/connectors/source/MySQL-CDC.md
b/docs/en/connectors/source/MySQL-CDC.md
index cd8c92a932..c518937e83 100644
--- a/docs/en/connectors/source/MySQL-CDC.md
+++ b/docs/en/connectors/source/MySQL-CDC.md
@@ -476,7 +476,14 @@ sink {
### Read tables without a primary key
-For tables without a physical primary key, set `exactly_once = false` and
supply a unique column via `table-names-config.primaryKeys` when you need
stable row identity for downstream upserts.
+Pick the path that matches what the source table guarantees:
+
+- **Append-only workload** (no UPDATE/DELETE will ever be produced
downstream): keep
+ `exactly_once = false` and do not declare a primary key. The source falls
back to a best-effort
+ row identity. Without a usable key, the connector cannot apply UPDATE/DELETE
events safely.
+- **Unique non-primary column is available**: declare it via
`table-names-config.primaryKeys` and
+ set `exactly_once = true` so the snapshot and binlog phases both use the
configured key for
+ consistent row identity.
```hocon
env {
@@ -492,7 +499,13 @@ source {
password = "mysqlpw"
table-names = ["mysql_cdc.mysql_cdc_e2e_source_table_no_primary_key"]
url = "jdbc:mysql://mysql_cdc_e2e:3306/mysql_cdc"
- exactly_once = false
+ table-names-config = [
+ {
+ table = "mysql_cdc.mysql_cdc_e2e_source_table_no_primary_key"
+ primaryKeys = ["id"]
+ }
+ ]
+ exactly_once = true
}
}
```
diff --git a/docs/en/connectors/source/Oracle-CDC.md
b/docs/en/connectors/source/Oracle-CDC.md
index ea55214f28..df9bcb9b0e 100644
--- a/docs/en/connectors/source/Oracle-CDC.md
+++ b/docs/en/connectors/source/Oracle-CDC.md
@@ -446,7 +446,14 @@ source {
### Read tables without a primary key
-For tables without a physical primary key, set `exactly_once = false` and
supply a unique column via `table-names-config.primaryKeys` when you need
stable row identity for downstream upserts.
+Pick the path that matches what the source table guarantees:
+
+- **Append-only workload** (no UPDATE/DELETE will ever be produced
downstream): keep
+ `exactly_once = false` and do not declare a primary key. The source falls
back to a best-effort
+ row identity. Without a usable key, the connector cannot apply UPDATE/DELETE
events safely.
+- **Unique non-primary column is available**: declare it via
`table-names-config.primaryKeys` and
+ set `exactly_once = true` so the snapshot and redo-log phases both use the
configured key for
+ consistent row identity.
```hocon
source {
@@ -463,6 +470,7 @@ source {
primaryKeys = ["ID"]
}
]
+ exactly_once = true
}
}
```
diff --git a/docs/en/connectors/source/PostgreSQL-CDC.md
b/docs/en/connectors/source/PostgreSQL-CDC.md
index 8ea2ca1850..248a710b9e 100644
--- a/docs/en/connectors/source/PostgreSQL-CDC.md
+++ b/docs/en/connectors/source/PostgreSQL-CDC.md
@@ -225,7 +225,7 @@ Use `startup.mode = "snapshot-only"` when the job must
perform an initial snapsh
```hocon
env {
- execution.parallelism = 1
+ parallelism = 1
job.mode = "BATCH"
checkpoint.interval = 5000
}
@@ -262,7 +262,14 @@ In `snapshot-only` mode, the connector skips WAL streaming
entirely; configure `
### Read tables without a primary key
-For tables without a physical primary key, set `exactly_once = false` and
supply a unique column via `table-names-config.primaryKeys` when you need
stable row identity for downstream upserts.
+Pick the path that matches what the source table guarantees:
+
+- **Append-only workload** (no UPDATE/DELETE will ever be produced
downstream): keep
+ `exactly_once = false` and do not declare a primary key. The source falls
back to a best-effort
+ row identity. Without a usable key, the connector cannot apply UPDATE/DELETE
events safely.
+- **Unique non-primary column is available**: declare it via
`table-names-config.primaryKeys` and
+ set `exactly_once = true` so the snapshot and WAL phases both use the
configured key for
+ consistent row identity.
```hocon
source {
@@ -274,7 +281,13 @@ source {
table-names = ["postgres_cdc.inventory.full_types_no_primary_key"]
url =
"jdbc:postgresql://postgres_cdc_e2e:5432/postgres_cdc?loggerLevel=OFF"
decoding.plugin.name = "decoderbufs"
- exactly_once = false
+ table-names-config = [
+ {
+ table = "postgres_cdc.inventory.full_types_no_primary_key"
+ primaryKeys = ["id"]
+ }
+ ]
+ exactly_once = true
slot.name = "seatunnel_postgres_cdc"
}
}
diff --git a/docs/en/connectors/source/Prometheus.md
b/docs/en/connectors/source/Prometheus.md
index 7eff2ceafd..2515ba03de 100644
--- a/docs/en/connectors/source/Prometheus.md
+++ b/docs/en/connectors/source/Prometheus.md
@@ -178,7 +178,7 @@ source {
query = "rate(node_cpu_seconds_total{mode!=\"idle\"}[1m])"
query_type = "Range"
start = "2026-08-10T00:00:00Z"
- end = "now"
+ end = CURRENT_TIMESTAMP
step = "30s"
content_field = "$.data.result.*"
format = "json"
diff --git a/docs/zh/connectors/source/MySQL-CDC.md
b/docs/zh/connectors/source/MySQL-CDC.md
index a322cb91df..84d610fd9f 100644
--- a/docs/zh/connectors/source/MySQL-CDC.md
+++ b/docs/zh/connectors/source/MySQL-CDC.md
@@ -473,7 +473,10 @@ sink {
### 读取没有主键的表
-对于没有物理主键的表,将 `exactly_once` 设为 `false`,并通过 `table-names-config.primaryKeys`
提供一列作为下游 upsert 所需的稳定行标识。
+根据源表能够提供的保证来选择合适的路径:
+
+- **仅追加(append-only)场景**:源表不会产生 UPDATE/DELETE 事件,保持 `exactly_once = false`
且不声明主键,源端会退回到尽力而为的行标识。在没有可用主键的情况下,connector 无法安全地应用 UPDATE/DELETE 事件。
+- **存在唯一非主键列**:通过 `table-names-config.primaryKeys` 显式声明该列,并设置 `exactly_once =
true`,让快照阶段与 binlog 阶段都使用同一配置主键作为稳定的行标识。
```hocon
env {
@@ -489,12 +492,18 @@ source {
password = "mysqlpw"
table-names = ["mysql_cdc.mysql_cdc_e2e_source_table_no_primary_key"]
url = "jdbc:mysql://mysql_cdc_e2e:3306/mysql_cdc"
- exactly_once = false
+ table-names-config = [
+ {
+ table = "mysql_cdc.mysql_cdc_e2e_source_table_no_primary_key"
+ primaryKeys = ["id"]
+ }
+ ]
+ exactly_once = true
}
}
```
-如果没有可用的主键(无论配置的或物理的),connector 就无法安全地应用 UPDATE/DELETE
事件。仅在仅追加(append-only)场景或下游 sink 行为不依赖行标识时使用此模式。
+上述示例演示的是"逻辑主键"场景:源表本身没有物理主键,但通过 `table-names-config.primaryKeys`
显式声明了一列作为稳定行标识,并启用 `exactly_once = true`,让快照阶段与 binlog
阶段都使用同一逻辑主键。只有当被声明的列在源数据中确实保持唯一时,UPDATE/DELETE 才能被正确路由;如果源数据中存在重复值,行为将不再可靠。
### 从指定 Binlog 位置启动
diff --git a/docs/zh/connectors/source/Oracle-CDC.md
b/docs/zh/connectors/source/Oracle-CDC.md
index 213e6aece8..41c52fe2d4 100644
--- a/docs/zh/connectors/source/Oracle-CDC.md
+++ b/docs/zh/connectors/source/Oracle-CDC.md
@@ -445,7 +445,10 @@ source {
### 读取没有主键的表
-对于没有物理主键的表,将 `exactly_once` 设为 `false`,并通过 `table-names-config.primaryKeys`
提供一列作为下游 upsert 所需的稳定行标识。
+根据源表能够提供的保证来选择合适的路径:
+
+- **仅追加(append-only)场景**:源表不会产生 UPDATE/DELETE 事件,保持 `exactly_once = false`
且不声明主键,源端会退回到尽力而为的行标识。在没有可用主键的情况下,connector 无法安全地应用 UPDATE/DELETE 事件。
+- **存在唯一非主键列**:通过 `table-names-config.primaryKeys` 显式声明该列,并设置 `exactly_once =
true`,让快照阶段与 redo log 阶段都使用同一配置主键作为稳定的行标识。
```hocon
source {
@@ -462,6 +465,7 @@ source {
primaryKeys = ["ID"]
}
]
+ exactly_once = true
}
}
```
diff --git a/docs/zh/connectors/source/PostgreSQL-CDC.md
b/docs/zh/connectors/source/PostgreSQL-CDC.md
index f055672a39..cbcbffaca6 100644
--- a/docs/zh/connectors/source/PostgreSQL-CDC.md
+++ b/docs/zh/connectors/source/PostgreSQL-CDC.md
@@ -223,7 +223,7 @@ source {
```hocon
env {
- execution.parallelism = 1
+ parallelism = 1
job.mode = "BATCH"
checkpoint.interval = 5000
}
@@ -260,7 +260,10 @@ sink {
### 读取没有主键的表
-对于没有物理主键的表,将 `exactly_once` 设为 `false`,并通过 `table-names-config.primaryKeys`
提供一列作为下游 upsert 所需的稳定行标识。
+根据源表能够提供的保证来选择合适的路径:
+
+- **仅追加(append-only)场景**:源表不会产生 UPDATE/DELETE 事件,保持 `exactly_once = false`
且不声明主键,源端会退回到尽力而为的行标识。在没有可用主键的情况下,connector 无法安全地应用 UPDATE/DELETE 事件。
+- **存在唯一非主键列**:通过 `table-names-config.primaryKeys` 显式声明该列,并设置 `exactly_once =
true`,让快照阶段与 WAL 阶段都使用同一配置主键作为稳定的行标识。
```hocon
source {
@@ -272,7 +275,13 @@ source {
table-names = ["postgres_cdc.inventory.full_types_no_primary_key"]
url =
"jdbc:postgresql://postgres_cdc_e2e:5432/postgres_cdc?loggerLevel=OFF"
decoding.plugin.name = "decoderbufs"
- exactly_once = false
+ table-names-config = [
+ {
+ table = "postgres_cdc.inventory.full_types_no_primary_key"
+ primaryKeys = ["id"]
+ }
+ ]
+ exactly_once = true
slot.name = "seatunnel_postgres_cdc"
}
}