This is an automated email from the ASF dual-hosted git repository.

github-merge-queue[bot] pushed a commit to branch dev
in repository https://gitbox.apache.org/repos/asf/seatunnel.git


The following commit(s) were added to refs/heads/dev by this push:
     new f0c4590557 [Docs][Connector-V2] Improve Datahub Socket Web3j and 
ActiveMQ connector docs (#11788)
f0c4590557 is described below

commit f0c459055706e362b617212abf330c4ca3f9502e
Author: Daniel Carter <[email protected]>
AuthorDate: Sun Sep 6 09:58:47 2026 +0000

    [Docs][Connector-V2] Improve Datahub Socket Web3j and ActiveMQ connector 
docs (#11788)
    
    Co-authored-by: DanielCarter-stack 
<[email protected]>
---
 docs/en/connectors/sink/Activemq.md | 25 +++++++++++++++++++++++++
 docs/en/connectors/sink/Datahub.md  | 15 +++++++++++++++
 docs/en/connectors/sink/Socket.md   | 15 +++++++++++++++
 docs/en/connectors/source/Socket.md | 15 +++++++++++++++
 docs/en/connectors/source/Web3j.md  | 15 +++++++++++++++
 docs/zh/connectors/sink/Activemq.md | 19 +++++++++++++++++++
 docs/zh/connectors/sink/Datahub.md  | 15 +++++++++++++++
 docs/zh/connectors/sink/Socket.md   | 15 +++++++++++++++
 docs/zh/connectors/source/Socket.md | 15 +++++++++++++++
 docs/zh/connectors/source/Web3j.md  | 15 +++++++++++++++
 10 files changed, 164 insertions(+)

diff --git a/docs/en/connectors/sink/Activemq.md 
b/docs/en/connectors/sink/Activemq.md
index 072febf334..e5590e5f5c 100644
--- a/docs/en/connectors/sink/Activemq.md
+++ b/docs/en/connectors/sink/Activemq.md
@@ -4,6 +4,12 @@ import ChangeLog from '../changelog/connector-activemq.md';
 
 > ActiveMQ sink connector
 
+## Support Those Engines
+
+> Spark<br/>
+> Flink<br/>
+> SeaTunnel Zeta<br/>
+
 ## Description
 
 Write SeaTunnel rows to an ActiveMQ queue. Each row is serialized as a JSON 
text message. This is
@@ -108,6 +114,25 @@ sink {
 }
 ```
 
+## FAQ
+
+### Does ActiveMQ sink support topics as well as queues?
+
+No. The current sink writes only to JMS queues identified by `queue_name`; 
topic destinations are not supported by the connector factory. If you need 
publish/subscribe semantics, use a generic JMS connector or a separate 
ActiveMQ-targeted bridge that maps the upstream rows onto a topic — but stick 
to queues when you want to drive this connector directly.
+
+### How are `username` and `password` validated?
+
+They are optional. When set, both must be present (configuring one without the 
other fails the job). They take effect at the JMS connection factory level, 
which means they override any credentials already embedded in `uri`. If your 
broker requires an account, prefer the explicit `username`/`password` options 
over embedding them in the URL so they appear in job config logs instead of 
inside the connection string.
+
+### What message format does each row become?
+
+Each SeaTunnel row is serialized as one JSON text message sent to the 
configured `queue_name`. There is no `format` option; the JSON shape is fixed 
by the sink's serializer, so any consumer that wants a different encoding must 
decode the JSON body itself first.
+
+### Is exactly-once delivery supported?
+
+No. The sink is best-effort with bounded reconnect behavior driven by the 
underlying JMS client. Enable checkpointing at the job level if at-least-once 
replay from upstream is acceptable.
+
 ## Changelog
 
 <ChangeLog />
+
diff --git a/docs/en/connectors/sink/Datahub.md 
b/docs/en/connectors/sink/Datahub.md
index f6b9ea8b6e..e58bcca3da 100644
--- a/docs/en/connectors/sink/Datahub.md
+++ b/docs/en/connectors/sink/Datahub.md
@@ -174,6 +174,21 @@ sink {
 }
 ```
 
+## FAQ
+
+### Does DataHub sink support exactly-once delivery?
+
+No. The connector performs best-effort writes with bounded retries 
(`retryTimes`, default `3`). Failures beyond the retry budget are surfaced to 
the job rather than silently swallowed, so enable checkpointing at the job 
level if at-least-once replay from upstream is acceptable.
+
+### How does multi-table routing work?
+
+When the upstream source emits more than one table, set `topic` to a value 
containing the `${table}` placeholder (for example `topic = "${table}"`). Each 
input table is then routed to a DataHub topic that shares its name. 
`${table_name}` is still recognized as a deprecated alias — prefer `${table}` 
in new jobs. The connector does not automatically create the target topic; 
create it in the DataHub project beforehand and ensure its schema field names 
match the upstream SeaTunnel schema.
+
+### Why is `topic` a required field even for single-table jobs?
+
+DataHub writes are schema-bound: the sink serializes each row against the 
field names declared on the target topic. A missing or empty `topic` leaves the 
connector with nowhere to write, so it is validated at job start rather than 
guessed from the project.
+
 ## Changelog
 
 <ChangeLog />
+
diff --git a/docs/en/connectors/sink/Socket.md 
b/docs/en/connectors/sink/Socket.md
index 3d61cb7f32..7e6a46ca7c 100644
--- a/docs/en/connectors/sink/Socket.md
+++ b/docs/en/connectors/sink/Socket.md
@@ -89,6 +89,21 @@ nc -l -v 9999
 {"name":"jared","age":17}{"name":"jared","age":18}...
 ```
 
+## FAQ
+
+### Does Socket sink append any delimiter between records?
+
+No. The sink serializes each SeaTunnel row to JSON via 
`JsonSerializationSchema` and writes the bytes to the TCP stream with no 
separator at all — neither `\n` nor any other character. Multiple records 
travel as one continuous concatenated byte stream (for example 
`{"a":1}{"a":2}{"a":3}`). The peer must therefore use a streaming JSON parser 
(such as Jackson's `MappingIterator`), not a line-oriented parser, to split 
records.
+
+### What does `max_retries` control exactly?
+
+`max_retries` is the number of times the writer retries a failed send after 
the TCP connection is established (connection refused, broken pipe, write 
timeouts, etc.). Default is `3`. Set it to `-1` to retry indefinitely, or `0` 
to fail the record immediately on the first write failure.
+
+### Can several Socket sink writers run in parallel?
+
+Yes. Each writer opens its own TCP connection to `host:port`, so 
`env.parallelism` greater than 1 produces N concurrent connections to the same 
socket server. Make sure the receiver on the other side is designed to handle 
multiple clients, otherwise it will only see one client at a time.
+
 ## Changelog
 
 <ChangeLog />
+
diff --git a/docs/en/connectors/source/Socket.md 
b/docs/en/connectors/source/Socket.md
index 3ed34c0baf..5c15aa4978 100644
--- a/docs/en/connectors/source/Socket.md
+++ b/docs/en/connectors/source/Socket.md
@@ -137,6 +137,21 @@ sink {
 }
 ```
 
+## FAQ
+
+### Does Socket source checkpoint its read position?
+
+No. The source does not record a server-side offset or message sequence; after 
a restart it starts the next read from whatever the socket returns. Use Kafka, 
Pulsar, or another offset-tracking source if the job needs replayable or 
exactly-once behavior. Socket source is intended for local debugging or 
one-shot text streams.
+
+### How are empty lines and partial trailing records handled?
+
+In batch mode the reader consumes whatever is currently buffered on the 
socket, splits on `\n`, emits each complete line as one STRING row, and then 
emits the trailing partial line (if any) as a final row before finishing. Empty 
lines are *not* skipped — they become a row whose payload is the empty string. 
In streaming mode empty lines are emitted in real time the same way.
+
+### Can I parallelize Socket source?
+
+No. The reader binds to a single client connection to the configured 
`host:port` and uses one split, so the job's `env.parallelism` is effectively 
capped at 1 for this source. To scale, run independent jobs — each with its own 
`host`/`port` pair — instead of increasing parallelism within a single Socket 
source.
+
 ## Changelog
 
 <ChangeLog />
+
diff --git a/docs/en/connectors/source/Web3j.md 
b/docs/en/connectors/source/Web3j.md
index e71d2f68a0..90b1a3b77e 100644
--- a/docs/en/connectors/source/Web3j.md
+++ b/docs/en/connectors/source/Web3j.md
@@ -128,6 +128,21 @@ sink {
 }
 ```
 
+## FAQ
+
+### How is the polling rate controlled?
+
+The connector issues one HTTP `eth_blockNumber` RPC per poll, blocks until the 
provider returns, and emits the result as one row. The actual emit cadence is 
therefore tied to the configured provider's response time — it cannot be set 
independently from a config option. Pair the source with a downstream sink that 
uses checkpointing if you need replayable backpressure; the Web3j source itself 
has no rate-limit or throttle setting.
+
+### What payload does the `value` field contain?
+
+A JSON object with the schema `{"blockNumber": <number>, "timestamp": 
"<ISO-8601 UTC>"}`. The `blockNumber` is the latest block head returned by the 
provider, and `timestamp` is generated by the connector at the moment it 
observes the response. The connector never inspects the payload contents; if 
you need additional fields (transaction count, gas used, peer metadata, etc.) 
you must call a different RPC or post-process upstream using a SQL Transform.
+
+### Does Web3j source support authentication beyond the URL?
+
+No. The connector passes the configured `url` through to the underlying HTTP 
client as-is. Authentication is therefore whatever the provider accepts in the 
URL — usually an embedded API key for hosted services such as Infura or 
Alchemy. There is no separate header-based auth path, so do not put secrets 
that should be rotated frequently into the URL; instead issue a long-lived 
provider key and store it via your normal secret-management system.
+
 ## Changelog
 
 <ChangeLog />
+
diff --git a/docs/zh/connectors/sink/Activemq.md 
b/docs/zh/connectors/sink/Activemq.md
index fbb1cea36a..c9f045ce25 100644
--- a/docs/zh/connectors/sink/Activemq.md
+++ b/docs/zh/connectors/sink/Activemq.md
@@ -107,6 +107,25 @@ sink {
 }
 ```
 
+## FAQ
+
+### ActiveMQ Sink 支持 topic 吗?
+
+不支持。当前 Sink 只会写入由 `queue_name` 指定的 JMS 队列,连接器工厂并不支持 topic 
目的地。如果需要发布/订阅语义,请改用通用的 JMS 连接器或者其它专门面向 ActiveMQ topic 的桥接组件;要直接驱动本连接器,请继续用 
queue。
+
+### `username` 和 `password` 是怎么校验的?
+
+这两项都是可选的。设置的时候必须同时设置(只设其中一个会让任务启动失败)。它们作用于 JMS 连接工厂这一层,因此会覆盖已经嵌在 `uri` 里的凭证。如果 
Broker 需要鉴权,建议显式配置 `username`/`password` 而不是把它们写在 URL 
里——这样凭证会出现在任务配置日志里,而不是埋在连接串里。
+
+### 每条数据最终变成什么格式的消息?
+
+每行 SeaTunnel 数据会被序列化为一条 JSON 文本消息写到配置的 `queue_name` 里。连接器没有 `format` 配置项,JSON 
结构由 Sink 序列化器固定,所以对端如果想要其它编码,需要自己先解码 JSON body。
+
+### 支持精确一次写入吗?
+
+不支持。Sink 是尽力而为模式,重连行为由底层 JMS 客户端决定。如果可以接受来自上游的至少一次重放,请在任务级别启用 checkpoint。
+
 ## 变更日志
 
 <ChangeLog />
+
diff --git a/docs/zh/connectors/sink/Datahub.md 
b/docs/zh/connectors/sink/Datahub.md
index 6051682a87..338ded00ae 100644
--- a/docs/zh/connectors/sink/Datahub.md
+++ b/docs/zh/connectors/sink/Datahub.md
@@ -166,6 +166,21 @@ sink {
 }
 ```
 
+## 常见问题
+
+### DataHub Sink 是否支持精确一次写入?
+
+不支持。连接器按最大重试次数(`retryTimes`,默认 
`3`)执行尽力而为的写入;超出重试预算后失败会抛给上层任务而不是默默吞掉。如果可以接受来自上游的至少一次重放,请在任务级别启用 checkpoint。
+
+### 多表写入是如何路由的?
+
+当上游 source 输出多张表时,把 `topic` 设置成包含 `${table}` 占位符的值(例如 `topic = 
"${table}"`),每张输入表就会被写到同名 DataHub topic。`${table_name}` 仍然作为已废弃的别名可以识别——新任务建议使用 
`${table}`。该连接器不会自动创建目标 topic,请提前在 DataHub 项目中创建好,并保证其字段名与上游 SeaTunnel schema 
一致。
+
+### 为什么单表任务 `topic` 也是必填?
+
+DataHub 写入是 schema 绑定的:sink 会按目标 topic 上声明的字段名序列化每行数据。如果 `topic` 
缺失或为空,连接器就找不到写入目标,因此会在任务启动阶段校验,而不是根据 project 自动猜测。
+
 ## 变更日志
 
 <ChangeLog />
+
diff --git a/docs/zh/connectors/sink/Socket.md 
b/docs/zh/connectors/sink/Socket.md
index ac6af3e326..3e21cc3971 100644
--- a/docs/zh/connectors/sink/Socket.md
+++ b/docs/zh/connectors/sink/Socket.md
@@ -87,6 +87,21 @@ nc -l -v 9999
 {"name":"jared","age":17}{"name":"jared","age":18}...
 ```
 
+## 常见问题
+
+### Socket Sink 会在记录之间追加分隔符吗?
+
+不会。Sink 用 `JsonSerializationSchema` 把每行序列化为 JSON,然后直接写到 TCP 
流,*不会*追加任何分隔符——既不会追加 `\n`,也不会追加任何其它字符。多条记录会作为一条连续的、拼接在一起的字节流传出去(例如 
`{"a":1}{"a":2}{"a":3}`)。所以对端必须使用流式 JSON 解析器(例如 Jackson 的 
`MappingIterator`),而不是按行解析的解析器来切分记录。
+
+### `max_retries` 到底控制什么?
+
+`max_retries` 是 Writer 在 TCP 连接已经建立后,发送失败时重试的次数(连接被拒绝、管道破裂、写超时等场景)。默认值为 
`3`。设置为 `-1` 表示无限重试,设置为 `0` 表示第一次写失败就立即抛错。
+
+### Socket Sink 能并行写吗?
+
+可以。每个 Writer 会各自建立一条到 `host:port` 的 TCP 连接,因此 `env.parallelism` 大于 1 时会向同一个 
socket server 同时打开 N 条连接。请确认对端能够接受多客户端连接,否则只会同时处理一个客户端。
+
 ## 变更日志
 
 <ChangeLog />
+
diff --git a/docs/zh/connectors/source/Socket.md 
b/docs/zh/connectors/source/Socket.md
index 74d7adc7d1..bf5d8b7c36 100644
--- a/docs/zh/connectors/source/Socket.md
+++ b/docs/zh/connectors/source/Socket.md
@@ -133,6 +133,21 @@ sink {
 }
 ```
 
+## 常见问题
+
+### Socket Source 会对读取位点做 checkpoint 吗?
+
+不会。连接器不会记录服务端偏移量或消息序号,重启后会从 socket 当前能读到的内容继续读下一次。如果任务需要可重放或精确一次,请改用 
Kafka、Pulsar 等具备位点跟踪能力的 Source。Socket Source 主要面向本地调试或一次性文本流。
+
+### 空行和末尾不完整的记录会怎样处理?
+
+批处理模式下,读取器把 socket 上当前缓冲的内容一次性读出来,按 `\n` 切分,每条完整的行作为一条 STRING 
数据发出去;如果末尾还有一段不完整的行,也会作为最后一行发出再结束任务。空行*不会*被跳过——它会变成一条内容为空字符串的数据。流处理模式下空行也是同样按原样即时发出去。
+
+### Socket Source 能不能并行?
+
+不能。读取器只会和配置的 `host:port` 建立一条客户端连接,使用单个 split,因此该 Source 在任务里的 
`env.parallelism` 实际上被限制为 1。要提高吞吐,请改跑多个独立任务(每个任务使用各自的 `host`/`port`),而不是在同一个 
Socket Source 里调高并行度。
+
 ## 变更日志
 
 <ChangeLog />
+
diff --git a/docs/zh/connectors/source/Web3j.md 
b/docs/zh/connectors/source/Web3j.md
index bc90c7adaf..f142778a3b 100644
--- a/docs/zh/connectors/source/Web3j.md
+++ b/docs/zh/connectors/source/Web3j.md
@@ -124,6 +124,21 @@ sink {
 }
 ```
 
+## 常见问题
+
+### 轮询节奏是怎么控制的?
+
+连接器每次轮询都会发一次 HTTP `eth_blockNumber` RPC,等到 provider 
返回结果后再把这一条数据发出来。因此实际的发数据节奏取决于所配置 provider 的响应速度——无法用一个独立的配置项直接设定。如果需要可重放的背压,请让 
Source 配合一个会做 checkpoint 的下游 Sink;Web3j Source 自身没有限流或节流配置。
+
+### `value` 字段里到底装了什么?
+
+一个结构固定的 JSON 对象:`{"blockNumber": <number>, "timestamp": "<ISO-8601 
UTC>"}`。`blockNumber` 是 provider 最新返回的区块号,`timestamp` 
是连接器在拿到响应那一刻生成的时间戳。连接器不会解析 `value` 的内容;如果还需要交易数、Gas、节点元信息等更多字段,请改用其它 RPC 或在下游通过 
SQL Transform 处理。
+
+### Web3j Source 除了 URL 之外还支持别的鉴权方式吗?
+
+不支持。连接器只会把配置的 `url` 原样传给底层的 HTTP 客户端。所以鉴权也只能是 URL 里能承载的那种——通常是 hosted 
服务(Infura / Alchemy 等)直接嵌在 URL 里的 API Key。连接器没有基于 Header 的鉴权路径,请不要把需要频繁轮换的密钥写在 
URL 里;直接使用一个长期的 provider Key,并通过你们平时的密钥管理系统托管。
+
 ## 变更日志
 
 <ChangeLog />
+

Reply via email to