This is an automated email from the ASF dual-hosted git repository.
davidzollo pushed a commit to branch dev
in repository https://gitbox.apache.org/repos/asf/seatunnel.git
The following commit(s) were added to refs/heads/dev by this push:
new 5d34f3bb8c [Docs][Connector-V2] Improve DingTalk Slack Sentry
GoogleSheets and OpenMldb connector docs (#11692)
5d34f3bb8c is described below
commit 5d34f3bb8cf9a665add314307e6ea3b35fdf2797
Author: Daniel Carter <[email protected]>
AuthorDate: Tue Aug 11 22:38:43 2026 +0800
[Docs][Connector-V2] Improve DingTalk Slack Sentry GoogleSheets and
OpenMldb connector docs (#11692)
Co-authored-by: DanielCarter-stack
<[email protected]>
---
docs/en/connectors/sink/DingTalk.md | 84 ++++++++++++++++-----
docs/en/connectors/sink/Sentry.md | 120 +++++++++++++++++++++++-------
docs/en/connectors/sink/Slack.md | 92 ++++++++++++++++++-----
docs/en/connectors/source/GoogleSheets.md | 114 +++++++++++++++++++++-------
docs/en/connectors/source/OpenMldb.md | 111 ++++++++++++++++++++-------
docs/zh/connectors/sink/DingTalk.md | 81 +++++++++++++++-----
docs/zh/connectors/sink/Sentry.md | 112 ++++++++++++++++++++++------
docs/zh/connectors/sink/Slack.md | 88 ++++++++++++++++++----
docs/zh/connectors/source/GoogleSheets.md | 108 ++++++++++++++++++++-------
docs/zh/connectors/source/OpenMldb.md | 103 ++++++++++++++++++-------
10 files changed, 794 insertions(+), 219 deletions(-)
diff --git a/docs/en/connectors/sink/DingTalk.md
b/docs/en/connectors/sink/DingTalk.md
index 2ef3eeaf1b..b5310ad05f 100644
--- a/docs/en/connectors/sink/DingTalk.md
+++ b/docs/en/connectors/sink/DingTalk.md
@@ -2,7 +2,7 @@ import ChangeLog from '../changelog/connector-dingtalk.md';
# DingTalk
-> DinkTalk sink connector
+> DingTalk sink connector
## Support Those Engines
@@ -13,43 +13,93 @@ import ChangeLog from '../changelog/connector-dingtalk.md';
## Key features
- [ ] [exactly-once](../../introduction/concepts/connector-v2-features.md)
+- [ ] [cdc](../../introduction/concepts/connector-v2-features.md)
+- [ ] [support multiple table
write](../../introduction/concepts/connector-v2-features.md)
## Description
-A sink plugin which use DingTalk robot send message
+A sink plugin that sends SeaTunnel rows to a DingTalk group chat through a
DingTalk custom robot
+webhook. The connector identifier used in job configuration is `DingTalk`. For
each row, the connector
+signs the request with the configured robot secret and posts a message to the
DingTalk robot address.
-## Options
+## Data Type Mapping
-| name | type | required | default value |
-|----------------|--------|----------|---------------|
-| url | String | yes | - |
-| secret | String | yes | - |
-| common-options | | no | - |
+The DingTalk connector serializes each row through `SeaTunnelRow.toString()`
and posts the resulting
+plain-text message to the DingTalk robot. There is no per-field JSON structure
on the wire — the entire
+row becomes a single text payload regardless of the underlying field types.
+
+## Sink Options
+
+| name | type | required | default value | description
|
+|--------------------|--------|----------|---------------|----------------------------------------------------------------------------------------------------------|
+| url | String | Yes | - | DingTalk robot
webhook URL, format `https://oapi.dingtalk.com/robot/send?access_token=XXXXXX`.
|
+| secret | String | Yes | - | DingTalk robot
secret used to sign the request.
|
+| common-options | | no | - | Sink plugin common
parameters, please refer to [Sink Common
Options](../common-options/sink-common-options.md) for details. |
### url [String]
-DingTalk robot address format is
https://oapi.dingtalk.com/robot/send?access_token=XXXXXX(String)
+DingTalk robot address format is
`https://oapi.dingtalk.com/robot/send?access_token=XXXXXX`. The `access_token`
+is the robot token created in the DingTalk group settings.
### secret [String]
-DingTalk robot secret (String)
+DingTalk robot secret used to sign messages sent to the robot defined in
`url`. The connector signs
+messages with the configured secret so DingTalk can verify the request source.
The secret must match
+the one bound to the robot configured in `url`. The signed client is created
lazily once per writer
+and reused for the writer's lifetime; the signature is not recomputed on every
individual write.
### common options
-Sink plugin common parameters, please refer to [Sink Common
Options](../common-options/sink-common-options.md) for details
+Sink plugin common parameters, please refer to [Sink Common
Options](../common-options/sink-common-options.md) for details.
+
+## Task Example
+
+### Simple
+
+Send rows to a DingTalk group through a configured robot.
+
+```hocon
+sink {
+ DingTalk {
+ url =
"https://oapi.dingtalk.com/robot/send?access_token=ec646cccd028d978a7156ceeac5b625ebd94f586ea0743fa501c100007890"
+ secret =
"SEC093249eef7aa57d4388aa635f678930c63db3d28b2829d5b2903fc1e5c10000"
+ }
+}
+```
+
+### With upstream source
-## Example
+A typical end-to-end job that reads rows from a fake source and forwards them
to DingTalk.
```hocon
+env {
+ parallelism = 1
+ job.mode = "BATCH"
+}
+
+source {
+ FakeSource {
+ schema = {
+ fields {
+ id = int
+ name = string
+ score = double
+ }
+ }
+ rows = [
+ { kind = "INSERT", fields = [1, "alice", 9.5] }
+ ]
+ }
+}
+
sink {
- DingTalk {
-
url="https://oapi.dingtalk.com/robot/send?access_token=ec646cccd028d978a7156ceeac5b625ebd94f586ea0743fa501c100007890"
- secret="SEC093249eef7aa57d4388aa635f678930c63db3d28b2829d5b2903fc1e5c10000"
- }
+ DingTalk {
+ url =
"https://oapi.dingtalk.com/robot/send?access_token=ec646cccd028d978a7156ceeac5b625ebd94f586ea0743fa501c100007890"
+ secret =
"SEC093249eef7aa57d4388aa635f678930c63db3d28b2829d5b2903fc1e5c10000"
+ }
}
```
## Changelog
<ChangeLog />
-
diff --git a/docs/en/connectors/sink/Sentry.md
b/docs/en/connectors/sink/Sentry.md
index 9f5a3ead74..44cd64a95a 100644
--- a/docs/en/connectors/sink/Sentry.md
+++ b/docs/en/connectors/sink/Sentry.md
@@ -2,66 +2,100 @@ import ChangeLog from '../changelog/connector-sentry.md';
# Sentry
-## Description
+> Sentry sink connector
-Write SeaTunnel rows to Sentry as messages. Each row is sent through the
Sentry SDK by calling
-`Sentry.captureMessage(row.toString())`.
+## Support Those Engines
+
+> Spark<br/>
+> Flink<br/>
+> SeaTunnel Zeta<br/>
## Key features
- [ ] [exactly-once](../../introduction/concepts/connector-v2-features.md)
+- [ ] [cdc](../../introduction/concepts/connector-v2-features.md)
+- [ ] [support multiple table
write](../../introduction/concepts/connector-v2-features.md)
-## Options
+## Description
-| name | type | required | default value |
description |
-|-----------------------------|---------|----------|---------------|-------------|
-| dsn | string | yes | - | Sentry
DSN used by the SDK. |
-| env | string | no | - | Sentry
environment name. |
-| release | string | no | - | Sentry
release value. |
-| cacheDirPath | string | no | - | Cache
directory for offline Sentry events. |
-| enableExternalConfiguration | boolean | no | - | Whether
the Sentry SDK can load external configuration. |
-| maxCacheItems | int | no | - | Maximum
number of cached events. |
-| flushTimeoutMillis | long | no | - | Time in
milliseconds to wait while flushing pending events. |
-| maxQueueSize | int | no | - | Maximum
queue size before events are flushed to disk. |
-| common-options | | no | - | Sink
plugin common parameters. |
+Write SeaTunnel rows to Sentry as messages. Each row is sent through the
Sentry SDK by calling
+`Sentry.captureMessage(row.toString())`. The connector is useful for
forwarding SeaTunnel events
+into Sentry for alerting and aggregation alongside other application events.
+
+## Data Type Mapping
+
+All row values are converted with `row.toString()` before they are passed to
the Sentry SDK, so the
+Sentry message payload is always a string regardless of the underlying field
types.
+
+| SeaTunnel Data Type | Sentry Message Format |
+|---------------------|-----------------------|
+| string | String |
+| tinyint / smallint / int / bigint | String (toString) |
+| float / double | String (toString) |
+| boolean | String (toString) |
+| date / time / timestamp | String (toString) |
+| bytes / array / map / row | String (toString) |
+
+## Sink Options
+
+| name | type | required | default value |
description
|
+|-----------------------------|---------|----------|---------------|---------------------------------------------------------------------------------------------------|
+| dsn | string | yes | - | Sentry
DSN used by the SDK to send events.
|
+| env | string | no | - | Sentry
environment name, attached to every event.
|
+| release | string | no | - | Sentry
release value, attached to every event.
|
+| cacheDirPath | string | no | - | Cache
directory used by the Sentry SDK to buffer offline events before they are sent.
|
+| enableExternalConfiguration | boolean | no | - | Whether
the Sentry SDK can load external configuration such as `sentry.properties`.
|
+| maxCacheItems | int | no | - | Maximum
number of cached events. Defaults to `30` in the SDK when not set.
|
+| flushTimeoutMillis | long | no | - | Time in
milliseconds to wait while flushing pending events.
|
+| maxQueueSize | int | no | - | Maximum
queue size before events are flushed to disk.
|
+| common-options | | no | - | Sink
plugin common parameters, please refer to [Sink Common
Options](../common-options/sink-common-options.md) for details. |
### dsn [string]
-The DSN tells the SDK where to send the events to.
+The DSN tells the SDK where to send the events to. Format is the standard
Sentry DSN, e.g.
+`https://<publicKey>@<host>/<projectId>`.
### env [string]
-specify the environment
+Specify the Sentry environment name (for example `prod`, `staging`). The value
is attached to every
+event captured through this sink.
### release [string]
-specify the release
+Specify the Sentry release value (for example `[email protected]`). The value is
attached to every event
+captured through this sink.
### cacheDirPath [string]
-the cache dir path for caching offline events
+The cache directory path for buffering offline events. Set this to a writable
local directory when
+the sink may run in environments where the Sentry server is not always
reachable.
### enableExternalConfiguration [boolean]
-if loading properties from external sources is enabled.
+If loading properties from external sources (such as `sentry.properties` on
the classpath) is enabled.
+Set this to `true` to let the Sentry SDK pick up environment-specific
configuration files.
### maxCacheItems [number]
-The max cache items for capping the number of events Default is 30
+The maximum number of cached events before the SDK starts dropping older ones.
Defaults to `30` when
+not set.
### flushTimeoutMillis [long]
-Controls how many milliseconds to wait while flushing pending events.
+Controls how many milliseconds to wait while flushing pending events when the
writer closes.
### maxQueueSize [number]
-Max queue size before flushing events/envelopes to the disk
+Maximum queue size before events are flushed to disk. Increase this when the
sink produces events
+faster than the network can drain them.
### common options
-Sink plugin common parameters, please refer to [Sink Common
Options](../common-options/sink-common-options.md) for details
+Sink plugin common parameters, please refer to [Sink Common
Options](../common-options/sink-common-options.md) for details.
+
+## Task Example
-## Example
+### Simple
```hocon
sink {
@@ -75,6 +109,42 @@ sink {
}
```
+### With upstream source
+
+A typical end-to-end job that forwards rows from a fake source to Sentry.
+
+```hocon
+env {
+ parallelism = 1
+ job.mode = "BATCH"
+}
+
+source {
+ FakeSource {
+ schema = {
+ fields {
+ event = string
+ severity = string
+ }
+ }
+ rows = [
+ { kind = "INSERT", fields = ["service-restart", "warning"] }
+ ]
+ }
+}
+
+sink {
+ Sentry {
+ dsn = "https://[email protected]:9999/6"
+ env = "prod"
+ release = "[email protected]"
+ enableExternalConfiguration = false
+ maxCacheItems = 1000
+ flushTimeoutMillis = 15000
+ }
+}
+```
+
## Changelog
<ChangeLog />
diff --git a/docs/en/connectors/sink/Slack.md b/docs/en/connectors/sink/Slack.md
index 98c9abed9f..5ffdff7641 100644
--- a/docs/en/connectors/sink/Slack.md
+++ b/docs/en/connectors/sink/Slack.md
@@ -14,26 +14,48 @@ import ChangeLog from '../changelog/connector-slack.md';
- [ ] [exactly-once](../../introduction/concepts/connector-v2-features.md)
- [ ] [cdc](../../introduction/concepts/connector-v2-features.md)
+- [ ] [support multiple table
write](../../introduction/concepts/connector-v2-features.md)
## Description
-Used to send SeaTunnel rows to a Slack channel. Both streaming and batch jobs
are supported.
-
-> The connector sends the row values as one comma-separated Slack message. For
example, a row with values
-> `huan` and `17` is sent as `huan,17`.
+Used to send SeaTunnel rows to a Slack channel. Both streaming and batch jobs
are supported. The
+connector first uses the configured OAuth token to look up the channel id,
then posts each row as a
+comma-separated message to that channel through Slack's Web API.
## Data Type Mapping
-All field values are converted to strings before they are sent to Slack.
+The Slack connector converts every field of a row to a string with
`String.valueOf(value)` and joins
+them with commas into a single plain-text message — there is no per-field JSON
structure on the wire,
+so the connector can post any SeaTunnel row regardless of the underlying type.
+
+## Sink Options
+
+| name | type | required | default value | description
|
+|-------------------|--------|----------|---------------|-------------------------------------------------------------------------------------------------------------------|
+| webhooks_url | String | Yes | - | Slack incoming
webhook URL. The connector checks for this option during initialization; the
message write path uses `oauth_token` and `slack_channel` to post via the Slack
Web API. |
+| oauth_token | String | Yes | - | Slack OAuth token
used to look up channels and post messages through the Slack Web API.
|
+| slack_channel | String | Yes | - | Slack channel name
where rows are posted. The connector resolves this to a channel id via the
OAuth token. |
+| common-options | | no | - | Sink plugin common
parameters, please refer to [Sink Common
Options](../common-options/sink-common-options.md) for details. |
+
+### webhooks_url [String]
+
+The Slack incoming webhook URL configured on the target Slack workspace. The
connector checks for this
+option during initialization; the message write path uses `oauth_token` and
`slack_channel` together
+with the Slack Web API to look up the channel id and post the row.
+
+### oauth_token [String]
-## Options
+Slack OAuth token with at least `chat:write` and `channels:read` (or
equivalent) scopes. The token is used
+to call the `conversations.list` and `chat.postMessage` APIs.
-| Name | Type | Required | Default |
Description |
-|----------------|--------|----------|---------|-------------------------------------------------------------------------------------------------------------|
-| webhooks_url | String | Yes | - | Slack webhook URL.
|
-| oauth_token | String | Yes | - | Slack OAuth token used to
list channels and post messages.
|
-| slack_channel | String | Yes | - | Slack channel name for data
writes.
|
-| common-options | | no | - | Sink plugin common
parameters, please refer to [Sink Common
Options](../common-options/sink-common-options.md) for details |
+### slack_channel [String]
+
+Slack channel name where rows are posted. The connector will resolve the
channel name to a channel id
+through the Slack Web API. The OAuth token must be able to access this channel.
+
+### common options
+
+Sink plugin common parameters, please refer to [Sink Common
Options](../common-options/sink-common-options.md) for details.
## Task Example
@@ -41,14 +63,50 @@ All field values are converted to strings before they are
sent to Slack.
```hocon
sink {
- Slack {
- webhooks_url =
"https://hooks.slack.com/services/xxxxxxxxxxxx/xxxxxxxxxxxx/xxxxxxxxxxxxxxxx"
- oauth_token = "xoxp-xxxxxxxxxx-xxxxxxxx-xxxxxxxxx-xxxxxxxxxxx"
- slack_channel = "seatunnel-alerts"
- }
+ Slack {
+ webhooks_url =
"https://hooks.slack.com/services/xxxxxxxxxxxx/xxxxxxxxxxxx/xxxxxxxxxxxxxxxx"
+ oauth_token = "xoxp-xxxxxxxxxx-xxxxxxxx-xxxxxxxxx-xxxxxxxxxxx"
+ slack_channel = "seatunnel-alerts"
+ }
}
```
+### With upstream source
+
+A simple batch job that forwards rows from a fake source to Slack.
+
+```hocon
+env {
+ parallelism = 1
+ job.mode = "BATCH"
+}
+
+source {
+ FakeSource {
+ schema = {
+ fields {
+ user = string
+ age = int
+ }
+ }
+ rows = [
+ { kind = "INSERT", fields = ["huan", 17] }
+ ]
+ }
+}
+
+sink {
+ Slack {
+ webhooks_url =
"https://hooks.slack.com/services/xxxxxxxxxxxx/xxxxxxxxxxxx/xxxxxxxxxxxxxxxx"
+ oauth_token = "xoxp-xxxxxxxxxx-xxxxxxxx-xxxxxxxxx-xxxxxxxxxxx"
+ slack_channel = "seatunnel-alerts"
+ }
+}
+```
+
+The connector sends the row values as one comma-separated Slack message, so
the example above produces
+`huan,17` in the configured channel.
+
## Changelog
<ChangeLog />
diff --git a/docs/en/connectors/source/GoogleSheets.md
b/docs/en/connectors/source/GoogleSheets.md
index 5af8d1429c..b52ce2ab26 100644
--- a/docs/en/connectors/source/GoogleSheets.md
+++ b/docs/en/connectors/source/GoogleSheets.md
@@ -6,7 +6,15 @@ import ChangeLog from
'../changelog/connector-google-sheets.md';
## Description
-Used to read data from GoogleSheets.
+Used to read data from Google Sheets through the Google Sheets API. The
connector reads a configured
+range from a sheet using a Google Cloud service account and converts each row
into a SeaTunnel record
+based on the user-defined schema.
+
+## Support Those Engines
+
+> Spark<br/>
+> Flink<br/>
+> SeaTunnel Zeta<br/>
## Key features
@@ -21,56 +29,110 @@ Used to read data from GoogleSheets.
- [ ] csv
- [ ] json
-## Options
+## Data Type Mapping
+
+The Google Sheets API does not expose per-cell types — every cell comes back
as an untyped raw value.
+The connector casts each cell according to the user-declared `schema` option,
so the resulting
+SeaTunnel type is driven entirely by your schema, not by any type detected
from the sheet itself.
+Cells with values that cannot be cast to the configured schema field will
cause the connector to fail
+the row.
+
+| Google Sheets Cell | SeaTunnel Data Type (after schema cast) |
+|--------------------|----------------------------------------|
+| string | string / numeric / boolean / date |
+| number | int / long / float / double |
+| boolean | boolean |
+| date | date / time / timestamp |
+
+## Source Options
-| name | type | required | default value |
-|---------------------|--------|----------|---------------|
-| service_account_key | string | yes | - |
-| sheet_id | string | yes | - |
-| sheet_name | string | yes | - |
-| range | string | yes | - |
-| schema | config | no | - |
+| name | type | required | default value | description
|
+|---------------------|--------|----------|---------------|---------------------------------------------------------------------------------------------------|
+| service_account_key | string | yes | - | Google Cloud
service account credentials. Must be provided as a Base64-encoded JSON string.
|
+| sheet_id | string | yes | - | The sheet id of
the Google Sheets URL, for example
`1VI0DvyZK-NIdssSdsDSsSSSC-_-rYMi7ppJiI_jhE`. |
+| sheet_name | string | yes | - | The name of the
sheet (tab) inside the Google Sheets document to read from.
|
+| range | string | yes | - | The A1 notation
range to read from the sheet, for example `A1:C3` or `Sheet1!A1:D100`.
|
+| schema | config | no | - | The schema of the
rows emitted by the connector. See [Schema
Feature](../../introduction/concepts/schema-feature.md). |
### service_account_key [string]
-google cloud service account, base64 required
+The Base64-encoded JSON content of a Google Cloud service account key file.
The service account must
+have access to the target Google Sheets document (share the sheet with the
service account email).
### sheet_id [string]
-sheet id in a Google Sheets URL
+The id of the Google Sheets document. It is the long identifier between `/d/`
and `/edit` in the
+sheet's URL.
### sheet_name [string]
-the name of the sheet you want to import
+The name of the sheet (tab) inside the Google Sheets document to read from,
for example `Sheet1`.
### range [string]
-the range of the sheet you want to import
+The A1 notation range to read from the sheet, for example `A1:C3` to read a
fixed area or `Sheet1!A:D`
+to read entire columns from a specific sheet.
### schema [config]
#### fields [config]
-The schema fields of upstream data. Please refer to [Schema
Feature](../../introduction/concepts/schema-feature.md).
+The schema fields of upstream data. The connector reads each cell as a string
and casts it to the
+declared field type. Please refer to [Schema
Feature](../../introduction/concepts/schema-feature.md) for
+the available types.
+
+## Task Example
+
+### Simple
+
+```hocon
+source {
+ GoogleSheets {
+ service_account_key = "seatunnel-test"
+ sheet_id = "1VI0DvyZK-NIdssSdsDSsSSSC-_-rYMi7ppJiI_jhE"
+ sheet_name = "sheets01"
+ range = "A1:C3"
+ schema = {
+ fields {
+ a = int
+ b = string
+ c = string
+ }
+ }
+ }
+}
+```
-## Example
+### With downstream sink
-simple:
+Read a sheet and print the rows through the Console sink.
```hocon
-GoogleSheets {
- service_account_key = "seatunnel-test"
- sheet_id = "1VI0DvyZK-NIdssSdsDSsSSSC-_-rYMi7ppJiI_jhE"
- sheet_name = "sheets01"
- range = "A1:C3"
- schema = {
- fields {
- a = int
- b = string
- c = string
+env {
+ parallelism = 1
+ job.mode = "BATCH"
+}
+
+source {
+ GoogleSheets {
+ service_account_key = "seatunnel-test"
+ sheet_id = "1VI0DvyZK-NIdssSdsDSsSSSC-_-rYMi7ppJiI_jhE"
+ sheet_name = "sheets01"
+ range = "A1:C100"
+ schema = {
+ fields {
+ a = int
+ b = string
+ c = string
+ }
}
}
}
+
+sink {
+ Console {
+ }
+}
```
## Changelog
diff --git a/docs/en/connectors/source/OpenMldb.md
b/docs/en/connectors/source/OpenMldb.md
index 2070b765e5..2881eadcc6 100644
--- a/docs/en/connectors/source/OpenMldb.md
+++ b/docs/en/connectors/source/OpenMldb.md
@@ -4,9 +4,17 @@ import ChangeLog from '../changelog/connector-openmldb.md';
> OpenMldb source connector
+## Support Those Engines
+
+> Spark<br/>
+> Flink<br/>
+> SeaTunnel Zeta<br/>
+
## Description
-Used to read data from OpenMldb.
+Used to read data from OpenMLDB. The connector executes the configured SQL
statement against
+OpenMLDB and turns the result rows into SeaTunnel records. Both standalone and
cluster deployment
+modes are supported.
## Key features
@@ -17,62 +25,85 @@ Used to read data from OpenMldb.
- [ ] [parallelism](../../introduction/concepts/connector-v2-features.md)
- [ ] [support user-defined
split](../../introduction/concepts/connector-v2-features.md)
-## Options
-
-| name | type | required | default value | description |
-|-----------------|---------|----------|---------------|-------------|
-| cluster_mode | boolean | yes | - | Whether to connect to
OpenMLDB in cluster mode. |
-| sql | string | yes | - | SQL statement to read
data. |
-| database | string | yes | - | Database name. |
-| host | string | no | - | Required when
`cluster_mode` is `false`. |
-| port | int | no | - | Required when
`cluster_mode` is `false`. |
-| zk_host | string | no | - | Required when
`cluster_mode` is `true`. |
-| zk_path | string | no | - | Required when
`cluster_mode` is `true`. |
-| session_timeout | int | no | 10000 | OpenMLDB session
timeout in milliseconds. |
-| request_timeout | int | no | 60000 | OpenMLDB request
timeout in milliseconds. |
-| common-options | | no | - | Source plugin common
parameters. |
+## Data Type Mapping
+
+OpenMLDB types are mapped to SeaTunnel types according to the result schema of
the configured `sql`
+statement. Columns whose types are not natively understood by SeaTunnel will
cause the read to fail with
+an `UNSUPPORTED_DATA_TYPE` error.
+
+| OpenMLDB Data Type | SeaTunnel Data Type |
+|--------------------|---------------------|
+| bool | boolean |
+| smallint | smallint |
+| int | int |
+| bigint | bigint |
+| float / double | float / double |
+| string / varchar | string |
+| date | date |
+| timestamp | timestamp |
+
+## Source Options
+
+| name | type | required | default value | description
|
+|-----------------|---------|----------|---------------|----------------------------------------------------------------------------------------|
+| cluster_mode | boolean | yes | - | Whether to connect to
OpenMLDB in cluster mode. Set to `false` for standalone mode. |
+| sql | string | yes | - | SQL statement to
execute against OpenMLDB. Column names and types follow the result. |
+| database | string | yes | - | The OpenMLDB database
name to connect to. |
+| host | string | no | - | Required when
`cluster_mode` is `false`. Host of the standalone OpenMLDB server. |
+| port | int | no | - | Required when
`cluster_mode` is `false`. Port of the standalone OpenMLDB server. |
+| zk_host | string | no | - | Required when
`cluster_mode` is `true`. ZooKeeper host list of the OpenMLDB cluster. |
+| zk_path | string | no | - | Required when
`cluster_mode` is `true`. ZooKeeper path of the OpenMLDB cluster. |
+| session_timeout | int | no | 10000 | OpenMLDB session
timeout in milliseconds. |
+| request_timeout | int | no | 60000 | OpenMLDB request
timeout in milliseconds. |
+| common-options | | no | - | Source plugin common
parameters, please refer to [Source Common
Options](../common-options/source-common-options.md) for details. |
### cluster_mode [boolean]
-Whether to connect to OpenMLDB in cluster mode. When it is `false`, configure
`host` and `port`. When it is `true`, configure `zk_host` and `zk_path`.
+Whether to connect to OpenMLDB in cluster mode. When it is `false`, configure
`host` and `port`.
+When it is `true`, configure `zk_host` and `zk_path`.
### sql [string]
-Sql statement
+The SQL statement to execute against OpenMLDB. The result set columns become
the schema of the
+emitted SeaTunnel rows.
### database [string]
-Database name
+The OpenMLDB database name to connect to. The configured database must exist
on the target
+OpenMLDB instance.
### host [string]
-OpenMldb host, only supported on OpenMldb single mode
+OpenMLDB host. Only used when `cluster_mode` is `false` (standalone mode).
### port [int]
-OpenMldb port, only supported on OpenMldb single mode
+OpenMLDB port. Only used when `cluster_mode` is `false` (standalone mode).
### zk_host [string]
-Zookeeper host, only supported on OpenMldb cluster mode
+ZooKeeper host list for the OpenMLDB cluster, for example
`zk-1:2181,zk-2:2181,zk-3:2181`. Only used
+when `cluster_mode` is `true`.
### zk_path [string]
-Zookeeper path, only supported on OpenMldb cluster mode
+ZooKeeper path of the OpenMLDB cluster, for example `/openmldb`. Only used
when `cluster_mode` is `true`.
### session_timeout [int]
-OpenMLDB session timeout in milliseconds.
+OpenMLDB session timeout in milliseconds. Defaults to `10000` (10 seconds).
### request_timeout [int]
-OpenMLDB request timeout in milliseconds.
+OpenMLDB request timeout in milliseconds. Defaults to `60000` (60 seconds).
### common options
-Source plugin common parameters, please refer to [Source Common
Options](../common-options/source-common-options.md) for details
+Source plugin common parameters, please refer to [Source Common
Options](../common-options/source-common-options.md) for details.
+
+## Task Example
-## Example
+### Standalone mode
```hocon
source {
@@ -86,7 +117,7 @@ source {
}
```
-Cluster mode example:
+### Cluster mode
```hocon
source {
@@ -100,6 +131,32 @@ source {
}
```
+### With downstream sink
+
+A typical end-to-end job that reads from OpenMLDB and prints the rows through
the Console sink.
+
+```hocon
+env {
+ parallelism = 1
+ job.mode = "BATCH"
+}
+
+source {
+ OpenMldb {
+ host = "172.17.0.2"
+ port = 6527
+ sql = "select id, name from demo_table1"
+ database = "demo_db"
+ cluster_mode = false
+ }
+}
+
+sink {
+ Console {
+ }
+}
+```
+
## Changelog
<ChangeLog />
diff --git a/docs/zh/connectors/sink/DingTalk.md
b/docs/zh/connectors/sink/DingTalk.md
index 39b27565b9..72ced19f36 100644
--- a/docs/zh/connectors/sink/DingTalk.md
+++ b/docs/zh/connectors/sink/DingTalk.md
@@ -2,7 +2,7 @@ import ChangeLog from '../changelog/connector-dingtalk.md';
# 钉钉
-> 钉钉 数据接收器
+> 钉钉数据接收器
## 支持的引擎
@@ -13,43 +13,90 @@ import ChangeLog from '../changelog/connector-dingtalk.md';
## 主要特性
- [ ] [精确一次](../../introduction/concepts/connector-v2-features.md)
-- [ ] [timer flush](../../introduction/concepts/connector-v2-features.md)
+- [ ] [cdc](../../introduction/concepts/connector-v2-features.md)
+- [ ] [支持多表写入](../../introduction/concepts/connector-v2-features.md)
## 描述
-一个使用钉钉机器人发送消息的Sink插件。
+通过钉钉自定义机器人 Webhook,将 SeaTunnel 行数据发送到钉钉群聊的接收器插件。作业配置中使用的连接器标识为
`DingTalk`。每一行数据都会使用配置的机器人密钥进行签名,然后发送到钉钉机器人地址。
-## Options
+## 数据类型映射
-| 名称 | 类型 | 是否必须 | 默认值 |
-|----------------|--------|------|-----|
-| url | String | 是 | - |
-| secret | String | 是 | - |
-| common-options | | 否 | - |
+钉钉连接器会把每一行通过 `SeaTunnelRow.toString()` 序列化为纯文本,并作为一条消息发送给钉钉机器人。
+线上传输的是单一文本消息,不存在按字段区分的 JSON 结构 —— 不论源字段类型是什么,整行都会被转换为
+一条文本消息。
+
+## 接收器选项
+
+| 名称 | 类型 | 是否必须 | 默认值 | 描述
|
+|---------------|--------|----------|--------|-----------------------------------------------------------------------------------------------|
+| url | String | 是 | - | 钉钉机器人 Webhook 地址,格式
`https://oapi.dingtalk.com/robot/send?access_token=XXXXXX`。 |
+| secret | String | 是 | - | 用于对请求进行签名的钉钉机器人密钥。
|
+| common-options| | 否 | - | Sink 插件通用参数,详见 [Sink
常见选项](../common-options/sink-common-options.md)。 |
### url [String]
-钉钉机器人地址格式为 https://oapi.dingtalk.com/robot/send?access_token=XXXXXX(String)
+钉钉机器人地址格式为 `https://oapi.dingtalk.com/robot/send?access_token=XXXXXX`,其中
`access_token`
+是钉钉群机器人设置中生成的令牌。
### secret [String]
-钉钉机器人的密钥 (String)
+钉钉机器人密钥,用于对发往 `url` 中机器人的消息进行签名。连接器使用该密钥为消息生成签名,以便
+钉钉端校验请求来源。该密钥必须与 `url` 中机器人绑定的密钥保持一致。签名客户端在写入器首次发送时
+按需创建一次,并在该写入器生命周期内复用,不会对每条消息重新计算签名。
### common options
-Sink插件的通用参数,请参考 [Sink Common
Options](../common-options/sink-common-options.md) 了解详情
+Sink 插件通用参数,请参考 [Sink 常见选项](../common-options/sink-common-options.md) 了解详情。
## 任务示例
+### 简单示例
+
+通过已配置的机器人将行数据发送到钉钉群。
+
```hocon
sink {
- DingTalk {
-
url="https://oapi.dingtalk.com/robot/send?access_token=ec646cccd028d978a7156ceeac5b625ebd94f586ea0743fa501c100007890"
- secret="SEC093249eef7aa57d4388aa635f678930c63db3d28b2829d5b2903fc1e5c10000"
- }
+ DingTalk {
+ url =
"https://oapi.dingtalk.com/robot/send?access_token=ec646cccd028d978a7156ceeac5b625ebd94f586ea0743fa501c100007890"
+ secret =
"SEC093249eef7aa57d4388aa635f678930c63db3d28b2829d5b2903fc1e5c10000"
+ }
+}
+```
+
+### 配合上游源使用
+
+一个典型的端到端作业,从 fake 源读取数据并转发到钉钉。
+
+```hocon
+env {
+ parallelism = 1
+ job.mode = "BATCH"
+}
+
+source {
+ FakeSource {
+ schema = {
+ fields {
+ id = int
+ name = string
+ score = double
+ }
+ }
+ rows = [
+ { kind = "INSERT", fields = [1, "alice", 9.5] }
+ ]
+ }
+}
+
+sink {
+ DingTalk {
+ url =
"https://oapi.dingtalk.com/robot/send?access_token=ec646cccd028d978a7156ceeac5b625ebd94f586ea0743fa501c100007890"
+ secret =
"SEC093249eef7aa57d4388aa635f678930c63db3d28b2829d5b2903fc1e5c10000"
+ }
}
```
## 变更日志
-<ChangeLog />
\ No newline at end of file
+<ChangeLog />
diff --git a/docs/zh/connectors/sink/Sentry.md
b/docs/zh/connectors/sink/Sentry.md
index fa48c134bf..9bc50c8aea 100644
--- a/docs/zh/connectors/sink/Sentry.md
+++ b/docs/zh/connectors/sink/Sentry.md
@@ -2,65 +2,95 @@ import ChangeLog from '../changelog/connector-sentry.md';
# Sentry
-## 描述
+> Sentry 数据接收器
+
+## 支持的引擎
-将 SeaTunnel 行数据作为消息写入 Sentry。每一行会通过 Sentry SDK 以
`Sentry.captureMessage(row.toString())` 的方式发送。
+> Spark<br/>
+> Flink<br/>
+> SeaTunnel Zeta<br/>
## 关键特性
- [ ] [精确一次](../../introduction/concepts/connector-v2-features.md)
+- [ ] [cdc](../../introduction/concepts/connector-v2-features.md)
+- [ ] [支持多表写入](../../introduction/concepts/connector-v2-features.md)
+
+## 描述
+
+将 SeaTunnel 行数据作为消息写入 Sentry。每一行都会通过 Sentry SDK 调用
+`Sentry.captureMessage(row.toString())` 进行发送。该连接器适合将 SeaTunnel 中的事件统一转发到
+Sentry,与其他业务事件一起做告警和聚合分析。
+
+## 数据类型映射
+
+所有行字段值在传入 Sentry SDK 之前都会通过 `row.toString()` 转成字符串,因此无论源字段类型
+是什么,最终发送给 Sentry 的消息载荷始终是字符串。
+
+| SeaTunnel 数据类型 | Sentry 消息格式 |
+|--------------------|-----------------|
+| string | String |
+| tinyint / smallint / int / bigint | String (toString) |
+| float / double | String (toString) |
+| boolean | String (toString) |
+| date / time / timestamp | String (toString) |
+| bytes / array / map / row | String (toString) |
## 选项
-| 名称 | 类型 | 必需 | 默认值 | 描述 |
-|-----------------------------|---------|------|--------|------|
-| dsn | string | 是 | - | Sentry SDK 使用的 DSN。 |
-| env | string | 否 | - | Sentry 环境名称。 |
-| release | string | 否 | - | Sentry release 值。 |
-| cacheDirPath | string | 否 | - | 离线事件缓存目录。 |
-| enableExternalConfiguration | boolean | 否 | - | 是否允许 Sentry SDK
加载外部配置。 |
-| maxCacheItems | int | 否 | - | 最大缓存事件数量。 |
-| flushTimeoutMillis | long | 否 | - | 刷新待发送事件时的等待时间,单位毫秒。 |
-| maxQueueSize | int | 否 | - | 事件刷新到磁盘前的最大队列大小。 |
-| common-options | | 否 | - | 接收器插件通用参数。 |
+| 名称 | 类型 | 必需 | 默认值 | 描述
|
+|-----------------------------|---------|------|--------|-------------------------------------------------------------------------------------------------|
+| dsn | string | 是 | - | Sentry SDK 使用的 DSN。
|
+| env | string | 否 | - | Sentry
环境名称,会附加到每一条事件上。 |
+| release | string | 否 | - | Sentry release
值,会附加到每一条事件上。 |
+| cacheDirPath | string | 否 | - | Sentry SDK
用于缓存离线事件的目录。 |
+| enableExternalConfiguration | boolean | 否 | - | 是否允许 Sentry SDK
从外部(例如 `sentry.properties`)加载配置。 |
+| maxCacheItems | int | 否 | - | 最大缓存事件数量。SDK 默认值为
`30`。 |
+| flushTimeoutMillis | long | 否 | - | 刷新待发送事件时的等待时间,单位毫秒。
|
+| maxQueueSize | int | 否 | - | 事件刷新到磁盘前的最大队列大小。
|
+| common-options | | 否 | - | 接收器插件通用参数,详见 [Sink
常见选项](../common-options/sink-common-options.md)。 |
### dsn [string]
-DSN告诉SDK将事件发送到何处.
+DSN 告诉 SDK 将事件发送到哪里。格式为标准 Sentry DSN,例如
+`https://<publicKey>@<host>/<projectId>`。
### env [string]
-指定环境
+指定 Sentry 环境名称(例如 `prod`、`staging`),会附加到该接收器捕获的每一条事件上。
### release [string]
-指定版本
+指定 Sentry release 值(例如 `[email protected]`),会附加到该接收器捕获的每一条事件上。
### cacheDirPath [string]
-缓存脱机事件的缓存目录路径
+用于缓存离线事件的目录。当接收器所在环境无法保证 Sentry 服务始终可达时,请配置为本地可写目录。
### enableExternalConfiguration [boolean]
-如果启用了从外部源加载属性.
+是否启用从外部源(例如 classpath 中的 `sentry.properties`)加载配置。设置为 `true` 后,SDK 会
+自动加载环境特定的配置文件。
### maxCacheItems [number]
-用于限制事件数量的最大缓存项默认值为30
+最大缓存事件数量,超过后会丢弃旧事件。不设置时 SDK 默认为 `30`。
### flushTimeoutMillis [long]
-刷新待发送事件时的等待时间,单位毫秒。
+刷新待发送事件时的等待时间,单位毫秒。用于在写入器关闭时控制阻塞时长。
### maxQueueSize [number]
-将事件/信封刷新到磁盘之前的最大队列大小
+事件刷新到磁盘前的最大队列大小。当事件产生速度快于网络发送速度时,可以适当调大该值。
### common options
-接收器插件常用参数,详见 [Sink 常见选项](../common-options/sink-common-options.md)
+接收器插件通用参数,详见 [Sink 常见选项](../common-options/sink-common-options.md)。
+
+## 任务示例
-## 示例
+### 简单示例
```hocon
sink {
@@ -74,6 +104,42 @@ sink {
}
```
+### 配合上游源使用
+
+将 fake 源产生的行数据转发到 Sentry 的典型端到端作业。
+
+```hocon
+env {
+ parallelism = 1
+ job.mode = "BATCH"
+}
+
+source {
+ FakeSource {
+ schema = {
+ fields {
+ event = string
+ severity = string
+ }
+ }
+ rows = [
+ { kind = "INSERT", fields = ["service-restart", "warning"] }
+ ]
+ }
+}
+
+sink {
+ Sentry {
+ dsn = "https://[email protected]:9999/6"
+ env = "prod"
+ release = "[email protected]"
+ enableExternalConfiguration = false
+ maxCacheItems = 1000
+ flushTimeoutMillis = 15000
+ }
+}
+```
+
## 变更日志
<ChangeLog />
diff --git a/docs/zh/connectors/sink/Slack.md b/docs/zh/connectors/sink/Slack.md
index 257b401c66..bb6fb286b0 100644
--- a/docs/zh/connectors/sink/Slack.md
+++ b/docs/zh/connectors/sink/Slack.md
@@ -14,26 +14,46 @@ import ChangeLog from '../changelog/connector-slack.md';
- [ ] [精确一次](../../introduction/concepts/connector-v2-features.md)
- [ ] [cdc](../../introduction/concepts/connector-v2-features.md)
+- [ ] [支持多表写入](../../introduction/concepts/connector-v2-features.md)
## 描述
-用于将 SeaTunnel 行数据发送到 Slack 频道。流处理和批处理作业都支持。
-
-> 连接器会把一行中的字段值拼成一条用逗号分隔的 Slack 消息。例如,字段值为 `huan` 和 `17` 时,
-> 发送内容为 `huan,17`。
+用于将 SeaTunnel 行数据发送到 Slack 频道,支持流处理和批处理作业。连接器首先使用配置的 OAuth 令牌
+查找频道 ID,然后通过 Slack Web API 将每一行以逗号分隔的消息发布到该频道。
## 数据类型映射
-所有字段值在发送到 Slack 前都会转换为字符串。
+Slack 连接器会把一行中的每个字段通过 `String.valueOf(value)` 转为字符串,再用逗号拼接成一条纯文本
+消息 —— 线上传输的是单一文本消息,不存在按字段区分的 JSON 结构,因此连接器可以发布任意类型的
+SeaTunnel 行。
## 选项
-| 名称 | 类型 | 必需 | 默认值 | 描述 |
-|----------------|--------|------|--------|------|
-| webhooks_url | String | 是 | - | Slack webhook URL。 |
-| oauth_token | String | 是 | - | 用于列出频道并发送消息的 Slack OAuth 令牌。 |
-| slack_channel | String | 是 | - | 写入数据的 Slack 频道名称。 |
-| common-options | | 否 | - | 接收器插件通用参数,详见 [Sink
常见选项](../common-options/sink-common-options.md)。 |
+| 名称 | 类型 | 必需 | 默认值 | 描述
|
+|------------------|--------|------|--------|---------------------------------------------------------------------------------------------------|
+| webhooks_url | String | 是 | - | Slack 传入 Webhook
URL,连接器在初始化时会校验该选项;消息发送路径使用 `oauth_token`、`slack_channel` 通过 Slack Web API
发布消息。 |
+| oauth_token | String | 是 | - | 用于查询频道和发送消息的 Slack OAuth 令牌。
|
+| slack_channel | String | 是 | - | 行数据发送到的 Slack 频道名称,连接器会通过 OAuth
令牌将其解析为频道 ID。 |
+| common-options | | 否 | - | 接收器插件通用参数,详见 [Sink
常见选项](../common-options/sink-common-options.md)。 |
+
+### webhooks_url [String]
+
+目标 Slack 工作空间中配置的传入 Webhook URL。连接器在初始化时会校验该选项;消息发送路径使用
+`oauth_token` 和 `slack_channel` 配合 Slack Web API 来解析频道 ID 并发布消息。
+
+### oauth_token [String]
+
+至少需要 `chat:write` 和 `channels:read`(或同等)权限的 Slack OAuth 令牌。该令牌用于调用
+`conversations.list` 和 `chat.postMessage` 接口。
+
+### slack_channel [String]
+
+行数据要发送到的 Slack 频道名称。连接器会通过 Slack Web API 将频道名解析为频道 ID。OAuth 令牌
+必须能访问该频道。
+
+### common options
+
+接收器插件通用参数,请参考 [Sink 常见选项](../common-options/sink-common-options.md) 了解详情。
## 任务示例
@@ -41,14 +61,50 @@ import ChangeLog from '../changelog/connector-slack.md';
```hocon
sink {
- Slack {
- webhooks_url =
"https://hooks.slack.com/services/xxxxxxxxxxxx/xxxxxxxxxxxx/xxxxxxxxxxxxxxxx"
- oauth_token = "xoxp-xxxxxxxxxx-xxxxxxxx-xxxxxxxxx-xxxxxxxxxxx"
- slack_channel = "seatunnel-alerts"
- }
+ Slack {
+ webhooks_url =
"https://hooks.slack.com/services/xxxxxxxxxxxx/xxxxxxxxxxxx/xxxxxxxxxxxxxxxx"
+ oauth_token = "xoxp-xxxxxxxxxx-xxxxxxxx-xxxxxxxxx-xxxxxxxxxxx"
+ slack_channel = "seatunnel-alerts"
+ }
}
```
+### 配合上游源使用
+
+将 fake 源产生的行数据转发到 Slack 的简单批处理作业。
+
+```hocon
+env {
+ parallelism = 1
+ job.mode = "BATCH"
+}
+
+source {
+ FakeSource {
+ schema = {
+ fields {
+ user = string
+ age = int
+ }
+ }
+ rows = [
+ { kind = "INSERT", fields = ["huan", 17] }
+ ]
+ }
+}
+
+sink {
+ Slack {
+ webhooks_url =
"https://hooks.slack.com/services/xxxxxxxxxxxx/xxxxxxxxxxxx/xxxxxxxxxxxxxxxx"
+ oauth_token = "xoxp-xxxxxxxxxx-xxxxxxxx-xxxxxxxxx-xxxxxxxxxxx"
+ slack_channel = "seatunnel-alerts"
+ }
+}
+```
+
+连接器会把一行中的字段值拼成一条用逗号分隔的 Slack 消息,因此上面的示例会在配置的频道中产生
+`huan,17` 这条消息。
+
## 变更日志
<ChangeLog />
diff --git a/docs/zh/connectors/source/GoogleSheets.md
b/docs/zh/connectors/source/GoogleSheets.md
index dfd26ce501..694a4cb43b 100644
--- a/docs/zh/connectors/source/GoogleSheets.md
+++ b/docs/zh/connectors/source/GoogleSheets.md
@@ -6,7 +6,14 @@ import ChangeLog from
'../changelog/connector-google-sheets.md';
## 描述
-用于从GoogleSheets读取数据.
+用于通过 Google Sheets API 从 Google 表格中读取数据。连接器使用 Google Cloud 服务账号凭据读取指定
+范围内的内容,并根据用户定义的 schema 把每一行转换为 SeaTunnel 记录。
+
+## 支持的引擎
+
+> Spark<br/>
+> Flink<br/>
+> SeaTunnel Zeta<br/>
## 关键特性
@@ -21,56 +28,105 @@ import ChangeLog from
'../changelog/connector-google-sheets.md';
- [ ] csv
- [ ] json
-## 选项
+## 数据类型映射
+
+Google Sheets API 不会返回单元格本身的数据类型 —— 每个单元格读取的都是一个无类型的原始值。连接器
+会按照用户声明的 `schema` 选项对单元格进行类型转换,因此最终输出的 SeaTunnel 类型完全由 schema 决定,
+与表格中单元格的原生类型无关。单元格值无法转换为目标 schema 类型时,对应行会失败。
+
+| Google Sheets 单元格 | SeaTunnel 数据类型(按 schema 转换后) |
+|----------------------|---------------------------------------|
+| string | string / 数值 / boolean / 日期 |
+| number | int / long / float / double |
+| boolean | boolean |
+| date | date / time / timestamp |
+
+## 源选项
-| 名称 | 类型 | 必需 | 默认值 |
-|---------------------|--------|----------|---------------|
-| service_account_key | string | 是 | - |
-| sheet_id | string | 是 | - |
-| sheet_name | string | 是 | - |
-| range | string | 是 | - |
-| schema | config | 否 | - |
+| 名称 | 类型 | 必需 | 默认值 | 描述
|
+|------------------------|--------|------|--------|-----------------------------------------------------------------------------------------------|
+| service_account_key | string | 是 | - | Google Cloud 服务账号凭据,必须使用
Base64 编码后的 JSON 字符串。 |
+| sheet_id | string | 是 | - | Google 表格的 id,即表格 URL 中
`/d/` 与 `/edit` 之间的长字符串。 |
+| sheet_name | string | 是 | - | 要读取的工作表(标签页)名称,例如 `Sheet1`。
|
+| range | string | 是 | - | 要读取的 A1 表示法范围,例如 `A1:C3` 或
`Sheet1!A1:D100`。 |
+| schema | config | 否 | - | 上游数据的字段定义,详见 [Schema
特性](../../introduction/concepts/schema-feature.md)。 |
### service_account_key [string]
-谷歌云服务帐户,需要base64编码
+Google Cloud 服务账号密钥 JSON 文件经过 Base64 编码后的内容。服务账号必须有目标 Google 表格的访问
+权限(请将服务账号邮箱加入表格的共享列表)。
### sheet_id [string]
-Google表格URL中的表格id
+要读取的 Google 表格 id,即表格 URL 中 `/d/` 与 `/edit` 之间的长字符串。
### sheet_name [string]
-要导入的工作表的名称
+要读取的工作表(标签页)名称,例如 `Sheet1`。
### range [string]
-要导入的 sheet 页的范围
+要读取的 A1 表示法范围,例如 `A1:C3` 用于读取固定区域,或 `Sheet1!A:D` 用于读取整列。
### schema [config]
#### fields [config]
-上游数据的字段。更多详情请参考 [Schema 特性](../../introduction/concepts/schema-feature.md)。
+上游数据的字段定义。连接器会把每个单元格作为字符串读取,然后按声明的类型进行转换。可用的类型请参考
+[Schema 特性](../../introduction/concepts/schema-feature.md)。
+
+## 任务示例
+
+### 简单示例
+
+```hocon
+source {
+ GoogleSheets {
+ service_account_key = "seatunnel-test"
+ sheet_id = "1VI0DvyZK-NIdssSdsDSsSSSC-_-rYMi7ppJiI_jhE"
+ sheet_name = "sheets01"
+ range = "A1:C3"
+ schema = {
+ fields {
+ a = int
+ b = string
+ c = string
+ }
+ }
+ }
+}
+```
-## 示例
+### 配合下游接收器
-简单示例:
+读取一个工作表并通过 Console 接收器打印读取的行数据。
```hocon
-GoogleSheets {
- service_account_key = "seatunnel-test"
- sheet_id = "1VI0DvyZK-NIdssSdsDSsSSSC-_-rYMi7ppJiI_jhE"
- sheet_name = "sheets01"
- range = "A1:C3"
- schema = {
- fields {
- a = int
- b = string
- c = string
+env {
+ parallelism = 1
+ job.mode = "BATCH"
+}
+
+source {
+ GoogleSheets {
+ service_account_key = "seatunnel-test"
+ sheet_id = "1VI0DvyZK-NIdssSdsDSsSSSC-_-rYMi7ppJiI_jhE"
+ sheet_name = "sheets01"
+ range = "A1:C100"
+ schema = {
+ fields {
+ a = int
+ b = string
+ c = string
+ }
}
}
}
+
+sink {
+ Console {
+ }
+}
```
## 变更日志
diff --git a/docs/zh/connectors/source/OpenMldb.md
b/docs/zh/connectors/source/OpenMldb.md
index 01e4aae04d..03143227a3 100644
--- a/docs/zh/connectors/source/OpenMldb.md
+++ b/docs/zh/connectors/source/OpenMldb.md
@@ -4,9 +4,16 @@ import ChangeLog from '../changelog/connector-openmldb.md';
> OpenMldb 源连接器
+## 支持的引擎
+
+> Spark<br/>
+> Flink<br/>
+> SeaTunnel Zeta<br/>
+
## 描述
-用于从 OpenMldb 读取数据.
+用于从 OpenMLDB 读取数据。连接器会执行配置的 SQL 语句并把结果转换为 SeaTunnel 记录,同时支持
+单机版和集群版两种部署模式。
## 关键特性
@@ -17,62 +24,82 @@ import ChangeLog from '../changelog/connector-openmldb.md';
- [ ] [并行度](../../introduction/concepts/connector-v2-features.md)
- [ ] [支持用户自定义分片](../../introduction/concepts/connector-v2-features.md)
+## 数据类型映射
+
+OpenMLDB 类型会按照所配置 `sql` 语句的结果集映射为 SeaTunnel 类型。SeaTunnel 不原生支持的类型会直接
+导致读取失败,并抛出 `UNSUPPORTED_DATA_TYPE` 错误。
+
+| OpenMLDB 数据类型 | SeaTunnel 数据类型 |
+|-------------------|--------------------|
+| bool | boolean |
+| smallint | smallint |
+| int | int |
+| bigint | bigint |
+| float / double | float / double |
+| string / varchar | string |
+| date | date |
+| timestamp | timestamp |
+
## 选项
-| 名称 | 类型 | 必需 | 默认值 | 描述 |
-|-----------------|---------|------|--------|------|
-| cluster_mode | boolean | 是 | - | 是否以 OpenMLDB 集群模式连接。 |
-| sql | string | 是 | - | 用于读取数据的 SQL 语句。 |
-| database | string | 是 | - | 数据库名称。 |
-| host | string | 否 | - | 当 `cluster_mode` 为 `false` 时必填。 |
-| port | int | 否 | - | 当 `cluster_mode` 为 `false` 时必填。 |
-| zk_host | string | 否 | - | 当 `cluster_mode` 为 `true` 时必填。 |
-| zk_path | string | 否 | - | 当 `cluster_mode` 为 `true` 时必填。 |
-| session_timeout | int | 否 | 10000 | OpenMLDB 会话超时时间,单位毫秒。 |
-| request_timeout | int | 否 | 60000 | OpenMLDB 请求超时时间,单位毫秒。 |
-| common-options | | 否 | - | 源插件通用参数。 |
+| 名称 | 类型 | 必需 | 默认值 | 描述
|
+|-----------------|---------|------|--------|---------------------------------------------------------------------------------------------------|
+| cluster_mode | boolean | 是 | - | 是否以 OpenMLDB 集群模式连接。`false`
表示单机模式,`true` 表示集群模式。 |
+| sql | string | 是 | - | 用于读取数据的 SQL 语句,列名和类型按结果集定义。
|
+| database | string | 是 | - | 要连接的 OpenMLDB 数据库名称。
|
+| host | string | 否 | - | 当 `cluster_mode` 为 `false`
时必填,OpenMLDB 单机版主机地址。 |
+| port | int | 否 | - | 当 `cluster_mode` 为 `false`
时必填,OpenMLDB 单机版端口。 |
+| zk_host | string | 否 | - | 当 `cluster_mode` 为 `true`
时必填,OpenMLDB 集群对应的 ZooKeeper 地址列表。 |
+| zk_path | string | 否 | - | 当 `cluster_mode` 为 `true`
时必填,OpenMLDB 集群在 ZooKeeper 上的路径,例如 `/openmldb`。 |
+| session_timeout | int | 否 | 10000 | OpenMLDB 会话超时时间,单位毫秒。
|
+| request_timeout | int | 否 | 60000 | OpenMLDB 请求超时时间,单位毫秒。
|
+| common-options | | 否 | - | 源插件通用参数,详见 [Source
常见选项](../common-options/source-common-options.md)。 |
### cluster_mode [boolean]
-是否以 OpenMLDB 集群模式连接。为 `false` 时配置 `host` 和 `port`;为 `true` 时配置 `zk_host` 和
`zk_path`。
+是否以 OpenMLDB 集群模式连接。为 `false` 时配置 `host` 和 `port`;为 `true` 时配置 `zk_host`
+和 `zk_path`。
### sql [string]
-Sql 语句
+针对 OpenMLDB 执行的 SQL 语句,结果集的列会成为连接器输出行的字段。
### database [string]
-数据库名称
+要连接的 OpenMLDB 数据库名称,配置的数据库必须在目标 OpenMLDB 实例上存在。
### host [string]
-OpenMldb主机,仅支持OpenMldb单模
+OpenMLDB 主机,仅在 `cluster_mode` 为 `false`(单机模式)下使用。
### port [int]
-OpenMldb端口,仅支持OpenMldb单模
+OpenMLDB 端口,仅在 `cluster_mode` 为 `false`(单机模式)下使用。
### zk_host [string]
-Zookeeper主机,仅在OpenMldb集群模式下受支持
+OpenMLDB 集群对应的 ZooKeeper 地址列表,例如 `zk-1:2181,zk-2:2181,zk-3:2181`,仅在
+`cluster_mode` 为 `true` 时使用。
### zk_path [string]
-Zookeeper路径,仅在OpenMldb集群模式下受支持
+OpenMLDB 集群在 ZooKeeper 上的路径,例如 `/openmldb`,仅在 `cluster_mode` 为 `true` 时使用。
### session_timeout [int]
-OpenMLDB 会话超时时间,单位毫秒。
+OpenMLDB 会话超时时间,单位毫秒,默认 `10000`(10 秒)。
### request_timeout [int]
-OpenMLDB 请求超时时间,单位毫秒。
+OpenMLDB 请求超时时间,单位毫秒,默认 `60000`(60 秒)。
### common options
-源插件常用参数, 详见 [Source Common
Options](../common-options/source-common-options.md)
+源插件通用参数,详见 [Source 常见选项](../common-options/source-common-options.md)。
+
+## 任务示例
-## 示例
+### 单机模式
```hocon
source {
@@ -86,7 +113,7 @@ source {
}
```
-集群模式示例:
+### 集群模式
```hocon
source {
@@ -100,6 +127,32 @@ source {
}
```
+### 配合下游接收器
+
+从 OpenMLDB 读取数据并通过 Console 接收器打印的典型端到端作业。
+
+```hocon
+env {
+ parallelism = 1
+ job.mode = "BATCH"
+}
+
+source {
+ OpenMldb {
+ host = "172.17.0.2"
+ port = 6527
+ sql = "select id, name from demo_table1"
+ database = "demo_db"
+ cluster_mode = false
+ }
+}
+
+sink {
+ Console {
+ }
+}
+```
+
## 变更日志
<ChangeLog />