DanielCarter-stack opened a new pull request, #11787:
URL: https://github.com/apache/seatunnel/pull/11787
## Summary
This PR is a documentation-only update that improves the reference pages for
five V2 connectors: **Clickhouse**, **Doris**, **Elasticsearch**, **Iceberg**,
and **Druid**. Both the English (`docs/en/connectors/{source,sink}/...`) and
Chinese (`docs/zh/connectors/{source,sink}/...`) pages are updated.
The five connectors were picked because none of them have been touched by a
PR within the last 7 days, and none of them were modified by DanielCarter-stack
on the current `dev` branch. Each connector pair was verified against the
connector's Java `Option` declarations before being changed.
## Changes by connector
### Clickhouse (source + sink, en + zh)
Source:
- Added explicit `Source Options` rows for `table_path`, `sql`,
`filter_query`, `partition_list`, `split.size`, and `batch_size` — these were
either missing from the option table or only appeared inside the `table_list`
sub-table. Default values are written in code-fence form (`Integer.MAX_VALUE`,
`1024`, `ZoneId.systemDefault()`) to mirror what `ClickhouseBaseOptions` and
`ClickhouseSourceOptions` actually declare.
- Tightened the `table_list` sub-table descriptions to match the actual
option keys (`split_size` is only valid inside `table_list`; `split.size` is
only valid at the outer source level).
- Added a `FAQ` section with three entries: the `host`/`host:port` format
question, the multi-table read pattern, and the behaviour when both
`table_path` and `sql` are configured.
- Mirrored the option table and FAQ into the Chinese page.
Sink:
- Added a missing `Support Those Engines` section so the source and sink
pages have a consistent top-of-page structure.
### Doris (source, en + zh)
- Restructured the option table to add `database` and `table` rows (required
for the flattened single-table path that the doc already describes in prose but
did not list explicitly). Added `doris.exec.mem.limit` (declared in
`DorisSourceOptions`) to both the base table and the `table_list` sub-table.
- Removed the redundant standalone paragraph that explained
`doris.request.retriesdoris.deserialize.queue.size` and folded the explanation
into the option's `Description` cell. The historical key name is preserved
verbatim — it is the actual runtime option key in
`DorisSourceOptions.DORIS_DESERIALIZE_QUEUE_SIZE` and was not invented in the
docs.
- Renamed the inner sub-table heading from `Table list configuration:` to
`Table list configuration (when using \`table_list\`):` to make the scope
explicit.
- Added a `FAQ` section explaining the historical key name, the multi-table
read pattern, and when to tune `doris.request.tablet.size`. Mirrored all of the
above into the Chinese page.
### Elasticsearch (source + sink, en + zh)
Both source and sink option tables previously listed every option without a
description column. This is the single biggest readability issue in the
existing docs, so both tables were rewritten with a `description` column that
explains every option:
Source option table now describes `hosts`, `auth_type`, `username`,
`password`, the three `auth.api_key*` fields, `index`, `index_list`, `source`,
`query`, `search_type`, `search_api_type`, `sql_query`, `scroll_time`,
`scroll_size`, the four TLS options, `array_column`, `pit_keep_alive`,
`pit_batch_size`, `runtime_fields`, `slice_max`, and `common-options`.
Sink option table now describes `hosts`, `index`, `schema_save_mode`,
`data_save_mode`, `index_type`, `primary_keys`, `key_delimiter`, the auth
fields, `max_retry_count`, `max_batch_size`, the four TLS options,
`common-options`, `vectorization_fields`, `vector_dimensions`, and
`multi_table_sink_replica`. Both tables cross-link to the `schema_save_mode`
and `data_save_mode` anchor sections below.
Source-only:
- Added a `Support Those Engines` section that was missing on the source
page.
- Marked `support multiple table read` in the Key Features list (the
`index_list` option supports this).
All of the above is mirrored into the Chinese
`docs/zh/connectors/source/Elasticsearch.md` and
`docs/zh/connectors/sink/Elasticsearch.md` pages with Chinese descriptions.
### Iceberg (source + sink, en + zh)
Source:
- Marked `support multiple table read` in Key Features — `table_list` is
supported by `IcebergSourceFactory` and is already documented but the Key
Features row was missing.
- Added a `FAQ` section with four entries: multi-table reads, the difference
between `start_snapshot_id` and `use_snapshot_id`, when to prefer
`start_snapshot_timestamp`, and how Kerberos authentication works.
Sink:
- Marked `exactly-once` in Key Features. The sink implements
`SinkAggregatedCommitter` via `IcebergSink`, so 2PC exactly-once is available
and the checkbox was incorrectly absent.
- Added a `FAQ` section with five entries: enabling upsert mode, the
`iceberg.table.partition-keys` format, committing to a non-default branch via
`iceberg.table.commit-branch`, when `custom_sql` is required, and Kerberos
configuration.
All Chinese pages mirror the same Key Features and FAQ additions.
### Druid (sink, en + zh)
- Added a `Support Those Engines` section that was missing on the page.
- Added a `FAQ` section with four entries: CDC support, supported SeaTunnel
data types, how flushing works, and multi-table routing via `${table_name}`.
## Notes
- No source, config, or SPI code is touched.
- All option names, default values, and descriptions were cross-checked
against the corresponding `*Options.java` files (`ClickhouseBaseOptions`,
`ClickhouseSourceOptions`, `DorisSourceOptions`, `ElasticsearchSourceOptions`,
`IcebergSourceOptions`, `IcebergSinkOptions`, `DruidSinkOptions`) where
applicable.
- No connector touched in this PR was modified by another PR on the `dev`
branch within the last 7 days.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]