DanielCarter-stack opened a new pull request, #11787:
URL: https://github.com/apache/seatunnel/pull/11787

   ## Summary
   
   This PR is a documentation-only update that improves the reference pages for 
five V2 connectors: **Clickhouse**, **Doris**, **Elasticsearch**, **Iceberg**, 
and **Druid**. Both the English (`docs/en/connectors/{source,sink}/...`) and 
Chinese (`docs/zh/connectors/{source,sink}/...`) pages are updated.
   
   The five connectors were picked because none of them have been touched by a 
PR within the last 7 days, and none of them were modified by DanielCarter-stack 
on the current `dev` branch. Each connector pair was verified against the 
connector's Java `Option` declarations before being changed.
   
   ## Changes by connector
   
   ### Clickhouse (source + sink, en + zh)
   
   Source:
   - Added explicit `Source Options` rows for `table_path`, `sql`, 
`filter_query`, `partition_list`, `split.size`, and `batch_size` — these were 
either missing from the option table or only appeared inside the `table_list` 
sub-table. Default values are written in code-fence form (`Integer.MAX_VALUE`, 
`1024`, `ZoneId.systemDefault()`) to mirror what `ClickhouseBaseOptions` and 
`ClickhouseSourceOptions` actually declare.
   - Tightened the `table_list` sub-table descriptions to match the actual 
option keys (`split_size` is only valid inside `table_list`; `split.size` is 
only valid at the outer source level).
   - Added a `FAQ` section with three entries: the `host`/`host:port` format 
question, the multi-table read pattern, and the behaviour when both 
`table_path` and `sql` are configured.
   - Mirrored the option table and FAQ into the Chinese page.
   
   Sink:
   - Added a missing `Support Those Engines` section so the source and sink 
pages have a consistent top-of-page structure.
   
   ### Doris (source, en + zh)
   
   - Restructured the option table to add `database` and `table` rows (required 
for the flattened single-table path that the doc already describes in prose but 
did not list explicitly). Added `doris.exec.mem.limit` (declared in 
`DorisSourceOptions`) to both the base table and the `table_list` sub-table.
   - Removed the redundant standalone paragraph that explained 
`doris.request.retriesdoris.deserialize.queue.size` and folded the explanation 
into the option's `Description` cell. The historical key name is preserved 
verbatim — it is the actual runtime option key in 
`DorisSourceOptions.DORIS_DESERIALIZE_QUEUE_SIZE` and was not invented in the 
docs.
   - Renamed the inner sub-table heading from `Table list configuration:` to 
`Table list configuration (when using \`table_list\`):` to make the scope 
explicit.
   - Added a `FAQ` section explaining the historical key name, the multi-table 
read pattern, and when to tune `doris.request.tablet.size`. Mirrored all of the 
above into the Chinese page.
   
   ### Elasticsearch (source + sink, en + zh)
   
   Both source and sink option tables previously listed every option without a 
description column. This is the single biggest readability issue in the 
existing docs, so both tables were rewritten with a `description` column that 
explains every option:
   
   Source option table now describes `hosts`, `auth_type`, `username`, 
`password`, the three `auth.api_key*` fields, `index`, `index_list`, `source`, 
`query`, `search_type`, `search_api_type`, `sql_query`, `scroll_time`, 
`scroll_size`, the four TLS options, `array_column`, `pit_keep_alive`, 
`pit_batch_size`, `runtime_fields`, `slice_max`, and `common-options`.
   
   Sink option table now describes `hosts`, `index`, `schema_save_mode`, 
`data_save_mode`, `index_type`, `primary_keys`, `key_delimiter`, the auth 
fields, `max_retry_count`, `max_batch_size`, the four TLS options, 
`common-options`, `vectorization_fields`, `vector_dimensions`, and 
`multi_table_sink_replica`. Both tables cross-link to the `schema_save_mode` 
and `data_save_mode` anchor sections below.
   
   Source-only:
   - Added a `Support Those Engines` section that was missing on the source 
page.
   - Marked `support multiple table read` in the Key Features list (the 
`index_list` option supports this).
   
   All of the above is mirrored into the Chinese 
`docs/zh/connectors/source/Elasticsearch.md` and 
`docs/zh/connectors/sink/Elasticsearch.md` pages with Chinese descriptions.
   
   ### Iceberg (source + sink, en + zh)
   
   Source:
   - Marked `support multiple table read` in Key Features — `table_list` is 
supported by `IcebergSourceFactory` and is already documented but the Key 
Features row was missing.
   - Added a `FAQ` section with four entries: multi-table reads, the difference 
between `start_snapshot_id` and `use_snapshot_id`, when to prefer 
`start_snapshot_timestamp`, and how Kerberos authentication works.
   
   Sink:
   - Marked `exactly-once` in Key Features. The sink implements 
`SinkAggregatedCommitter` via `IcebergSink`, so 2PC exactly-once is available 
and the checkbox was incorrectly absent.
   - Added a `FAQ` section with five entries: enabling upsert mode, the 
`iceberg.table.partition-keys` format, committing to a non-default branch via 
`iceberg.table.commit-branch`, when `custom_sql` is required, and Kerberos 
configuration.
   
   All Chinese pages mirror the same Key Features and FAQ additions.
   
   ### Druid (sink, en + zh)
   
   - Added a `Support Those Engines` section that was missing on the page.
   - Added a `FAQ` section with four entries: CDC support, supported SeaTunnel 
data types, how flushing works, and multi-table routing via `${table_name}`.
   
   ## Notes
   
   - No source, config, or SPI code is touched.
   - All option names, default values, and descriptions were cross-checked 
against the corresponding `*Options.java` files (`ClickhouseBaseOptions`, 
`ClickhouseSourceOptions`, `DorisSourceOptions`, `ElasticsearchSourceOptions`, 
`IcebergSourceOptions`, `IcebergSinkOptions`, `DruidSinkOptions`) where 
applicable.
   - No connector touched in this PR was modified by another PR on the `dev` 
branch within the last 7 days.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to