davidzollo opened a new pull request, #11748:
URL: https://github.com/apache/seatunnel/pull/11748
## Summary
A support-bot analytics export flagged 51 Q&A pairs (out of 1,459) as
"uncertain" across 32 conversation threads. As maintainer, I reviewed all of
them, verified every candidate documentation gap directly against current
source code (not just against what the bot said), and screened out anything
that turned out to already be documented, be a genuine feature gap rather than
a doc gap, or be too narrow/third-party-specific to generalize.
Most candidates turned out to already be fixed in `dev` (this repo evolves
fast): Hudi/Hive Metastore sync behavior, REST API pause/resume semantics, and
MySQL CDC's `TRUNCATE TABLE` limitation are all already documented. Six gaps
were real and are fixed in this PR, each backed by a source-code citation:
- **JDBC source** (`docs/{en,zh}/connectors/source/Jdbc.md`): documents that
whether `table_path`/`use_regex` also matches database views (not just base
tables) depends on the dialect's internal table-listing query —
MySQL/PostgreSQL include views, SQL Server/Oracle/Dameng exclude them — and
there's no config option to control this explicitly. Verified against
`JdbcCatalogUtils.java` and each dialect's `Catalog` implementation.
- **Avro format** (`docs/{en,zh}/connectors/formats/avro.md`): documents
that Confluent Schema Registry wire-format Avro (5-byte header: magic byte +
schema ID) is not supported — this format only decodes plain/embedded-schema
Avro. Verified against `AvroDeserializationSchema.java` and
`KafkaSourceConfig.java`; contrasted with Protobuf's existing
`strip_schema_registry_header` option, which has no Avro equivalent.
- **Debezium JSON format**
(`docs/{en,zh}/connectors/formats/debezium-json.md`): adds a "See Also" section
cross-linking the sibling Canal/Maxwell/OGG JSON CDC formats, which already
exist as separate doc pages but weren't discoverable from the Debezium page.
- **`job.retry.times`**
(`docs/{en,zh}/introduction/configuration/JobEnvConfig.md`): documents that the
retry counter accumulates for the life of the pipeline and is **not** reset by
an intermediate successful recovery — e.g. with `job.retry.times = 5`, a
fail→retry→recover on attempt #3 followed by a later failure only has 2 retries
left (#4, #5), not a fresh budget of 5. Verified against `SubPlan.java`'s
`pipelineRestoreNum` handling (incremented in `prepareRestorePipeline()`, never
reset on success, single instance for the pipeline's lifetime). The one
exception — an active-master failover, which rebuilds the plan and its counter
— is also documented.
- **Telemetry** (`docs/{en,zh}/engines/zeta/telemetry.md`): marks the
"Thread Pool Status" and "Job info detail" metric sections as master-only,
matching the scope note already used elsewhere on the same page for other
master-only metric groups (`JobMetricExports`/`JobThreadPoolStatusExports` both
gate on `isMaster()`).
- **REST API** (`docs/{en,zh}/engines/zeta/rest-api-v2.md`,
`rest-api-job-lifecycle.md`): documents `restoreMode`/`restoreSourceJobId` on
**both** `/submit-job` and `/submit-job/upload` — both endpoints already
support these params via the shared `submitJobInternal`/`resolveRestoreMode`
path (added by #11421), but the parameter tables only listed
`jobId`/`jobName`/`isStartWithSavePoint`. Adds an upload+restore curl example
to both the reference and the lifecycle cookbook. Also fixes a pre-existing
copy-paste bug in the zh doc where the upload endpoint's `<summary>` showed
`/submit-job` instead of `/submit-job/upload`.
## Test plan
- [x] Read every one of the 51 flagged rows and their full bot-generated
answers before triaging.
- [x] For every candidate gap, verified the underlying behavior directly in
source (not just trusting the bot's claim) — file:line citations are in the PR
description above and were cross-checked against the current `dev` HEAD.
- [x] Confirmed no duplicate PR exists for this work (`gh pr list --author
davidzollo`).
- [x] `git diff --stat` confirms only the 14 intended doc files changed, no
unrelated files.
- [x] Spotless: this repo's `spotless-maven-plugin` config (`pom.xml`) only
formats `<java>` sources (googleJavaFormat/importOrder/removeUnusedImports) —
there is no markdown/prettier formatter in the repo. Since this PR touches only
`.md` files, `spotless:apply` is a no-op and was not run against the full
multi-module reactor.
- [x] Verified all new Docusaurus anchor links
(`#6-job-recovery-and-restart`, `#6-作业恢复与重启`,
`#submit-a-job-by-upload-config-file`, `#提交作业来源上传配置文件`) against the existing
slug convention already used elsewhere in the same files (e.g.
`separated-cluster-deployment.md#44-history-job-expiry-configuration`).
- [x] Verified all 14 edited files have balanced Docusaurus admonition
blocks (`:::note`/`:::`) and even code-fence counts.
- No compile/test/E2E run needed — this PR contains no code changes, only
documentation.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]