zhangshenghang opened a new pull request, #12159:
URL: https://github.com/apache/seatunnel/pull/12159
## Purpose
Some JDBC dialects return a zero or negative estimate from
`queryApproximateRowCnt`:
- AnalyticDB MySQL external tables, which do not have accurate table
statistics.
- MySQL tables whose `information_schema` statistics have not been updated.
- PostgreSQL tables that have never been `ANALYZE`d.
With `approximateRowCnt <= 0` the previous code path went through the
uneven-distribution branch and `splitUnevenlySizedChunks`, which issues one
extra boundary query per chunk. For large uneven tables this is slow and
produces a noisy split log.
This change adds a bounded range-based split that uses the measured `min` /
`max` of the split key directly, falling back to a single full-table split when
the range is unsafe to measure or would produce too many chunks.
## Changes
- `DynamicChunkSplitter`:
- When `queryApproximateRowCnt` returns `0` or a negative value, call the
new `splitEvenlySizedChunksByRange` instead of the uneven-distribution path.
- `splitEvenlySizedChunksByRange` generates chunks of `chunkSize` keys
anchored at the measured `min`, stopping when the next boundary would exceed
`max`, the chunk boundary stops advancing, or arithmetic overflows.
- `isRangeChunkFallbackSafe` caps the estimated chunk count at the
existing `split.sample-sharding.threshold` and refuses to fall back when `min`,
`max` is an unsupported numeric pair (for example `NaN`).
- `ObjectUtils#plus`: handle `Byte` and `Short` operands with
`Math.addExact` and an explicit range check, so the new fallback works for all
numeric split key types supported by `ObjectUtils#minus`.
- `DynamicChunkSplitterTest`: use `ObjectUtils.compare` in the chunk
ordering check, since the new test exercises `Byte` and `Short` ranges; add
`testSplitEvenlySizedChunksByRangeWhenApproximateRowCountUnavailable` covering
the happy path, oversized sparse ranges, unsafe numeric ranges, and the
chunkSize / maxChunkCount argument validation.
## Validation
```
./mvnw -pl seatunnel-connectors-v2/connector-jdbc \
-Dtest=DynamicChunkSplitterTest \
-Dcheckstyle.skip -Dspotless.check.skip test
```
Result: `Tests run: 6, Failures: 0, Errors: 0, Skipped: 0`
## Impact
- Behavior change: when `approximateRowCnt <= 0` for a numeric split key,
the source now uses a single range-based split plan rather than N boundary
probes. Read SQL is unchanged; only the number and shape of splits differ.
Tables with valid row count estimates are unaffected.
- No new public API, no configuration change. The existing
`split.sample-sharding.threshold` knob is reused as the chunk-count cap.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]