goutamadwant opened a new issue, #12239:
URL: https://github.com/apache/seatunnel/issues/12239
### Search before asking
- [x] I searched existing issues and pull requests and found no duplicate
implementation of this focused S3 source capability.
### Description
Add metadata-only S3 source validation to the existing `--dry-run connect`
workflow. At the reproduced baseline, `S3FileSourceFactory` does not implement
`SupportSourceDryRunValidation`, so the core connectivity validator skips this
source. This is an enhancement gap, not a report that ordinary S3 jobs are
broken.
Users should be able to check whether their configured S3 endpoint and
credentials can access the selected object or prefix before submitting a job,
without downloading records, recursively enumerating files, or modifying
storage.
#### Reproduction
Baseline: `apache/seatunnel` revision
`75fd4ed4b2e63579a57b461285997951490c4fe7`.
Add the following regression to the existing `S3FileSourceFactoryTest`,
importing `SupportSourceDryRunValidation` from
`org.apache.seatunnel.api.table.factory`:
```java
@Test
void shouldSupportConnectivityDryRun() {
assertTrue(
new S3FileSourceFactory() instanceof
SupportSourceDryRunValidation,
"S3File source connectivity is skipped when the factory does not
expose the dry-run SPI");
}
```
With the required reactor dependencies built, run:
```bash
mvn -o -q -pl seatunnel-connectors-v2/connector-file/connector-file-s3 \
-Dtest=S3FileSourceFactoryTest#shouldSupportConnectivityDryRun test
```
On original production code, this produced one test and one assertion
failure (`expected: <true> but was: <false>`) on both Java 8 and Java 11. The
factory capability check, together with the existing core dispatch path,
establishes why connectivity is skipped. It does not demonstrate a failed
production job.
#### Proposed behavior and acceptance criteria
- Implement the existing source validation SPI without changing ordinary
source/sink execution or defaults.
- Initially support single-table `s3a://bucket` sources with an absolute
path, explicit inline `schema.fields` or `schema.columns`, text/CSV/JSON/XML,
`parse_partition_from_path=false`, and no `read_columns`.
- Reuse current Hadoop S3A endpoint, credential-provider chain, proxy,
path-style and bucket-override configuration; introduce no dependency upgrade
or separate credential resolver.
- Check an exact object with HEAD. On object 404, perform one prefix LIST
with delimiter `/` and `maxKeys=1`. Do not fetch contents, paginate, upload,
delete or initialize a shared filesystem.
- Fail missing batch prefixes and denied requests. Accept a successfully
listed empty continuous prefix because files may arrive later. Missing buckets
must fail in either mode; an accessible empty bucket root may pass.
- Cap validation connect/socket timeouts at five seconds, preserve smaller
positive values and disable request retries. These are not a global deadline
for DNS or credential initialization.
- Reject unsupported schema/format configurations explicitly rather than
reporting validation success. Document the limits in English and Chinese.
#### Before and after
Before: S3 source connectivity is skipped by connect dry-run.
After: supported S3 configurations receive bounded metadata checks and
actionable failures before submission. Unsupported configurations fail with an
explanation. Normal jobs retain their existing behavior.
#### Prepared implementation and validation
Prepared commit: `d66fb63f93fcaed5744d575bdfc2e6b98a97c67e`.
- Java 8 and Java 11: 68 connector unit tests and 13 factory integration
tests passed on each, with zero failures/errors/skips. The final runs used
Corretto 8u504 and Temurin 11.0.32.1.
- Integration coverage includes six real MinIO tests and seven HTTP-wire
tests for metadata-only requests, bounded listing, permissions, missing
buckets, empty prefixes, timeout handling and cleanup after success, request
failure and endpoint setup failure.
- Root formatting and full `./mvnw -q -DskipTests verify` passed on Java 11
for the branch and unchanged baseline. This is compilation/package
verification, not execution of every Java test.
- A separate baseline starter-command test attempt discovered zero tests and
is not included in the passing totals.
- [Fork
CI](https://github.com/goutamadwant/seatunnel/actions/runs/34320638878) is not
fully green: the Java 8 SFTP job failed before tests when the Maven wrapper
download received HTTP 403. This is separate from the passing local S3
validation.
The check does not establish object-content readability, file-format
correctness, worker credentials, IAM/KMS coverage, or successful engine-job
execution. The initial scope excludes multi-table configuration, file-derived
schemas, projection/partition inference, legacy `s3n`, SSE-C, S3-specific
credential-store preprocessing, S3Guard, multipart purge and custom S3 client
factories. Hadoop's existing partial custom-provider constructor cleanup
limitation is not changed.
### Usage Scenario
Preflight checks for pipelines reading S3 objects or continuously arriving
files, especially where endpoint, credential or path mistakes would otherwise
be found after job submission. The existing command remains `--dry-run
connect`; the feature adds S3 source coverage rather than a new submission mode.
### Related issues
- Existing source connectivity SPI: #11186.
- Separate Hadoop/AWS SDK migration: #11648. This implementation targets the
current stack and would need porting and revalidation if that migration lands
first.
- [Prepared compare
branch](https://github.com/apache/seatunnel/compare/dev...goutamadwant:feature/s3-source-connectivity).
### Are you willing to submit a PR?
- [x] Yes, I am willing to submit a PR. The implementation is prepared on
the linked compare branch.
### Code of Conduct
- [x] I agree to follow this project's [Code of
Conduct](https://www.apache.org/foundation/policies/conduct).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]