goutamadwant opened a new issue, #12239:
URL: https://github.com/apache/seatunnel/issues/12239

   ### Search before asking
   
   - [x] I searched existing issues and pull requests and found no duplicate 
implementation of this focused S3 source capability.
   
   ### Description
   
   Add metadata-only S3 source validation to the existing `--dry-run connect` 
workflow. At the reproduced baseline, `S3FileSourceFactory` does not implement 
`SupportSourceDryRunValidation`, so the core connectivity validator skips this 
source. This is an enhancement gap, not a report that ordinary S3 jobs are 
broken.
   
   Users should be able to check whether their configured S3 endpoint and 
credentials can access the selected object or prefix before submitting a job, 
without downloading records, recursively enumerating files, or modifying 
storage.
   
   #### Reproduction
   
   Baseline: `apache/seatunnel` revision 
`75fd4ed4b2e63579a57b461285997951490c4fe7`.
   
   Add the following regression to the existing `S3FileSourceFactoryTest`, 
importing `SupportSourceDryRunValidation` from 
`org.apache.seatunnel.api.table.factory`:
   
   ```java
   @Test
   void shouldSupportConnectivityDryRun() {
       assertTrue(
               new S3FileSourceFactory() instanceof 
SupportSourceDryRunValidation,
               "S3File source connectivity is skipped when the factory does not 
expose the dry-run SPI");
   }
   ```
   
   With the required reactor dependencies built, run:
   
   ```bash
   mvn -o -q -pl seatunnel-connectors-v2/connector-file/connector-file-s3 \
     -Dtest=S3FileSourceFactoryTest#shouldSupportConnectivityDryRun test
   ```
   
   On original production code, this produced one test and one assertion 
failure (`expected: <true> but was: <false>`) on both Java 8 and Java 11. The 
factory capability check, together with the existing core dispatch path, 
establishes why connectivity is skipped. It does not demonstrate a failed 
production job.
   
   #### Proposed behavior and acceptance criteria
   
   - Implement the existing source validation SPI without changing ordinary 
source/sink execution or defaults.
   - Initially support single-table `s3a://bucket` sources with an absolute 
path, explicit inline `schema.fields` or `schema.columns`, text/CSV/JSON/XML, 
`parse_partition_from_path=false`, and no `read_columns`.
   - Reuse current Hadoop S3A endpoint, credential-provider chain, proxy, 
path-style and bucket-override configuration; introduce no dependency upgrade 
or separate credential resolver.
   - Check an exact object with HEAD. On object 404, perform one prefix LIST 
with delimiter `/` and `maxKeys=1`. Do not fetch contents, paginate, upload, 
delete or initialize a shared filesystem.
   - Fail missing batch prefixes and denied requests. Accept a successfully 
listed empty continuous prefix because files may arrive later. Missing buckets 
must fail in either mode; an accessible empty bucket root may pass.
   - Cap validation connect/socket timeouts at five seconds, preserve smaller 
positive values and disable request retries. These are not a global deadline 
for DNS or credential initialization.
   - Reject unsupported schema/format configurations explicitly rather than 
reporting validation success. Document the limits in English and Chinese.
   
   #### Before and after
   
   Before: S3 source connectivity is skipped by connect dry-run.
   
   After: supported S3 configurations receive bounded metadata checks and 
actionable failures before submission. Unsupported configurations fail with an 
explanation. Normal jobs retain their existing behavior.
   
   #### Prepared implementation and validation
   
   Prepared commit: `d66fb63f93fcaed5744d575bdfc2e6b98a97c67e`.
   
   - Java 8 and Java 11: 68 connector unit tests and 13 factory integration 
tests passed on each, with zero failures/errors/skips. The final runs used 
Corretto 8u504 and Temurin 11.0.32.1.
   - Integration coverage includes six real MinIO tests and seven HTTP-wire 
tests for metadata-only requests, bounded listing, permissions, missing 
buckets, empty prefixes, timeout handling and cleanup after success, request 
failure and endpoint setup failure.
   - Root formatting and full `./mvnw -q -DskipTests verify` passed on Java 11 
for the branch and unchanged baseline. This is compilation/package 
verification, not execution of every Java test.
   - A separate baseline starter-command test attempt discovered zero tests and 
is not included in the passing totals.
   - [Fork 
CI](https://github.com/goutamadwant/seatunnel/actions/runs/34320638878) is not 
fully green: the Java 8 SFTP job failed before tests when the Maven wrapper 
download received HTTP 403. This is separate from the passing local S3 
validation.
   
   The check does not establish object-content readability, file-format 
correctness, worker credentials, IAM/KMS coverage, or successful engine-job 
execution. The initial scope excludes multi-table configuration, file-derived 
schemas, projection/partition inference, legacy `s3n`, SSE-C, S3-specific 
credential-store preprocessing, S3Guard, multipart purge and custom S3 client 
factories. Hadoop's existing partial custom-provider constructor cleanup 
limitation is not changed.
   
   ### Usage Scenario
   
   Preflight checks for pipelines reading S3 objects or continuously arriving 
files, especially where endpoint, credential or path mistakes would otherwise 
be found after job submission. The existing command remains `--dry-run 
connect`; the feature adds S3 source coverage rather than a new submission mode.
   
   ### Related issues
   
   - Existing source connectivity SPI: #11186.
   - Separate Hadoop/AWS SDK migration: #11648. This implementation targets the 
current stack and would need porting and revalidation if that migration lands 
first.
   - [Prepared compare 
branch](https://github.com/apache/seatunnel/compare/dev...goutamadwant:feature/s3-source-connectivity).
   
   ### Are you willing to submit a PR?
   
   - [x] Yes, I am willing to submit a PR. The implementation is prepared on 
the linked compare branch.
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to