[
https://issues.apache.org/jira/browse/FLINK-40942?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated FLINK-40942:
-----------------------------------
Labels: pull-request-available (was: )
> NativeS3FileSystem#getFileStatus fails with "Key cannot be empty" for
> bucket-root paths
> ---------------------------------------------------------------------------------------
>
> Key: FLINK-40942
> URL: https://issues.apache.org/jira/browse/FLINK-40942
> Project: Flink
> Issue Type: Bug
> Components: Connectors / FileSystem
> Affects Versions: 2.3.0
> Reporter: João Boto
> Priority: Major
> Labels: pull-request-available
>
> h3. Problem
> In flink-s3-fs-native, getFileStatus() / exists() on a bucket-root path
> (s3://bucket or s3://bucket/) throws instead of returning a directory status.
> NativeS3AccessHelper#extractKey strips the leading "/" from the URI path, so
> both forms yield an empty key. getFileStatus then builds
> HeadObjectRequest.builder().bucket(bucketName).key("") and the AWS SDK
> v2 request marshaller rejects it client-side, before any HTTP call:
>
> {code:java}
> SdkClientException: Unable to marshall request to JSON: Key cannot be empty.
> at
> ...services.s3.transform.HeadObjectRequestMarshaller.marshall(HeadObjectRequestMarshaller.java:53)
> at ...services.s3.DefaultS3Client.headObject(DefaultS3Client.java:8133)
> at
> org.apache.flink.fs.s3native.NativeS3FileSystem.getFileStatus(NativeS3FileSystem.java:206)
> at org.apache.flink.core.fs.IFileSystem.exists(IFileSystem.java:276)
> ...
> Caused by: java.lang.IllegalArgumentException: Key cannot be empty.
> at ...utils.Validate.notEmpty(Validate.java:314)
> at
> ...protocols.core.PathMarshaller$GreedyLeadingSlashPathMarshaller.marshall(PathMarshaller.java:83){code}
> SdkClientException is not an S3Exception, so neither the NoSuchKeyException
> nor the S3Exception handler in getFileStatus catches it, and it reaches the
> caller.
> h3. Impact
> This is a regression for users migrating from flink-s3-fs-hadoop: Hadoop S3A
> treats the bucket root as an always-existing directory. Any caller that
> checks the existence of a root path breaks. We hit it through Apache Paimon's
> HiveCatalog, whose FileIO.get(warehouse) -> checkAccess() calls
> exists(warehouse) at catalog creation, with warehouse = s3a://<bucket>. Every
> job using that catalog failed to start. The Flink docs also show bucket-root
> paths such as checkpoint dirs of the form s3://<bucket>/.
> h3. Reproduce
>
> {code:java}
> Path root = new Path("s3://my-bucket");
> root.getFileSystem().exists(root); // throws SdkClientException{code}
> h3. Suggested fix
> In getFileStatus, when extractKey(path) is empty, return a directory status
> for the bucket root without calling HeadObject (optionally after a HeadBucket
> check, to surface missing-bucket/permission errors), matching S3A's root
> handling.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)