João Boto created FLINK-40942:
---------------------------------

             Summary: NativeS3FileSystem#getFileStatus fails with "Key cannot 
be empty" for bucket-root paths
                 Key: FLINK-40942
                 URL: https://issues.apache.org/jira/browse/FLINK-40942
             Project: Flink
          Issue Type: Bug
          Components: Connectors / FileSystem
    Affects Versions: 2.3.0
            Reporter: João Boto


h3. Problem

In flink-s3-fs-native, getFileStatus() / exists() on a bucket-root path 
(s3://bucket or s3://bucket/) throws instead of returning a directory status.

NativeS3AccessHelper#extractKey strips the leading "/" from the URI path, so 
both forms yield an empty key. getFileStatus then builds 
HeadObjectRequest.builder().bucket(bucketName).key("") and the AWS SDK
  v2 request marshaller rejects it client-side, before any HTTP call:

 
{code:java}
 SdkClientException: Unable to marshall request to JSON: Key cannot be empty.
    at 
...services.s3.transform.HeadObjectRequestMarshaller.marshall(HeadObjectRequestMarshaller.java:53)
    at ...services.s3.DefaultS3Client.headObject(DefaultS3Client.java:8133)
    at 
org.apache.flink.fs.s3native.NativeS3FileSystem.getFileStatus(NativeS3FileSystem.java:206)
    at org.apache.flink.core.fs.IFileSystem.exists(IFileSystem.java:276)
    ...
  Caused by: java.lang.IllegalArgumentException: Key cannot be empty.
    at ...utils.Validate.notEmpty(Validate.java:314)
    at 
...protocols.core.PathMarshaller$GreedyLeadingSlashPathMarshaller.marshall(PathMarshaller.java:83){code}
SdkClientException is not an S3Exception, so neither the NoSuchKeyException nor 
the S3Exception handler in getFileStatus catches it, and it reaches the caller.
h3. Impact

This is a regression for users migrating from flink-s3-fs-hadoop: Hadoop S3A 
treats the bucket root as an always-existing directory. Any caller that checks 
the existence of a root path breaks. We hit it through Apache Paimon's 
HiveCatalog, whose FileIO.get(warehouse) -> checkAccess() calls 
exists(warehouse) at catalog creation, with warehouse = s3a://<bucket>. Every 
job using that catalog failed to start. The Flink docs also show bucket-root 
paths such as checkpoint dirs of the form s3://<bucket>/.
h3. Reproduce

 
{code:java}
  Path root = new Path("s3://my-bucket");
  root.getFileSystem().exists(root);   // throws SdkClientException{code}
h3. Suggested fix

In getFileStatus, when extractKey(path) is empty, return a directory status for 
the bucket root without calling HeadObject (optionally after a HeadBucket 
check, to surface missing-bucket/permission errors), matching S3A's root 
handling.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to