[ 
https://issues.apache.org/jira/browse/FLINK-40942?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated FLINK-40942:
-----------------------------------
    Labels: pull-request-available  (was: )

> NativeS3FileSystem#getFileStatus fails with "Key cannot be empty" for 
> bucket-root paths
> ---------------------------------------------------------------------------------------
>
>                 Key: FLINK-40942
>                 URL: https://issues.apache.org/jira/browse/FLINK-40942
>             Project: Flink
>          Issue Type: Bug
>          Components: Connectors / FileSystem
>    Affects Versions: 2.3.0
>            Reporter: João Boto
>            Priority: Major
>              Labels: pull-request-available
>
> h3. Problem
> In flink-s3-fs-native, getFileStatus() / exists() on a bucket-root path 
> (s3://bucket or s3://bucket/) throws instead of returning a directory status.
> NativeS3AccessHelper#extractKey strips the leading "/" from the URI path, so 
> both forms yield an empty key. getFileStatus then builds 
> HeadObjectRequest.builder().bucket(bucketName).key("") and the AWS SDK
>   v2 request marshaller rejects it client-side, before any HTTP call:
>  
> {code:java}
>  SdkClientException: Unable to marshall request to JSON: Key cannot be empty.
>     at 
> ...services.s3.transform.HeadObjectRequestMarshaller.marshall(HeadObjectRequestMarshaller.java:53)
>     at ...services.s3.DefaultS3Client.headObject(DefaultS3Client.java:8133)
>     at 
> org.apache.flink.fs.s3native.NativeS3FileSystem.getFileStatus(NativeS3FileSystem.java:206)
>     at org.apache.flink.core.fs.IFileSystem.exists(IFileSystem.java:276)
>     ...
>   Caused by: java.lang.IllegalArgumentException: Key cannot be empty.
>     at ...utils.Validate.notEmpty(Validate.java:314)
>     at 
> ...protocols.core.PathMarshaller$GreedyLeadingSlashPathMarshaller.marshall(PathMarshaller.java:83){code}
> SdkClientException is not an S3Exception, so neither the NoSuchKeyException 
> nor the S3Exception handler in getFileStatus catches it, and it reaches the 
> caller.
> h3. Impact
> This is a regression for users migrating from flink-s3-fs-hadoop: Hadoop S3A 
> treats the bucket root as an always-existing directory. Any caller that 
> checks the existence of a root path breaks. We hit it through Apache Paimon's 
> HiveCatalog, whose FileIO.get(warehouse) -> checkAccess() calls 
> exists(warehouse) at catalog creation, with warehouse = s3a://<bucket>. Every 
> job using that catalog failed to start. The Flink docs also show bucket-root 
> paths such as checkpoint dirs of the form s3://<bucket>/.
> h3. Reproduce
>  
> {code:java}
>   Path root = new Path("s3://my-bucket");
>   root.getFileSystem().exists(root);   // throws SdkClientException{code}
> h3. Suggested fix
> In getFileStatus, when extractKey(path) is empty, return a directory status 
> for the bucket root without calling HeadObject (optionally after a HeadBucket 
> check, to surface missing-bucket/permission errors), matching S3A's root 
> handling.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to