João Boto created FLINK-40942:
---------------------------------
Summary: NativeS3FileSystem#getFileStatus fails with "Key cannot
be empty" for bucket-root paths
Key: FLINK-40942
URL: https://issues.apache.org/jira/browse/FLINK-40942
Project: Flink
Issue Type: Bug
Components: Connectors / FileSystem
Affects Versions: 2.3.0
Reporter: João Boto
h3. Problem
In flink-s3-fs-native, getFileStatus() / exists() on a bucket-root path
(s3://bucket or s3://bucket/) throws instead of returning a directory status.
NativeS3AccessHelper#extractKey strips the leading "/" from the URI path, so
both forms yield an empty key. getFileStatus then builds
HeadObjectRequest.builder().bucket(bucketName).key("") and the AWS SDK
v2 request marshaller rejects it client-side, before any HTTP call:
{code:java}
SdkClientException: Unable to marshall request to JSON: Key cannot be empty.
at
...services.s3.transform.HeadObjectRequestMarshaller.marshall(HeadObjectRequestMarshaller.java:53)
at ...services.s3.DefaultS3Client.headObject(DefaultS3Client.java:8133)
at
org.apache.flink.fs.s3native.NativeS3FileSystem.getFileStatus(NativeS3FileSystem.java:206)
at org.apache.flink.core.fs.IFileSystem.exists(IFileSystem.java:276)
...
Caused by: java.lang.IllegalArgumentException: Key cannot be empty.
at ...utils.Validate.notEmpty(Validate.java:314)
at
...protocols.core.PathMarshaller$GreedyLeadingSlashPathMarshaller.marshall(PathMarshaller.java:83){code}
SdkClientException is not an S3Exception, so neither the NoSuchKeyException nor
the S3Exception handler in getFileStatus catches it, and it reaches the caller.
h3. Impact
This is a regression for users migrating from flink-s3-fs-hadoop: Hadoop S3A
treats the bucket root as an always-existing directory. Any caller that checks
the existence of a root path breaks. We hit it through Apache Paimon's
HiveCatalog, whose FileIO.get(warehouse) -> checkAccess() calls
exists(warehouse) at catalog creation, with warehouse = s3a://<bucket>. Every
job using that catalog failed to start. The Flink docs also show bucket-root
paths such as checkpoint dirs of the form s3://<bucket>/.
h3. Reproduce
{code:java}
Path root = new Path("s3://my-bucket");
root.getFileSystem().exists(root); // throws SdkClientException{code}
h3. Suggested fix
In getFileStatus, when extractKey(path) is empty, return a directory status for
the bucket root without calling HeadObject (optionally after a HeadBucket
check, to surface missing-bucket/permission errors), matching S3A's root
handling.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)