Aditya1404Sal opened a new issue, #865:
URL: https://github.com/apache/arrow-rs-object-store/issues/865

   **Is your feature request related to a problem or challenge? Please describe 
what you are trying to do.**
   
   `object_store` has no bucket-level API. Every store is built for one bucket
   (`AmazonS3Builder::with_bucket_name` and the GCS and Azure equivalents), and 
every `ObjectStore`
   method works on paths inside it. There is no way to create that bucket, 
delete it, or check whether
   it exists.
   
   I'm working on wasmCloud's blobstore support. wasmCloud implements the
   
[`wasi:blobstore`](https://github.com/WebAssembly/wasi-blobstore/blob/main/wit/blobstore.wit)
   interface, which requires `create-container`, `delete-container` and 
`container-exists`. We want one
   `object_store`-backed implementation for S3, GCS and Azure
   
([wasmCloud/wasmCloud#4977](https://github.com/wasmCloud/wasmCloud/issues/4977)).
 The maintainers
   there (Including me) want bucket creation to be supported rather than 
provisioned out of band. Without it, a
   consumer has two choices:
   - give up part of that interface; or
   - add a provider SDK next to `object_store` for three calls, which undercuts 
running the same
     binary against every cloud.
   
   Multi-tenant services that create a bucket or container per tenant at 
onboarding, and remove it at
   offboarding, run into the same gap.
   
   I couldn't find an existing issue for this.
   
   **Describe the solution you'd like**
   
   A separate, opt-in trait, not new methods on `ObjectStore`. It follows the 
pattern of `Signer` and
   `MultipartStore`, and it is implemented only by `AmazonS3`, 
`GoogleCloudStorage` and
   `MicrosoftAzure`. Existing `ObjectStore` implementations are unaffected.
   
   ```rust
   #[async_trait]
   pub trait BucketStore: Send + Sync + fmt::Debug + 'static {
       /// Create the bucket this store is configured for.
       async fn create_bucket(&self) -> Result<()>;
       /// Delete the bucket this store is configured for.
       async fn delete_bucket(&self) -> Result<()>;
       /// Return whether the bucket this store is configured for exists.
       async fn bucket_exists(&self) -> Result<bool>;
   }
   ```
   
   The methods act on the store's own bucket, because the builders already tie 
the endpoint, region
   and addressing style to one bucket. Callers that work with bucket names 
build one store per bucket
   with `builder.clone().with_bucket_name(name)`.
   
   The semantics reuse existing `Error` variants:
   - Creating an existing bucket returns `AlreadyExists`.
   - Deleting a missing bucket returns `NotFound`.
   - Deleting a non-empty S3 or GCS bucket returns `Generic` with the 
provider's message. A 409 on
     delete means "not empty", not "already exists".
   - `bucket_exists` returns `Ok(false)` only on a 404. A 403 stays an error, 
because on S3 it can
     also mean the bucket belongs to another account.
   
   Provider differences are documented:
   - S3 in `us-east-1` returns success when re-creating a bucket you own.
   - Azure deletes containers asynchronously.
   - S3 Express One Zone directory buckets return `NotSupported`.
   
   GCS needs a project ID to create a bucket, and `GoogleCloudStorageBuilder` 
has none today. So this
   also adds `with_project_id` and a `google_project_id` config key. The 
project ID is taken from
   `GOOGLE_PROJECT_ID`, then `GOOGLE_CLOUD_PROJECT`, then the service-account 
key. Only
   `create_bucket` uses it.
   
   The bucket operations check the configured name before building a request. 
While writing that, I
   noticed that the existing object operations put bucket and container names 
into request URLs as-is
   too: S3 interpolates the name, and Azure's URL handling drops `.` and `..` 
path segments. That's
   harmless while names come from configuration. But a caller that derives 
bucket names from user or
   tenant input, like the per-tenant case above, currently has to validate them 
itself. Validating in
   the builders for all operations could be a follow-up.
   
   **Describe alternatives you've considered**
   
   1. **Methods on `ObjectStore` with `NotSupported` defaults.** This grows the 
core trait with API
      that is a runtime error for most implementations, including those outside 
this crate.
   2. **A name-parameterized or account-scoped API** (`create_bucket(&self, 
name)`). It is closer to
      the provider APIs and to name-based consumers, and it would allow 
`list_buckets`. But it adds a
      second way to construct stores that doesn't fit today's per-bucket 
builders.
   3. **`bucket_info() -> Result<Option<BucketMeta>>` instead of 
`bucket_exists`.** This leaves room
      for metadata such as creation time, which providers expose unevenly.
   
   **Additional context**
   
   I have an implementation, about 1,500 lines including tests, and will open a 
PR that references
   this issue. I'm happy to split it into the trait plus S3, then GCS with the 
project-ID option, then
   Azure, if that is easier to review.
   
   It passes against LocalStack, fake-gcs-server and Azurite. A wasmCloud
   `wasi:blobstore` backend built on it passes on the same emulators for all 
three providers. It hasn't
   been run against real AWS, GCS or Azure accounts yet.
   
   Related gaps I ran into, none of them changed by this proposal:
   - **Cross-bucket copy** (#297). `wasi:blobstore` copies objects between 
containers. Today that
     means a get followed by a put.
   - **No bucket creation time.** `container.info()` wants one. Only GCS 
exposes it per bucket, which
     is alternative 3.
   - **List errors lose their status code** (#851). A list-based existence 
probe can't tell a missing
     bucket from a forbidden one.
   - **One unparseable key fails a listing.** Listing parses keys with 
`Path::parse`, so a single key
     it rejects (for example `a//b` written by another tool) fails the whole 
page. Such a bucket can't
     be emptied through `object_store` before `delete_bucket`.
   - **S3 ignores the bucket name with a custom endpoint.** With an explicit 
endpoint and
     virtual-hosted-style requests, S3 uses the endpoint as-is. Stores built 
for different bucket names
     then all address the same bucket, and the new bucket operations do too. 
This is documented on the
     `AmazonS3` implementation.
   - **A range past the object's end is a generic error.** A range starting at 
or beyond the object's
     length comes back as `Error::Generic` rather than a distinct error, so 
callers clamping ranges
     need a `head` first.
   - **Wrong emulator variable in `CONTRIBUTING.md`.** Its Azure instructions 
set
     `AZURE_USE_EMULATOR`, which `from_env` ignores; the key is 
`AZURE_STORAGE_USE_EMULATOR`.
   
   Open questions, assuming this is in scope for the crate:
   1. Should the operations act on the store's own bucket, as proposed, or take 
a bucket name (see
      alternative 2)?
   2. `bucket_exists() -> bool`, or `bucket_info() -> Option<BucketMeta>` 
(alternative 3)?
   3. Does the GCS project-ID option look right: the key name, 
`GOOGLE_PROJECT_ID` taking precedence
      over `GOOGLE_CLOUD_PROJECT`, and the fallback to the service-account key?
   
   Transparency note : This Issue was drafted using the Help of AI, The 
Contents of the Issue have been vetted but might sound slightly robotic -- 
apologies for the same


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to