Aditya1404Sal opened a new issue, #865: URL: https://github.com/apache/arrow-rs-object-store/issues/865
**Is your feature request related to a problem or challenge? Please describe what you are trying to do.** `object_store` has no bucket-level API. Every store is built for one bucket (`AmazonS3Builder::with_bucket_name` and the GCS and Azure equivalents), and every `ObjectStore` method works on paths inside it. There is no way to create that bucket, delete it, or check whether it exists. I'm working on wasmCloud's blobstore support. wasmCloud implements the [`wasi:blobstore`](https://github.com/WebAssembly/wasi-blobstore/blob/main/wit/blobstore.wit) interface, which requires `create-container`, `delete-container` and `container-exists`. We want one `object_store`-backed implementation for S3, GCS and Azure ([wasmCloud/wasmCloud#4977](https://github.com/wasmCloud/wasmCloud/issues/4977)). The maintainers there (Including me) want bucket creation to be supported rather than provisioned out of band. Without it, a consumer has two choices: - give up part of that interface; or - add a provider SDK next to `object_store` for three calls, which undercuts running the same binary against every cloud. Multi-tenant services that create a bucket or container per tenant at onboarding, and remove it at offboarding, run into the same gap. I couldn't find an existing issue for this. **Describe the solution you'd like** A separate, opt-in trait, not new methods on `ObjectStore`. It follows the pattern of `Signer` and `MultipartStore`, and it is implemented only by `AmazonS3`, `GoogleCloudStorage` and `MicrosoftAzure`. Existing `ObjectStore` implementations are unaffected. ```rust #[async_trait] pub trait BucketStore: Send + Sync + fmt::Debug + 'static { /// Create the bucket this store is configured for. async fn create_bucket(&self) -> Result<()>; /// Delete the bucket this store is configured for. async fn delete_bucket(&self) -> Result<()>; /// Return whether the bucket this store is configured for exists. async fn bucket_exists(&self) -> Result<bool>; } ``` The methods act on the store's own bucket, because the builders already tie the endpoint, region and addressing style to one bucket. Callers that work with bucket names build one store per bucket with `builder.clone().with_bucket_name(name)`. The semantics reuse existing `Error` variants: - Creating an existing bucket returns `AlreadyExists`. - Deleting a missing bucket returns `NotFound`. - Deleting a non-empty S3 or GCS bucket returns `Generic` with the provider's message. A 409 on delete means "not empty", not "already exists". - `bucket_exists` returns `Ok(false)` only on a 404. A 403 stays an error, because on S3 it can also mean the bucket belongs to another account. Provider differences are documented: - S3 in `us-east-1` returns success when re-creating a bucket you own. - Azure deletes containers asynchronously. - S3 Express One Zone directory buckets return `NotSupported`. GCS needs a project ID to create a bucket, and `GoogleCloudStorageBuilder` has none today. So this also adds `with_project_id` and a `google_project_id` config key. The project ID is taken from `GOOGLE_PROJECT_ID`, then `GOOGLE_CLOUD_PROJECT`, then the service-account key. Only `create_bucket` uses it. The bucket operations check the configured name before building a request. While writing that, I noticed that the existing object operations put bucket and container names into request URLs as-is too: S3 interpolates the name, and Azure's URL handling drops `.` and `..` path segments. That's harmless while names come from configuration. But a caller that derives bucket names from user or tenant input, like the per-tenant case above, currently has to validate them itself. Validating in the builders for all operations could be a follow-up. **Describe alternatives you've considered** 1. **Methods on `ObjectStore` with `NotSupported` defaults.** This grows the core trait with API that is a runtime error for most implementations, including those outside this crate. 2. **A name-parameterized or account-scoped API** (`create_bucket(&self, name)`). It is closer to the provider APIs and to name-based consumers, and it would allow `list_buckets`. But it adds a second way to construct stores that doesn't fit today's per-bucket builders. 3. **`bucket_info() -> Result<Option<BucketMeta>>` instead of `bucket_exists`.** This leaves room for metadata such as creation time, which providers expose unevenly. **Additional context** I have an implementation, about 1,500 lines including tests, and will open a PR that references this issue. I'm happy to split it into the trait plus S3, then GCS with the project-ID option, then Azure, if that is easier to review. It passes against LocalStack, fake-gcs-server and Azurite. A wasmCloud `wasi:blobstore` backend built on it passes on the same emulators for all three providers. It hasn't been run against real AWS, GCS or Azure accounts yet. Related gaps I ran into, none of them changed by this proposal: - **Cross-bucket copy** (#297). `wasi:blobstore` copies objects between containers. Today that means a get followed by a put. - **No bucket creation time.** `container.info()` wants one. Only GCS exposes it per bucket, which is alternative 3. - **List errors lose their status code** (#851). A list-based existence probe can't tell a missing bucket from a forbidden one. - **One unparseable key fails a listing.** Listing parses keys with `Path::parse`, so a single key it rejects (for example `a//b` written by another tool) fails the whole page. Such a bucket can't be emptied through `object_store` before `delete_bucket`. - **S3 ignores the bucket name with a custom endpoint.** With an explicit endpoint and virtual-hosted-style requests, S3 uses the endpoint as-is. Stores built for different bucket names then all address the same bucket, and the new bucket operations do too. This is documented on the `AmazonS3` implementation. - **A range past the object's end is a generic error.** A range starting at or beyond the object's length comes back as `Error::Generic` rather than a distinct error, so callers clamping ranges need a `head` first. - **Wrong emulator variable in `CONTRIBUTING.md`.** Its Azure instructions set `AZURE_USE_EMULATOR`, which `from_env` ignores; the key is `AZURE_STORAGE_USE_EMULATOR`. Open questions, assuming this is in scope for the crate: 1. Should the operations act on the store's own bucket, as proposed, or take a bucket name (see alternative 2)? 2. `bucket_exists() -> bool`, or `bucket_info() -> Option<BucketMeta>` (alternative 3)? 3. Does the GCS project-ID option look right: the key name, `GOOGLE_PROJECT_ID` taking precedence over `GOOGLE_CLOUD_PROJECT`, and the fallback to the service-account key? Transparency note : This Issue was drafted using the Help of AI, The Contents of the Issue have been vetted but might sound slightly robotic -- apologies for the same -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
