HesandaLiyanage opened a new pull request, #3193: URL: https://github.com/apache/james-project/pull/3193
JIRA: https://issues.apache.org/jira/browse/JAMES-4231 ### Context & Problem High email ingest volumes in Apache James deployments utilizing S3/MinIO object storage result in hundreds of millions of small S3 objects (average body/header size ~20–50KB). Operating at this granularity introduces high S3 API request fees (PUT/GET/LIST), elevated bucket listing latency during storage maintenance and garbage collection, and reduced overall object storage throughput. ### Solution Overview This PR implements autonomous S3 object compaction for Apache James, packing historical standalone blobs into large, immutable chunk files (~100MB target size) with slot-level virtual addressing and transparent single-request HTTP ranged read support. ### Key Components & Architecture 1. **Chunk Format (`ChunkFormat`):** - Single binary chunk containing concatenated slots and a trailer footer: `[Header (16B)][Slot 0]...[Slot N-1][Footer][Footer Size (4B)][End Magic 0x454F4643]` - **Per-slot independent Zstandard compression:** Each slot has its own Zstd codec (`content-encoding=zstd\n`), allowing single-slot ranged reads without decompressing or buffering the full 100MB chunk. - **Suffix-indexed 64KB footer:** Suffix range read (`readRange(..., -65536, -1)`) fetches chunk slot metadata in $O(1)$ without scanning payloads. - Fully documented in `src/adr/0076-s3-object-compaction.md`. 2. **Virtual Slot Addressing & Ranged Reads (`ChunkedBlobStoreDAO` & `ChunkId`):** - Format: `<family>_<generation>_chunk<hash>~<offset>~<limit>`. - `ChunkedBlobStoreDAO` transparently intercepts slot references and issues HTTP byte-range requests directly against the underlying storage connector (`readRange(bucket, chunkId, offset, offset + limit - 1)`). - Conforms strictly to `ChunkMarker` regex (`^\d+_\d+_chunk[A-Za-z0-9_-]{16,}(~\d+~\d+)?$`) to ensure bloom-filter GC never misclassifies chunks as unreferenced garbage. 3. **Compaction Pipeline & Crash Safety (`BlobCompactionAlgorithm`):** - **Initial Compaction (`initialCompact`):** 1. Candidate scanning windowed into batches of 1,000 blobs (`DEFAULT_CANDIDATE_BATCH_SIZE = 1000`) so candidate byte arrays are packed and freed window-by-window without buffering the generation's payloads in heap. 2. Chunk assembled and saved to raw S3 storage. 3. Metadata references updated in Cassandra (`messageV3`, `messageIdToImapUid`, `messageIdTable`). 4. Original standalone blobs deleted from S3 only for candidates whose references updated successfully. If an update fails mid-batch, deletion is skipped for that candidate, guaranteeing zero dangling references. - **GC Compaction (`gcCompact`):** - Inspects existing chunks via footer-only ranged reads (metadata-only). - Orphan chunks (100% dead slots) deleted with 0 payload bytes read. - Chunks exceeding dead ratio rewritten; small adjacent chunks merged. Live slots streamed individually via ranged reads, bounding GC heap to $O(\text{maxSlotSize})$. - **Self-Healing Repairer (`CassandraBlobIdRepairer`):** - On missing slot / 404, falls back to original blob ID from `blobReferenceMapping` and restores Cassandra references. Also heals any transient inconsistency window across Cassandra tables. 4. **Configuration Guards & Scope:** - Compaction automatically disables itself with an informative log when client-side AES encryption or whole-blob compression is active (AES cipher blocks HTTP range slicing). - Generation-scoped: triggered via `DELETE /blobs?scope=compaction&generation=<gen>&family=<fam>`. 5. **Memory Characteristics & Operational Ceiling:** - Candidate payload heap: strictly bounded by $O(\min(N \times \text{avgBlobSize}, \text{chunkTargetSize}))$. - GC payload heap: strictly bounded by $O(\text{maxSlotSize})$ (~1MB). - Reference mapping ceiling: $O(\text{totalLiveGenerationReferences} \times \approx 200\text{ bytes})$ (~200MB heap for 1M live references; ~2GB heap for 10M). Documented in class javadoc; future follow-up can introduce partition-paged lookups. ### Verification & Tests - `server/blob/blob-compaction`: 47 tests passed (100%), including allocation-bound proof test and task serialization. - `server/blob/blob-storage-strategy`: 78 tests passed (100%). - `mailbox/cassandra` (`CassandraBlobId*IntegrationTest`): 3/3 passed with real Cassandra 5.0.9 testcontainers. - `server/blob/blob-s3`: Ranged read and MinIO S3 end-to-end compaction integration tests passed. - `server/container/guice/distributed`: 43 tests passed (100%). - Checkstyle: 0 errors across all 7 modified modules. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
