fallintoplace opened a new pull request, #2143:
URL: https://github.com/apache/iceberg-go/pull/2143

   **What**
   - Sort Puffin blob indexes when reading all blobs.
   
   **Why**
   - Choosing read order only needs offsets. Copying full metadata into the 
ordering slice adds overhead.
   
   **Implementation**
   - Sort `[]int` with `slices.SortFunc` and look up metadata by index.
   - Keep offset-order reads, footer-order results and defensive metadata 
copies.
   - Add a benchmark using real in-memory Puffin files.
   
   **Benchmark**
   Apple M1 Pro, darwin/arm64, Go 1.26.3. Median of 5 runs, 300ms each, 
`-cpu=1`. Full `ReadAllBlobs`, 64-byte payloads, writer-order offsets.
   
   | Blobs | Before µs/op | After µs/op | B/op before → after | Allocs/op 
before → after |
   | --- | ---: | ---: | ---: | ---: |
   | 16 | 5.0 | 5.0 | 10408 → 8576 | 85 → 82 |
   | 256 | 86.9 | 77.5 | 162856 → 137472 | 1285 → 1282 |
   | 1024 | 325.1 | 312.0 | 640424 → 550144 | 5125 → 5122 |
   
   At 1,024 blobs, the temporary ordering slice saves 90,112 bytes per call. 
Storage I/O can dominate overall time.
   
   ```sh
   go test ./puffin -run '^$' -bench '^BenchmarkReadAllBlobs$' -benchmem 
-benchtime=300ms -count=5 -cpu=1
   ```
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to