zenfenan opened a new issue, #3657:
URL: https://github.com/apache/iceberg-python/issues/3657

   ### Feature Request / Improvement
   
   When scanning a table with row filter, `_open_manifest` currently 
deserializes all live manifest entries via 
`ManifestEntry#fetch_manifest_entry()` and then applies partition and metrics 
filters as a second pass. This materialises the full list of entries before 
discarding non-matching ones.
   
   The intermediate list and second iteration pass add overhead that scales 
with total entry count, not matched entry count.
   
   **Proposal:**
   
   Add `ManifestFile#prune_manifest_entry()` that applies the partition/metrics 
evaluation as the entries are deserialized, avoiding the intermediate list 
allocation for non-matching entries. Update `_open_manifest()` to use it.
   
   **Benchmark results** (across multiple tables and filter expression):
   
   | Total Manifest Entries | Entries Matching Filter | Two-pass Time | Fused 
Time | Improvement |
   
|-----------------------|------------------------|----------------|-------------|--------------|
   | 116k | 78 (0.07%) | 9.13s | 7.18s | -21.3% |
   | 438k | 1 (~0%) | 35.12s | 27.81s | -20.8% |
   | 823k | 4,962 (0.6%) | 54.38s | 41.38s | -23.9% |
   | 18k | 269 (1.5%) | 1.23s | 0.90s | -27.0% |
   | 8k | 185 (2.2%) | 0.57s | 0.40s | -30.2% |
   
   No behavioral change - correctness validated across all runs.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to