gripleaf commented on PR #323:
URL: https://github.com/apache/paimon-cpp/pull/323#issuecomment-5617814817

   We evaluated this optimization using the same query load on a Paimon table 
with 90 partitions and approximately 16K buckets per partition. The workload 
performs point lookups using `WHERE rowkey = value`, with the bucket inferred 
from the predicate.
   
   **Before**
   
   Manifest reading still spends substantial CPU decoding Avro records into 
Arrow arrays before filtering candidates. 
`AvroDirectDecoder::DecodeAvroToBuilder` accounts for **58.65%** of sampled 
CPU, and `PrepareBucketRead` does not appear in the profile.
   
   <img width="3942" height="1182" alt="image" 
src="https://github.com/user-attachments/assets/f32aa413-4329-43b0-b2fe-d0384df40420";
 />
   
   
   **After**
   
   The profile now includes `PrepareBucketRead`: the first pass reads only the 
fields required to identify candidate rows, and the second pass fully decodes 
only those candidates. Entries with historical bucket counts or schema IDs 
remain eligible for the existing compatibility checks.
   
   <img width="3942" height="1370" alt="image" 
src="https://github.com/user-attachments/assets/6e729e34-933b-4f1b-bfc4-79dea8a1c13f";
 />
   
   Under the same query load, estimated CPU consumption decreased substantially:
   
   | Call path (inclusive CPU) | Before (CPU cores) | After (CPU cores) | 
Reduction |
   |---|---:|---:|---:|
   | `ManifestFile::ReadInferredBucketEntries` | 30.74 | 2.50 | 91.9% |
   | `AvroFileBatchReader::NextBatch` | 27.83 | 2.31 | 91.7% |
   | `AvroDirectDecoder::DecodeAvroToBuilder` | 21.92 | 1.16 | 94.7% |
   
   These estimates are derived from the 99 Hz profiles and normalized by the 
available sampling duration: 39 seconds before and 40 seconds after. The rows 
represent overlapping call paths and should not be added together.
   
   The results demonstrate a substantial reduction in CPU spent decoding 
manifest metadata under the tested workload. They measure CPU consumption, 
rather than end-to-end query latency or a corresponding query speedup.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to