gripleaf commented on PR #323: URL: https://github.com/apache/paimon-cpp/pull/323#issuecomment-5617814817
We evaluated this optimization using the same query load on a Paimon table with 90 partitions and approximately 16K buckets per partition. The workload performs point lookups using `WHERE rowkey = value`, with the bucket inferred from the predicate. **Before** Manifest reading still spends substantial CPU decoding Avro records into Arrow arrays before filtering candidates. `AvroDirectDecoder::DecodeAvroToBuilder` accounts for **58.65%** of sampled CPU, and `PrepareBucketRead` does not appear in the profile. <img width="3942" height="1182" alt="image" src="https://github.com/user-attachments/assets/f32aa413-4329-43b0-b2fe-d0384df40420" /> **After** The profile now includes `PrepareBucketRead`: the first pass reads only the fields required to identify candidate rows, and the second pass fully decodes only those candidates. Entries with historical bucket counts or schema IDs remain eligible for the existing compatibility checks. <img width="3942" height="1370" alt="image" src="https://github.com/user-attachments/assets/6e729e34-933b-4f1b-bfc4-79dea8a1c13f" /> Under the same query load, estimated CPU consumption decreased substantially: | Call path (inclusive CPU) | Before (CPU cores) | After (CPU cores) | Reduction | |---|---:|---:|---:| | `ManifestFile::ReadInferredBucketEntries` | 30.74 | 2.50 | 91.9% | | `AvroFileBatchReader::NextBatch` | 27.83 | 2.31 | 91.7% | | `AvroDirectDecoder::DecodeAvroToBuilder` | 21.92 | 1.16 | 94.7% | These estimates are derived from the 99 Hz profiles and normalized by the available sampling duration: 39 seconds before and 40 seconds after. The rows represent overlapping call paths and should not be added together. The results demonstrate a substantial reduction in CPU spent decoding manifest metadata under the tested workload. They measure CPU consumption, rather than end-to-end query latency or a corresponding query speedup. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
