wangzhigang1999 opened a new pull request, #10143:
URL: https://github.com/apache/paimon/pull/10143

   ### Purpose
   
   Part of #10142.
   
   When projecting a few keys from a shared-shredding MAP, PyPaimon converts 
the field mapping to Python lists and searches candidate columns for each row 
and key. For large batches, this adds CPU overhead even after pruning unrelated 
columns.
   
   This PR optimizes selected-key MAP reconstruction for large batches by using 
Arrow operations to locate and extract values for supported scalar types. The 
vectorized path is enabled when the batch entering reconstruction contains at 
least 1,024 rows and the sum of candidate physical-column counts across 
requested keys is at most 16. Other inputs retain the row-wise path.
   
   Both paths preserve first-match selection, physical null values, null MAPs, 
sliced-array offsets, and overflow handling.
   
   ### Tests
   
   - Added regression coverage for nulls, missing keys, sliced arrays, 
overflow, malformed mappings, and both sides of the activation thresholds. 
Passed 114 tests and 48 subtests across seven related modules, plus 424 
independent reconstruction cases each on PyArrow 6.0.1 and 19.0.1.
   - Passed repository Flake8 and diff checks. Pyright reports no new 
diagnostics and no diagnostics in the changed functions.
   
   **Benchmark datasets**
   
   The benchmark measures PyPaimon reader scans and MAP projection. Labels and 
attributes from the following datasets were written to shared-shredding Paimon 
tables stored as Parquet. Both versions read the same files and fully consume 
the projected results.
   
   | Dataset | Source and MAP contents | Size |
   | --- | --- | ---: |
   | Prometheus labels | A frozen one-hour snapshot from the [PromLabs 
demo](https://demo.promlabs.com/); one row per sample, with `__name__` in a 
separate column and other labels stored as a MAP | 5,198 series; 1,151,842 rows 
|
   | OTel resource attributes | The first two consecutive row groups of public 
OpenTelemetry demo logs from [ClickHouse 
TextBench](https://github.com/ClickHouse/TextBench); `ResourceAttributes` 
contains service and Kubernetes metadata | 1,236,992 rows |
   | OTel event attributes | `LogAttributes` from the same log records, 
including file paths, output streams and user IDs | 1,236,992 rows |
   | Amazon product attributes | The `details` field from the All Beauty 
product metadata in [Amazon Reviews 
2023](https://amazon-reviews-2023.github.io/) | 112,590 rows |
   | GitHub event attributes | Event `payload` from the complete 2025-01-01 
12:00 UTC hour in [GH Archive](https://www.gharchive.org/) | 223,540 rows |
   
   The OTel inputs are demo-generated logs; both attribute presets use the same 
records. Prometheus samples retain series order. Nested Amazon and GitHub 
attributes are encoded as JSON strings, preserving empty MAPs and original 
record order.
   
   Synthetic controls cover 14 scalar types with stable or moving key-to-column 
mappings: 28 tables of 32,771 rows each. Four additional service-layout and 
overflow tables contain 32,768 or 65,536 rows each. These controls exercise 
value types, mapping layouts, nulls and fallback paths.
   
   **Benchmark cases**
   
   Each of the five public/demo presets runs the following 15 cases. Hot keys 
are selected before timing by descending row frequency, with key order breaking 
ties.
   
   | Case | Read operation |
   | --- | --- |
   | Plain column and full MAP, 2 cases | Read row ID `rid` alone, or `rid` 
plus the complete `tags` MAP, as controls for unchanged paths |
   | Hot keys, 3 cases | Read `rid` plus the 1, 4 or 8 most frequent MAP keys |
   | Rare, missing and mixed keys, 3 cases | Read the least frequent existing 
key, an absent key, or the most frequent key together with an absent key |
   | Filter then project, 1 case | Filter on the plain column with `rid < 10% 
of the row count`, then project the most frequent key; no MAP-value predicate 
pushdown |
   | Batch sizes, 4 cases | Project the four most frequent keys with requested 
batches of 32, 256, 1,024 or 4,096 rows |
   | Wide projection and batch consumption, 2 cases | Project eight hot keys 
with batch size 1,024, or consume the same static table's four-key projection 
through the batch reader |
   
   The metrics four-key case reads `job`, `instance`, `id` and `le`. Resource 
logs use `service.name`, `k8s.namespace.name`, `k8s.node.name` and 
`k8s.pod.name`. The event-log single-key case reads `log.file.path`; its 
four-key case adds `log.iostream`, `logtag` and `userId`. Key projection scans 
input rows without filtering out rows that lack the requested keys.
   
   These 75 cases plus 84 scalar and 18 layout comparisons make up the 177 
formal comparisons. Parquet row groups can cap actual batches: requesting 4,096 
rows can still yield 1,024-row batches. Actual batch sizes are recorded in the 
reports.
   
   **Benchmark results**
   
   The comparison uses the unoptimized baseline at `9b630346bd3b`, Python 
3.11.13, PyArrow 19.0.1, an Intel Xeon Platinum 8575C, and one thread each for 
Arrow CPU and I/O operations. The following results are median read times from 
three measured runs after warmup, on local Parquet files with warm OS cache. 
Result validation is excluded from timing.
   
   | Workload | Before | After | Read time reduction |
   | --- | ---: | ---: | ---: |
   | Prometheus labels, 4 keys, 1.15M rows | 7.247 s | 1.462 s | 79.8% |
   | OTel resource attributes, 4 keys, 1.24M rows | 8.611 s | 2.497 s | 71.0% |
   | OTel event attributes, 1 key, 1.24M rows | 8.413 s | 2.872 s | 65.9% |
   | OTel event attributes, 4 keys (row-wise fallback) | 9.889 s | 9.665 s | 
2.3% |
   | Prometheus labels, batch size 256 (row-wise fallback) | 8.069 s | 8.168 s 
| -1.2% |
   | Prometheus complete MAP (unchanged path) | 9.725 s | 9.784 s | -0.6% |
   
   All 1,440 reader executions in the groups above passed correctness and 
integrity checks. Under the tested layouts and environment, 93 of the 177 
comparisons reduced median read time by more than 5% and 84 stayed within ±5%; 
none regressed by more than 5%.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to