wangzhigang1999 opened a new pull request, #10143: URL: https://github.com/apache/paimon/pull/10143
### Purpose Part of #10142. When projecting a few keys from a shared-shredding MAP, PyPaimon converts the field mapping to Python lists and searches candidate columns for each row and key. For large batches, this adds CPU overhead even after pruning unrelated columns. This PR optimizes selected-key MAP reconstruction for large batches by using Arrow operations to locate and extract values for supported scalar types. The vectorized path is enabled when the batch entering reconstruction contains at least 1,024 rows and the sum of candidate physical-column counts across requested keys is at most 16. Other inputs retain the row-wise path. Both paths preserve first-match selection, physical null values, null MAPs, sliced-array offsets, and overflow handling. ### Tests - Added regression coverage for nulls, missing keys, sliced arrays, overflow, malformed mappings, and both sides of the activation thresholds. Passed 114 tests and 48 subtests across seven related modules, plus 424 independent reconstruction cases each on PyArrow 6.0.1 and 19.0.1. - Passed repository Flake8 and diff checks. Pyright reports no new diagnostics and no diagnostics in the changed functions. **Benchmark datasets** The benchmark measures PyPaimon reader scans and MAP projection. Labels and attributes from the following datasets were written to shared-shredding Paimon tables stored as Parquet. Both versions read the same files and fully consume the projected results. | Dataset | Source and MAP contents | Size | | --- | --- | ---: | | Prometheus labels | A frozen one-hour snapshot from the [PromLabs demo](https://demo.promlabs.com/); one row per sample, with `__name__` in a separate column and other labels stored as a MAP | 5,198 series; 1,151,842 rows | | OTel resource attributes | The first two consecutive row groups of public OpenTelemetry demo logs from [ClickHouse TextBench](https://github.com/ClickHouse/TextBench); `ResourceAttributes` contains service and Kubernetes metadata | 1,236,992 rows | | OTel event attributes | `LogAttributes` from the same log records, including file paths, output streams and user IDs | 1,236,992 rows | | Amazon product attributes | The `details` field from the All Beauty product metadata in [Amazon Reviews 2023](https://amazon-reviews-2023.github.io/) | 112,590 rows | | GitHub event attributes | Event `payload` from the complete 2025-01-01 12:00 UTC hour in [GH Archive](https://www.gharchive.org/) | 223,540 rows | The OTel inputs are demo-generated logs; both attribute presets use the same records. Prometheus samples retain series order. Nested Amazon and GitHub attributes are encoded as JSON strings, preserving empty MAPs and original record order. Synthetic controls cover 14 scalar types with stable or moving key-to-column mappings: 28 tables of 32,771 rows each. Four additional service-layout and overflow tables contain 32,768 or 65,536 rows each. These controls exercise value types, mapping layouts, nulls and fallback paths. **Benchmark cases** Each of the five public/demo presets runs the following 15 cases. Hot keys are selected before timing by descending row frequency, with key order breaking ties. | Case | Read operation | | --- | --- | | Plain column and full MAP, 2 cases | Read row ID `rid` alone, or `rid` plus the complete `tags` MAP, as controls for unchanged paths | | Hot keys, 3 cases | Read `rid` plus the 1, 4 or 8 most frequent MAP keys | | Rare, missing and mixed keys, 3 cases | Read the least frequent existing key, an absent key, or the most frequent key together with an absent key | | Filter then project, 1 case | Filter on the plain column with `rid < 10% of the row count`, then project the most frequent key; no MAP-value predicate pushdown | | Batch sizes, 4 cases | Project the four most frequent keys with requested batches of 32, 256, 1,024 or 4,096 rows | | Wide projection and batch consumption, 2 cases | Project eight hot keys with batch size 1,024, or consume the same static table's four-key projection through the batch reader | The metrics four-key case reads `job`, `instance`, `id` and `le`. Resource logs use `service.name`, `k8s.namespace.name`, `k8s.node.name` and `k8s.pod.name`. The event-log single-key case reads `log.file.path`; its four-key case adds `log.iostream`, `logtag` and `userId`. Key projection scans input rows without filtering out rows that lack the requested keys. These 75 cases plus 84 scalar and 18 layout comparisons make up the 177 formal comparisons. Parquet row groups can cap actual batches: requesting 4,096 rows can still yield 1,024-row batches. Actual batch sizes are recorded in the reports. **Benchmark results** The comparison uses the unoptimized baseline at `9b630346bd3b`, Python 3.11.13, PyArrow 19.0.1, an Intel Xeon Platinum 8575C, and one thread each for Arrow CPU and I/O operations. The following results are median read times from three measured runs after warmup, on local Parquet files with warm OS cache. Result validation is excluded from timing. | Workload | Before | After | Read time reduction | | --- | ---: | ---: | ---: | | Prometheus labels, 4 keys, 1.15M rows | 7.247 s | 1.462 s | 79.8% | | OTel resource attributes, 4 keys, 1.24M rows | 8.611 s | 2.497 s | 71.0% | | OTel event attributes, 1 key, 1.24M rows | 8.413 s | 2.872 s | 65.9% | | OTel event attributes, 4 keys (row-wise fallback) | 9.889 s | 9.665 s | 2.3% | | Prometheus labels, batch size 256 (row-wise fallback) | 8.069 s | 8.168 s | -1.2% | | Prometheus complete MAP (unchanged path) | 9.725 s | 9.784 s | -0.6% | All 1,440 reader executions in the groups above passed correctness and integrity checks. Under the tested layouts and environment, 93 of the 177 comparisons reduced median read time by more than 5% and 84 stayed within ±5%; none regressed by more than 5%. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
