wangzhigang1999 opened a new pull request, #10147: URL: https://github.com/apache/paimon/pull/10147
### Purpose Part of #10142. When projecting selected keys from an ordinary MAP column, PyPaimon converts the full key array to Python objects and searches each row for each requested key. This adds CPU overhead even when the query selects only one key. Use Arrow's `map_lookup(..., "first")` in `assemble_normal_map_selected_keys` for string and large-string keys. Preserve first-match semantics, including a null first value, parent MAP nulls, and ORC temporal restoration. Keep the existing path for Arrow versions without this kernel, unsupported value builders, and other key types. ### Tests - Five added unit tests pass on PyArrow 6.0.1, 19.0.1 and 23.0.1, covering duplicate keys, null/empty MAPs, missing keys, sliced arrays, scalar and nested values, ORC time restoration, non-string keys, and the extension-value fallback. - 452 benchmark reader executions pass value/type/null checks and source/table integrity checks, including warmups and 32 baseline-versus-baseline calibration executions. - 16 additional full-reader queries pass a separate Python first-match oracle, including all 112,590 Amazon rows. - Ruff lint and formatting checks, project Flake8, and `git diff --check` pass. License checks passed during implementation. Pyright reports no new diagnostics; the module retains eight existing diagnostics that match the baseline. Benchmark against `9b630346bd3b52463aa57c12c8df0a69ccd0a1a2`: Python 3.11.13, PyArrow 19.0.1, local Parquet, one Arrow/BLAS thread. Each side runs in a fresh process, with alternating order, one warmup and five measured repetitions. The table below shows medians for the full Amazon dataset written with ordinary MAP storage. | Query | Before | After | Time reduction | |---|---:|---:|---:| | Project 1 key | 465.86 ms | 100.97 ms | 78.3% | | Project 8 keys | 708.86 ms | 263.53 ms | 62.8% | | Project 4 keys, batch size 32 | 1296.95 ms | 784.19 ms | 39.5% | | Read the complete MAP (control) | 62.37 ms | 62.31 ms | Within 1% | Across the Amazon dataset and three synthetic MAP layouts, the 30 selected-key comparisons reduce median time by 28.6%–87.9%; the four complete-MAP controls remain within 1%. Baseline-versus-baseline calibration differs by less than 1%. Timing covers planning, reading and Arrow materialization, excluding process startup, imports, catalog initialization and correctness checks. These are warmed local-file results; OSS and cold-storage performance are outside this measurement. These measurements precede the final formatting and replacement of the kernel existence check with `getattr`. The unit tests were rerun after that cleanup, and again on the rebased PR branch with PyArrow 23.0.1. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
