wangzhigang1999 opened a new pull request, #10147:
URL: https://github.com/apache/paimon/pull/10147

   ### Purpose
   
   Part of #10142.
   
   When projecting selected keys from an ordinary MAP column, PyPaimon converts 
the full key array to Python objects and searches each row for each requested 
key. This adds CPU overhead even when the query selects only one key.
   
   Use Arrow's `map_lookup(..., "first")` in 
`assemble_normal_map_selected_keys` for string and large-string keys. Preserve 
first-match semantics, including a null first value, parent MAP nulls, and ORC 
temporal restoration. Keep the existing path for Arrow versions without this 
kernel, unsupported value builders, and other key types.
   
   ### Tests
   
   - Five added unit tests pass on PyArrow 6.0.1, 19.0.1 and 23.0.1, covering 
duplicate keys, null/empty MAPs, missing keys, sliced arrays, scalar and nested 
values, ORC time restoration, non-string keys, and the extension-value fallback.
   - 452 benchmark reader executions pass value/type/null checks and 
source/table integrity checks, including warmups and 32 
baseline-versus-baseline calibration executions.
   - 16 additional full-reader queries pass a separate Python first-match 
oracle, including all 112,590 Amazon rows.
   - Ruff lint and formatting checks, project Flake8, and `git diff --check` 
pass. License checks passed during implementation. Pyright reports no new 
diagnostics; the module retains eight existing diagnostics that match the 
baseline.
   
   Benchmark against `9b630346bd3b52463aa57c12c8df0a69ccd0a1a2`: Python 
3.11.13, PyArrow 19.0.1, local Parquet, one Arrow/BLAS thread. Each side runs 
in a fresh process, with alternating order, one warmup and five measured 
repetitions. The table below shows medians for the full Amazon dataset written 
with ordinary MAP storage.
   
   | Query | Before | After | Time reduction |
   |---|---:|---:|---:|
   | Project 1 key | 465.86 ms | 100.97 ms | 78.3% |
   | Project 8 keys | 708.86 ms | 263.53 ms | 62.8% |
   | Project 4 keys, batch size 32 | 1296.95 ms | 784.19 ms | 39.5% |
   | Read the complete MAP (control) | 62.37 ms | 62.31 ms | Within 1% |
   
   Across the Amazon dataset and three synthetic MAP layouts, the 30 
selected-key comparisons reduce median time by 28.6%–87.9%; the four 
complete-MAP controls remain within 1%. Baseline-versus-baseline calibration 
differs by less than 1%.
   
   Timing covers planning, reading and Arrow materialization, excluding process 
startup, imports, catalog initialization and correctness checks. These are 
warmed local-file results; OSS and cold-storage performance are outside this 
measurement.
   
   These measurements precede the final formatting and replacement of the 
kernel existence check with `getattr`. The unit tests were rerun after that 
cleanup, and again on the rebased PR branch with PyArrow 23.0.1.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to