wangzhigang1999 opened a new pull request, #10148:
URL: https://github.com/apache/paimon/pull/10148

   ### Purpose
   
   Part of #10142.
   
   Full shared-shredding MAP reconstruction converts each field ID in 
`__field_mapping` through Arrow's `to_pylist()`. For wide mappings, this 
conversion accounts for a substantial part of the read time.
   
   Convert the integer buffer through a NumPy view, then reconstruct mapping 
rows from their offsets and validity. Use this path for batches of at least 32 
rows with physical value columns and no null mapping children; retain the 
existing path for other inputs. The shared-shredding schema defines mapping 
children as int32, independently of the MAP value type.
   
   This changes only `assemble_shared_shredding_map`. Value handling, entry 
order, duplicate entries and ID validation remain unchanged for valid 
shared-shredding schemas. Ordinary MAP and shared-shredding selected-key 
projection are outside this PR's scope.
   
   Local reader measurements used four frozen append-only Parquet tables, 
Python 3.11.13 / PyArrow 19.0.1 on Linux x86-64 (Xeon Platinum 8575C), with 
Arrow CPU/I/O threads and reader parallelism set to 1. Each case used separate 
processes, alternating baseline/candidate order, one warmup and three measured 
runs per side under a shared timing lock. Times below are medians for planning, 
full consumption and materialization.
   
   | Dataset | Batch | Baseline (s) | Candidate (s) | Time reduction | RSS 
baseline / candidate (MiB) |
   |---|---|---:|---:|---:|---:|
   | scalar_int64 | default | 0.238532 | 0.202746 | 15.00% | 210.3 / 210.3 |
   | scalar_int64 | 32 | 0.456362 | 0.427029 | 6.43% | 209.3 / 209.7 |
   | scalar_moving_string | default | 0.187024 | 0.147425 | 21.17% | 208.7 / 
209.4 |
   | scalar_moving_string | 32 | 0.397368 | 0.367593 | 7.49% | 211.6 / 211.5 |
   | amazon | default | 0.989895 | 0.556463 | 43.79% | 277.4 / 275.4 |
   | amazon | 32 | 2.206143 | 1.808098 | 18.04% | 282.7 / 282.7 |
   | shared_sparse_unique_int64_w64 | default | 0.400907 | 0.172098 | 57.07% | 
221.1 / 221.3 |
   | shared_sparse_unique_int64_w64 | 32 | 0.853646 | 0.627120 | 26.54% | 219.5 
/ 217.7 |
   
   The scalar fixtures have 32,771 rows and four physical slots plus overflow; 
the string fixture changes key-to-slot assignments. Amazon has 112,590 rows. 
The sparse fixture has 16,387 rows, 64 slots, unique keys and about 90% empty 
MAPs.
   
   The measurements compare against `9b630346bd3b52463aa57c12c8df0a69ccd0a1a2`, 
whose shared-shredding reader file matches this PR's upstream base. The 
measured candidate predates removal of a redundant integer-type guard and the 
final offset-pairing cleanup with `zip`; the exact submitted revision has not 
been performance-benchmarked. These results cover the stated local workloads, 
not OSS or other layouts.
   
   In a separate mapping-only measurement at 1,024 rows and 64 slots, peak 
Python allocation increased from about 0.56 to 1.11 MiB. The temporary flat 
list is released after row reconstruction. Reader RSS above includes import 
overhead.
   
   ### Tests
   
   Added four focused tests covering LIST/LARGE_LIST slices, 31/32/33-row 
boundaries, null/empty MAPs, entry order and duplicates, nested values, 
unknown/negative IDs, large uint64 IDs, invalid mappings, missing physical 
columns and ORC time conversion.
   
   On the final code:
   
   - All four focused tests passed on Python 3.12 / PyArrow 23.0.1, including 
after moving the isolated commit onto the upstream base.
   - Ruff check, formatting of added code, project-configured Flake8 7.3.0 and 
`git diff --check` passed. Pyright reports no new diagnostics; eight existing 
production diagnostics match the baseline, and the new test file has none.
   
   ```sh
   PYTHONPATH=paimon-python python -m unittest 
pypaimon.tests.shared_shredding_full_restore_test -v
   ```
   
   Additional validation before the guard removal and offset-pairing cleanup:
   
   - 4 reader smoke runs and 64 formal reader runs passed independent 
correctness and integrity checks. The formal count includes 16 warmups and 48 
measured reads. All eight comparisons have three complete measured rounds per 
side; actual 32-row batches and tail batches match on both sides.
   - The focused and existing shared-shredding reader test modules passed on 
PyArrow 19.0.1: 16 tests and 15 subtests.
   - Earlier raw-input validation passed 396 checks on each of PyArrow 6.0.1 
and 19.0.1 across 28 frozen tables and 14 value types, including boundary 
slices and NaN.
   
   A separate duplicate-key sparse writer round-trip experiment failed 16 reads 
on both baseline and candidate. Those failures are excluded from the passing 
counts and performance table above. This PR preserves already-encoded duplicate 
entries and does not address that writer issue.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to