wangzhigang1999 opened a new pull request, #10148: URL: https://github.com/apache/paimon/pull/10148
### Purpose Part of #10142. Full shared-shredding MAP reconstruction converts each field ID in `__field_mapping` through Arrow's `to_pylist()`. For wide mappings, this conversion accounts for a substantial part of the read time. Convert the integer buffer through a NumPy view, then reconstruct mapping rows from their offsets and validity. Use this path for batches of at least 32 rows with physical value columns and no null mapping children; retain the existing path for other inputs. The shared-shredding schema defines mapping children as int32, independently of the MAP value type. This changes only `assemble_shared_shredding_map`. Value handling, entry order, duplicate entries and ID validation remain unchanged for valid shared-shredding schemas. Ordinary MAP and shared-shredding selected-key projection are outside this PR's scope. Local reader measurements used four frozen append-only Parquet tables, Python 3.11.13 / PyArrow 19.0.1 on Linux x86-64 (Xeon Platinum 8575C), with Arrow CPU/I/O threads and reader parallelism set to 1. Each case used separate processes, alternating baseline/candidate order, one warmup and three measured runs per side under a shared timing lock. Times below are medians for planning, full consumption and materialization. | Dataset | Batch | Baseline (s) | Candidate (s) | Time reduction | RSS baseline / candidate (MiB) | |---|---|---:|---:|---:|---:| | scalar_int64 | default | 0.238532 | 0.202746 | 15.00% | 210.3 / 210.3 | | scalar_int64 | 32 | 0.456362 | 0.427029 | 6.43% | 209.3 / 209.7 | | scalar_moving_string | default | 0.187024 | 0.147425 | 21.17% | 208.7 / 209.4 | | scalar_moving_string | 32 | 0.397368 | 0.367593 | 7.49% | 211.6 / 211.5 | | amazon | default | 0.989895 | 0.556463 | 43.79% | 277.4 / 275.4 | | amazon | 32 | 2.206143 | 1.808098 | 18.04% | 282.7 / 282.7 | | shared_sparse_unique_int64_w64 | default | 0.400907 | 0.172098 | 57.07% | 221.1 / 221.3 | | shared_sparse_unique_int64_w64 | 32 | 0.853646 | 0.627120 | 26.54% | 219.5 / 217.7 | The scalar fixtures have 32,771 rows and four physical slots plus overflow; the string fixture changes key-to-slot assignments. Amazon has 112,590 rows. The sparse fixture has 16,387 rows, 64 slots, unique keys and about 90% empty MAPs. The measurements compare against `9b630346bd3b52463aa57c12c8df0a69ccd0a1a2`, whose shared-shredding reader file matches this PR's upstream base. The measured candidate predates removal of a redundant integer-type guard and the final offset-pairing cleanup with `zip`; the exact submitted revision has not been performance-benchmarked. These results cover the stated local workloads, not OSS or other layouts. In a separate mapping-only measurement at 1,024 rows and 64 slots, peak Python allocation increased from about 0.56 to 1.11 MiB. The temporary flat list is released after row reconstruction. Reader RSS above includes import overhead. ### Tests Added four focused tests covering LIST/LARGE_LIST slices, 31/32/33-row boundaries, null/empty MAPs, entry order and duplicates, nested values, unknown/negative IDs, large uint64 IDs, invalid mappings, missing physical columns and ORC time conversion. On the final code: - All four focused tests passed on Python 3.12 / PyArrow 23.0.1, including after moving the isolated commit onto the upstream base. - Ruff check, formatting of added code, project-configured Flake8 7.3.0 and `git diff --check` passed. Pyright reports no new diagnostics; eight existing production diagnostics match the baseline, and the new test file has none. ```sh PYTHONPATH=paimon-python python -m unittest pypaimon.tests.shared_shredding_full_restore_test -v ``` Additional validation before the guard removal and offset-pairing cleanup: - 4 reader smoke runs and 64 formal reader runs passed independent correctness and integrity checks. The formal count includes 16 warmups and 48 measured reads. All eight comparisons have three complete measured rounds per side; actual 32-row batches and tail batches match on both sides. - The focused and existing shared-shredding reader test modules passed on PyArrow 19.0.1: 16 tests and 15 subtests. - Earlier raw-input validation passed 396 checks on each of PyArrow 6.0.1 and 19.0.1 across 28 frozen tables and 14 value types, including boundary slices and NaN. A separate duplicate-key sparse writer round-trip experiment failed 16 reads on both baseline and candidate. Those failures are excluded from the passing counts and performance table above. This PR preserves already-encoded duplicate entries and does not address that writer issue. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
