zhuxiangyi opened a new pull request, #9979:
URL: https://github.com/apache/paimon/pull/9979

   ## Purpose
   
   Chain branch scans are incomplete inputs to a later merge. Applying 
value-statistics pruning independently to snapshot/delta partitions can remove 
a newer update or deletion needed to suppress an older matching row. A residual 
predicate cannot repair the stale row, since the old value itself satisfies it. 
Partial updates can also lose a required version.
   
   Minimal example: snapshot `(k=1, seq=1, v=old)` plus delta `(k=1, seq=2, 
v=new)`. Reading the logical delta partition with `v=old` must return no rows. 
Before this change it can return the snapshot row.
   
   ## Changes
   
   - Project predicates inclusively onto trimmed primary keys before scanning 
inputs to a chain merge.
   - Preserve original row indices and reuse existing AND/OR projection 
semantics.
   - Keep the original filter for complete snapshot reads.
   - Leave logical partition/anchor selection, authorization, bucket filtering 
and merge-reader behavior unchanged.
   
   No new options or dependencies.
   
   ## Validation
   
   - Baseline regression reproduced; the first expanded 12 configurations 
failed by assertion, not harness error.
   - Final targeted Maven reactor: **132 tests passed** (9 Common + 123 Core), 
with normal quality checks enabled.
   - Table-level differential tests compare full row multisets against 
merge-before-filter results; a seeded version-history test additionally uses an 
independent expected-state map.
   - Cover Parquet/ORC, updates, NULL transitions, deletes with DV disabled, 
partial-update, multiple deltas/groups, no snapshot anchor, key-range 
splitting, safe key pruning, complete-snapshot pruning, schema reorder across 
files, logical partition predicates, key transforms and authorization-masked 
keys.
   - Production and changed tests compile with actual JDK 8; runtime validation 
used JDK 17.0.13.
   
   Command:
   ```sh
   mvn -pl paimon-core -am -DwildcardSuites=none -DfailIfNoTests=false \
     
-Dtest=PredicateProjectionConverterTest,ChainTableFileStoreTableTest,ChainTableDeletionVectorReadTest,FallbackReadFileStoreTableTest,ChainTableUtilsTest,ChainPartitionProjectorTest,ChainSplitTest,KeyValueFileStoreScanTest,MergeFileSplitReadTest
 test
   ```
   
   ## Performance tradeoff (not a no-regression claim)
   
   Local warm-cache benchmark: four groups, 16 key ranges, 2,048 rows per 
range, three versions (393,216 physical records; 131,072 final rows), 
separately prepared immutable Parquet and ORC fixtures. Baseline and fixed use 
the exact same fixture for each format. Two JVM rounds per variant, each with 3 
warmups + 11 measurements per workload and push/no-push mode. Actual 
implementation and fixture hashes are recorded; exact full rows are checked on 
every sample.
   
   - Fixed: all **672 samples** correct; baseline: 56 stale-value samples 
incorrect (32,768 stale rows instead of zero).
   - Stable key filtering still prunes **192 files to 12**, and 
complete-snapshot value filtering **64 to 32**.
   - Median fixed key query versus fixed no-push is about 2.51x faster for 
Parquet and 2.04x for ORC on this fixture.
   - Necessary cost: a value-only `v=absent` query can no longer discard all 
merge inputs independently. Median total time increased **31.59→105.31 ms 
(Parquet)** and **32.64→72.62 ms (ORC)**, with 0→192 planned files. Its old 
answer happens to be correct for this fixture, but the pruning is not safe in 
general.
   - Old `v=old` timings are not performance baselines because those answers 
are wrong.
   - Untouched unfiltered timings also varied across JVMs (+12.2% Parquet, 
-2.0% ORC). These single-machine small-file measurements are descriptive, not 
production throughput predictions or proof of no regression.
   
   Planned bytes mean selected file metadata size, not physical I/O. Read time 
includes residual evaluation and output materialization.
   
   ## Scope limitations
   
   A separate pre-existing cross-branch tombstone case with DV enabled failed 
even without a pushed predicate; this PR does not claim to fix it. DV update 
coverage and existing DV read tests pass. The pre-existing `withBucket(int)` 
no-op on this chain path is also outside scope; explicit `withBucketFilter` is 
tested.
   
   No Spark/Flink SQL integration matrix, remote object-storage benchmark or 
full repository suite was run. Benchmark scripts/results are local review 
evidence and are not added to production modules.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to