TheR1sing3un opened a new pull request, #10166:
URL: https://github.com/apache/paimon/pull/10166

   ### Purpose
   
   A primary-key table with `sequence.field=seq` can keep a stale row when a 
newer business version arrives before an older version in the same write 
buffer. Writing `(seq=100, value='high')` followed by `(seq=50, value='low')` 
currently persists `low`; committing the rows separately returns `high` because 
the read-side merge already honors the sequence field.
   
   Sort buffered rows by primary key, configured sequence fields and generated 
sequence number before folding equal keys. Keep the columnar Arrow sort, 
nulls-first ordering in both directions, and generated sequence numbers as the 
tie breaker. Validate sequence configuration and supported field types before 
buffering.
   
   ### Tests
   
   - 202 focused writer, merge-buffer, dispatch, aggregation, partial-update 
and sequence-field tests plus 6 subtests passed on Python 3.11 / PyArrow 19.0.1.
   - 42 new real-table cases cover batch and stream writers; one batch, 
separate Arrow batches, row writes and separate commits; ascending/descending 
order, nulls, ties, compound fields, partitioned keys, partial-update null 
filling, buffer rolling, projection and typed sequence fields. Replacing the 
sort with the original implementation causes 33 failures; 9 unaffected cases 
pass.
   - Combined with #10162 and unsupported-aggregation write validation: 135 
focused tests passed.
   - Changed-file flake8, Python 3.6 syntax parsing, license headers and `git 
diff --check` passed.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to