TheR1sing3un opened a new pull request, #10165:
URL: https://github.com/apache/paimon/pull/10165

   ### Purpose
   
   Aggregation tables configured with unsupported options can currently commit 
data using deduplicate semantics. For example, enabling 
`aggregation.remove-record-on-delete` still allows same-key values `10` and 
`20` to be written as `20`, despite the table declaring `sum`. A later Python 
read error cannot recover the discarded input, and another engine can read the 
incorrect file.
   
   Run the existing read-side option validation before constructing the 
aggregation table's write buffer. Unsupported retract options, sequence groups, 
aggregator identifiers and invalid sequence configurations fail before data is 
buffered. Explicitly false retract flags remain accepted.
   
   This complements the supported-aggregation writer implementation in #10162 
with write-side validation. The PR is based on master and can be reviewed 
independently; both changes touch the aggregation dispatch and its old 
fallback-warning test.
   
   ### Tests
   
   - 92 focused aggregation, partial-update, sequence-field, merge-buffer and 
dispatch tests passed on Python 3.11 / PyArrow 19.0.1.
   - Seven selected regressions fail with the original writer dispatch. Tests 
verify batch and stream rejection, no committed snapshot or Parquet data files 
after failure, and writable false-valued options.
   - Combined with #10162 and the sequence-field writer fix: 135 focused tests 
passed.
   - Changed-file flake8, Python 3.6 syntax parsing, license headers and `git 
diff --check` passed.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to