dwangatt opened a new pull request, #10170:
URL: https://github.com/apache/paimon/pull/10170
## Summary
This is a focused follow-up extracted from #9370, following the merged
partition-layout scan prerequisite (#10052).
This PR makes partition bucket layout a core write contract:
- adds `PartitionBucketMapping` to resolve the current bucket count for each
partition;
- adds write APIs that carry the routing bucket count:
`write(partition, bucket, totalBuckets, data)`;
- propagates the mapping into write restore so empty buckets recover the
partition's active layout;
- ensures core overwrite writes can route and stamp files using the target
partition layout;
- rejects the legacy `write(partition, bucket, data)` path when
`bucket.per-partition-count-enabled=true` on a partitioned table.
The legacy path only carries a bucket id and cannot establish which bucket
count was used for hashing. Rejecting it prevents a caller from silently
routing a row with the table-level bucket count while the target partition uses
a different layout.
## Scope
This PR intentionally does not expose complete engine support for the option:
- Flink routing and end-to-end integration will be added in a follow-up PR.
- Spark will reject writes to tables with per-partition bucket counts until
it has a safe routing implementation.
- Documentation and end-to-end tests will follow with the Flink integration.
## Verification
- `mvn -q -pl paimon-core -DskipTests compile`
- `mvn -q -pl paimon-core spotless:check`
- `git diff --check`
The focused restore test was updated to cover legacy-path rejection,
stale-layout rejection, and a successful explicit partition-layout write.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]