JingsongLi commented on PR #10278: URL: https://github.com/apache/paimon/pull/10278#issuecomment-5936175593
[P2] Preserve exact sequence-group field names during unrelated renames (`SchemaManagerUtils.java`, lines 728–739). Unlike bucket/sequence/clustering CSV readers, `PartialUpdateMergeFunction.sequenceGroupOrderingFields` and `sequenceGroupProtectedFields` split on comma without trimming; schema validation also resolves the ordering-field key by exact name. Leading/trailing spaces are allowed in quoted column names. Trimming these tokens therefore changes a valid field reference, even when that field is not being renamed. I reproduced this through a real catalog table with primary key `id`, columns ` ts`, ` value`, and `x`, and options `merge-engine=partial-update`, `fields. ts.sequence-group= value`. A write/commit succeeds. Renaming only `x`→`y` on this PR rewrites the existing option to `fields.ts.sequence-group=value` and fails with `Field ts can not be found in table schema`. The same committed-table rename with the merge-base `SchemaManagerUtils` succeeds and preserves the original sequence-group names. Trim only options whose actual parsing contract trims tokens. Keep sequence-group ordering/protected names exact, or introduce their parser-contract migration separately with compatibility coverage. Validation: 91 normal Maven tests passed, plus actual catalog rename/write/read checks for spaced bucket/sequence CSV, canonical clustering and its fallback. The committed-table boundary-space case above and the merge-base helper control establish the new regression. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
