JingsongLi commented on PR #9606: URL: https://github.com/apache/paimon/pull/9606#issuecomment-5747999915
Requirement fit: PIVOT. Implementation: CLEAN for the low-level copy primitive. The production scenario is real: the in-place `HiveMigrator` passes one `HiveCatalog.fileIO()` into `FileMetaUtils.construct`, which renames source Hive files into the target Paimon bucket, so a same-region cross-bucket migration currently needs a transfer path. The adaptive multipart sizing and copy-before-delete behavior address the previously reviewed correctness boundary. The remaining concern is ownership and credentials, not absence of end-to-end value. A single `OSSFileIO` can only use one configured credential/endpoint/security policy for both buckets; production migrations commonly have different source and target identities. The newer Flink clone path already models this correctly with separate source and target FileIOs and streams between them. Please move the capability to the migration/copy boundary (source FileIO + target FileIO), using the server-side OSS copy optimization only when both sides prove compatible, and retain the stream-copy fallback otherwise. That keeps generic `FileIO.rename` within one owner while supporting the real migration use case. I am not closing this PR because the user-visible need is demonstrated, but the current abstraction should pivot before merge. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
