JingsongLi commented on PR #9606:
URL: https://github.com/apache/paimon/pull/9606#issuecomment-5747999915

   Requirement fit: PIVOT. Implementation: CLEAN for the low-level copy 
primitive.
   
   The production scenario is real: the in-place `HiveMigrator` passes one 
`HiveCatalog.fileIO()` into `FileMetaUtils.construct`, which renames source 
Hive files into the target Paimon bucket, so a same-region cross-bucket 
migration currently needs a transfer path. The adaptive multipart sizing and 
copy-before-delete behavior address the previously reviewed correctness 
boundary.
   
   The remaining concern is ownership and credentials, not absence of 
end-to-end value. A single `OSSFileIO` can only use one configured 
credential/endpoint/security policy for both buckets; production migrations 
commonly have different source and target identities. The newer Flink clone 
path already models this correctly with separate source and target FileIOs and 
streams between them. Please move the capability to the migration/copy boundary 
(source FileIO + target FileIO), using the server-side OSS copy optimization 
only when both sides prove compatible, and retain the stream-copy fallback 
otherwise. That keeps generic `FileIO.rename` within one owner while supporting 
the real migration use case.
   
   I am not closing this PR because the user-visible need is demonstrated, but 
the current abstraction should pivot before merge.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to