azwanzuharimi commented on issue #3388: URL: https://github.com/apache/iceberg-python/issues/3388#issuecomment-5916823502
I am working on this and will open a PR this week. Both earlier PRs, #3336 and #3661, closed as stale before a review. Since then the file format writer API landed in #3119 and #3381. So a new PR must go through `FileFormatWriter` and not call `pq.ParquetWriter` directly. Plan: - Add a `length()` method to `FileFormatWriter`. It returns the bytes written so far, from `OutputStream.tell()`. This mirrors `FileAppender.length()` in Java. - In the streaming path, write each batch through the format writer. Roll to a new file when `length()` reaches `write.target-file-size-bytes`. This mirrors `RollingFileWriter` in Java. - Keep the `pa.Table` path unchanged. - Add tests that check file sizes on disk and that a large stream does not hold more than one batch in memory. @paultmathew tell me if you plan to revive #3336, so we do not do the work twice. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
