JingsongLi opened a new issue, #1048: URL: https://github.com/apache/paimon-rust/issues/1048
### Search before asking - [x] I searched existing issues and found no matching request. ### Motivation Java Paimon supports `scan.ignore-lost-files` and `scan.ignore-corrupt-files` for data-file reads. Rust currently only consults the lost-files option when falling back from a missing ROW auxiliary file to its primary file. Reading a missing or corrupt primary Parquet file still fails with the corresponding scan option enabled, preventing PyPaimon Native reads from honoring the same table options. ### Solution Apply file-scoped recovery to Paimon data-file creation and decoding, following Java `DataFileRecordReader`: - At reader creation, check the actual file's existence. Missing files require the lost-files option; present unreadable files require the corrupt-files option. - During batch decoding, only the corrupt-files option may end the current file. Preserve already returned batches and continue subsequent files. - For Data Evolution, end the column group if an active column file is skipped, matching `DataEvolutionFileReader` rather than inventing NULL values. - Preserve strict errors for resource limits, configuration, schema/metadata processing, and external BLOB payload resolution. - Keep both options disabled by default and retain all files. Java references: - https://github.com/apache/paimon/blob/master/paimon-core/src/main/java/org/apache/paimon/io/DataFileRecordReader.java - https://github.com/apache/paimon/blob/master/paimon-common/src/main/java/org/apache/paimon/reader/DataEvolutionFileReader.java ### Willingness to contribute - [x] I'm willing to submit a PR. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
