wangyong9999 opened a new issue, #10314: URL: https://github.com/apache/paimon/issues/10314
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon/issues) and found nothing similar. ### Motivation File index is pluggable through the `FileIndexerFactory` SPI, and a `FileIndexWriter` receives every row of a data file in file order. However, the writer is created by `FileIndexer#createWriter()` without knowing which data file it indexes, although `DataFileIndexWriter` is always constructed right after the data file path is decided. This blocks plugins whose output has to reference the data file, for example: - maintaining an external `key -> (data file, row position)` mapping to serve point lookups on append tables, where a single row can then be read by position with the Parquet page index; - recording per-file lineage or audit information in an external system. Current workarounds are fragile: - passing the path through a `ThreadLocal`, which depends on internal construction order; - wrapping the file format to get `FileAwareFormatWriter#setFile`, which changes the data file extension and requires registering the custom format on every reader; - resolving files in a `CommitCallback`, which needs an extra token-to-file indirection and delays availability until commit. Paimon already exposes the file location at write time elsewhere: `FileAwareFormatWriter#setFile` for format writers, and `TableWrite#withBlobConsumer` (#7074), which reports the uri, offset and length of each blob as it is written. File index writers are the remaining writers without this information. ### Solution Add a small context class and a default method, fully backward compatible: ```java public class FileIndexWriterContext { public Path dataFilePath(); // the data file indexed by the writer public long schemaId(); // the schema the data file is written with } public interface FileIndexer { FileIndexWriter createWriter(); default FileIndexWriter createWriter(FileIndexWriterContext context) { return createWriter(); } } ``` The schema id is included because Paimon's own reader resolves a data file through `DataFileMeta#schemaId` to handle schema evolution; an external reference that bypasses the manifest has to carry it. `DataFileIndexWriter` passes the context to all index writers, including nested map index writers. Every call site provides it: `RowDataFileWriter`, `KeyValueDataFileWriter`, `KeyValueClusteringFileWriter` and `FileIndexProcessor`, so `rewrite_file_index` also provides it for existing files. - Built-in indexes and the index file format are unchanged. - Existing plugins behave the same unless they override the new method. ### Anything else? _No response_ ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
