wangyong9999 opened a new issue, #10314:
URL: https://github.com/apache/paimon/issues/10314

   ### Search before asking
   
   - [x] I searched in the [issues](https://github.com/apache/paimon/issues) 
and found nothing similar.
   
   
   ### Motivation
   
   File index is pluggable through the `FileIndexerFactory` SPI, and a 
`FileIndexWriter` receives every row of a data file in file order. However, the 
writer is created by `FileIndexer#createWriter()` without knowing which data 
file it indexes, although `DataFileIndexWriter` is always constructed right 
after the data file path is decided.
   
   This blocks plugins whose output has to reference the data file, for example:
   - maintaining an external `key -> (data file, row position)` mapping to 
serve point lookups on append tables, where a single row can then be read by 
position with the Parquet page index;
   - recording per-file lineage or audit information in an external system.
   
   Current workarounds are fragile:
   - passing the path through a `ThreadLocal`, which depends on internal 
construction order;
   - wrapping the file format to get `FileAwareFormatWriter#setFile`, which 
changes the data file extension and requires registering the custom format on 
every reader;
   - resolving files in a `CommitCallback`, which needs an extra token-to-file 
indirection and delays availability until commit.
   
   Paimon already exposes the file location at write time elsewhere: 
`FileAwareFormatWriter#setFile` for format writers, and 
`TableWrite#withBlobConsumer` (#7074), which reports the uri, offset and length 
of each blob as it is written. File index writers are the remaining writers 
without this information.
   
   ### Solution
   
   Add a small context class and a default method, fully backward compatible:
   
   ```java
   public class FileIndexWriterContext {
       public Path dataFilePath();  // the data file indexed by the writer
       public long schemaId();      // the schema the data file is written with
   }
   
   public interface FileIndexer {
       FileIndexWriter createWriter();
   
       default FileIndexWriter createWriter(FileIndexWriterContext context) {
           return createWriter();
       }
   }
   ```
   
   The schema id is included because Paimon's own reader resolves a data file 
through `DataFileMeta#schemaId` to handle schema evolution; an external 
reference that bypasses the manifest has to carry it.
   
   `DataFileIndexWriter` passes the context to all index writers, including 
nested map index writers. Every call site provides it: `RowDataFileWriter`, 
`KeyValueDataFileWriter`, `KeyValueClusteringFileWriter` and 
`FileIndexProcessor`, so `rewrite_file_index` also provides it for existing 
files.
   
   - Built-in indexes and the index file format are unchanged.
   - Existing plugins behave the same unless they override the new method.
   
   ### Anything else?
   
   _No response_
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to