ChaomingZhangCN opened a new issue, #197:
URL: https://github.com/apache/paimon-cpp/issues/197

   ### Search before asking
   
   - [x] I searched in the 
[issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.
   
   
   ### Motivation
   
   Apache Paimon Java has introduced the VECTOR<T, N> data type based on 
[PIP-40](https://cwiki.apache.org/confluence/spaces/PAIMON/pages/399279132/PIP-40+Introduce+a+new+Vector+data+type).
 It supports schema representation, regular data-file storage such as Parquet, 
dedicated vector storage, reads, writes, and Data Evolution.
   
   Paimon C++ currently recognizes .vector. file names for some file-level 
bookkeeping, but it does not yet provide a VECTOR logical type, Arrow mapping, 
serialization, storage, or end-to-end read and write support.
   
   This issue implements the VECTOR roadmap item tracked in 
[#186](https://github.com/apache/paimon-cpp/issues/186).
   
   ### Solution
   
   Introduce `VECTOR<T, N>` support incrementally, while keeping the schema and 
storage behavior compatible with Apache Paimon Java.
   
   #### Phase 1: Schema and regular Parquet storage
   
   - [ ] Add `VECTOR<T, N>` to the Paimon C++ logical type system.
   - [ ] Implement schema JSON serialization and deserialization compatible 
with Paimon Java.
   - [ ] Map `VECTOR<T, N>` to Arrow `FixedSizeList<T, N>`.
   - [ ] Support the element types defined by PIP-40:
     - `BOOLEAN`
     - `TINYINT`
     - `SMALLINT`
     - `INT`
     - `BIGINT`
     - `FLOAT`
     - `DOUBLE`
   - [ ] Validate that the dimension is positive and fixed.
   - [ ] Validate that the written vector length equals `N`.
   - [ ] Reject null vector elements.
   - [ ] Support reading and writing VECTOR columns in regular Parquet data 
files.
   - [ ] Add end-to-end append-table tests.
   - [ ] Add Java/C++ schema and file compatibility tests.
   
   #### Phase 2: Data Evolution
   
   - [ ] Support adding and dropping VECTOR columns.
   - [ ] Support reading files written with previous table schemas.
   - [ ] Reject incompatible dimension changes, such as `VECTOR<FLOAT, 3>` to 
`VECTOR<FLOAT, 5>`.
   - [ ] Reject VECTOR columns as primary keys, partition keys, or sorting keys.
   - [ ] Add Data Evolution integration tests.
   
   #### Phase 3: Dedicated vector storage
   
   - [ ] Support the vector file format configuration used by Paimon Java.
   - [ ] Support reading and writing dedicated `*.vector.vortex` files.
   - [ ] Integrate vector files with row tracking and Data Evolution.
   - [ ] Integrate vector files with scan planning, file commits, and conflict 
handling.
   - [ ] Add Java, Python, and C++ Vortex compatibility tests.
   
   ### Initial scope
   
   The initial implementation can focus on Phase 1, providing a usable 
end-to-end vertical slice through schema representation, Arrow mapping, and 
regular Parquet reads and writes.
   
   Dedicated Vortex vector storage and Data Evolution can be delivered through 
follow-up pull requests under this issue.
   
   The following items are not required for the initial implementation:
   
   - ORC VECTOR support
   - Vector indexes or similarity search
   - Changing vector dimensions through schema evolution
   - VECTOR values inside shared-shredding MAP columns
   - Element types not supported by Apache Paimon Java
   
   PIP-40 should only be considered fully supported after all phases are 
complete. Completing Phase 1 means that regular Parquet VECTOR storage is 
supported, but does not imply support for dedicated Vortex vector files.
   
   ### Anything else?
   
   Related roadmap: [#186](https://github.com/apache/paimon-cpp/issues/186)
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to