ChaomingZhangCN opened a new issue, #197: URL: https://github.com/apache/paimon-cpp/issues/197
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar. ### Motivation Apache Paimon Java has introduced the VECTOR<T, N> data type based on [PIP-40](https://cwiki.apache.org/confluence/spaces/PAIMON/pages/399279132/PIP-40+Introduce+a+new+Vector+data+type). It supports schema representation, regular data-file storage such as Parquet, dedicated vector storage, reads, writes, and Data Evolution. Paimon C++ currently recognizes .vector. file names for some file-level bookkeeping, but it does not yet provide a VECTOR logical type, Arrow mapping, serialization, storage, or end-to-end read and write support. This issue implements the VECTOR roadmap item tracked in [#186](https://github.com/apache/paimon-cpp/issues/186). ### Solution Introduce `VECTOR<T, N>` support incrementally, while keeping the schema and storage behavior compatible with Apache Paimon Java. #### Phase 1: Schema and regular Parquet storage - [ ] Add `VECTOR<T, N>` to the Paimon C++ logical type system. - [ ] Implement schema JSON serialization and deserialization compatible with Paimon Java. - [ ] Map `VECTOR<T, N>` to Arrow `FixedSizeList<T, N>`. - [ ] Support the element types defined by PIP-40: - `BOOLEAN` - `TINYINT` - `SMALLINT` - `INT` - `BIGINT` - `FLOAT` - `DOUBLE` - [ ] Validate that the dimension is positive and fixed. - [ ] Validate that the written vector length equals `N`. - [ ] Reject null vector elements. - [ ] Support reading and writing VECTOR columns in regular Parquet data files. - [ ] Add end-to-end append-table tests. - [ ] Add Java/C++ schema and file compatibility tests. #### Phase 2: Data Evolution - [ ] Support adding and dropping VECTOR columns. - [ ] Support reading files written with previous table schemas. - [ ] Reject incompatible dimension changes, such as `VECTOR<FLOAT, 3>` to `VECTOR<FLOAT, 5>`. - [ ] Reject VECTOR columns as primary keys, partition keys, or sorting keys. - [ ] Add Data Evolution integration tests. #### Phase 3: Dedicated vector storage - [ ] Support the vector file format configuration used by Paimon Java. - [ ] Support reading and writing dedicated `*.vector.vortex` files. - [ ] Integrate vector files with row tracking and Data Evolution. - [ ] Integrate vector files with scan planning, file commits, and conflict handling. - [ ] Add Java, Python, and C++ Vortex compatibility tests. ### Initial scope The initial implementation can focus on Phase 1, providing a usable end-to-end vertical slice through schema representation, Arrow mapping, and regular Parquet reads and writes. Dedicated Vortex vector storage and Data Evolution can be delivered through follow-up pull requests under this issue. The following items are not required for the initial implementation: - ORC VECTOR support - Vector indexes or similarity search - Changing vector dimensions through schema evolution - VECTOR values inside shared-shredding MAP columns - Element types not supported by Apache Paimon Java PIP-40 should only be considered fully supported after all phases are complete. Completing Phase 1 means that regular Parquet VECTOR storage is supported, but does not imply support for dedicated Vortex vector files. ### Anything else? Related roadmap: [#186](https://github.com/apache/paimon-cpp/issues/186) ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
