lxy-9602 opened a new issue, #186:
URL: https://github.com/apache/paimon-cpp/issues/186

   ### Search before asking
   
   - [x] I searched in the 
[issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.
   
   
   ### Motivation
   
   Looking ahead, Paimon C++ will focus on several areas:
   
   - Storage and retrieval optimizations for vertical workloads, including 
efficient MAP storage, richer data types such as VECTOR, and complete index 
support.
   - Making newly written data queryable in real time through pluggable 
real-time writes and memory/disk union reads.
   - Following up on important capabilities from the broader Paimon community, 
including changelog production, Format Table, and extensions to Data Evolution.
   
   
   ### Solution
   
   #### 1. Storage and retrieval optimizations for vertical workloads
   
   - [ ] Support shared-shredding columnar storage for `MAP<STRING, T>` based 
on 
[PIP-43](https://cwiki.apache.org/confluence/spaces/PAIMON/pages/430408347/PIP-43+Columnar+Storage+Optimization+for+MAP+Type+in+Paimon),
 including adaptive physical columns, field mappings, overflow storage, and 
end-to-end reads and writes.
   - [ ] Support richer data types, starting with `VECTOR<T, N>` based on 
[PIP-40](https://cwiki.apache.org/confluence/spaces/PAIMON/pages/399279132/PIP-40+Introduce+a+new+Vector+data+type),
 including schema representation, storage, reads, writes, and Data Evolution.
   - [ ] Complete File Index support by adding index generation to the existing 
read path.
   - [ ] Continue improving storage layout, predicate pushdown, column pruning, 
point lookups, and index-assisted retrieval for workload-specific scenarios.
   
   #### 2. Making real-time data queryable
   
   - [ ] Support pluggable real-time writes and memory/disk union reads based 
on 
[PIP-46](https://cwiki.apache.org/confluence/spaces/PAIMON/pages/444334302/PIP-46+Support+pluggable+real-time+writes+and+memory+disk+union+reads+for+Paimon+C),
 tracked by [#158](https://github.com/apache/paimon-cpp/issues/158).
   - [ ] Make data held by an active writer queryable before it is committed 
into a snapshot.
   - [ ] Guarantee consistent query results without missing or duplicate rows 
while writing, committing, and reclaiming memory segments concurrently.
   - [ ] Support append tables first, followed by primary-key tables, 
deletion-vector mode, and additional production scenarios.
   
   #### 3. Following up on Paimon community capabilities
   
   - [ ] Support changelog production, including `input`, `lookup`, and 
`full-compaction` modes, tracked by 
[#174](https://github.com/apache/paimon-cpp/issues/174).
   - [ ] Support Format Table reads and writes for Hive-style file directories, 
tracked by [#170](https://github.com/apache/paimon-cpp/issues/170).
   - [ ] Extend Data Evolution to work with deletion vectors, tracked by 
[#169](https://github.com/apache/paimon-cpp/issues/169).
   - [ ] Extend Data Evolution to primary-key tables and support compaction 
across evolved field groups.
   - [ ] Continue following relevant Paimon features and make them available 
through native Paimon C++ APIs where appropriate.
   
   Each roadmap item should be delivered through focused issues and reviewable 
pull requests, with unit tests, integration tests, documentation, and 
compatibility validation where applicable.
   
   ### Anything else?
   
   _No response_
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to