Hi Xinyu, Thanks for the proposal. Overall, the proposed direction aligns well with our use cases.
For our low-latency serving workloads, the higher-priority scenario is primary-key tables with merge-on-read semantics. Newly ingested records need to be queryable before the next snapshot commit and merged with committed disk data under Paimon’s existing primary-key semantics. We also have append-table use cases that require union reads across committed and real-time data. The proposed MemIndexer provides a suitable common framework, with append and primary-key semantics handled by their respective readers. I have one clarification about commit and recovery. Does the design assume only one in-flight PrepareCommit and snapshot commit per real-time writer? My understanding is that each PrepareCommit seals the current building segment and prepares the returned sealed segment. Snapshot progress advances only after the corresponding commit succeeds, and recovery then restores the next _OFFSET from the committed progress and replays uncommitted input. If multiple sealed segments can be prepared concurrently, are their commits serialized so that each partition-bucket’s committed _OFFSET advances only across a contiguous committed prefix? Best, Jianfeng On Tue, 28 Jul 2026 15:40:51 +0800 (CST), “刘欣瑀” [email protected] wrote: Hi everyone, I’d like to start a discussion on adding pluggable in-process real-time write and union-read support to Paimon C++. Problem Paimon uses committed snapshots as stable visibility boundaries. Records buffered by a writer cannot be queried through the standard reader until they are written to data files and committed into a snapshot. Some embedded and serving workloads require lower read latency: newly ingested records should be queryable before the next snapshot commit; ingestion should continue while PrepareCommit writes previously buffered records; queries should consistently combine committed disk data with uncommitted real-time data; append and primary-key tables should retain their existing Paimon semantics. Proposed Solution We propose an opt-in, process-local MemIndexer plugin integrated with the Paimon C++ write and read pipelines. The main ideas are: Paimon assigns an independent _OFFSET for each row within its partition-bucket. Each MemIndexer manages building and sealed segments. PrepareCommit seals the current segment and immediately opens a new one for continued ingestion. Sealed data is read from the plugin through commit readers and written by existing Paimon rolling writers. Per-partition-bucket _OFFSET progress is persisted through snapshot metadata. TableScan creates a RealtimeSplit containing disk splits and a pinned memory read view. Append tables concatenate disk and memory readers. Primary-key tables adapt memory readers into the existing Paimon merge pipeline. The plugin may use memory, local spill files, or specialized indexes internally, but Paimon continues to own table semantics, file formats, rolling policies, manifests, snapshots, and commit messages. The feature is disabled by default, so existing disk-only write and read paths are unaffected. Main Interfaces The proposal introduces several extension points: MemIndexer::Write MemIndexer::SealForCommit MemIndexer::AcquireReadView MemIndexer::CreateQueryReaders MemIndexer::CreateCommitReaders MemIndexer::Reclaim RealtimeContext internal RealtimeSplit planning PIP Document The full proposal, including lifecycle, public interfaces, append union read, and primary-key merge integration, is available here: https://cwiki.apache.org/confluence/spaces/PAIMON/pages/444334302/PIP-46+Support+pluggable+real-time+writes+and+memory+disk+union+reads+for+Paimon+C Looking for Feedback I’d appreciate community feedback on: The overall MemIndexer plugin boundary. The building/sealed segment lifecycle during PrepareCommit. The _OFFSET assignment and snapshot progress model. The integration with existing primary-key merge semantics. Any compatibility or implementation concerns. Looking forward to the discussion! Best regards, Xinyu
