Hi Yuxia,

Thank you for proposing PIP-39 and for the detailed design. We believe 
improving data freshness is very important for Paimon, and the pluggable Stream 
Store approach provides a promising way to combine Paimon’s historical storage 
capabilities with a low-latency real-time data layer.

Previously, PIP-46 proposed a similar real-time framework that provides native, 
pluggable support for writing and querying real-time data inside Paimon. The 
real-time state is process-local and recovery currently relies on replaying 
data from the upstream system based on the durable offsets recorded in Paimon 
snapshots. A pluggable RealtimeStore interface allows customized storage 
implementations.

The initial implementation of this framework has been completed in paimon-cpp, 
and integration work with several query engines has already started.

Since PIP-39 and PIP-46 share several important concepts, we hope the two 
designs can reuse a common protocol where possible. In particular, it would be 
useful to align on:
- The representation and serialization of real-time splits and commit message 
(may with realtime progress). Per-partition and per-bucket offsets, including 
committed, tiered, and readable progress.
- The consistency boundary between a Paimon snapshot and the remaining 
real-time data.
- Schema, row-kind, and primary-key merge semantics across the lake and stream 
layers.

We would also like to explore whether Fluss could provide a paimon-cpp 
RealtimeStore plugin. Since paimon-cpp already provides some of the required 
interfaces and real-time read/write capabilities, we believe starting the 
integration with paimon-cpp could be a simpler and faster way to validate and 
deliver this architecture. Such an integration could also help the two proposal 
gradually align their data formats, split and offset protocols, and lake-stream 
consistency semantics. This would allow the existing paimon-cpp real-time 
framework to use Fluss as its durable real-time store, providing stronger 
support for task restart and failover while keeping the integration pluggable. 
We would be happy to collaborate on the common split and offset protocol, as 
well as the Fluss-based RealtimeStore integration.

Thanks again for the proposal. Looking forward to further discussion.

Best regards,
Xinyu Liu

At 2026-09-28 17:39:13, "yuxia" <[email protected]> wrote:
>Hi Paimon community, 
>
>I would like to start a discussion on PIP-39: Introduce Lake-Stream Mode to 
>Provide Second-Level Data Freshness for Paimon Tables. 
>
>This PIP proposes a pluggable Stream Store integration for Paimon, with Fluss 
>as the first implementation. Existing Paimon tables can enable lake-stream 
>mode in place: Paimon continues to manage historical data, while the Stream 
>Store provides second-level data freshness. 
>
>The proposal defines the read and write semantics, the enablement and 
>disablement procedures, and the corresponding Flink and Spark connector 
>changes. 
>
>Please find the proposal here: 
>https://cwiki.apache.org/confluence/spaces/PAIMON/pages/398000424/PIP-39+Introduce+Lake-Stream+Mode+to+Provide+Second-Level+Data+Freshness+for+Paimon+Tables
> 
>
>Feedback and suggestions are welcome. 
>
>Best regards, 
>Yuxia 

Reply via email to