Hi, Xinyu.

Thanks for sharing the details and current progress of PIP-46.

>From my understanding, PIP-39 and PIP-46 do not conflict and have little 
>direct overlap, because they address different layers and can be discussed 
>independently.

PIP-39 defines the table-level read and write semantics for lake-stream mode. 
The goal is to allow each engine to recognize when a Paimon table is operating 
in lake-stream mode and consistently follow the behavior defined by PIP-39. The 
Paimon connectors do not implement the Stream Store’s internal read or write 
paths. Instead, they discover the configured Stream Store and delegate the 
actual operations to its native source and sink implementations. Therefore, 
PIP-39 defines the engine-facing integration contract and observable table 
behavior, rather than the internal implementation of a Stream Store.

As I understand it, PIP-46 focuses on the implementation of native real-time 
capabilities inside Paimon and paimon-cpp, including real-time state 
management, split and commit message formats, offset tracking, and recovery. 
These are implementation details of the real-time read and write path and are 
outside the scope of PIP-39.

A PIP-46 RealtimeStore implementation could be one way for an engine to provide 
the real-time side of a lake-stream table. Similarly, a Fluss-backed paimon-cpp 
RealtimeStore could be a useful integration to explore. However, I think this 
can be discussed separately and should not be a prerequisite for PIP-39.

If a concrete integration later requires Paimon and a Stream Store to exchange 
splits, offsets, or commit messages, we can discuss and align on a common 
protocol at that point. For now, I believe PIP-39 and PIP-46 can proceed 
independently.

Best regards,
Yuxia

----- 原始邮件 -----
发件人: "刘欣瑀" <[email protected]>
收件人: "dev" <[email protected]>
发送时间: 星期一, 2026年 9 月 28日 下午 6:48:45
主题: Re:[DISCUSS] PIP-39: Introduce Lake-Stream Mode to Provide Second-Level 
Data Freshness for Paimon Tables

Hi Yuxia,

Thank you for proposing PIP-39 and for the detailed design. We believe 
improving data freshness is very important for Paimon, and the pluggable Stream 
Store approach provides a promising way to combine Paimon’s historical storage 
capabilities with a low-latency real-time data layer.

Previously, PIP-46 proposed a similar real-time framework that provides native, 
pluggable support for writing and querying real-time data inside Paimon. The 
real-time state is process-local and recovery currently relies on replaying 
data from the upstream system based on the durable offsets recorded in Paimon 
snapshots. A pluggable RealtimeStore interface allows customized storage 
implementations.

The initial implementation of this framework has been completed in paimon-cpp, 
and integration work with several query engines has already started.

Since PIP-39 and PIP-46 share several important concepts, we hope the two 
designs can reuse a common protocol where possible. In particular, it would be 
useful to align on:
- The representation and serialization of real-time splits and commit message 
(may with realtime progress). Per-partition and per-bucket offsets, including 
committed, tiered, and readable progress.
- The consistency boundary between a Paimon snapshot and the remaining 
real-time data.
- Schema, row-kind, and primary-key merge semantics across the lake and stream 
layers.

We would also like to explore whether Fluss could provide a paimon-cpp 
RealtimeStore plugin. Since paimon-cpp already provides some of the required 
interfaces and real-time read/write capabilities, we believe starting the 
integration with paimon-cpp could be a simpler and faster way to validate and 
deliver this architecture. Such an integration could also help the two proposal 
gradually align their data formats, split and offset protocols, and lake-stream 
consistency semantics. This would allow the existing paimon-cpp real-time 
framework to use Fluss as its durable real-time store, providing stronger 
support for task restart and failover while keeping the integration pluggable. 
We would be happy to collaborate on the common split and offset protocol, as 
well as the Fluss-based RealtimeStore integration.

Thanks again for the proposal. Looking forward to further discussion.

Best regards,
Xinyu Liu

At 2026-09-28 17:39:13, "yuxia" <[email protected]> wrote:
>Hi Paimon community, 
>
>I would like to start a discussion on PIP-39: Introduce Lake-Stream Mode to 
>Provide Second-Level Data Freshness for Paimon Tables. 
>
>This PIP proposes a pluggable Stream Store integration for Paimon, with Fluss 
>as the first implementation. Existing Paimon tables can enable lake-stream 
>mode in place: Paimon continues to manage historical data, while the Stream 
>Store provides second-level data freshness. 
>
>The proposal defines the read and write semantics, the enablement and 
>disablement procedures, and the corresponding Flink and Spark connector 
>changes. 
>
>Please find the proposal here: 
>https://cwiki.apache.org/confluence/spaces/PAIMON/pages/398000424/PIP-39+Introduce+Lake-Stream+Mode+to+Provide+Second-Level+Data+Freshness+for+Paimon+Tables
> 
>
>Feedback and suggestions are welcome. 
>
>Best regards, 
>Yuxia

Reply via email to