Hi, Yang Zhang

Thanks for driving this FIP.+1 from my side — Spark Ecosystem will give users 
more tiering options.

>  PoC pull request: https://github.com/apache/fluss/pull/3099

Would it be possible to first split out the extraction of common tiering code 
(e.g., TieringSplitGenerator, TieringCoordinator, TieringCommitter) into a 
shared module as a separate PR?This would make the main PR easier to review.


Best regards,
Junbo Wang

> 2026年5月8日 14:36,yuxia <[email protected]> 写道:
> 
> Hi, Yang Zhang
> Thanks for driving this FIP. Left minor comments:
> 
> 1: New configuration options
> Do we really need introduce/expose these configuration options? According to 
> my experience, the deault value always be fine, and hardly users will want to 
> change it. So, I tend 
> to not expose/introduce them unless we have a strong reason to expose them. 
> 
> 2: About code reuse
> In the Module Structure section, several proposed classes (e.g., 
> TieringSplitGenerator, TieringCoordinator, TieringCommitter) already have 
> counterparts in the existing Flink tiering module (fluss-flink-common). 
> Rather than duplicating this logic into a new Spark-specific module, I'd 
> recommend extracting the engine-agnostic components into a shared module 
> (e.g.,fluss-lake-tiering-common or moving them down to fluss-client).
> 
> Best regards,
> Yuxia
> 
> ----- 原始邮件 -----
> 发件人: "Yang Zhang" <[email protected]>
> 收件人: "dev" <[email protected]>
> 发送时间: 星期三, 2026年 5 月 06日 下午 3:15:26
> 主题: [SPAM][DISCUSS] Proposal: Add Spark as Tiering Compute Engine
> 
> Hi everyone,
> 
> I'd like to propose adding Spark as a supported compute engine for Fluss's
> data tiering service.
> Below is the motivation, key design decisions, and future roadmap.
> 
> Motivation
> 
> Spark is the most widely adopted compute engine in data platforms,
> supported across all major cloud platforms (Databricks, EMR, Dataproc) and
> used extensively in production environments.
> Adding Spark tiering brings these key benefits:
> - Broader accessibility: Users with existing Spark infrastructure can
> enable tiering without introducing additional compute frameworks, lowering
> the barrier to adoption.
> - Natural fit: Tiering is inherently a periodic batch operation (heartbeat
> -> poll -> read -> write -> commit), which maps cleanly to Spark's
> RDD-based batch parallelism. Each tiering round processes a snapshot or log
> segment in parallel across buckets, making it a good match for Spark's
> execution model.
> - Ecosystem completeness: Fluss 0.9 already supports Spark catalog, read,
> and write. Fluss 1.0 plans Union Read for Spark. Adding tiering completes
> the full lifecycle from a single Spark cluster.
> 
> Key Highlights
> 
> - Driver-Loop Architecture: A long-running driver loop orchestrates
> heartbeat -> split generation -> RDD job -> commit -> report.
> - Compatibility: Pure additive feature, no changes to existing APIs or
> Flink tiering, all 3 lake plugins (Paimon, Iceberg, Lance) work unmodified.
> 
> Future Work
> 
> - Multi-table concurrent tiering: Currently one table per loop iteration;
> parallel tiering will improve throughput and reduce latency.
> - Spark DataSource V2 integration: Replace the low-level RDD pattern with
> Spark's DataSource V2 or custom operators for better scheduler integration,
> metrics, and fault tolerance.
> - Decouple from datalake format: A single tiering service should tier
> tables to different lake formats simultaneously, rather than being bound to
> one format per instance.
> 
> You can find the full proposal document here:
> https://cwiki.apache.org/confluence/display/FLUSS/FIP-42%3A+Add+Spark+as+Tiering+Compute+Engine
> 
> You can find the PoC pull request here:
> https://github.com/apache/fluss/pull/3099
> 
> I'd appreciate your feedback on the design direction. Looking forward to
> the discussion!
> 
> Best regards,
> Yang Zhang

Reply via email to