Thanks Yang,

The proposal looks good to me in general. I'm fine to start a vote
when above comments are addressed.

Best,
Jark

On Fri, 15 May 2026 at 08:59, Junbo Wang <[email protected]> wrote:
>
> Hi, Yang Zhang
>
> Thanks for driving this FIP.+1 from my side — Spark Ecosystem will give users 
> more tiering options.
>
> >  PoC pull request: https://github.com/apache/fluss/pull/3099
>
> Would it be possible to first split out the extraction of common tiering code 
> (e.g., TieringSplitGenerator, TieringCoordinator, TieringCommitter) into a 
> shared module as a separate PR?This would make the main PR easier to review.
>
>
> Best regards,
> Junbo Wang
>
> > 2026年5月8日 14:36,yuxia <[email protected]> 写道:
> >
> > Hi, Yang Zhang
> > Thanks for driving this FIP. Left minor comments:
> >
> > 1: New configuration options
> > Do we really need introduce/expose these configuration options? According 
> > to my experience, the deault value always be fine, and hardly users will 
> > want to change it. So, I tend
> > to not expose/introduce them unless we have a strong reason to expose them.
> >
> > 2: About code reuse
> > In the Module Structure section, several proposed classes (e.g., 
> > TieringSplitGenerator, TieringCoordinator, TieringCommitter) already have 
> > counterparts in the existing Flink tiering module (fluss-flink-common). 
> > Rather than duplicating this logic into a new Spark-specific module, I'd 
> > recommend extracting the engine-agnostic components into a shared module 
> > (e.g.,fluss-lake-tiering-common or moving them down to fluss-client).
> >
> > Best regards,
> > Yuxia
> >
> > ----- 原始邮件 -----
> > 发件人: "Yang Zhang" <[email protected]>
> > 收件人: "dev" <[email protected]>
> > 发送时间: 星期三, 2026年 5 月 06日 下午 3:15:26
> > 主题: [SPAM][DISCUSS] Proposal: Add Spark as Tiering Compute Engine
> >
> > Hi everyone,
> >
> > I'd like to propose adding Spark as a supported compute engine for Fluss's
> > data tiering service.
> > Below is the motivation, key design decisions, and future roadmap.
> >
> > Motivation
> >
> > Spark is the most widely adopted compute engine in data platforms,
> > supported across all major cloud platforms (Databricks, EMR, Dataproc) and
> > used extensively in production environments.
> > Adding Spark tiering brings these key benefits:
> > - Broader accessibility: Users with existing Spark infrastructure can
> > enable tiering without introducing additional compute frameworks, lowering
> > the barrier to adoption.
> > - Natural fit: Tiering is inherently a periodic batch operation (heartbeat
> > -> poll -> read -> write -> commit), which maps cleanly to Spark's
> > RDD-based batch parallelism. Each tiering round processes a snapshot or log
> > segment in parallel across buckets, making it a good match for Spark's
> > execution model.
> > - Ecosystem completeness: Fluss 0.9 already supports Spark catalog, read,
> > and write. Fluss 1.0 plans Union Read for Spark. Adding tiering completes
> > the full lifecycle from a single Spark cluster.
> >
> > Key Highlights
> >
> > - Driver-Loop Architecture: A long-running driver loop orchestrates
> > heartbeat -> split generation -> RDD job -> commit -> report.
> > - Compatibility: Pure additive feature, no changes to existing APIs or
> > Flink tiering, all 3 lake plugins (Paimon, Iceberg, Lance) work unmodified.
> >
> > Future Work
> >
> > - Multi-table concurrent tiering: Currently one table per loop iteration;
> > parallel tiering will improve throughput and reduce latency.
> > - Spark DataSource V2 integration: Replace the low-level RDD pattern with
> > Spark's DataSource V2 or custom operators for better scheduler integration,
> > metrics, and fault tolerance.
> > - Decouple from datalake format: A single tiering service should tier
> > tables to different lake formats simultaneously, rather than being bound to
> > one format per instance.
> >
> > You can find the full proposal document here:
> > https://cwiki.apache.org/confluence/display/FLUSS/FIP-42%3A+Add+Spark+as+Tiering+Compute+Engine
> >
> > You can find the PoC pull request here:
> > https://github.com/apache/fluss/pull/3099
> >
> > I'd appreciate your feedback on the design direction. Looking forward to
> > the discussion!
> >
> > Best regards,
> > Yang Zhang
>

Reply via email to