Hi Anton,
We will also support union read in StarRocks in Fluss C++ (
https://github.com/StarRocks/starrocks/issues/75785), and Yangyang(
https://github.com/naivedogger) and I will do it. This way, our Rust client
ecosystem can become more than just a demo.

Best,
Hongshun

On Fri, Jul 3, 2026 at 3:42 PM Anton Borisov <[email protected]> wrote:

> Hi all,
>
> Thanks Forward, Hongshun and Junbo.
>
> On the release framing: I agree we should avoid making "0.3.0" the
> public target name once fluss-rust is consolidated into the main Fluss
> repo. It was a useful shorthand while fluss-rust had its own
> versioning, but after FIP-40 it is probably clearer to talk about the
> next post-consolidation Fluss release scope.
> My reading is that consensus is forming around the following shape, so
> let's define the roadmap like this for the time being:
>
> Theme 1  DataFusion integration adapter
> This looks like the main must-have item. There seems to be agreement
> that it should be a standalone adapter over the Rust core, not a
> bundled engine, and that Arrow/DataFusion versions should be pinned
> consistently across the workspace.
>
> Theme 2  Lake + log union read
> This also looks like a must-have scope. Paimon-first seems to be the
> pragmatic path, with the lake-side dependencies isolated in a separate
> crate. We should be careful with the boundary semantics for PK tables,
> especially how newer log records are applied over the lake snapshot.
>
> Theme 3  Predicate pushdown
> There also seems to be agreement that predicate pushdown is
> high-value, especially for both the DataFusion adapter and gateway
> read paths. I would still keep it as stretch unless based on the
> current capacity and amount of contributors. In the interim, we can
> push predicates into the lake side where supported and apply the
> remaining filters over the real-time log tail.
>
> Theme 4  Gateway
> For Gateway, I think the discussion is converging toward treating it
> as a separate  FIP from the core fluss-rust roadmap, especially if we
> split the plain REST API and SQL surface.
>
> The auth/authz questions Hongshun raised are important and probably
> belong in that Gateway design rather than being hidden inside the rust
> roadmap. My initial take is that gateway-side authentication can be
> independent, but authorization should stay aligned with Fluss
> server-side ACL semantics. Whether the gateway uses a single service
> identity or delegated user identities is a real design choice and
> affects connection management, so it deserves explicit discussion in
> the Gateway FIP.
>
> On the production-readiness gap Hongshun mentioned: I agree this is
> important, but I would separate it from the analytical surface itself.
> The goal is to close these gaps, while first bringing the core
> capabilities into place. I would treat this as a separate
> production-readiness track/umbrella issue as part of the
> post-consolidation work. Once fluss-rust is consolidated into the main
> repo, we should also test it regularly against the current Fluss
> build, so compatibility and stability issues are caught earlier.
>
> For today’s community meeting, I added the fluss-rust roadmap topic. I
> can summarize the dev@ discussion so far and keep it clearly framed as
> discussion status, not a final decision. Junbo, it would be good if
> you could also briefly cover the prototype /lake kernel /gateway parts
> you explored in code. It will be useful to understand the scope and
> the work we can integrate back.
>
> -- Anton
>
> пт, 3 июл. 2026 г. в 07:42, Junbo Wang <[email protected]>:
> >
> > Hi Hongshun,
> >
> > Good point — fluss-rust is currently at 0.1.0 with master tracking
> 0.2.0, so "0.3.0" was just our shorthand here. Once FIP-40 lands and
> fluss-rust moves into the fluss repo, it would naturally follow fluss
> versioning — likely the release after 1.0. I'll update the framing to
> reflect that.
> >
> >
> > Best regards,
> > Junbo Wang
> >
> > > 2026年7月3日 14:16,Hongshun Wang <[email protected]> 写道:
> > >
> > > Hi Junbo,
> > > What's 0.3.0? Since fluss-rust will be merged into fluss repo, maybe
> fluss
> > > release-1.0?
> > >
> > > Best,
> > > Hongshun
> > >
> > > On Fri, Jul 3, 2026 at 11:27 AM Junbo Wang <[email protected]>
> wrote:
> > >
> > >> Thanks Forward and Hongshun for the thoughtful input.
> > >>
> > >> On ForwardXu's question about the target release — I've been thinking
> a
> > >> bit about the 0.3.0 scope, and would like to share my personal take,
> mostly
> > >> as a starting point for discussion.
> > >>
> > >> A possible 0.3.0 scope
> > >>
> > >> FIP-40: consolidate fluss-rust into apache/fluss. thanks to Anton for
> > >> already opening the PR-3401 <
> https://github.com/apache/fluss/pull/3401>
> > >> DataFusion integration adapter (Theme 1), as a fluss-rust submodule.
> > >> Lake + log union read, Paimon-first (Theme 2), as a separate
> submodule so
> > >> lake-side dependencies stay isolated from the core. S
> > >> erver-side filter pushdown (Theme 3) — would be great to have if
> capacity
> > >> allows, but I'd suggest keeping it optional rather than a hard
> requirement
> > >> for 0.3.0.
> > >>
> > >> Gateway — perhaps outside 0.3.0
> > >>
> > >> Once fluss-rust is consolidated into apache/fluss, it might make
> sense for
> > >> the gateway to live as its own module depending on fluss-rust, rather
> than
> > >> being tied to the 0.3.0 release train. That would also give the
> auth/authz
> > >> questions Hongshun raised some room to be designed in a dedicated FIP.
> > >>
> > >> Just my personal thinking — happy to adjust based on what others feel
> is
> > >> right.
> > >>
> > >>
> > >> Best regards,
> > >> Junbo Wang
> > >>
> > >>> 2026年7月2日 11:47,Forward Xu <[email protected]> 写道:
> > >>>
> > >>> Hi Anton,
> > >>>
> > >>> Thanks for putting this together, and thanks to everyone who landed
> the
> > >>> previous roadmap. Framing the next phase around "fluss-rust as a
> > >>> first-class analytical query surface" makes a lot of sense to me — it
> > >>> builds naturally on the analytical primitives we just shipped. A few
> > >>> thoughts, roughly following your structure.
> > >>>
> > >>> *Theme 1 — DataFusion integration adapter [MUST-HAVE]* +1. I strongly
> > >> agree
> > >>> it should be an *adapter over the Rust core*, not a bundled engine.
> > >> Keeping
> > >>> fluss-datafusion as a standalone crate exposing TableProvider +
> > >>> CatalogProvider/SchemaProvider keeps the dependency direction clean
> and
> > >>> lets any DataFusion-based consumer (or the gateway) opt in. Mapping
> > >>> projection/filter/limit pushdown onto the existing access paths (PK
> > >>> equality → lookup, bucket-key prefix → prefix lookup, LIMIT → bounded
> > >> scan,
> > >>> else log scan) is the right approach. One thing worth nailing down
> early
> > >> is
> > >>> the exactness contract we report back to DataFusion (
> > >>> TableProviderFilterPushDown::Exact vs Inexact), since getting that
> wrong
> > >>> silently drops or double-applies filters.
> > >>>
> > >>> *Theme 2 — Lake + log union read (Paimon-first) [MUST-HAVE]* +1, and
> > >>> Paimon-first is the right call — reusing paimon-rust by wrapping an
> > >>> existing reader is low-risk, and the lake-snapshot-records-log-offset
> > >>> design gives us a clean stitch point. Starting bounded-batch and
> > >> deferring
> > >>> streaming also de-risks correctness of the log-offset boundary
> before we
> > >>> add continuous reads. The main correctness edge I'd like to see
> covered
> > >> is
> > >>> the boundary semantics for PK tables when applying newer log records
> over
> > >>> the lake snapshot (dedup/ordering at the tiered offset).
> > >>>
> > >>> *Theme 3 — Foundations / server-side filter pushdown [STRETCH]* Agree
> > >> this
> > >>> is a stretch, but it's the item with the highest leverage since it
> prunes
> > >>> both lake and log scans and unblocks the gateway's filter pushdown.
> > >>> Cross-checking against the Java PredicateConverter semantics is
> important
> > >>> so both clients behave identically. If capacity allows, I'd lean
> toward
> > >>> pulling at least the simple value-predicate cases forward, since
> Themes 1
> > >>> and 4 both benefit.
> > >>>
> > >>> *Theme 4 — Gateway [STRETCH]* The "REST is the no-regret floor, SQL
> is
> > >> the
> > >>> ceiling deferred until Themes 1–2" framing is a good way to think
> about
> > >> it.
> > >>> Keeping it a thin, stateless HTTP frontend over the existing client
> > >>> primitives (one endpoint per primitive) sounds right and keeps scope
> > >>> contained.
> > >>>
> > >>> *Open questions*
> > >>>
> > >>>  - *Themes 1–2 must-have, rest stretch — reasonable?* Yes, ordering
> > >> looks
> > >>>  right to me. Nothing obviously missing.
> > >>>  - *Paimon first, Iceberg later?* Agree, provided the union-read
> path is
> > >>>  written against a lake abstraction so Iceberg/Lance slot in the same
> > >> way
> > >>>  rather than requiring a rewrite.
> > >>>  - *Pin one arrow/DataFusion version across core, union-read/lake,
> and
> > >>>  DF?* Strong +1 — we should pin a single Arrow/DataFusion version
> > >>>  workspace-wide; version skew between these crates is a common and
> > >> painful
> > >>>  source of breakage.
> > >>>  - *Separate crates: lake kernel, DataFusion adapter, gateway?* +1 to
> > >>>  separate crates. It matches the "adapter, not bundled engine"
> > >> principle and
> > >>>  keeps the DataFusion dependency out of the core for consumers that
> > >> don't
> > >>>  need it.
> > >>>
> > >>> One meta-question: do we have a rough sense of the target release
> > >> (0.3.0?)
> > >>> these must-haves land in, so we can size the Theme 3/4 stretch work
> > >>> accordingly?
> > >>>
> > >>> Thanks again — happy to help on the DataFusion adapter side.
> > >>>
> > >>> Best,
> > >>>
> > >>> ForwardXu
> > >>>
> > >>> Junbo Wang <[email protected]> 于2026年7月1日周三 22:41写道:
> > >>>
> > >>>>> I'd also like us to focus on closing the gap with the Java client
> on
> > >>>> predicate pushdown.
> > >>>> Agreed — predicate pushdown is a great optimization, and the read
> path
> > >>>> will benefit significantly from it.
> > >>>>
> > >>>>
> > >>>>
> > >>>> One thing that was missing in upstream paimon-rust was an
> > >>>> object-store-safe existence check in the filesystem catalog. It
> checked
> > >>>> whether database/table directories “exist”, which works on local
> > >>>> filesystems but can fail on S3/OSS because those directories are
> just
> > >>>> prefixes, not real objects. I forked it to fix that, so lake reads
> can
> > >>>> reliably open Paimon tables from object storage.
> > >>>>
> > >>>> Thanks again, Anton, for putting this roadmap together! I think we
> could
> > >>>> share it at the July 3rd community meeting. I've also been working
> on
> > >> some
> > >>>> designs around lake kernel read and the Fluss Gateway REST API —
> happy
> > >> to
> > >>>> discuss those as well if time allows.
> > >>>>
> > >>>>
> > >>>> Best regards,
> > >>>> Junbo Wang
> > >>>>
> > >>>>> 2026年6月25日 07:23,Anton Borisov <[email protected]> 写道:
> > >>>>>
> > >>>>> Hi Yuxia and Junbo,
> > >>>>>
> > >>>>> Yuxia, on PK changelog reads across the bindings: I have already
> > >>>>> prepared record-mode CDC for Rust, Python and C++. It's small, and
> > >>>>> the PK union read in Theme 2 needs it anyway.
> > >>>>>
> > >>>>> Junbo, I looked through your prototype branch. It is clear that a
> lot
> > >>>>> of Themes 1, 2 and 4 already exist there as working code and it
> does
> > >>>>> seem to answer two of the open questions: the three-way crate split
> > >>>>> (lake kernel / adapter / gateway), and pinning a single
> > >> Arrow/DataFusion
> > >>>>> version at the workspace level. +1 to both.
> > >>>>>
> > >>>>> Two separate FIPs with detailed design for REST/SQL gateway parts -
> > >>>> makes sense.
> > >>>>>
> > >>>>> One thing I noticed and want to discuss:
> > >>>>> The "lake read" currently depends on a fork of paimon-rust, can you
> > >> share
> > >>>>> what was missing in paimon-rust?
> > >>>>>
> > >>>>> I'd also like us to focus on closing the gap with the Java client
> on
> > >>>>> predicate pushdown.
> > >>>>> In the interim we can push predicates into the lake (Paimon already
> > >>>>> supports it) and apply them as a filter pass over the real-time log
> > >>>>> tail.
> > >>>>>
> > >>>>> Good to hear you want to drive this forward, given the prototype,
> that
> > >>>>> makes sense to me. I am happy to review the design and following
> PRs.
> > >>>>>
> > >>>>> Looking forward to the write-up with your findings.
> > >>>>>
> > >>>>> -- Anton
> > >>>>>
> > >>>>> вт, 23 июн. 2026 г. в 15:10, Junbo Wang <[email protected]>:
> > >>>>>>
> > >>>>>> Thanks Anton for kicking this off — funnily enough, I'd been
> arriving
> > >>>> at very similar conclusions while prototyping on my side, so let me
> > >> share a
> > >>>> bit of what came out of those experiments.
> > >>>>>>
> > >>>>>>> THEME 1: DATAFUSION INTEGRATION ADAPTER [MUST-HAVE]
> > >>>>>>
> > >>>>>> Agree. Having the DataFusion integration live in fluss-rust would
> make
> > >>>> it reusable across more downstream projects. One thing we might
> want to
> > >>>> keep in mind: it would probably be worth aligning the DataFusion and
> > >> Arrow
> > >>>> versions with pg-datafusion so the whole stack stays on a consistent
> > >> set of
> > >>>> versions.
> > >>>>>>
> > >>>>>>> THEME 2: LAKE + LOG UNION READ (PAIMON-FIRST) [MUST-HAVE]
> > >>>>>>
> > >>>>>> Agree. I'd lean toward putting this in a separate fluss-lake
> module
> > >>>> that exposes a client capable of doing the union read over the lake
> and
> > >> the
> > >>>> Fluss log. That way the lake-side dependencies stay isolated in
> their
> > >> own
> > >>>> module and don't bleed into the core fluss-rust code. One thing that
> > >> might
> > >>>> be worth flagging here: the lake side already supports predicate
> > >> pushdown,
> > >>>> while Fluss doesn't yet — so we'd probably think about how the union
> > >> read
> > >>>> handles that asymmetry in the interim.
> > >>>>>>
> > >>>>>> THEME 4: GATEWAY [STRETCH] — A thin HTTP frontend over the client
> > >>>>>>
> > >>>>>> I wonder if we could expand on this one a bit. My suggestion
> would be
> > >>>> to split the SQL REST surface and the plain REST API into two
> separate
> > >>>> FIPs. A read/write REST API on its own — somewhat in the spirit of
> > >>>> kafka-rest, covering reads, writes, and metadata management — would
> > >> already
> > >>>> be a quick win for broadening Fluss's data ingestion/access reach,
> and
> > >> it
> > >>>> doesn't depend on anything else. The SQL REST surface, on the other
> > >> hand,
> > >>>> has to wait on Themes 1–2, so it might make sense to split them and
> ship
> > >>>> the plain REST API first.
> > >>>>>>
> > >>>>>>
> > >>>>>>
> > >>>>>> Also, it might make sense to have the gateway as a separate crate
> > >>>> inside the Fluss project itself — especially since fluss-rust will
> be
> > >>>> moving into the Fluss repo anyway.
> > >>>>>>
> > >>>>>> On a personal note, this is something I'm really excited about and
> > >>>> would love to help push forward. I'll share a write-up of what came
> out
> > >> of
> > >>>> my demo experiments with the community over the next few weeks.
> > >>>>>>
> > >>>>>>
> > >>>>>> Best regards,
> > >>>>>> Junbo Wang
> > >>>>>>
> > >>>>>>> 2026年6月23日 20:55,Yuxia Luo <[email protected]> 写道:
> > >>>>>>>
> > >>>>>>> Thanks for putting this together. +1 on the overall framing. My
> > >>>> answers to the open questions, plus one addition:
> > >>>>>>>
> > >>>>>>> Open questions:
> > >>>>>>> - Themes 1-2 as must-have, rest stretch: agree, the ordering
> makes
> > >>>> sense to me.
> > >>>>>>> - Paimon first, Iceberg later: +1. Reusing paimon-rust is the
> > >>>> pragmatic path, and the union-read design generalizes to Iceberg the
> > >> same
> > >>>> way later.
> > >>>>>>> - Pin one arrow/DataFusion version across core, union-read/lake,
> and
> > >>>> DF: yes, we should pin a single version. Letting these drift will
> cause
> > >>>> painful Arrow type/ABI mismatches across the crate boundary, so a
> shared
> > >>>> pinned version (bumped deliberately, not per-crate) is worth the
> > >> discipline.
> > >>>>>>> - Separate crates (lake kernel, DataFusion adapter, gateway): +1
> on
> > >>>> separate crates. It keeps the dependency surface clean - the
> adapter and
> > >>>> gateway are optional consumers, and the lake kernel shouldn't drag
> > >>>> DataFusion into anyone who only needs the core read path.
> > >>>>>>>
> > >>>>>>> One addition I'd like to propose for the scope:
> > >>>>>>> PK-table CHANGELOG read, and expose it across the bindings -
> > >>>> specifically C++ and Python - so non-Rust CDC consumers can
> subscribe.
> > >>>>>>>
> > >>>>>>> On 2026/06/22 14:49:52 Anton Borisov wrote:
> > >>>>>>>> Hi all,
> > >>>>>>>>
> > >>>>>>>> With the previous roadmap wrapping up - complex types
> > >>>>>>>> (Array/Row/Map/nesting), limit and prefix scan, schema-aware KV
> > >>>>>>>> decoding, the metrics framework, and write optimisations,
> thanks to
> > >>>>>>>> everyone who contributed and reviewed, I'd like to open
> discussion
> > >> on
> > >>>>>>>> the roadmap.
> > >>>>>>>>
> > >>>>>>>> I propose framing around a single goal: make fluss-rust a
> > >> first-class
> > >>>>>>>> analytical query surface - a Rust-native SQL/DataFrame path over
> > >> Fluss
> > >>>>>>>> (DataFusion, and through it Polars/DuckDB/gateways), reading the
> > >> lake
> > >>>>>>>> tier at scale, built on the analytical primitives we just
> landed.
> > >>>>>>>>
> > >>>>>>>> As before, I grouped the items into themes with an initial
> > >> must-have /
> > >>>>>>>> stretch positioning. Please push back where you disagree.
> > >>>>>>>>
> > >>>>>>>> THEME 1: DATAFUSION INTEGRATION ADAPTER [MUST-HAVE]
> > >>>>>>>> A standalone fluss-datafusion crate exposing TableProvider +
> Catalog
> > >>>>>>>> over the existing client, so SELECT ... FROM . works from any
> > >>>>>>>> DataFusion-based engine. It should be framed as an integration
> > >> adapter
> > >>>>>>>> over the Rust core, not a bundled engine. A gateway (FIP-32) or
> any
> > >>>>>>>> analytical consumer can use it or not.
> > >>>>>>>>
> > >>>>>>>> Scope: TableProvider + CatalogProvider/SchemaProvider backed by
> the
> > >>>>>>>> metadata path.
> > >>>>>>>>
> > >>>>>>>> Map DataFusion's projection/filter/limit pushdown onto the
> client's
> > >>>>>>>> existing access paths: a full primary-key equality becomes a
> > >> lookup, a
> > >>>>>>>> bucket-key prefix becomes a prefix lookup, LIMIT becomes a
> bounded
> > >>>>>>>> scan, otherwise a log scan. Filters an access path fully
> satisfies
> > >> are
> > >>>>>>>> reported exact, the rest are left for DataFusion to apply
> > >> (server-side
> > >>>>>>>> filter pushdown is Theme 3).
> > >>>>>>>>
> > >>>>>>>> THEME 2: LAKE + LOG UNION READ (PAIMON-FIRST) [MUST-HAVE]
> > >>>>>>>> A client-side read that presents a lake-enabled table as one
> table.
> > >>>>>>>> Its history is tiered to the lake (Paimon) and recent writes are
> > >> still
> > >>>>>>>> in the Fluss log, the lake snapshot records the log offset it
> > >> covers,
> > >>>>>>>> so we read the lake up to that offset and the Fluss log past it,
> > >> then
> > >>>>>>>> combine.
> > >>>>>>>> Scope:
> > >>>>>>>> - Per bucket: the lake snapshot plus the Fluss log past the
> tiered
> > >>>> offset.
> > >>>>>>>> - Log tables: concatenate the two. PK tables: apply the newer
> log
> > >>>>>>>> records to get current state.
> > >>>>>>>> - Paimon first, reusing paimon-rust, Iceberg/Lance later the
> same
> > >>>>>>>> way. The DataFusion adapter reads lake-enabled tables through
> this
> > >>>>>>>> path.
> > >>>>>>>> - Bounded batch first, streaming later
> > >>>>>>>>
> > >>>>>>>> Reusing paimon-rust means wrapping an existing reader, so it's
> > >> rather
> > >>>>>>>> feasible and straightforward.
> > >>>>>>>>
> > >>>>>>>> THEME 3: FOUNDATIONS [STRETCH]
> > >>>>>>>> Server-side value-predicate (filter) pushdown. The proto field
> > >> exists,
> > >>>>>>>> but the client currently sends none. This prunes lake/log scans
> and
> > >>>>>>>> unblocks the gateway's filter pushdown (the open question I
> raised
> > >> for
> > >>>>>>>> 0.2.0), cross-checked against the Java PredicateConverter
> semantics.
> > >>>>>>>>
> > >>>>>>>> THEME 4: GATEWAY [STRETCH]
> > >>>>>>>> A thin HTTP frontend over the client we already have
> independent of
> > >>>>>>>> Themes 1–3. It's how non-Rust/non-SQL clients (TS/JS,
> microservices,
> > >>>>>>>> curl) write to Fluss and do simple key access.
> > >>>>>>>> - Write: upsert/append, delete, create/drop tables, metadata.
> The
> > >>>>>>>> mature client write path essentially.
> > >>>>>>>> - Read: point lookup, prefix lookup, bounded log / CDC poll -
> one
> > >>>>>>>> endpoint per existing primitive, bounded and stateless.
> > >>>>>>>> The tradeoff: no joins, filters, or aggregations - that's the
> SQL
> > >>>>>>>> surface, and it needs the adapter and probably will be an
> ongoing
> > >>>>>>>> effort to optimise later.
> > >>>>>>>> - If Themes 1–2 ship, we unlock SQL frontends on top -
> PostgreSQL
> > >>>>>>>> via datafusion-postgres, as one example
> > >>>>>>>> In short: REST is the no-regret floor, SQL is the ceiling
> deferred
> > >>>>>>>> until Themes 1–2.
> > >>>>>>>>
> > >>>>>>>> PARALLEL TRACK: ELIXIR + BINDING EXPOSURE
> > >>>>>>>> Continue Elixir parity and expose newly-landed primitives
> (prefix
> > >>>>>>>> lookup, etc.) across Python / C++ / Elixir.
> > >>>>>>>>
> > >>>>>>>> GOVERNANCE TRACK
> > >>>>>>>> FIP-40: Consolidate apache/fluss-rust into apache/fluss - PR is
> > >> ready,
> > >>>>>>>> looks pretty good, waiting for a good moment to rebase/merge.
> > >>>>>>>>
> > >>>>>>>> OPEN QUESTIONS
> > >>>>>>>> - Themes 1–2 must-have, the rest stretch. Reasonable? Anything
> > >>>>>>>> missing or mis-ordered?
> > >>>>>>>> - Paimon first, Iceberg later?
> > >>>>>>>> - Pin one arrow/DataFusion version across core,
> union-read/lake, and
> > >>>> DF?
> > >>>>>>>> - Separate crates: lake kernel, DataFusion adapter, gateway?
> > >>>>>>>>
> > >>>>>>>> Looking forward to feedback.
> > >>>>>>>>
> > >>>>>>>> -- Anton
> > >>>>>>>>
> > >>>>>>
> > >>>>
> > >>>>
> > >>
> > >>
> >
>

Reply via email to