Hi all,

Thanks, Junbo, Anton and Leonard, for summarizing and aligning the roadmap.

+1 to splitting the work into focused FIPs for Union Read, the DataFusion
Adapter, and Gateway. I think this will help keep the scope clear and make
the components easier to evolve independently.

I'm looking forward to the upcoming FIP discussions and am happy to help
review and contribute where I can.

Thanks!

Best,

Yangyang

Leonard Xu <[email protected]> 于2026年7月14日周二 12:02写道:

> Hi Anton and all,
>
>
>
> Thanks for the summary. This matches my understanding as well.
>
>
>
> +1 to splitting the roadmap into focused FIPs: Lake + Log Union Read,
>
> DataFusion Adapter, and Gateway as a separate follow-up topic. This keeps
> the
>
> scope clear and the modules independently reusable.
>
>
>
> It is also great to see that there are contributors interested in the Union
>
> Read work, DataFusion integration, and the Gateway direction. I think we
> can
>
> use Junbo's Union Read/lake kernel draft as the starting point, and let the
>
> Gateway discussion continue in parallel as a separate FIP when the scope is
>
> ready.
>
>
>
> I'm happy to help review and coordinate the related FIPs.
>
>
>
> Best,
>
> Leonard
>
>
>
> > 2026 7月 11 12:16 上午,Anton Borisov <[email protected]> 写道:
> >
> > Hi Leonard and all,
> >
> > Thanks, this matches my reading, and I don’t think anything material is
> missing.
> >
> > I agree we should split the roadmap into focused FIPs:
> >
> > Lake + Log Union Read -- Paimon first, but format-neutral at the API
> boundary.
> > DataFusion Adapter -- a separate integration layer, with Arrow and
> > DataFusion versions pinned consistently.
> > Gateway -- a separate follow-up FIP, given the broader REST/SQL and
> > auth/authz design.
> >
> > This keeps the lake kernel, DataFusion adapter, and gateway modular
> > and independently reusable.
> >
> > Predicate pushdown remains high-value stretch scope. First-class
> > StarRocks support should also be treated as a high-priority stretch
> > goal: we should prioritise the reviews and APIs needed to support it
> > well, without committing the integration itself as a roadmap
> > deliverable given the outlined dependencies and available capacity.
> >
> > With FIP-40, which merged fluss-rust into the main repository, now
> > landed, I suggest we move the Union Read and DataFusion Adapter work
> > into focused FIP discussions and agree their boundaries on dev@.
> >
> > Junbo has already shared an initial draft covering the lake kernel and
> > union-read direction [1].
> > I suggest we use it as the starting point and invite everyone
> > interested to review and refine the scope before moving it into a
> > formal FIP. Thanks, Junbo, for putting this together.
> >
> > The DataFusion proposal can proceed in parallel, with a clear boundary
> > between the two.
> > StarRocks can serve as an important validation consumer for the
> > underlying APIs, while Gateway continues as an independent design
> > track.
> >
> >
> > Thanks to everyone who contributed to the discussion and helped
> > converge the scope.
> > I’m happy to help coordinate the proposals and keep the proposals
> > aligned as they evolve.
> >
> > -- Anton
> >
> > [1]
> https://docs.google.com/document/d/1FE56807gsaYxiDMKAuqKavnvJrAmlzblDNnFNZM7msE/edit?tab=t.0#heading=h.2qxre6j4t17p
> >
> > ср, 8 июл. 2026 г. в 11:18, Leonard Xu <[email protected]>:
> >>
> >> Hi Anton and all,
> >>
> >> Thanks again for starting this discussion and for summarizing the
> roadmap
> >> direction so far.
> >>
> >> There have also been some follow-up discussions in Slack channel /
> community
> >> meeting. To avoid losing that context outside dev@, maybe it would be
> >> helpful to bring the latest consensus and open questions back to this
> >> thread.
> >>
> >> My rough reading is that the discussion is converging around:
> >>
> >> - DataFusion adapter and Lake + Log union read as the main must-have
> items.
> >> - Paimon first for union read.
> >> - Predicate pushdown as high-value but probably still stretch.
> >> - Gateway as a separate FIP/topic, especially considering auth/authz.
> >> - Keeping Arrow/DataFusion versions pinned consistently.
> >> - Keeping lake kernel / DataFusion adapter / gateway modular.
> >>
> >> Anton, since you kindly started this roadmap discussion, would you be
> willing to
> >> help summarize the latest consensus and suggest the next concrete step?
> >> One possible next step could be to split the roadmap into smaller FIPs,
> >> for example:
> >>
> >> - Lake + Log Union Read in fluss-rust
> >> - DataFusion Integration Adapter for Fluss
> >> - Gateway REST/SQL surface as a separate follow-up FIP
> >>
> >> WDYT? And please correct me if I missed or misrepresented anything.
> >>
> >> Best,
> >> Leonard
> >>
> >>> 2026 7月 3 6:22 下午,Anton Borisov <[email protected]> 写道:
> >>>
> >>> Hi Yangyang,
> >>>
> >>> Welcome, and great to hear you’re interested in contributing.
> >>>
> >>> StarRocks sounds like a very useful real-world consumer for this work.
> >>> Since you and Hongshun will be looking at the native connector side,
> >>> it would be helpful if you could share the main gaps you hit in the
> >>> current C++ bindings as you go: API shape, missing primitives,
> >>> performance constraints, or anything that feels awkward to integrate.
> >>>
> >>> That should give us a good feedback loop between the Fluss side and
> >>> the StarRocks integration without overfitting the Fluss API to one
> >>> consumer.
> >>> Happy to help review and discuss the API/design as it develops.
> >>>
> >>> -- Anton
> >>>
> >>> пт, 3 июл. 2026 г. в 09:58, Jim Hu <[email protected]>:
> >>>>
> >>>> Hi all,
> >>>>
> >>>>
> >>>> I'm Yangyang, and I'm relatively new to the Fluss community. Hongshun
> and I
> >>>> are planning to work on the union read support in the fluss-rust C++
> >>>> bindings and bring it into StarRocks as a native connector
> >>>> (StarRocks/starrocks#75785).
> >>>>
> >>>>
> >>>> Thanks Anton for the great advice! As a newcomer, I'd really
> appreciate any
> >>>> guidance or suggestions from the community.
> >>>>
> >>>>
> >>>> Looking forward to contributing!
> >>>>
> >>>>
> >>>> Best,
> >>>>
> >>>> Yangyang
> >>>>
> >>>> Anton Borisov <[email protected]> 于2026年7月3日周五 16:35写道:
> >>>>
> >>>>> Hi Hongshun,
> >>>>>
> >>>>> Thanks, this sounds like a very useful landing case.
> >>>>> Since StarRocks support relies on the same union-read / Paimon lake
> >>>>> path covered by Theme 2, it fits the direction well. From the
> >>>>> fluss-rust side, I think the important part is to define the common
> >>>>> union-read semantics clearly, and then make sure the Rust client and
> >>>>> C++ bindings expose the right primitives for integrations like
> >>>>> StarRocks to build on.
> >>>>>
> >>>>> So this looks like a good validation of the Theme 2 direction, and
> >>>>> also a useful reminder that the lake/log union read should not be
> >>>>> DataFusion-specific.
> >>>>>
> >>>>> -- Anton
> >>>>>
> >>>>> пт, 3 июл. 2026 г. в 09:10, Hongshun Wang <[email protected]>:
> >>>>>>
> >>>>>> Hi Anton,
> >>>>>> We will also support union read in StarRocks in Fluss C++ (
> >>>>>> https://github.com/StarRocks/starrocks/issues/75785), and Yangyang(
> >>>>>> https://github.com/naivedogger) and I will do it. This way, our
> Rust
> >>>>> client
> >>>>>> ecosystem can become more than just a demo.
> >>>>>>
> >>>>>> Best,
> >>>>>> Hongshun
> >>>>>>
> >>>>>> On Fri, Jul 3, 2026 at 3:42 PM Anton Borisov <[email protected]>
> >>>>> wrote:
> >>>>>>
> >>>>>>> Hi all,
> >>>>>>>
> >>>>>>> Thanks Forward, Hongshun and Junbo.
> >>>>>>>
> >>>>>>> On the release framing: I agree we should avoid making "0.3.0" the
> >>>>>>> public target name once fluss-rust is consolidated into the main
> Fluss
> >>>>>>> repo. It was a useful shorthand while fluss-rust had its own
> >>>>>>> versioning, but after FIP-40 it is probably clearer to talk about
> the
> >>>>>>> next post-consolidation Fluss release scope.
> >>>>>>> My reading is that consensus is forming around the following
> shape, so
> >>>>>>> let's define the roadmap like this for the time being:
> >>>>>>>
> >>>>>>> Theme 1  DataFusion integration adapter
> >>>>>>> This looks like the main must-have item. There seems to be
> agreement
> >>>>>>> that it should be a standalone adapter over the Rust core, not a
> >>>>>>> bundled engine, and that Arrow/DataFusion versions should be pinned
> >>>>>>> consistently across the workspace.
> >>>>>>>
> >>>>>>> Theme 2  Lake + log union read
> >>>>>>> This also looks like a must-have scope. Paimon-first seems to be
> the
> >>>>>>> pragmatic path, with the lake-side dependencies isolated in a
> separate
> >>>>>>> crate. We should be careful with the boundary semantics for PK
> tables,
> >>>>>>> especially how newer log records are applied over the lake
> snapshot.
> >>>>>>>
> >>>>>>> Theme 3  Predicate pushdown
> >>>>>>> There also seems to be agreement that predicate pushdown is
> >>>>>>> high-value, especially for both the DataFusion adapter and gateway
> >>>>>>> read paths. I would still keep it as stretch unless based on the
> >>>>>>> current capacity and amount of contributors. In the interim, we can
> >>>>>>> push predicates into the lake side where supported and apply the
> >>>>>>> remaining filters over the real-time log tail.
> >>>>>>>
> >>>>>>> Theme 4  Gateway
> >>>>>>> For Gateway, I think the discussion is converging toward treating
> it
> >>>>>>> as a separate  FIP from the core fluss-rust roadmap, especially if
> we
> >>>>>>> split the plain REST API and SQL surface.
> >>>>>>>
> >>>>>>> The auth/authz questions Hongshun raised are important and probably
> >>>>>>> belong in that Gateway design rather than being hidden inside the
> rust
> >>>>>>> roadmap. My initial take is that gateway-side authentication can be
> >>>>>>> independent, but authorization should stay aligned with Fluss
> >>>>>>> server-side ACL semantics. Whether the gateway uses a single
> service
> >>>>>>> identity or delegated user identities is a real design choice and
> >>>>>>> affects connection management, so it deserves explicit discussion
> in
> >>>>>>> the Gateway FIP.
> >>>>>>>
> >>>>>>> On the production-readiness gap Hongshun mentioned: I agree this is
> >>>>>>> important, but I would separate it from the analytical surface
> itself.
> >>>>>>> The goal is to close these gaps, while first bringing the core
> >>>>>>> capabilities into place. I would treat this as a separate
> >>>>>>> production-readiness track/umbrella issue as part of the
> >>>>>>> post-consolidation work. Once fluss-rust is consolidated into the
> main
> >>>>>>> repo, we should also test it regularly against the current Fluss
> >>>>>>> build, so compatibility and stability issues are caught earlier.
> >>>>>>>
> >>>>>>> For today’s community meeting, I added the fluss-rust roadmap
> topic. I
> >>>>>>> can summarize the dev@ discussion so far and keep it clearly
> framed as
> >>>>>>> discussion status, not a final decision. Junbo, it would be good if
> >>>>>>> you could also briefly cover the prototype /lake kernel /gateway
> parts
> >>>>>>> you explored in code. It will be useful to understand the scope and
> >>>>>>> the work we can integrate back.
> >>>>>>>
> >>>>>>> -- Anton
> >>>>>>>
> >>>>>>> пт, 3 июл. 2026 г. в 07:42, Junbo Wang <[email protected]>:
> >>>>>>>>
> >>>>>>>> Hi Hongshun,
> >>>>>>>>
> >>>>>>>> Good point — fluss-rust is currently at 0.1.0 with master tracking
> >>>>>>> 0.2.0, so "0.3.0" was just our shorthand here. Once FIP-40 lands
> and
> >>>>>>> fluss-rust moves into the fluss repo, it would naturally follow
> fluss
> >>>>>>> versioning — likely the release after 1.0. I'll update the framing
> to
> >>>>>>> reflect that.
> >>>>>>>>
> >>>>>>>>
> >>>>>>>> Best regards,
> >>>>>>>> Junbo Wang
> >>>>>>>>
> >>>>>>>>> 2026年7月3日 14:16,Hongshun Wang <[email protected]> 写道:
> >>>>>>>>>
> >>>>>>>>> Hi Junbo,
> >>>>>>>>> What's 0.3.0? Since fluss-rust will be merged into fluss repo,
> >>>>> maybe
> >>>>>>> fluss
> >>>>>>>>> release-1.0?
> >>>>>>>>>
> >>>>>>>>> Best,
> >>>>>>>>> Hongshun
> >>>>>>>>>
> >>>>>>>>> On Fri, Jul 3, 2026 at 11:27 AM Junbo Wang <[email protected]
> >
> >>>>>>> wrote:
> >>>>>>>>>
> >>>>>>>>>> Thanks Forward and Hongshun for the thoughtful input.
> >>>>>>>>>>
> >>>>>>>>>> On ForwardXu's question about the target release — I've been
> >>>>> thinking
> >>>>>>> a
> >>>>>>>>>> bit about the 0.3.0 scope, and would like to share my personal
> >>>>> take,
> >>>>>>> mostly
> >>>>>>>>>> as a starting point for discussion.
> >>>>>>>>>>
> >>>>>>>>>> A possible 0.3.0 scope
> >>>>>>>>>>
> >>>>>>>>>> FIP-40: consolidate fluss-rust into apache/fluss. thanks to
> Anton
> >>>>> for
> >>>>>>>>>> already opening the PR-3401 <
> >>>>>>> https://github.com/apache/fluss/pull/3401>
> >>>>>>>>>> DataFusion integration adapter (Theme 1), as a fluss-rust
> >>>>> submodule.
> >>>>>>>>>> Lake + log union read, Paimon-first (Theme 2), as a separate
> >>>>>>> submodule so
> >>>>>>>>>> lake-side dependencies stay isolated from the core. S
> >>>>>>>>>> erver-side filter pushdown (Theme 3) — would be great to have if
> >>>>>>> capacity
> >>>>>>>>>> allows, but I'd suggest keeping it optional rather than a hard
> >>>>>>> requirement
> >>>>>>>>>> for 0.3.0.
> >>>>>>>>>>
> >>>>>>>>>> Gateway — perhaps outside 0.3.0
> >>>>>>>>>>
> >>>>>>>>>> Once fluss-rust is consolidated into apache/fluss, it might make
> >>>>>>> sense for
> >>>>>>>>>> the gateway to live as its own module depending on fluss-rust,
> >>>>> rather
> >>>>>>> than
> >>>>>>>>>> being tied to the 0.3.0 release train. That would also give the
> >>>>>>> auth/authz
> >>>>>>>>>> questions Hongshun raised some room to be designed in a
> dedicated
> >>>>> FIP.
> >>>>>>>>>>
> >>>>>>>>>> Just my personal thinking — happy to adjust based on what others
> >>>>> feel
> >>>>>>> is
> >>>>>>>>>> right.
> >>>>>>>>>>
> >>>>>>>>>>
> >>>>>>>>>> Best regards,
> >>>>>>>>>> Junbo Wang
> >>>>>>>>>>
> >>>>>>>>>>> 2026年7月2日 11:47,Forward Xu <[email protected]> 写道:
> >>>>>>>>>>>
> >>>>>>>>>>> Hi Anton,
> >>>>>>>>>>>
> >>>>>>>>>>> Thanks for putting this together, and thanks to everyone who
> >>>>> landed
> >>>>>>> the
> >>>>>>>>>>> previous roadmap. Framing the next phase around "fluss-rust as
> a
> >>>>>>>>>>> first-class analytical query surface" makes a lot of sense to
> me
> >>>>> — it
> >>>>>>>>>>> builds naturally on the analytical primitives we just shipped.
> A
> >>>>> few
> >>>>>>>>>>> thoughts, roughly following your structure.
> >>>>>>>>>>>
> >>>>>>>>>>> *Theme 1 — DataFusion integration adapter [MUST-HAVE]* +1. I
> >>>>> strongly
> >>>>>>>>>> agree
> >>>>>>>>>>> it should be an *adapter over the Rust core*, not a bundled
> >>>>> engine.
> >>>>>>>>>> Keeping
> >>>>>>>>>>> fluss-datafusion as a standalone crate exposing TableProvider +
> >>>>>>>>>>> CatalogProvider/SchemaProvider keeps the dependency direction
> >>>>> clean
> >>>>>>> and
> >>>>>>>>>>> lets any DataFusion-based consumer (or the gateway) opt in.
> >>>>> Mapping
> >>>>>>>>>>> projection/filter/limit pushdown onto the existing access paths
> >>>>> (PK
> >>>>>>>>>>> equality → lookup, bucket-key prefix → prefix lookup, LIMIT →
> >>>>> bounded
> >>>>>>>>>> scan,
> >>>>>>>>>>> else log scan) is the right approach. One thing worth nailing
> >>>>> down
> >>>>>>> early
> >>>>>>>>>> is
> >>>>>>>>>>> the exactness contract we report back to DataFusion (
> >>>>>>>>>>> TableProviderFilterPushDown::Exact vs Inexact), since getting
> >>>>> that
> >>>>>>> wrong
> >>>>>>>>>>> silently drops or double-applies filters.
> >>>>>>>>>>>
> >>>>>>>>>>> *Theme 2 — Lake + log union read (Paimon-first) [MUST-HAVE]*
> +1,
> >>>>> and
> >>>>>>>>>>> Paimon-first is the right call — reusing paimon-rust by
> wrapping
> >>>>> an
> >>>>>>>>>>> existing reader is low-risk, and the
> >>>>> lake-snapshot-records-log-offset
> >>>>>>>>>>> design gives us a clean stitch point. Starting bounded-batch
> and
> >>>>>>>>>> deferring
> >>>>>>>>>>> streaming also de-risks correctness of the log-offset boundary
> >>>>>>> before we
> >>>>>>>>>>> add continuous reads. The main correctness edge I'd like to see
> >>>>>>> covered
> >>>>>>>>>> is
> >>>>>>>>>>> the boundary semantics for PK tables when applying newer log
> >>>>> records
> >>>>>>> over
> >>>>>>>>>>> the lake snapshot (dedup/ordering at the tiered offset).
> >>>>>>>>>>>
> >>>>>>>>>>> *Theme 3 — Foundations / server-side filter pushdown [STRETCH]*
> >>>>> Agree
> >>>>>>>>>> this
> >>>>>>>>>>> is a stretch, but it's the item with the highest leverage since
> >>>>> it
> >>>>>>> prunes
> >>>>>>>>>>> both lake and log scans and unblocks the gateway's filter
> >>>>> pushdown.
> >>>>>>>>>>> Cross-checking against the Java PredicateConverter semantics is
> >>>>>>> important
> >>>>>>>>>>> so both clients behave identically. If capacity allows, I'd
> lean
> >>>>>>> toward
> >>>>>>>>>>> pulling at least the simple value-predicate cases forward,
> since
> >>>>>>> Themes 1
> >>>>>>>>>>> and 4 both benefit.
> >>>>>>>>>>>
> >>>>>>>>>>> *Theme 4 — Gateway [STRETCH]* The "REST is the no-regret floor,
> >>>>> SQL
> >>>>>>> is
> >>>>>>>>>> the
> >>>>>>>>>>> ceiling deferred until Themes 1–2" framing is a good way to
> think
> >>>>>>> about
> >>>>>>>>>> it.
> >>>>>>>>>>> Keeping it a thin, stateless HTTP frontend over the existing
> >>>>> client
> >>>>>>>>>>> primitives (one endpoint per primitive) sounds right and keeps
> >>>>> scope
> >>>>>>>>>>> contained.
> >>>>>>>>>>>
> >>>>>>>>>>> *Open questions*
> >>>>>>>>>>>
> >>>>>>>>>>> - *Themes 1–2 must-have, rest stretch — reasonable?* Yes,
> >>>>> ordering
> >>>>>>>>>> looks
> >>>>>>>>>>> right to me. Nothing obviously missing.
> >>>>>>>>>>> - *Paimon first, Iceberg later?* Agree, provided the union-read
> >>>>>>> path is
> >>>>>>>>>>> written against a lake abstraction so Iceberg/Lance slot in the
> >>>>> same
> >>>>>>>>>> way
> >>>>>>>>>>> rather than requiring a rewrite.
> >>>>>>>>>>> - *Pin one arrow/DataFusion version across core,
> >>>>> union-read/lake,
> >>>>>>> and
> >>>>>>>>>>> DF?* Strong +1 — we should pin a single Arrow/DataFusion
> version
> >>>>>>>>>>> workspace-wide; version skew between these crates is a common
> >>>>> and
> >>>>>>>>>> painful
> >>>>>>>>>>> source of breakage.
> >>>>>>>>>>> - *Separate crates: lake kernel, DataFusion adapter, gateway?*
> >>>>> +1 to
> >>>>>>>>>>> separate crates. It matches the "adapter, not bundled engine"
> >>>>>>>>>> principle and
> >>>>>>>>>>> keeps the DataFusion dependency out of the core for consumers
> >>>>> that
> >>>>>>>>>> don't
> >>>>>>>>>>> need it.
> >>>>>>>>>>>
> >>>>>>>>>>> One meta-question: do we have a rough sense of the target
> release
> >>>>>>>>>> (0.3.0?)
> >>>>>>>>>>> these must-haves land in, so we can size the Theme 3/4 stretch
> >>>>> work
> >>>>>>>>>>> accordingly?
> >>>>>>>>>>>
> >>>>>>>>>>> Thanks again — happy to help on the DataFusion adapter side.
> >>>>>>>>>>>
> >>>>>>>>>>> Best,
> >>>>>>>>>>>
> >>>>>>>>>>> ForwardXu
> >>>>>>>>>>>
> >>>>>>>>>>> Junbo Wang <[email protected]> 于2026年7月1日周三 22:41写道:
> >>>>>>>>>>>
> >>>>>>>>>>>>> I'd also like us to focus on closing the gap with the Java
> >>>>> client
> >>>>>>> on
> >>>>>>>>>>>> predicate pushdown.
> >>>>>>>>>>>> Agreed — predicate pushdown is a great optimization, and the
> >>>>> read
> >>>>>>> path
> >>>>>>>>>>>> will benefit significantly from it.
> >>>>>>>>>>>>
> >>>>>>>>>>>>
> >>>>>>>>>>>>
> >>>>>>>>>>>> One thing that was missing in upstream paimon-rust was an
> >>>>>>>>>>>> object-store-safe existence check in the filesystem catalog.
> It
> >>>>>>> checked
> >>>>>>>>>>>> whether database/table directories “exist”, which works on
> local
> >>>>>>>>>>>> filesystems but can fail on S3/OSS because those directories
> are
> >>>>>>> just
> >>>>>>>>>>>> prefixes, not real objects. I forked it to fix that, so lake
> >>>>> reads
> >>>>>>> can
> >>>>>>>>>>>> reliably open Paimon tables from object storage.
> >>>>>>>>>>>>
> >>>>>>>>>>>> Thanks again, Anton, for putting this roadmap together! I
> think
> >>>>> we
> >>>>>>> could
> >>>>>>>>>>>> share it at the July 3rd community meeting. I've also been
> >>>>> working
> >>>>>>> on
> >>>>>>>>>> some
> >>>>>>>>>>>> designs around lake kernel read and the Fluss Gateway REST
> API —
> >>>>>>> happy
> >>>>>>>>>> to
> >>>>>>>>>>>> discuss those as well if time allows.
> >>>>>>>>>>>>
> >>>>>>>>>>>>
> >>>>>>>>>>>> Best regards,
> >>>>>>>>>>>> Junbo Wang
> >>>>>>>>>>>>
> >>>>>>>>>>>>> 2026年6月25日 07:23,Anton Borisov <[email protected]> 写道:
> >>>>>>>>>>>>>
> >>>>>>>>>>>>> Hi Yuxia and Junbo,
> >>>>>>>>>>>>>
> >>>>>>>>>>>>> Yuxia, on PK changelog reads across the bindings: I have
> >>>>> already
> >>>>>>>>>>>>> prepared record-mode CDC for Rust, Python and C++. It's
> small,
> >>>>> and
> >>>>>>>>>>>>> the PK union read in Theme 2 needs it anyway.
> >>>>>>>>>>>>>
> >>>>>>>>>>>>> Junbo, I looked through your prototype branch. It is clear
> >>>>> that a
> >>>>>>> lot
> >>>>>>>>>>>>> of Themes 1, 2 and 4 already exist there as working code and
> it
> >>>>>>> does
> >>>>>>>>>>>>> seem to answer two of the open questions: the three-way crate
> >>>>> split
> >>>>>>>>>>>>> (lake kernel / adapter / gateway), and pinning a single
> >>>>>>>>>> Arrow/DataFusion
> >>>>>>>>>>>>> version at the workspace level. +1 to both.
> >>>>>>>>>>>>>
> >>>>>>>>>>>>> Two separate FIPs with detailed design for REST/SQL gateway
> >>>>> parts -
> >>>>>>>>>>>> makes sense.
> >>>>>>>>>>>>>
> >>>>>>>>>>>>> One thing I noticed and want to discuss:
> >>>>>>>>>>>>> The "lake read" currently depends on a fork of paimon-rust,
> >>>>> can you
> >>>>>>>>>> share
> >>>>>>>>>>>>> what was missing in paimon-rust?
> >>>>>>>>>>>>>
> >>>>>>>>>>>>> I'd also like us to focus on closing the gap with the Java
> >>>>> client
> >>>>>>> on
> >>>>>>>>>>>>> predicate pushdown.
> >>>>>>>>>>>>> In the interim we can push predicates into the lake (Paimon
> >>>>> already
> >>>>>>>>>>>>> supports it) and apply them as a filter pass over the
> >>>>> real-time log
> >>>>>>>>>>>>> tail.
> >>>>>>>>>>>>>
> >>>>>>>>>>>>> Good to hear you want to drive this forward, given the
> >>>>> prototype,
> >>>>>>> that
> >>>>>>>>>>>>> makes sense to me. I am happy to review the design and
> >>>>> following
> >>>>>>> PRs.
> >>>>>>>>>>>>>
> >>>>>>>>>>>>> Looking forward to the write-up with your findings.
> >>>>>>>>>>>>>
> >>>>>>>>>>>>> -- Anton
> >>>>>>>>>>>>>
> >>>>>>>>>>>>> вт, 23 июн. 2026 г. в 15:10, Junbo Wang <
> [email protected]
> >>>>>> :
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>> Thanks Anton for kicking this off — funnily enough, I'd been
> >>>>>>> arriving
> >>>>>>>>>>>> at very similar conclusions while prototyping on my side, so
> >>>>> let me
> >>>>>>>>>> share a
> >>>>>>>>>>>> bit of what came out of those experiments.
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>>> THEME 1: DATAFUSION INTEGRATION ADAPTER [MUST-HAVE]
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>> Agree. Having the DataFusion integration live in fluss-rust
> >>>>> would
> >>>>>>> make
> >>>>>>>>>>>> it reusable across more downstream projects. One thing we
> might
> >>>>>>> want to
> >>>>>>>>>>>> keep in mind: it would probably be worth aligning the
> >>>>> DataFusion and
> >>>>>>>>>> Arrow
> >>>>>>>>>>>> versions with pg-datafusion so the whole stack stays on a
> >>>>> consistent
> >>>>>>>>>> set of
> >>>>>>>>>>>> versions.
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>>> THEME 2: LAKE + LOG UNION READ (PAIMON-FIRST) [MUST-HAVE]
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>> Agree. I'd lean toward putting this in a separate fluss-lake
> >>>>>>> module
> >>>>>>>>>>>> that exposes a client capable of doing the union read over the
> >>>>> lake
> >>>>>>> and
> >>>>>>>>>> the
> >>>>>>>>>>>> Fluss log. That way the lake-side dependencies stay isolated
> in
> >>>>>>> their
> >>>>>>>>>> own
> >>>>>>>>>>>> module and don't bleed into the core fluss-rust code. One
> thing
> >>>>> that
> >>>>>>>>>> might
> >>>>>>>>>>>> be worth flagging here: the lake side already supports
> predicate
> >>>>>>>>>> pushdown,
> >>>>>>>>>>>> while Fluss doesn't yet — so we'd probably think about how the
> >>>>> union
> >>>>>>>>>> read
> >>>>>>>>>>>> handles that asymmetry in the interim.
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>> THEME 4: GATEWAY [STRETCH] — A thin HTTP frontend over the
> >>>>> client
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>> I wonder if we could expand on this one a bit. My suggestion
> >>>>>>> would be
> >>>>>>>>>>>> to split the SQL REST surface and the plain REST API into two
> >>>>>>> separate
> >>>>>>>>>>>> FIPs. A read/write REST API on its own — somewhat in the
> spirit
> >>>>> of
> >>>>>>>>>>>> kafka-rest, covering reads, writes, and metadata management —
> >>>>> would
> >>>>>>>>>> already
> >>>>>>>>>>>> be a quick win for broadening Fluss's data ingestion/access
> >>>>> reach,
> >>>>>>> and
> >>>>>>>>>> it
> >>>>>>>>>>>> doesn't depend on anything else. The SQL REST surface, on the
> >>>>> other
> >>>>>>>>>> hand,
> >>>>>>>>>>>> has to wait on Themes 1–2, so it might make sense to split
> them
> >>>>> and
> >>>>>>> ship
> >>>>>>>>>>>> the plain REST API first.
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>> Also, it might make sense to have the gateway as a separate
> >>>>> crate
> >>>>>>>>>>>> inside the Fluss project itself — especially since fluss-rust
> >>>>> will
> >>>>>>> be
> >>>>>>>>>>>> moving into the Fluss repo anyway.
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>> On a personal note, this is something I'm really excited
> >>>>> about and
> >>>>>>>>>>>> would love to help push forward. I'll share a write-up of what
> >>>>> came
> >>>>>>> out
> >>>>>>>>>> of
> >>>>>>>>>>>> my demo experiments with the community over the next few
> weeks.
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>> Best regards,
> >>>>>>>>>>>>>> Junbo Wang
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>>>> 2026年6月23日 20:55,Yuxia Luo <[email protected]> 写道:
> >>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>> Thanks for putting this together. +1 on the overall
> framing.
> >>>>> My
> >>>>>>>>>>>> answers to the open questions, plus one addition:
> >>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>> Open questions:
> >>>>>>>>>>>>>>> - Themes 1-2 as must-have, rest stretch: agree, the
> ordering
> >>>>>>> makes
> >>>>>>>>>>>> sense to me.
> >>>>>>>>>>>>>>> - Paimon first, Iceberg later: +1. Reusing paimon-rust is
> the
> >>>>>>>>>>>> pragmatic path, and the union-read design generalizes to
> >>>>> Iceberg the
> >>>>>>>>>> same
> >>>>>>>>>>>> way later.
> >>>>>>>>>>>>>>> - Pin one arrow/DataFusion version across core,
> >>>>> union-read/lake,
> >>>>>>> and
> >>>>>>>>>>>> DF: yes, we should pin a single version. Letting these drift
> >>>>> will
> >>>>>>> cause
> >>>>>>>>>>>> painful Arrow type/ABI mismatches across the crate boundary,
> so
> >>>>> a
> >>>>>>> shared
> >>>>>>>>>>>> pinned version (bumped deliberately, not per-crate) is worth
> the
> >>>>>>>>>> discipline.
> >>>>>>>>>>>>>>> - Separate crates (lake kernel, DataFusion adapter,
> >>>>> gateway): +1
> >>>>>>> on
> >>>>>>>>>>>> separate crates. It keeps the dependency surface clean - the
> >>>>>>> adapter and
> >>>>>>>>>>>> gateway are optional consumers, and the lake kernel shouldn't
> >>>>> drag
> >>>>>>>>>>>> DataFusion into anyone who only needs the core read path.
> >>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>> One addition I'd like to propose for the scope:
> >>>>>>>>>>>>>>> PK-table CHANGELOG read, and expose it across the bindings
> -
> >>>>>>>>>>>> specifically C++ and Python - so non-Rust CDC consumers can
> >>>>>>> subscribe.
> >>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>> On 2026/06/22 14:49:52 Anton Borisov wrote:
> >>>>>>>>>>>>>>>> Hi all,
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> With the previous roadmap wrapping up - complex types
> >>>>>>>>>>>>>>>> (Array/Row/Map/nesting), limit and prefix scan,
> >>>>> schema-aware KV
> >>>>>>>>>>>>>>>> decoding, the metrics framework, and write optimisations,
> >>>>>>> thanks to
> >>>>>>>>>>>>>>>> everyone who contributed and reviewed, I'd like to open
> >>>>>>> discussion
> >>>>>>>>>> on
> >>>>>>>>>>>>>>>> the roadmap.
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> I propose framing around a single goal: make fluss-rust a
> >>>>>>>>>> first-class
> >>>>>>>>>>>>>>>> analytical query surface - a Rust-native SQL/DataFrame
> path
> >>>>> over
> >>>>>>>>>> Fluss
> >>>>>>>>>>>>>>>> (DataFusion, and through it Polars/DuckDB/gateways),
> >>>>> reading the
> >>>>>>>>>> lake
> >>>>>>>>>>>>>>>> tier at scale, built on the analytical primitives we just
> >>>>>>> landed.
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> As before, I grouped the items into themes with an initial
> >>>>>>>>>> must-have /
> >>>>>>>>>>>>>>>> stretch positioning. Please push back where you disagree.
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> THEME 1: DATAFUSION INTEGRATION ADAPTER [MUST-HAVE]
> >>>>>>>>>>>>>>>> A standalone fluss-datafusion crate exposing
> TableProvider +
> >>>>>>> Catalog
> >>>>>>>>>>>>>>>> over the existing client, so SELECT ... FROM . works from
> >>>>> any
> >>>>>>>>>>>>>>>> DataFusion-based engine. It should be framed as an
> >>>>> integration
> >>>>>>>>>> adapter
> >>>>>>>>>>>>>>>> over the Rust core, not a bundled engine. A gateway
> >>>>> (FIP-32) or
> >>>>>>> any
> >>>>>>>>>>>>>>>> analytical consumer can use it or not.
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> Scope: TableProvider + CatalogProvider/SchemaProvider
> >>>>> backed by
> >>>>>>> the
> >>>>>>>>>>>>>>>> metadata path.
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> Map DataFusion's projection/filter/limit pushdown onto the
> >>>>>>> client's
> >>>>>>>>>>>>>>>> existing access paths: a full primary-key equality
> becomes a
> >>>>>>>>>> lookup, a
> >>>>>>>>>>>>>>>> bucket-key prefix becomes a prefix lookup, LIMIT becomes a
> >>>>>>> bounded
> >>>>>>>>>>>>>>>> scan, otherwise a log scan. Filters an access path fully
> >>>>>>> satisfies
> >>>>>>>>>> are
> >>>>>>>>>>>>>>>> reported exact, the rest are left for DataFusion to apply
> >>>>>>>>>> (server-side
> >>>>>>>>>>>>>>>> filter pushdown is Theme 3).
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> THEME 2: LAKE + LOG UNION READ (PAIMON-FIRST) [MUST-HAVE]
> >>>>>>>>>>>>>>>> A client-side read that presents a lake-enabled table as
> one
> >>>>>>> table.
> >>>>>>>>>>>>>>>> Its history is tiered to the lake (Paimon) and recent
> >>>>> writes are
> >>>>>>>>>> still
> >>>>>>>>>>>>>>>> in the Fluss log, the lake snapshot records the log offset
> >>>>> it
> >>>>>>>>>> covers,
> >>>>>>>>>>>>>>>> so we read the lake up to that offset and the Fluss log
> >>>>> past it,
> >>>>>>>>>> then
> >>>>>>>>>>>>>>>> combine.
> >>>>>>>>>>>>>>>> Scope:
> >>>>>>>>>>>>>>>> - Per bucket: the lake snapshot plus the Fluss log past
> the
> >>>>>>> tiered
> >>>>>>>>>>>> offset.
> >>>>>>>>>>>>>>>> - Log tables: concatenate the two. PK tables: apply the
> >>>>> newer
> >>>>>>> log
> >>>>>>>>>>>>>>>> records to get current state.
> >>>>>>>>>>>>>>>> - Paimon first, reusing paimon-rust, Iceberg/Lance later
> the
> >>>>>>> same
> >>>>>>>>>>>>>>>> way. The DataFusion adapter reads lake-enabled tables
> >>>>> through
> >>>>>>> this
> >>>>>>>>>>>>>>>> path.
> >>>>>>>>>>>>>>>> - Bounded batch first, streaming later
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> Reusing paimon-rust means wrapping an existing reader, so
> >>>>> it's
> >>>>>>>>>> rather
> >>>>>>>>>>>>>>>> feasible and straightforward.
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> THEME 3: FOUNDATIONS [STRETCH]
> >>>>>>>>>>>>>>>> Server-side value-predicate (filter) pushdown. The proto
> >>>>> field
> >>>>>>>>>> exists,
> >>>>>>>>>>>>>>>> but the client currently sends none. This prunes lake/log
> >>>>> scans
> >>>>>>> and
> >>>>>>>>>>>>>>>> unblocks the gateway's filter pushdown (the open question
> I
> >>>>>>> raised
> >>>>>>>>>> for
> >>>>>>>>>>>>>>>> 0.2.0), cross-checked against the Java PredicateConverter
> >>>>>>> semantics.
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> THEME 4: GATEWAY [STRETCH]
> >>>>>>>>>>>>>>>> A thin HTTP frontend over the client we already have
> >>>>>>> independent of
> >>>>>>>>>>>>>>>> Themes 1–3. It's how non-Rust/non-SQL clients (TS/JS,
> >>>>>>> microservices,
> >>>>>>>>>>>>>>>> curl) write to Fluss and do simple key access.
> >>>>>>>>>>>>>>>> - Write: upsert/append, delete, create/drop tables,
> >>>>> metadata.
> >>>>>>> The
> >>>>>>>>>>>>>>>> mature client write path essentially.
> >>>>>>>>>>>>>>>> - Read: point lookup, prefix lookup, bounded log / CDC
> poll
> >>>>> -
> >>>>>>> one
> >>>>>>>>>>>>>>>> endpoint per existing primitive, bounded and stateless.
> >>>>>>>>>>>>>>>> The tradeoff: no joins, filters, or aggregations - that's
> >>>>> the
> >>>>>>> SQL
> >>>>>>>>>>>>>>>> surface, and it needs the adapter and probably will be an
> >>>>>>> ongoing
> >>>>>>>>>>>>>>>> effort to optimise later.
> >>>>>>>>>>>>>>>> - If Themes 1–2 ship, we unlock SQL frontends on top -
> >>>>>>> PostgreSQL
> >>>>>>>>>>>>>>>> via datafusion-postgres, as one example
> >>>>>>>>>>>>>>>> In short: REST is the no-regret floor, SQL is the ceiling
> >>>>>>> deferred
> >>>>>>>>>>>>>>>> until Themes 1–2.
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> PARALLEL TRACK: ELIXIR + BINDING EXPOSURE
> >>>>>>>>>>>>>>>> Continue Elixir parity and expose newly-landed primitives
> >>>>>>> (prefix
> >>>>>>>>>>>>>>>> lookup, etc.) across Python / C++ / Elixir.
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> GOVERNANCE TRACK
> >>>>>>>>>>>>>>>> FIP-40: Consolidate apache/fluss-rust into apache/fluss -
> >>>>> PR is
> >>>>>>>>>> ready,
> >>>>>>>>>>>>>>>> looks pretty good, waiting for a good moment to
> >>>>> rebase/merge.
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> OPEN QUESTIONS
> >>>>>>>>>>>>>>>> - Themes 1–2 must-have, the rest stretch. Reasonable?
> >>>>> Anything
> >>>>>>>>>>>>>>>> missing or mis-ordered?
> >>>>>>>>>>>>>>>> - Paimon first, Iceberg later?
> >>>>>>>>>>>>>>>> - Pin one arrow/DataFusion version across core,
> >>>>>>> union-read/lake, and
> >>>>>>>>>>>> DF?
> >>>>>>>>>>>>>>>> - Separate crates: lake kernel, DataFusion adapter,
> gateway?
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> Looking forward to feedback.
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>>> -- Anton
> >>>>>>>>>>>>>>>>
> >>>>>>>>>>>>>>
> >>>>>>>>>>>>
> >>>>>>>>>>>>
> >>>>>>>>>>
> >>>>>>>>>>
> >>>>>>>>
> >>>>>>>
> >>>>>
> >>
>
>

Reply via email to