Hello,

I really like the "Fluss as embedding pipeline" direction.

The core pattern of ingestion -> Flink batching + embedding -> writing back
to Fluss -> Lance Tiering is already possible today, albeit without
modality awareness.
This PR on Lance Quickstart demonstrates exactly that (minus batching):
Flink reads from Fluss log table A, calls an embedding model (streaming),
then writes into Fluss log table B. We can update the PR to use batched
calls.

https://github.com/apache/fluss/pull/2716/changes#diff-e549694d35816df1240d10f4597fa20c6df050bd845998f248ce5f883782d93dR391-R398

That said, I think there's room to architect Fluss more intentionally for
this use case. Reading from a log table, generating embeddings, and writing
to another log table is something Kafka can be used for - the Lance tiering
is the differentiator, but it's not enough on its own.
One concrete improvement worth exploring: allowing post-append enrichment
on log tables (e.g. populating a row's vector column in-place). This would
significantly reduce data movement between tables.

+1 on collecting this in a document to help us brainstorm further - happy
to contribute to that.

Best regards
Keith

On Wed, Apr 8, 2026 at 6:55 AM Giannis Polyzos <[email protected]>
wrote:

> Exciting indeed 😄
>
> Actually, you are correct, that "multimodal ingestion" is currently a
> storage/transport capability (BYTES for content, ARRAY for pre-computed
> vectors) rather than type-system-level semantic awareness. A native VECTOR
> type would make the design intent explicit. I see paimon has a proposal for
> a Vector type, which might be something relevant to Fluss as well
>
> https://cwiki.apache.org/confluence/display/PAIMON/PIP-40%3A+Introduce+a+new+Vector+data+type
>
> On freshness: Fluss reduces every bottleneck except embedding model
> latency, which it cannot eliminate but can mitigate via micro-batch
> embedding in a Flink job. The hot log layer also enables immediate
> raw-content retrieval before vector indexes are updated, which has a
> standalone value for hybrid retrieval in context engineering.
>
> The lowest-complexity path forward could be "Fluss as embedding pipeline"
> pattern: ingest -> Flink batching + embedding UDF or AI Functions -> write
> vectors back to Fluss -> Lance (or even Paimon, assuming it goes down that
> direction) tiering. This consolidates what currently requires Kafka +
> external embedding service + vector DB into a single system. The missing
> piece is making this pattern well-documented and ergonomic (ideally with a
> VECTOR type).
>
> It might be worth collecting all this information in a document to better
> help us brainstorm, but I think this might be a good first approach.
>
> Best,
> Giannis
>
> On Wed, Apr 8, 2026 at 1:17 AM Keith Lee <[email protected]> wrote:
>
> > Hi Giannis, dev,
> >
> > Thank you for following up and your input. I agree in general on not
> > fixating on the details of technical implementation. Adding my
> observations
> > here.
> >
> > > Currently, Fluss supports the ingestion of multi-modal data and
> > tiering on the Lance format
> > > ingestion of multi-modal data
> > > fast serving so that it can be used for context engineering use case
> >
> > I am not certain that Fluss currently supports ingestion of multimodal
> > data. Or, at least, it is not aware of image and video on the type
> system /
> > metadata level. We do have to think about how ingestion and serving will
> > look like here if we decide to defer vector processing e.g. will a hot
> > layer for multi modal data be useful for context engineering if there’s
> no
> > vector query capability?
> >
> > The discussion we had left me thinking around the aspect of using Fluss
> as
> > hot layer on top of existing format / infra used for multi-modal context
> > engineering. Specifically, it’d be important for us to understand
> > 1. what is the de-facto average / worse case data freshness (vector index
> > freshness?) achievable with existing format / tools?
> > 2. will adding Fluss on top of existing format / infra actually help
> > improve data freshness (vector index freshness)? I imagine that vector
> > embedding might be the bottleneck (Fluss will need to call a model to get
> > vector embedding)
> > 3. Can Fluss help in a different way e.g. achieve similar data freshness
> at
> > lower complexity / cost? E.g. Fluss performing vector embedding by
> batching
> > and calling an embedding model (locally or cloud)
> >
> > Exciting discussions!
> >
> > Best regards
> > Keith
> >
> >
> >
> > On Mon, 6 Apr 2026 at 08:52, Giannis Polyzos <[email protected]>
> > wrote:
> >
> > > Hi devs,
> > >
> > > Following up on our discussions and Fluss direction on Vector data
> > > support, i
> > > wanted to leave here my two cents.
> > >
> > > I wanna start by saying that im trying to follow-up with a few
> companies
> > > that work with vectors - like Yelp and Booking to understand their use
> > > cases and ideally get some feedback from them to better help us shape
> > this
> > > direction. Currently, Fluss supports the ingestion of multi-modal data
> > and
> > > tiering on the Lance format.. Seems like Paimon will also invest
> towards
> > > that direction.
> > > So I think a good first step for Fluss in that direction would be to
> act
> > as
> > > a streaming storage layer that can support:
> > > 1. The ingestion of multi-modal data
> > > 2. Fast serving of that data so it can be used for context engineering
> > use
> > > cases
> > > 3. Continue its support and enhancement on paimon and Lance format -
> for
> > > example supporting the Primary Key table there.
> > >
> > > I think for now these would be some good first steps, considering there
> > is
> > > already ground work there, the Lance format seems to be getting some
> good
> > > community adoption.
> > > So my suggestion would be to use the above as guideliness and not spend
> > too
> > > much time now at processing vectors and defer that to integrations, for
> > > example a LanceDB integration and then as we collect more feedback
> > > re-iterate.
> > >
> > > Another thing that may be good to think about is how users can
> integrate
> > > existing unstructured data --- think legal documents that already live
> on
> > > S3 or other object storage -- and make fluss aware of them for serving
> > them
> > > again as part of some context engineering jobs.
> > > https://fluss.apache.org/blog/fluss-for-ai/
> > > I think that what we have in Fluss for AI is already a compelling story
> > and
> > > allow fluss to act as a centralized data repository for all types of
> > data,
> > > so lets focus on that as a first step.
> > >
> > > Let me know your thoughts, and if there are more suggestions and
> > proposal I
> > > would be eager to hear your thoughts.
> > >
> > > Best,
> > > Giannis
> > >
> > > On Mon, Mar 9, 2026 at 5:20 PM Lorenzo Affetti <
> > > [email protected]> wrote:
> > >
> > > > Thanks guys for the valuable feedback.
> > > >
> > > > I will put this on the table with Wangcheng and Giannis Polyzos (I
> know
> > > he
> > > > has quite a vision for the future of Fluss for AI:
> > > > https://fluss.apache.org/blog/fluss-for-ai/).
> > > > So that we can come up with a roadmap and put that under the
> discussion
> > > > thread on Github.
> > > >
> > > > Thrilled!
> > > >
> > > > On Mon, Mar 2, 2026 at 1:35 PM ForwardXu <[email protected]> wrote:
> > > >
> > > >> Hi all,
> > > >> I think it makes perfect sense to create a dedicated roadmap for
> Lance
> > > >> support. This will help us clarify our priorities and ensure we can
> > > deliver
> > > >> more comprehensive support, including advanced features like complex
> > > data
> > > >> types and blob types, among others.
> > > >> Looking forward to discussing this further on Slack.
> > > >>
> > > >> Best,
> > > >> Forwardxu
> > > >>
> > > >> 原始邮件
> > > >> ------------------------------
> > > >> 发件人:Lorenzo Affetti via dev <[email protected]>
> > > >> 发件时间:2026年3月2日 18:48
> > > >> 收件人:dev <[email protected]>
> > > >> 抄送:forwardxu <[email protected]>, Lorenzo Affetti <
> > > >> [email protected]>
> > > >> 主题:Re: Analysis of Lance storage format support
> > > >>
> > > >> Hello! Thanks for wrapping this up!
> > > >>
> > > >> I do understand both Cheng and Keith.
> > > >> For sure Lance support should be on par with other lake formats. If
> > > >> something is not supported, there should be a concrete reason why
> > (apart
> > > >> from a lack of resources :) ).
> > > >>
> > > >> Still, input from the Lance community would be essential for
> > > >> understanding evolution areas of the support itself.
> > > >>
> > > >> For this item, I would take an approach similar to what Mehul did
> for
> > > >> Iceberg support.
> > > >> I think there is a lack of a roadmap for Lance support in 2026.
> > > >>
> > > >> Having a roadmap doesn't actually mean we will accomplish
> everything,
> > > but,
> > > >> it signals that we understand the problem space and have an idea of
> > the
> > > >> sequence of actions to take.
> > > >>
> > > >> @cheng, I think you are the de-facto owner of the Lance module.
> > > >> Would it make sense to dedicate some of our resources to discuss
> this
> > > via
> > > >> Slack and start drafting a roadmap?
> > > >>
> > > >> On Sun, Mar 1, 2026 at 2:11 PM Keith Lee <
> [email protected]
> > >
> > > >> wrote:
> > > >>
> > > >> > Hello Cheng,
> > > >> >
> > > >> > Good call. I agree that gathering input from Lance community will
> be
> > > >>
> > > >> > beneficial to inform integration of features such as vector
> search,
> > > vector
> > > >> > indexing and hybrid search.
> > > >> >
> > > >>
> > > >> > However, the issues I’ve outlined only meant to cover the scope of
> > > bringing
> > > >> > current fluss lance integration up to parity to other lakehouses
> > like
> > > >> > paimon or iceberg e.g. batch or union read without lance feature
> > such
> > > as
> > > >>
> > > >> > vector search. As such, I believe these can be decoupled and we
> can
> > > have a
> > > >>
> > > >> > separate effort, gathering input from lance community and FIP
> > > proposal for
> > > >> > integrating vector search into feature such as union read.
> > > >> >
> > > >> > Let me know what your thoughts are on this. Thank you!
> > > >> >
> > > >> > Best regards
> > > >> > Keith Lee
> > > >> >
> > > >> >
> > > >> > On Sun, 1 Mar 2026 at 10:20, Cheng Wang <[email protected]> wrote:
> > > >> >
> > > >> > > Hello Keith,
> > > >> > >
> > > >> > >
> > > >>
> > > >> > > Regarding our plan to implement union read for Lance using
> Flink,
> > > might
> > > >> > it
> > > >> > > be beneficial to first gather input from the Lance community?
> > > >> > Understanding
> > > >>
> > > >> > > the primary scenarios where union read would help in the machine
> > > learning
> > > >> > > scenario, along with the most popular execution engine in Lance
> > > >> > ecosystem,
> > > >> > > could ensure we're building the right integration to maximize
> its
> > > >> > adoption.
> > > >> > >
> > > >> > >
> > > >> > >
> > > >> > >
> > > >> > > Regards,
> > > >> > > Cheng Wang
> > > >> > >
> > > >> > >
> > > >> > >
> > > >> > > &nbsp;
> > > >> > >
> > > >> > >
> > > >> > >
> > > >> > >
> > > >> > > ------------------&nbsp;Original&nbsp;------------------
> > > >> > > From:
> > > >> > >                                                   "dev"
> > > >> > >
>  <
> > > >> > > [email protected]&gt;;
> > > >> > > Date:&nbsp;Sat, Feb 28, 2026 11:20 PM
> > > >> > > To:&nbsp;"dev"<[email protected]&gt;;
> > > >> > > Cc:&nbsp;"Cheng Wang"<[email protected]&gt;;"forwardxu"<
> > > >> > > [email protected]&gt;;
> > > >> > > Subject:&nbsp;Re: Analysis of Lance storage format support
> > > >> > >
> > > >> > >
> > > >> > >
> > > >> > > This is extremely helpful, thanks for putting this together.
> > > >> > >
> > > >>
> > > >> > > Maybe we can create an umbrella ticket on GitHub to keep track
> on
> > > these
> > > >> > and
> > > >> > > open individual tasks, for tracking.
> > > >> > >
> > > >> > > Best,
> > > >> > > Giannis
> > > >> > >
> > > >> > > On Sat, 28 Feb 2026 at 3:52 PM, Keith Lee <
> > > >> [email protected]
> > > >> > &gt;
> > > >> > > wrote:
> > > >> > >
> > > >> > > &gt; Hello,
> > > >> > > &gt;
> > > >>
> > > >> > > &gt; As discussed on community sync yesterday on analysing where
> > we
> > > are
> > > >> > at
> > > >> > > the
> > > >> > > &gt; moment in terms of Lance format support.
> > > >> > > &gt; Here are my findings as part of working on Lance QuickStart
> > > >> > > documentation
> > > >> > > &gt; [1]. Lance lake tiering works in general, however there are
> > > some
> > > >> > gaps
> > > >> > > that
> > > >>
> > > >> > > &gt; to be addressed to bring Lance format support in parity
> with
> > > Paimon
> > > >> > /
> > > >> > > &gt; Iceberg.
> > > >> > > &gt;
> > > >>
> > > >> > > &gt; - (Merged) Support for Arrow FixedSizeList to enable
> pylance
> > > native
> > > >> > > vector
> > > >> > > &gt; search [2]
> > > >> > > &gt; - (In progress) Support Flink SQL Union Read query against
> > > Lance
> > > >> > > table [3]
> > > >> > > &gt; - (Open) Support Flink SQL batch query against Lance table
> > [4]
> > > >> > > &gt; - (Blocked) Primary Key table support - I believe this is
> > still
> > > >> > > blocking on
> > > >> > > &gt; Lance format support for delete API [5]
> > > >> > > &gt;
> > > >> > > &gt; Finally there is also a gap in the ability of performing
> > vector
> > > >> > > search on
> > > >> > > &gt; hot data / via union read. After discussion with Mehul,
> > native
> > > >> > vector
> > > >> > > &gt; indexing on hot data in Fluss would be a separate, bigger
> > > effort
> > > >> > that
> > > >> > > we
> > > >> > > &gt; can evolve towards if there's demand for it.
> > > >> > > &gt;
> > > >> > > &gt; Appreciate feedback here from Cheng, Forward and anyone
> else
> > > with
> > > >>
> > > >> > > &gt; familiarity around this area as I have only started dipping
> > my
> > > toes
> > > >> > > into
> > > >> > > &gt; Lance.
> > > >> > > &gt;
> > > >> > > &gt; *Additionally, if anyone wants to help contributing in this
> > > area,
> > > >> > > please
> > > >> > > &gt; reach out. *
> > > >> > > &gt;
> > > >> > > &gt; Best regards
> > > >> > > &gt; Keith Lee
> > > >> > > &gt;
> > > >> > > &gt; Reference
> > > >> > > &gt; [1] https://github.com/apache/fluss/pull/2716
> > > >> > > &gt; [2] https://github.com/apache/fluss/issues/2706
> > > >> > > &gt; [3] https://github.com/apache/fluss/issues/2715
> > > >> > > &gt; [4] https://github.com/apache/fluss/issues/2751
> > > >> > > &gt; [5] https://github.com/lance-format/lance/issues/3961
> > > >> > > &gt;
> > > >> >
> > > >>
> > > >>
> > > >> --
> > > >> Lorenzo Affetti
> > > >> Senior Software Engineer @ Flink Team
> > > >> Ververica <http://www.ververica.com>
> > > >>
> > > >>
> > > >>
> > > >
> > > > --
> > > > Lorenzo Affetti
> > > > Senior Software Engineer @ Flink Team
> > > > Ververica <http://www.ververica.com>
> > > >
> > >
> >
>

Reply via email to