Hello, I really like the "Fluss as embedding pipeline" direction.
The core pattern of ingestion -> Flink batching + embedding -> writing back to Fluss -> Lance Tiering is already possible today, albeit without modality awareness. This PR on Lance Quickstart demonstrates exactly that (minus batching): Flink reads from Fluss log table A, calls an embedding model (streaming), then writes into Fluss log table B. We can update the PR to use batched calls. https://github.com/apache/fluss/pull/2716/changes#diff-e549694d35816df1240d10f4597fa20c6df050bd845998f248ce5f883782d93dR391-R398 That said, I think there's room to architect Fluss more intentionally for this use case. Reading from a log table, generating embeddings, and writing to another log table is something Kafka can be used for - the Lance tiering is the differentiator, but it's not enough on its own. One concrete improvement worth exploring: allowing post-append enrichment on log tables (e.g. populating a row's vector column in-place). This would significantly reduce data movement between tables. +1 on collecting this in a document to help us brainstorm further - happy to contribute to that. Best regards Keith On Wed, Apr 8, 2026 at 6:55 AM Giannis Polyzos <[email protected]> wrote: > Exciting indeed 😄 > > Actually, you are correct, that "multimodal ingestion" is currently a > storage/transport capability (BYTES for content, ARRAY for pre-computed > vectors) rather than type-system-level semantic awareness. A native VECTOR > type would make the design intent explicit. I see paimon has a proposal for > a Vector type, which might be something relevant to Fluss as well > > https://cwiki.apache.org/confluence/display/PAIMON/PIP-40%3A+Introduce+a+new+Vector+data+type > > On freshness: Fluss reduces every bottleneck except embedding model > latency, which it cannot eliminate but can mitigate via micro-batch > embedding in a Flink job. The hot log layer also enables immediate > raw-content retrieval before vector indexes are updated, which has a > standalone value for hybrid retrieval in context engineering. > > The lowest-complexity path forward could be "Fluss as embedding pipeline" > pattern: ingest -> Flink batching + embedding UDF or AI Functions -> write > vectors back to Fluss -> Lance (or even Paimon, assuming it goes down that > direction) tiering. This consolidates what currently requires Kafka + > external embedding service + vector DB into a single system. The missing > piece is making this pattern well-documented and ergonomic (ideally with a > VECTOR type). > > It might be worth collecting all this information in a document to better > help us brainstorm, but I think this might be a good first approach. > > Best, > Giannis > > On Wed, Apr 8, 2026 at 1:17 AM Keith Lee <[email protected]> wrote: > > > Hi Giannis, dev, > > > > Thank you for following up and your input. I agree in general on not > > fixating on the details of technical implementation. Adding my > observations > > here. > > > > > Currently, Fluss supports the ingestion of multi-modal data and > > tiering on the Lance format > > > ingestion of multi-modal data > > > fast serving so that it can be used for context engineering use case > > > > I am not certain that Fluss currently supports ingestion of multimodal > > data. Or, at least, it is not aware of image and video on the type > system / > > metadata level. We do have to think about how ingestion and serving will > > look like here if we decide to defer vector processing e.g. will a hot > > layer for multi modal data be useful for context engineering if there’s > no > > vector query capability? > > > > The discussion we had left me thinking around the aspect of using Fluss > as > > hot layer on top of existing format / infra used for multi-modal context > > engineering. Specifically, it’d be important for us to understand > > 1. what is the de-facto average / worse case data freshness (vector index > > freshness?) achievable with existing format / tools? > > 2. will adding Fluss on top of existing format / infra actually help > > improve data freshness (vector index freshness)? I imagine that vector > > embedding might be the bottleneck (Fluss will need to call a model to get > > vector embedding) > > 3. Can Fluss help in a different way e.g. achieve similar data freshness > at > > lower complexity / cost? E.g. Fluss performing vector embedding by > batching > > and calling an embedding model (locally or cloud) > > > > Exciting discussions! > > > > Best regards > > Keith > > > > > > > > On Mon, 6 Apr 2026 at 08:52, Giannis Polyzos <[email protected]> > > wrote: > > > > > Hi devs, > > > > > > Following up on our discussions and Fluss direction on Vector data > > > support, i > > > wanted to leave here my two cents. > > > > > > I wanna start by saying that im trying to follow-up with a few > companies > > > that work with vectors - like Yelp and Booking to understand their use > > > cases and ideally get some feedback from them to better help us shape > > this > > > direction. Currently, Fluss supports the ingestion of multi-modal data > > and > > > tiering on the Lance format.. Seems like Paimon will also invest > towards > > > that direction. > > > So I think a good first step for Fluss in that direction would be to > act > > as > > > a streaming storage layer that can support: > > > 1. The ingestion of multi-modal data > > > 2. Fast serving of that data so it can be used for context engineering > > use > > > cases > > > 3. Continue its support and enhancement on paimon and Lance format - > for > > > example supporting the Primary Key table there. > > > > > > I think for now these would be some good first steps, considering there > > is > > > already ground work there, the Lance format seems to be getting some > good > > > community adoption. > > > So my suggestion would be to use the above as guideliness and not spend > > too > > > much time now at processing vectors and defer that to integrations, for > > > example a LanceDB integration and then as we collect more feedback > > > re-iterate. > > > > > > Another thing that may be good to think about is how users can > integrate > > > existing unstructured data --- think legal documents that already live > on > > > S3 or other object storage -- and make fluss aware of them for serving > > them > > > again as part of some context engineering jobs. > > > https://fluss.apache.org/blog/fluss-for-ai/ > > > I think that what we have in Fluss for AI is already a compelling story > > and > > > allow fluss to act as a centralized data repository for all types of > > data, > > > so lets focus on that as a first step. > > > > > > Let me know your thoughts, and if there are more suggestions and > > proposal I > > > would be eager to hear your thoughts. > > > > > > Best, > > > Giannis > > > > > > On Mon, Mar 9, 2026 at 5:20 PM Lorenzo Affetti < > > > [email protected]> wrote: > > > > > > > Thanks guys for the valuable feedback. > > > > > > > > I will put this on the table with Wangcheng and Giannis Polyzos (I > know > > > he > > > > has quite a vision for the future of Fluss for AI: > > > > https://fluss.apache.org/blog/fluss-for-ai/). > > > > So that we can come up with a roadmap and put that under the > discussion > > > > thread on Github. > > > > > > > > Thrilled! > > > > > > > > On Mon, Mar 2, 2026 at 1:35 PM ForwardXu <[email protected]> wrote: > > > > > > > >> Hi all, > > > >> I think it makes perfect sense to create a dedicated roadmap for > Lance > > > >> support. This will help us clarify our priorities and ensure we can > > > deliver > > > >> more comprehensive support, including advanced features like complex > > > data > > > >> types and blob types, among others. > > > >> Looking forward to discussing this further on Slack. > > > >> > > > >> Best, > > > >> Forwardxu > > > >> > > > >> 原始邮件 > > > >> ------------------------------ > > > >> 发件人:Lorenzo Affetti via dev <[email protected]> > > > >> 发件时间:2026年3月2日 18:48 > > > >> 收件人:dev <[email protected]> > > > >> 抄送:forwardxu <[email protected]>, Lorenzo Affetti < > > > >> [email protected]> > > > >> 主题:Re: Analysis of Lance storage format support > > > >> > > > >> Hello! Thanks for wrapping this up! > > > >> > > > >> I do understand both Cheng and Keith. > > > >> For sure Lance support should be on par with other lake formats. If > > > >> something is not supported, there should be a concrete reason why > > (apart > > > >> from a lack of resources :) ). > > > >> > > > >> Still, input from the Lance community would be essential for > > > >> understanding evolution areas of the support itself. > > > >> > > > >> For this item, I would take an approach similar to what Mehul did > for > > > >> Iceberg support. > > > >> I think there is a lack of a roadmap for Lance support in 2026. > > > >> > > > >> Having a roadmap doesn't actually mean we will accomplish > everything, > > > but, > > > >> it signals that we understand the problem space and have an idea of > > the > > > >> sequence of actions to take. > > > >> > > > >> @cheng, I think you are the de-facto owner of the Lance module. > > > >> Would it make sense to dedicate some of our resources to discuss > this > > > via > > > >> Slack and start drafting a roadmap? > > > >> > > > >> On Sun, Mar 1, 2026 at 2:11 PM Keith Lee < > [email protected] > > > > > > >> wrote: > > > >> > > > >> > Hello Cheng, > > > >> > > > > >> > Good call. I agree that gathering input from Lance community will > be > > > >> > > > >> > beneficial to inform integration of features such as vector > search, > > > vector > > > >> > indexing and hybrid search. > > > >> > > > > >> > > > >> > However, the issues I’ve outlined only meant to cover the scope of > > > bringing > > > >> > current fluss lance integration up to parity to other lakehouses > > like > > > >> > paimon or iceberg e.g. batch or union read without lance feature > > such > > > as > > > >> > > > >> > vector search. As such, I believe these can be decoupled and we > can > > > have a > > > >> > > > >> > separate effort, gathering input from lance community and FIP > > > proposal for > > > >> > integrating vector search into feature such as union read. > > > >> > > > > >> > Let me know what your thoughts are on this. Thank you! > > > >> > > > > >> > Best regards > > > >> > Keith Lee > > > >> > > > > >> > > > > >> > On Sun, 1 Mar 2026 at 10:20, Cheng Wang <[email protected]> wrote: > > > >> > > > > >> > > Hello Keith, > > > >> > > > > > >> > > > > > >> > > > >> > > Regarding our plan to implement union read for Lance using > Flink, > > > might > > > >> > it > > > >> > > be beneficial to first gather input from the Lance community? > > > >> > Understanding > > > >> > > > >> > > the primary scenarios where union read would help in the machine > > > learning > > > >> > > scenario, along with the most popular execution engine in Lance > > > >> > ecosystem, > > > >> > > could ensure we're building the right integration to maximize > its > > > >> > adoption. > > > >> > > > > > >> > > > > > >> > > > > > >> > > > > > >> > > Regards, > > > >> > > Cheng Wang > > > >> > > > > > >> > > > > > >> > > > > > >> > > > > > >> > > > > > >> > > > > > >> > > > > > >> > > > > > >> > > ------------------ Original ------------------ > > > >> > > From: > > > >> > > "dev" > > > >> > > > < > > > >> > > [email protected]>; > > > >> > > Date: Sat, Feb 28, 2026 11:20 PM > > > >> > > To: "dev"<[email protected]>; > > > >> > > Cc: "Cheng Wang"<[email protected]>;"forwardxu"< > > > >> > > [email protected]>; > > > >> > > Subject: Re: Analysis of Lance storage format support > > > >> > > > > > >> > > > > > >> > > > > > >> > > This is extremely helpful, thanks for putting this together. > > > >> > > > > > >> > > > >> > > Maybe we can create an umbrella ticket on GitHub to keep track > on > > > these > > > >> > and > > > >> > > open individual tasks, for tracking. > > > >> > > > > > >> > > Best, > > > >> > > Giannis > > > >> > > > > > >> > > On Sat, 28 Feb 2026 at 3:52 PM, Keith Lee < > > > >> [email protected] > > > >> > > > > > >> > > wrote: > > > >> > > > > > >> > > > Hello, > > > >> > > > > > > >> > > > >> > > > As discussed on community sync yesterday on analysing where > > we > > > are > > > >> > at > > > >> > > the > > > >> > > > moment in terms of Lance format support. > > > >> > > > Here are my findings as part of working on Lance QuickStart > > > >> > > documentation > > > >> > > > [1]. Lance lake tiering works in general, however there are > > > some > > > >> > gaps > > > >> > > that > > > >> > > > >> > > > to be addressed to bring Lance format support in parity > with > > > Paimon > > > >> > / > > > >> > > > Iceberg. > > > >> > > > > > > >> > > > >> > > > - (Merged) Support for Arrow FixedSizeList to enable > pylance > > > native > > > >> > > vector > > > >> > > > search [2] > > > >> > > > - (In progress) Support Flink SQL Union Read query against > > > Lance > > > >> > > table [3] > > > >> > > > - (Open) Support Flink SQL batch query against Lance table > > [4] > > > >> > > > - (Blocked) Primary Key table support - I believe this is > > still > > > >> > > blocking on > > > >> > > > Lance format support for delete API [5] > > > >> > > > > > > >> > > > Finally there is also a gap in the ability of performing > > vector > > > >> > > search on > > > >> > > > hot data / via union read. After discussion with Mehul, > > native > > > >> > vector > > > >> > > > indexing on hot data in Fluss would be a separate, bigger > > > effort > > > >> > that > > > >> > > we > > > >> > > > can evolve towards if there's demand for it. > > > >> > > > > > > >> > > > Appreciate feedback here from Cheng, Forward and anyone > else > > > with > > > >> > > > >> > > > familiarity around this area as I have only started dipping > > my > > > toes > > > >> > > into > > > >> > > > Lance. > > > >> > > > > > > >> > > > *Additionally, if anyone wants to help contributing in this > > > area, > > > >> > > please > > > >> > > > reach out. * > > > >> > > > > > > >> > > > Best regards > > > >> > > > Keith Lee > > > >> > > > > > > >> > > > Reference > > > >> > > > [1] https://github.com/apache/fluss/pull/2716 > > > >> > > > [2] https://github.com/apache/fluss/issues/2706 > > > >> > > > [3] https://github.com/apache/fluss/issues/2715 > > > >> > > > [4] https://github.com/apache/fluss/issues/2751 > > > >> > > > [5] https://github.com/lance-format/lance/issues/3961 > > > >> > > > > > > >> > > > > >> > > > >> > > > >> -- > > > >> Lorenzo Affetti > > > >> Senior Software Engineer @ Flink Team > > > >> Ververica <http://www.ververica.com> > > > >> > > > >> > > > >> > > > > > > > > -- > > > > Lorenzo Affetti > > > > Senior Software Engineer @ Flink Team > > > > Ververica <http://www.ververica.com> > > > > > > > > > >
