Mehul, this is huge. Can you please participate in the next episodes? :) I think we should consolidate all these sharp conversations into a single document.
I can work on that with the help of my friend Claude and give you a Google doc, so that we can start directly commenting there and make edits maybe? On Thu, Apr 23, 2026 at 11:26 AM Mehul Batra <[email protected]> wrote: > Hi all, > > Wanted to start a separate thread (or sub-thread) on the multimodal / > unstructured data direction, since I think it deserves its own scoping > discussion rather than being folded into the VECTOR type work. > > Keith made the key observation earlier: current "multimodal ingestion" in > Fluss is really just BYTES columns with no type-system awareness of images, > video, or document formats. Giannis acknowledged this. My view is that we > should explicitly defer deep multimodal investment in the near term, for > the following reasons. > > *1. The VECTOR type work is the prerequisite, not a parallel track* > > Most production multimodal pipelines (text, images, audio) converge on the > same bottleneck: *embedding*. Whether the source is a PDF, a video frame, > or an audio clip, the valuable retrieval signal ends up as a vector. If we > do not get the VECTOR type right (model lineage, partial visibility, > hierarchy points raised in the other thread), multimodal support built on > top will inherit every one of those gaps, multiplied. > > Better to nail the vector foundation first, then layer modality-specific > extraction on top. > > *2. Multimodal ingestion is a pipeline problem, not a storage problem* > > What production systems actually need for multimodal is not "a place to > put bytes" object storage already solves that. What they need is: > > - File type detection and routing (PDF vs scanned image vs DOCX need > different extraction paths) > - Extraction metadata preservation (page numbers, bounding boxes, > timestamps lost at ingest is unrecoverable) > - Extraction failure handling as a first-class state (content present > but extraction failed is distinct from content missing) > - Blob reference vs inline storage decisions (inlining every video > frame into a streaming layer is impractical) > > None of this is a Fluss-shaped problem today. Most of it lives in the > extraction pipeline (Unstructured.io, Tika, dedicated OCR services, vision > models) that feeds the storage layer. Fluss adding multimodal awareness > without solving the pipeline layer means we inherit the hard parts without > the upstream tooling. > > *3. "Fluss supports multimodal" as a narrative risks overpromising* > > Today, storing BYTES is not meaningfully different from what Kafka + S3 > already offers. Claiming multimodal support on top of BYTES columns invites > comparison with systems that actually have semantic awareness (Milvus with > multimodal collections, LanceDB with typed content, vector-native > platforms). We will not win that comparison on the current foundation. > > The honest positioning, at least for now, is: *Fluss is the real-time > data backbone for vector workloads, and multimodal retrieval is delegated > to integrations until the extraction and type-system story is mature.* > > *4. Use-case validation is especially thin here* > > Giannis flagged that he is reaching out to companies like Yelp and Booking > to validate vector retrieval use cases. For multimodal specifically, we do > not yet have a clear picture of: > > - Which modalities matter most to Fluss users (text-from-docs? images? > video? audio?) > - Whether users want Fluss to own the extraction pipeline or just > ingest already-extracted content > - What blob sizes and retention patterns we are actually designing for > > Without this input, a multimodal type system risks being designed for a > use case no one has. > > I believe, we should defer deep multimodal investment until the VECTOR > type FIP ships and stabilizes, In the interim, document the current pattern > clearly: BYTES for raw content, external extraction pipeline (Flink + > Unstructured/Tika/OCR), vector written back to Fluss via the embedding > pipeline pattern, Collect multimodal use cases alongside the vector > use-case outreach Giannis is driving. Treat it as an input to a future FIP > rather than a parallel track now, once the VECTOR type is in place and real > user demand is characterized, revisit multimodal as its own scoped FIP > covering type-system awareness, extraction metadata, and blob reference > semantics. > > Happy to help scope the multimodal discussion. > > Best regards, > Mehul > > On Thu, Apr 23, 2026 at 2:45 PM Mehul Batra <[email protected]> > wrote: > >> Hi all, >> >> Adding four observations from running a RAG platform in production . >> Hopefully useful as input to the shared doc and the upcoming VECTOR type >> FIP. >> >> *1. Embedding latency dominates freshness, not storage* >> >> Keith raised the right question on whether Fluss on top of existing infra >> actually improves vector index freshness. From our deployment, end-to-end >> freshness from ingest to queryable vector is bounded by embedding >> throughput, not storage. On a embedding model like (BGE-M3) with fractional >> GPU on RayServe, embedding is the dominant cost by a wide margin. >> >> One additional wrinkle worth flagging: if users opt into semantic >> chunking (as opposed to fixed-length or section-based), the chunking step >> itself involves embedding calls to compute semantic boundaries. That stacks >> a second embedding pass before the retrieval embedding pass, compounding >> the latency. So the freshness bound depends on chunking strategy, not just >> the final embedding model. >> >> Implication: Fluss hot-layer serving improves raw content freshness >> (immediately useful for BM25/keyword retrieval), but vector index freshness >> is gated by the embedding step(s). Worth being precise about which >> freshness we are claiming to improve, because conflating the two oversells >> the story. >> >> *2. VECTOR type without model lineage is a production footgun* >> >> A vector column carrying only dimension metadata will break the first >> time someone upgrades an embedding model. Vectors from BGE-M3 and Qwen3 >> Embedding 0.6B can share dimensions but are not interchangeable. Any join, >> union read, or index built across them silently produces garbage. >> >> Suggestion: the VECTOR type should carry model identity (model name + >> version + normalization scheme), not just dimension. Without this, schema >> evolution becomes a data corruption event. Worth aligning with Paimon >> PIP-40 on this specifically before the API freezes. >> >> *3. Hierarchical chunking breaks the flat VECTOR assumption* >> >> Production RAG is not one vector per document. It is a tree of vectors at >> different granularities, typically parent chunks for context reconstruction >> and child chunks for retrieval precision. Different embedding models have >> different optimal child chunk sizes, which we tune per model. >> >> For the initial VECTOR type, fixed-length and section-based chunking are >> reasonable starting points and map cleanly onto a flat schema. I want to >> raise hierarchical chunking as a forward-looking concern so we do not paint >> ourselves into a corner. >> >> Two implications for Fluss once hierarchy is in scope: >> >> - Union read semantics get harder when the hot tier has child chunks >> but retrieval needs parent context reconstruction. >> - The type system needs to either understand parent-child >> relationships natively, or the storage layer needs conventions for >> co-locating them. >> - Hierarchical chunking is the standard small-to-big retrieval >> pattern embed small chunks for match precision, return parent chunks for >> LLM context, increases the overall precision >> >> Flagging this now so the flat-case design leaves room for it, rather than >> retrofitting later at higher cost. >> >> *4. Post-append enrichment needs partial row visibility semantics* >> >> Building on Keith's post-append enrichment idea: this is not just a log >> immutability question, it is a query semantics question. In production RAG, >> raw content is queryable immediately (for BM25/keyword) but vectors >> populate asynchronously. >> >> The system needs explicit partial row visibility so that queries can >> distinguish rows where the vector is NULL (keyword-only retrieval) from >> rows where the vector is populated (hybrid retrieval). Without this, hybrid >> retrieval either blocks on embedding or returns inconsistent results. This >> is a concrete design constraint for the FIP. >> >> Happy to contribute to the shared doc on any of these, particularly (2) >> and (3) since they directly shape the VECTOR type FIP. >> >> Best regards, >> Mehul >> >> On Fri, Apr 10, 2026 at 5:57 PM Lorenzo Affetti via dev < >> [email protected]> wrote: >> >>> Hi all, >>> >>> Apologies in advance for the length of this email — I considered >>> splitting >>> it into a series, but I figured a single consolidated reply would be >>> easier >>> to reference. >>> We may opt for adding this to our shared document on Drive. >>> >>> Thank you Keith for the sharp observations — especially around whether >>> the >>> hot layer actually improves freshness when embedding latency is the >>> bottleneck, and the post-append enrichment idea. I want to build on this >>> and try to frame things around concrete use cases, because I think the >>> retrieval pattern we need depends entirely on what we're solving for. >>> >>> I did my research and I decided to wrap everything up into 3 relevant use >>> cases focusing on retrieval patterns, what we have, and what's missing. >>> I also added a broader look at agentic workloads in section 2. >>> >>> Before I dive in, I want to put forward a framing that I think should >>> guide >>> us (coming out of our last discussion): >>> >>> """ >>> Vector databases answer "what's relevant?" >>> Fluss can answer "what's relevant right now, what just happened, and >>> what's >>> the current state?" — in one system. >>> For workloads that need shared state, event-driven coordination, and >>> fresh >>> context alongside semantic retrieval (especially agentic workloads), >>> Fluss is the real-time data backbone that vector DBs were never designed >>> to >>> be. >>> """ >>> >>> With that in mind, let's go! >>> >>> ## 1. Use Cases and Their Retrieval Patterns >>> >>> As I see it, there are three distinct use cases where Fluss can play a >>> role >>> in AI workloads, and each implies a different retrieval pattern: >>> >>> ### 1a. Real-time feature serving for ML inference >>> - Pattern: Point lookup by key (e.g. "give me the latest embedding for >>> user_id 12345") >>> - Fluss capability today: KV tables already support this. The value just >>> happens to be a vector. >>> - What's missing: A native VECTOR type to make the intent explicit. >>> Otherwise this works today. >>> >>> ### 1b. Keeping external vector DBs fresh (Fluss as CDC source) >>> - Pattern: Changelog subscription — downstream systems (Milvus, Qdrant, >>> LanceDB) consume the log to keep their indexes up to date. >>> - Fluss capability today: Log tables with CDC already support this. >>> - What's missing: Again, the VECTOR type for schema clarity, and >>> potentially optimized serialization for high-dimensional vectors. But >>> functionally, this works. >>> >>> ### 1c. Real-time RAG / context engineering >>> - Pattern: Similarity search — "find the 10 most relevant chunks to this >>> query from the last N minutes + historical data." >>> - Fluss capability today: This does not exist. Point lookup doesn't help >>> because you don't know which keys are relevant — that's what similarity >>> search tells you. >>> - What's missing: This is the big gap. It requires vector search >>> capability >>> on the hot layer (brute-force for small datasets as Lorenzo suggested), >>> union-read with the cold tier (Lance/Paimon), and a query interface to >>> express nearest-neighbor queries. >>> >>> Use cases 1a and 1b are achievable with relatively small investments >>> (VECTOR type + documentation). Use case 1c is the most compelling for the >>> AI narrative but requires the most new capability. >>> >>> >>> ## 2. Agentic Workloads — Where Fluss Can Truly Differentiate >>> >>> I've been thinking about where Fluss fits in the emerging agentic AI >>> landscape, and I believe this is where our architecture gives us a >>> genuine >>> edge over pure vector databases. Agents have fundamentally different data >>> patterns from traditional RAG — they're stateful, multi-step, often >>> collaborative, and they produce data as much as they consume it. >>> >>> Here's how I see the agent data needs mapping to Fluss's capabilities: >>> >>> ### 2a. Shared context across agents >>> When multiple agents collaborate (research agent + coding agent + review >>> agent on the same task), they need shared, evolving context. This isn't a >>> vector search problem — it's a real-time state synchronization problem. >>> This maps directly to Fluss's KV table with changelog: each agent writes >>> contributions keyed by task/step ID, others subscribe or do point >>> lookups. >>> This works today. >>> >>> ### 2b. Agent memory at multiple time horizons >>> - Working memory (current task state, intermediate results): hot, >>> mutable, >>> keyed state → Fluss KV table. Works today. >>> - Episodic memory (what worked in past similar tasks): semantic search >>> over >>> historical data → Log tiered to Lance/Paimon, needs similarity search on >>> cold tier. Partially works (storage/tiering yes, retrieval no). >>> - Organizational memory (policies, reference docs, past decisions): >>> classic >>> RAG corpus. Fluss keeps it fresh via CDC, but the heavy retrieval is >>> delegated to the cold tier or an external system. >>> >>> The key insight: vector search is only needed for episodic and >>> organizational memory, not for working memory. And working memory is the >>> highest-frequency, most latency-sensitive pattern — exactly where Fluss's >>> KV layer is strongest. >>> >>> ### 2c. Tool call and action logging >>> Agents make tool calls and take actions that need durable, ordered >>> logging >>> for observability, debugging, and learning. This is what Fluss's >>> append-only log was literally built for. Hot tier for real-time >>> monitoring, >>> cold tier (via tiering) for post-hoc analysis and fine-tuning data >>> extraction. >>> >>> ### 2d. Event-driven agent coordination >>> "When the research agent publishes findings, wake up the synthesis >>> agent." >>> This is event-driven orchestration via log subscription — native Fluss >>> capability. Kafka can do this too, but Fluss adds KV serving and tiering >>> in >>> the same system. >>> >>> ### 2e. Context window management >>> Agents need to decide what goes into the LLM context window at each step. >>> This requires fast retrieval with both recency awareness and semantic >>> relevance. A pure vector DB gives relevance but no recency guarantees. A >>> pure stream gives recency but no semantic retrieval. Fluss with hot-layer >>> brute-force + KV lookups + cold-tier indexed search could serve all three >>> needs. >>> >>> To summarize the gap analysis for agents: >>> >>> - Shared mutable state: Vector DBs are not designed for this. Fluss KV >>> tables are built for it. >>> - Real-time event reactions: Vector DBs have no pub/sub. Fluss log >>> subscription handles it. >>> - Ordered action history: Vector DBs aren't logs. Fluss is. >>> - Data freshness: Vector DBs depend on re-indexing latency. Fluss hot >>> layer >>> is immediately available. >>> - Semantic retrieval over history: This is where vector DBs excel and >>> where >>> Fluss has a gap today. >>> >>> Fluss is strong where vector DBs are weak, and weak where vector DBs are >>> strong. We don't need to beat Milvus at similarity search — we need to >>> complement it, or provide "good enough" search for the hot tier while >>> excelling at everything else. >>> >>> >>> ## 3. On the "All-in-One" Framing >>> >>> I want to be honest about something: even with the embedding pipeline >>> pattern (Fluss → Flink + embedding UDF → Fluss → Lance tiering), this is >>> not truly an all-in-one system. The architecture still involves Fluss >>> (storage), Flink (compute/orchestration), and an external embedding model >>> (Claude, OpenAI, local model). That's three systems minimum. >>> >>> What Fluss actually offers is a simplification of the pipeline topology — >>> from 4-5 systems (Kafka → embedding service → queue → vector DB → lake) >>> down to 2-3 (Fluss → Flink w/ embedding → Fluss w/ tiering). Fewer moving >>> parts, unified storage, one operational surface. That's genuinely >>> valuable, >>> but we should frame it as integration simplification rather than >>> consolidation into a single system. I believe this framing is more >>> credible >>> and still compelling. >>> >>> >>> ## 4. The VECTOR Type as Foundation >>> >>> Regardless of which use case we prioritize, the VECTOR type is the common >>> prerequisite. Giannis correctly pointed to Paimon's PIP-40 as a >>> reference. >>> I think this should be our first concrete deliverable — it's low risk, >>> high >>> signal, and unblocks everything else. Without it, we're storing vectors >>> as >>> BYTES or ARRAY, which works mechanically but tells neither the system nor >>> the user anything about what the data actually is. >>> >>> ## 5. On Post-Append Enrichment >>> >>> Keith, your idea about populating a vector column in-place rather than >>> writing to a separate table is one of the most compelling differentiators >>> I've heard in this discussion. A table that's immediately queryable for >>> raw >>> content and becomes vector-searchable shortly after — that's genuinely >>> different from what Kafka or any current streaming system offers. >>> >>> That said, I think this deserves its own focused discussion and >>> potentially >>> its own FIP, as it touches fundamental assumptions about log >>> immutability. >>> The most pragmatic path might be leveraging KV table upsert semantics >>> (write raw first, update with vector second), but there are more >>> ambitious >>> designs worth exploring. I'd suggest we scope this separately so it >>> doesn't >>> block the foundational work. >>> >>> >>> ## 6. Suggested Priorities >>> >>> 1. VECTOR type (FIP, aligning with Paimon PIP-40) — common prerequisite, >>> unblocks everything >>> 2. Embedding pipeline quickstart (building on PR #2716 with batched >>> calls) >>> — proves the story with minimal effort >>> 3. Use case gathering (Giannis's outreach to Yelp, Booking, etc.) — >>> validate which retrieval patterns and which agentic patterns matter to >>> real >>> users >>> 4. Agentic patterns documentation — KV tables for agent state, log tables >>> for coordination. This works today and is an untold story. >>> 5. Post-append enrichment design (separate FIP) — Keith's differentiating >>> idea, scoped independently >>> 6. Hot-layer similarity search + union-read — once use case 1c and the >>> agentic context engineering need are validated >>> >>> I think collecting all of this in a shared document would be a great next >>> step. Happy to contribute — especially on the agentic workload framing >>> and >>> the retrieval pattern analysis. >>> >>> Looking forward to the discussion. >>> >>> On Wed, Apr 8, 2026 at 1:20 PM Keith Lee <[email protected]> wrote: >>> >>> > Hello, >>> > >>> > I really like the "Fluss as embedding pipeline" direction. >>> > >>> > The core pattern of ingestion -> Flink batching + embedding -> writing >>> back >>> > to Fluss -> Lance Tiering is already possible today, albeit without >>> > modality awareness. >>> > This PR on Lance Quickstart demonstrates exactly that (minus batching): >>> > Flink reads from Fluss log table A, calls an embedding model >>> (streaming), >>> > then writes into Fluss log table B. We can update the PR to use batched >>> > calls. >>> > >>> > >>> > >>> https://github.com/apache/fluss/pull/2716/changes#diff-e549694d35816df1240d10f4597fa20c6df050bd845998f248ce5f883782d93dR391-R398 >>> > >>> > That said, I think there's room to architect Fluss more intentionally >>> for >>> > this use case. Reading from a log table, generating embeddings, and >>> writing >>> > to another log table is something Kafka can be used for - the Lance >>> tiering >>> > is the differentiator, but it's not enough on its own. >>> > One concrete improvement worth exploring: allowing post-append >>> enrichment >>> > on log tables (e.g. populating a row's vector column in-place). This >>> would >>> > significantly reduce data movement between tables. >>> > >>> > +1 on collecting this in a document to help us brainstorm further - >>> happy >>> > to contribute to that. >>> > >>> > Best regards >>> > Keith >>> > >>> > On Wed, Apr 8, 2026 at 6:55 AM Giannis Polyzos <[email protected]> >>> > wrote: >>> > >>> > > Exciting indeed 😄 >>> > > >>> > > Actually, you are correct, that "multimodal ingestion" is currently a >>> > > storage/transport capability (BYTES for content, ARRAY for >>> pre-computed >>> > > vectors) rather than type-system-level semantic awareness. A native >>> > VECTOR >>> > > type would make the design intent explicit. I see paimon has a >>> proposal >>> > for >>> > > a Vector type, which might be something relevant to Fluss as well >>> > > >>> > > >>> > >>> https://cwiki.apache.org/confluence/display/PAIMON/PIP-40%3A+Introduce+a+new+Vector+data+type >>> > > >>> > > On freshness: Fluss reduces every bottleneck except embedding model >>> > > latency, which it cannot eliminate but can mitigate via micro-batch >>> > > embedding in a Flink job. The hot log layer also enables immediate >>> > > raw-content retrieval before vector indexes are updated, which has a >>> > > standalone value for hybrid retrieval in context engineering. >>> > > >>> > > The lowest-complexity path forward could be "Fluss as embedding >>> pipeline" >>> > > pattern: ingest -> Flink batching + embedding UDF or AI Functions -> >>> > write >>> > > vectors back to Fluss -> Lance (or even Paimon, assuming it goes down >>> > that >>> > > direction) tiering. This consolidates what currently requires Kafka + >>> > > external embedding service + vector DB into a single system. The >>> missing >>> > > piece is making this pattern well-documented and ergonomic (ideally >>> with >>> > a >>> > > VECTOR type). >>> > > >>> > > It might be worth collecting all this information in a document to >>> better >>> > > help us brainstorm, but I think this might be a good first approach. >>> > > >>> > > Best, >>> > > Giannis >>> > > >>> > > On Wed, Apr 8, 2026 at 1:17 AM Keith Lee <[email protected]> wrote: >>> > > >>> > > > Hi Giannis, dev, >>> > > > >>> > > > Thank you for following up and your input. I agree in general on >>> not >>> > > > fixating on the details of technical implementation. Adding my >>> > > observations >>> > > > here. >>> > > > >>> > > > > Currently, Fluss supports the ingestion of multi-modal data and >>> > > > tiering on the Lance format >>> > > > > ingestion of multi-modal data >>> > > > > fast serving so that it can be used for context engineering use >>> case >>> > > > >>> > > > I am not certain that Fluss currently supports ingestion of >>> multimodal >>> > > > data. Or, at least, it is not aware of image and video on the type >>> > > system / >>> > > > metadata level. We do have to think about how ingestion and serving >>> > will >>> > > > look like here if we decide to defer vector processing e.g. will a >>> hot >>> > > > layer for multi modal data be useful for context engineering if >>> there’s >>> > > no >>> > > > vector query capability? >>> > > > >>> > > > The discussion we had left me thinking around the aspect of using >>> Fluss >>> > > as >>> > > > hot layer on top of existing format / infra used for multi-modal >>> > context >>> > > > engineering. Specifically, it’d be important for us to understand >>> > > > 1. what is the de-facto average / worse case data freshness (vector >>> > index >>> > > > freshness?) achievable with existing format / tools? >>> > > > 2. will adding Fluss on top of existing format / infra actually >>> help >>> > > > improve data freshness (vector index freshness)? I imagine that >>> vector >>> > > > embedding might be the bottleneck (Fluss will need to call a model >>> to >>> > get >>> > > > vector embedding) >>> > > > 3. Can Fluss help in a different way e.g. achieve similar data >>> > freshness >>> > > at >>> > > > lower complexity / cost? E.g. Fluss performing vector embedding by >>> > > batching >>> > > > and calling an embedding model (locally or cloud) >>> > > > >>> > > > Exciting discussions! >>> > > > >>> > > > Best regards >>> > > > Keith >>> > > > >>> > > > >>> > > > >>> > > > On Mon, 6 Apr 2026 at 08:52, Giannis Polyzos < >>> [email protected]> >>> > > > wrote: >>> > > > >>> > > > > Hi devs, >>> > > > > >>> > > > > Following up on our discussions and Fluss direction on Vector >>> data >>> > > > > support, i >>> > > > > wanted to leave here my two cents. >>> > > > > >>> > > > > I wanna start by saying that im trying to follow-up with a few >>> > > companies >>> > > > > that work with vectors - like Yelp and Booking to understand >>> their >>> > use >>> > > > > cases and ideally get some feedback from them to better help us >>> shape >>> > > > this >>> > > > > direction. Currently, Fluss supports the ingestion of multi-modal >>> > data >>> > > > and >>> > > > > tiering on the Lance format.. Seems like Paimon will also invest >>> > > towards >>> > > > > that direction. >>> > > > > So I think a good first step for Fluss in that direction would >>> be to >>> > > act >>> > > > as >>> > > > > a streaming storage layer that can support: >>> > > > > 1. The ingestion of multi-modal data >>> > > > > 2. Fast serving of that data so it can be used for context >>> > engineering >>> > > > use >>> > > > > cases >>> > > > > 3. Continue its support and enhancement on paimon and Lance >>> format - >>> > > for >>> > > > > example supporting the Primary Key table there. >>> > > > > >>> > > > > I think for now these would be some good first steps, considering >>> > there >>> > > > is >>> > > > > already ground work there, the Lance format seems to be getting >>> some >>> > > good >>> > > > > community adoption. >>> > > > > So my suggestion would be to use the above as guideliness and not >>> > spend >>> > > > too >>> > > > > much time now at processing vectors and defer that to >>> integrations, >>> > for >>> > > > > example a LanceDB integration and then as we collect more >>> feedback >>> > > > > re-iterate. >>> > > > > >>> > > > > Another thing that may be good to think about is how users can >>> > > integrate >>> > > > > existing unstructured data --- think legal documents that already >>> > live >>> > > on >>> > > > > S3 or other object storage -- and make fluss aware of them for >>> > serving >>> > > > them >>> > > > > again as part of some context engineering jobs. >>> > > > > https://fluss.apache.org/blog/fluss-for-ai/ >>> > > > > I think that what we have in Fluss for AI is already a compelling >>> > story >>> > > > and >>> > > > > allow fluss to act as a centralized data repository for all >>> types of >>> > > > data, >>> > > > > so lets focus on that as a first step. >>> > > > > >>> > > > > Let me know your thoughts, and if there are more suggestions and >>> > > > proposal I >>> > > > > would be eager to hear your thoughts. >>> > > > > >>> > > > > Best, >>> > > > > Giannis >>> > > > > >>> > > > > On Mon, Mar 9, 2026 at 5:20 PM Lorenzo Affetti < >>> > > > > [email protected]> wrote: >>> > > > > >>> > > > > > Thanks guys for the valuable feedback. >>> > > > > > >>> > > > > > I will put this on the table with Wangcheng and Giannis >>> Polyzos (I >>> > > know >>> > > > > he >>> > > > > > has quite a vision for the future of Fluss for AI: >>> > > > > > https://fluss.apache.org/blog/fluss-for-ai/). >>> > > > > > So that we can come up with a roadmap and put that under the >>> > > discussion >>> > > > > > thread on Github. >>> > > > > > >>> > > > > > Thrilled! >>> > > > > > >>> > > > > > On Mon, Mar 2, 2026 at 1:35 PM ForwardXu <[email protected]> >>> wrote: >>> > > > > > >>> > > > > >> Hi all, >>> > > > > >> I think it makes perfect sense to create a dedicated roadmap >>> for >>> > > Lance >>> > > > > >> support. This will help us clarify our priorities and ensure >>> we >>> > can >>> > > > > deliver >>> > > > > >> more comprehensive support, including advanced features like >>> > complex >>> > > > > data >>> > > > > >> types and blob types, among others. >>> > > > > >> Looking forward to discussing this further on Slack. >>> > > > > >> >>> > > > > >> Best, >>> > > > > >> Forwardxu >>> > > > > >> >>> > > > > >> 原始邮件 >>> > > > > >> ------------------------------ >>> > > > > >> 发件人:Lorenzo Affetti via dev <[email protected]> >>> > > > > >> 发件时间:2026年3月2日 18:48 >>> > > > > >> 收件人:dev <[email protected]> >>> > > > > >> 抄送:forwardxu <[email protected]>, Lorenzo Affetti < >>> > > > > >> [email protected]> >>> > > > > >> 主题:Re: Analysis of Lance storage format support >>> > > > > >> >>> > > > > >> Hello! Thanks for wrapping this up! >>> > > > > >> >>> > > > > >> I do understand both Cheng and Keith. >>> > > > > >> For sure Lance support should be on par with other lake >>> formats. >>> > If >>> > > > > >> something is not supported, there should be a concrete reason >>> why >>> > > > (apart >>> > > > > >> from a lack of resources :) ). >>> > > > > >> >>> > > > > >> Still, input from the Lance community would be essential for >>> > > > > >> understanding evolution areas of the support itself. >>> > > > > >> >>> > > > > >> For this item, I would take an approach similar to what Mehul >>> did >>> > > for >>> > > > > >> Iceberg support. >>> > > > > >> I think there is a lack of a roadmap for Lance support in >>> 2026. >>> > > > > >> >>> > > > > >> Having a roadmap doesn't actually mean we will accomplish >>> > > everything, >>> > > > > but, >>> > > > > >> it signals that we understand the problem space and have an >>> idea >>> > of >>> > > > the >>> > > > > >> sequence of actions to take. >>> > > > > >> >>> > > > > >> @cheng, I think you are the de-facto owner of the Lance >>> module. >>> > > > > >> Would it make sense to dedicate some of our resources to >>> discuss >>> > > this >>> > > > > via >>> > > > > >> Slack and start drafting a roadmap? >>> > > > > >> >>> > > > > >> On Sun, Mar 1, 2026 at 2:11 PM Keith Lee < >>> > > [email protected] >>> > > > > >>> > > > > >> wrote: >>> > > > > >> >>> > > > > >> > Hello Cheng, >>> > > > > >> > >>> > > > > >> > Good call. I agree that gathering input from Lance community >>> > will >>> > > be >>> > > > > >> >>> > > > > >> > beneficial to inform integration of features such as vector >>> > > search, >>> > > > > vector >>> > > > > >> > indexing and hybrid search. >>> > > > > >> > >>> > > > > >> >>> > > > > >> > However, the issues I’ve outlined only meant to cover the >>> scope >>> > of >>> > > > > bringing >>> > > > > >> > current fluss lance integration up to parity to other >>> lakehouses >>> > > > like >>> > > > > >> > paimon or iceberg e.g. batch or union read without lance >>> feature >>> > > > such >>> > > > > as >>> > > > > >> >>> > > > > >> > vector search. As such, I believe these can be decoupled >>> and we >>> > > can >>> > > > > have a >>> > > > > >> >>> > > > > >> > separate effort, gathering input from lance community and >>> FIP >>> > > > > proposal for >>> > > > > >> > integrating vector search into feature such as union read. >>> > > > > >> > >>> > > > > >> > Let me know what your thoughts are on this. Thank you! >>> > > > > >> > >>> > > > > >> > Best regards >>> > > > > >> > Keith Lee >>> > > > > >> > >>> > > > > >> > >>> > > > > >> > On Sun, 1 Mar 2026 at 10:20, Cheng Wang <[email protected]> >>> > wrote: >>> > > > > >> > >>> > > > > >> > > Hello Keith, >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> >>> > > > > >> > > Regarding our plan to implement union read for Lance using >>> > > Flink, >>> > > > > might >>> > > > > >> > it >>> > > > > >> > > be beneficial to first gather input from the Lance >>> community? >>> > > > > >> > Understanding >>> > > > > >> >>> > > > > >> > > the primary scenarios where union read would help in the >>> > machine >>> > > > > learning >>> > > > > >> > > scenario, along with the most popular execution engine in >>> > Lance >>> > > > > >> > ecosystem, >>> > > > > >> > > could ensure we're building the right integration to >>> maximize >>> > > its >>> > > > > >> > adoption. >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> > > Regards, >>> > > > > >> > > Cheng Wang >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> > > ------------------ Original ------------------ >>> > > > > >> > > From: >>> > > > > >> > > "dev" >>> > > > > >> > > >>> > > < >>> > > > > >> > > [email protected]>; >>> > > > > >> > > Date: Sat, Feb 28, 2026 11:20 PM >>> > > > > >> > > To: "dev"<[email protected]>; >>> > > > > >> > > Cc: "Cheng Wang"<[email protected]>;"forwardxu"< >>> > > > > >> > > [email protected]>; >>> > > > > >> > > Subject: Re: Analysis of Lance storage format support >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> > > >>> > > > > >> > > This is extremely helpful, thanks for putting this >>> together. >>> > > > > >> > > >>> > > > > >> >>> > > > > >> > > Maybe we can create an umbrella ticket on GitHub to keep >>> track >>> > > on >>> > > > > these >>> > > > > >> > and >>> > > > > >> > > open individual tasks, for tracking. >>> > > > > >> > > >>> > > > > >> > > Best, >>> > > > > >> > > Giannis >>> > > > > >> > > >>> > > > > >> > > On Sat, 28 Feb 2026 at 3:52 PM, Keith Lee < >>> > > > > >> [email protected] >>> > > > > >> > > >>> > > > > >> > > wrote: >>> > > > > >> > > >>> > > > > >> > > > Hello, >>> > > > > >> > > > >>> > > > > >> >>> > > > > >> > > > As discussed on community sync yesterday on analysing >>> > where >>> > > > we >>> > > > > are >>> > > > > >> > at >>> > > > > >> > > the >>> > > > > >> > > > moment in terms of Lance format support. >>> > > > > >> > > > Here are my findings as part of working on Lance >>> > QuickStart >>> > > > > >> > > documentation >>> > > > > >> > > > [1]. Lance lake tiering works in general, however >>> there >>> > are >>> > > > > some >>> > > > > >> > gaps >>> > > > > >> > > that >>> > > > > >> >>> > > > > >> > > > to be addressed to bring Lance format support in >>> parity >>> > > with >>> > > > > Paimon >>> > > > > >> > / >>> > > > > >> > > > Iceberg. >>> > > > > >> > > > >>> > > > > >> >>> > > > > >> > > > - (Merged) Support for Arrow FixedSizeList to enable >>> > > pylance >>> > > > > native >>> > > > > >> > > vector >>> > > > > >> > > > search [2] >>> > > > > >> > > > - (In progress) Support Flink SQL Union Read query >>> > against >>> > > > > Lance >>> > > > > >> > > table [3] >>> > > > > >> > > > - (Open) Support Flink SQL batch query against Lance >>> > table >>> > > > [4] >>> > > > > >> > > > - (Blocked) Primary Key table support - I believe >>> this is >>> > > > still >>> > > > > >> > > blocking on >>> > > > > >> > > > Lance format support for delete API [5] >>> > > > > >> > > > >>> > > > > >> > > > Finally there is also a gap in the ability of >>> performing >>> > > > vector >>> > > > > >> > > search on >>> > > > > >> > > > hot data / via union read. After discussion with >>> Mehul, >>> > > > native >>> > > > > >> > vector >>> > > > > >> > > > indexing on hot data in Fluss would be a separate, >>> bigger >>> > > > > effort >>> > > > > >> > that >>> > > > > >> > > we >>> > > > > >> > > > can evolve towards if there's demand for it. >>> > > > > >> > > > >>> > > > > >> > > > Appreciate feedback here from Cheng, Forward and >>> anyone >>> > > else >>> > > > > with >>> > > > > >> >>> > > > > >> > > > familiarity around this area as I have only started >>> > dipping >>> > > > my >>> > > > > toes >>> > > > > >> > > into >>> > > > > >> > > > Lance. >>> > > > > >> > > > >>> > > > > >> > > > *Additionally, if anyone wants to help contributing >>> in >>> > this >>> > > > > area, >>> > > > > >> > > please >>> > > > > >> > > > reach out. * >>> > > > > >> > > > >>> > > > > >> > > > Best regards >>> > > > > >> > > > Keith Lee >>> > > > > >> > > > >>> > > > > >> > > > Reference >>> > > > > >> > > > [1] https://github.com/apache/fluss/pull/2716 >>> > > > > >> > > > [2] https://github.com/apache/fluss/issues/2706 >>> > > > > >> > > > [3] https://github.com/apache/fluss/issues/2715 >>> > > > > >> > > > [4] https://github.com/apache/fluss/issues/2751 >>> > > > > >> > > > [5] >>> https://github.com/lance-format/lance/issues/3961 >>> > > > > >> > > > >>> > > > > >> > >>> > > > > >> >>> > > > > >> >>> > > > > >> -- >>> > > > > >> Lorenzo Affetti >>> > > > > >> Senior Software Engineer @ Flink Team >>> > > > > >> Ververica <http://www.ververica.com> >>> > > > > >> >>> > > > > >> >>> > > > > >> >>> > > > > > >>> > > > > > -- >>> > > > > > Lorenzo Affetti >>> > > > > > Senior Software Engineer @ Flink Team >>> > > > > > Ververica <http://www.ververica.com> >>> > > > > > >>> > > > > >>> > > > >>> > > >>> > >>> >>> >>> -- >>> Lorenzo Affetti >>> Senior Software Engineer @ Flink Team >>> Ververica <http://www.ververica.com> >>> >> -- Lorenzo Affetti Senior Software Engineer @ Flink Team Ververica <http://www.ververica.com>
