I created this design doc: https://docs.google.com/document/d/1Ww84Z1dtnOCNEv_fv49Nqz9zMVpc3nKc4a162_ELGus/edit?usp=sharing
I already shared it with @Keith Lee <https://apache-fluss.slack.com/team/U0A047XE0QH> @Giorgos Stamatakis <https://apache-fluss.slack.com/team/U09KZ3JJSHE> @Giannis Polyzos <https://apache-fluss.slack.com/team/U08C8MT54V8> @Jark Wu <https://apache-fluss.slack.com/team/U08CXD51A1E> @Mehul Batra <https://apache-fluss.slack.com/team/U08CQTKH3RR>. If anybody else wants, please request [image: :handshake:] I generated that with the help of Claude and feeding all the context from our e-mail conversations to sync minutes [image: :handshake:] I tried to keep it compressed, but the topic is wide [image: :sweat_smile:] You will see the separation in 3 inter-dependent tracks: - vector and multi-modal model - real-time context engine - Agent skills & MCP There is also a "roadmap" proposal. If we can validate that via comments on the document and converge it would be awesome so that we can either update/add to the official roadmap (if it makes sense). We can also continue in this thread for visibility. On Fri, Apr 24, 2026 at 9:01 AM Lorenzo Affetti < [email protected]> wrote: > Mehul, this is huge. > Can you please participate in the next episodes? :) > > I think we should consolidate all these sharp conversations into a single > document. > > I can work on that with the help of my friend Claude and give you a Google > doc, so that we can start directly commenting there and make edits maybe? > > > On Thu, Apr 23, 2026 at 11:26 AM Mehul Batra <[email protected]> > wrote: > >> Hi all, >> >> Wanted to start a separate thread (or sub-thread) on the multimodal / >> unstructured data direction, since I think it deserves its own scoping >> discussion rather than being folded into the VECTOR type work. >> >> Keith made the key observation earlier: current "multimodal ingestion" in >> Fluss is really just BYTES columns with no type-system awareness of images, >> video, or document formats. Giannis acknowledged this. My view is that we >> should explicitly defer deep multimodal investment in the near term, for >> the following reasons. >> >> *1. The VECTOR type work is the prerequisite, not a parallel track* >> >> Most production multimodal pipelines (text, images, audio) converge on >> the same bottleneck: *embedding*. Whether the source is a PDF, a video >> frame, or an audio clip, the valuable retrieval signal ends up as a vector. >> If we do not get the VECTOR type right (model lineage, partial visibility, >> hierarchy points raised in the other thread), multimodal support built on >> top will inherit every one of those gaps, multiplied. >> >> Better to nail the vector foundation first, then layer modality-specific >> extraction on top. >> >> *2. Multimodal ingestion is a pipeline problem, not a storage problem* >> >> What production systems actually need for multimodal is not "a place to >> put bytes" object storage already solves that. What they need is: >> >> - File type detection and routing (PDF vs scanned image vs DOCX need >> different extraction paths) >> - Extraction metadata preservation (page numbers, bounding boxes, >> timestamps lost at ingest is unrecoverable) >> - Extraction failure handling as a first-class state (content present >> but extraction failed is distinct from content missing) >> - Blob reference vs inline storage decisions (inlining every video >> frame into a streaming layer is impractical) >> >> None of this is a Fluss-shaped problem today. Most of it lives in the >> extraction pipeline (Unstructured.io, Tika, dedicated OCR services, vision >> models) that feeds the storage layer. Fluss adding multimodal awareness >> without solving the pipeline layer means we inherit the hard parts without >> the upstream tooling. >> >> *3. "Fluss supports multimodal" as a narrative risks overpromising* >> >> Today, storing BYTES is not meaningfully different from what Kafka + S3 >> already offers. Claiming multimodal support on top of BYTES columns invites >> comparison with systems that actually have semantic awareness (Milvus with >> multimodal collections, LanceDB with typed content, vector-native >> platforms). We will not win that comparison on the current foundation. >> >> The honest positioning, at least for now, is: *Fluss is the real-time >> data backbone for vector workloads, and multimodal retrieval is delegated >> to integrations until the extraction and type-system story is mature.* >> >> *4. Use-case validation is especially thin here* >> >> Giannis flagged that he is reaching out to companies like Yelp and >> Booking to validate vector retrieval use cases. For multimodal >> specifically, we do not yet have a clear picture of: >> >> - Which modalities matter most to Fluss users (text-from-docs? >> images? video? audio?) >> - Whether users want Fluss to own the extraction pipeline or just >> ingest already-extracted content >> - What blob sizes and retention patterns we are actually designing for >> >> Without this input, a multimodal type system risks being designed for a >> use case no one has. >> >> I believe, we should defer deep multimodal investment until the VECTOR >> type FIP ships and stabilizes, In the interim, document the current pattern >> clearly: BYTES for raw content, external extraction pipeline (Flink + >> Unstructured/Tika/OCR), vector written back to Fluss via the embedding >> pipeline pattern, Collect multimodal use cases alongside the vector >> use-case outreach Giannis is driving. Treat it as an input to a future FIP >> rather than a parallel track now, once the VECTOR type is in place and real >> user demand is characterized, revisit multimodal as its own scoped FIP >> covering type-system awareness, extraction metadata, and blob reference >> semantics. >> >> Happy to help scope the multimodal discussion. >> >> Best regards, >> Mehul >> >> On Thu, Apr 23, 2026 at 2:45 PM Mehul Batra <[email protected]> >> wrote: >> >>> Hi all, >>> >>> Adding four observations from running a RAG platform in production . >>> Hopefully useful as input to the shared doc and the upcoming VECTOR type >>> FIP. >>> >>> *1. Embedding latency dominates freshness, not storage* >>> >>> Keith raised the right question on whether Fluss on top of existing >>> infra actually improves vector index freshness. From our deployment, >>> end-to-end freshness from ingest to queryable vector is bounded by >>> embedding throughput, not storage. On a embedding model like (BGE-M3) with >>> fractional GPU on RayServe, embedding is the dominant cost by a wide margin. >>> >>> One additional wrinkle worth flagging: if users opt into semantic >>> chunking (as opposed to fixed-length or section-based), the chunking step >>> itself involves embedding calls to compute semantic boundaries. That stacks >>> a second embedding pass before the retrieval embedding pass, compounding >>> the latency. So the freshness bound depends on chunking strategy, not just >>> the final embedding model. >>> >>> Implication: Fluss hot-layer serving improves raw content freshness >>> (immediately useful for BM25/keyword retrieval), but vector index freshness >>> is gated by the embedding step(s). Worth being precise about which >>> freshness we are claiming to improve, because conflating the two oversells >>> the story. >>> >>> *2. VECTOR type without model lineage is a production footgun* >>> >>> A vector column carrying only dimension metadata will break the first >>> time someone upgrades an embedding model. Vectors from BGE-M3 and Qwen3 >>> Embedding 0.6B can share dimensions but are not interchangeable. Any join, >>> union read, or index built across them silently produces garbage. >>> >>> Suggestion: the VECTOR type should carry model identity (model name + >>> version + normalization scheme), not just dimension. Without this, schema >>> evolution becomes a data corruption event. Worth aligning with Paimon >>> PIP-40 on this specifically before the API freezes. >>> >>> *3. Hierarchical chunking breaks the flat VECTOR assumption* >>> >>> Production RAG is not one vector per document. It is a tree of vectors >>> at different granularities, typically parent chunks for context >>> reconstruction and child chunks for retrieval precision. Different >>> embedding models have different optimal child chunk sizes, which we tune >>> per model. >>> >>> For the initial VECTOR type, fixed-length and section-based chunking are >>> reasonable starting points and map cleanly onto a flat schema. I want to >>> raise hierarchical chunking as a forward-looking concern so we do not paint >>> ourselves into a corner. >>> >>> Two implications for Fluss once hierarchy is in scope: >>> >>> - Union read semantics get harder when the hot tier has child chunks >>> but retrieval needs parent context reconstruction. >>> - The type system needs to either understand parent-child >>> relationships natively, or the storage layer needs conventions for >>> co-locating them. >>> - Hierarchical chunking is the standard small-to-big retrieval >>> pattern embed small chunks for match precision, return parent chunks for >>> LLM context, increases the overall precision >>> >>> Flagging this now so the flat-case design leaves room for it, rather >>> than retrofitting later at higher cost. >>> >>> *4. Post-append enrichment needs partial row visibility semantics* >>> >>> Building on Keith's post-append enrichment idea: this is not just a log >>> immutability question, it is a query semantics question. In production RAG, >>> raw content is queryable immediately (for BM25/keyword) but vectors >>> populate asynchronously. >>> >>> The system needs explicit partial row visibility so that queries can >>> distinguish rows where the vector is NULL (keyword-only retrieval) from >>> rows where the vector is populated (hybrid retrieval). Without this, hybrid >>> retrieval either blocks on embedding or returns inconsistent results. This >>> is a concrete design constraint for the FIP. >>> >>> Happy to contribute to the shared doc on any of these, particularly (2) >>> and (3) since they directly shape the VECTOR type FIP. >>> >>> Best regards, >>> Mehul >>> >>> On Fri, Apr 10, 2026 at 5:57 PM Lorenzo Affetti via dev < >>> [email protected]> wrote: >>> >>>> Hi all, >>>> >>>> Apologies in advance for the length of this email — I considered >>>> splitting >>>> it into a series, but I figured a single consolidated reply would be >>>> easier >>>> to reference. >>>> We may opt for adding this to our shared document on Drive. >>>> >>>> Thank you Keith for the sharp observations — especially around whether >>>> the >>>> hot layer actually improves freshness when embedding latency is the >>>> bottleneck, and the post-append enrichment idea. I want to build on this >>>> and try to frame things around concrete use cases, because I think the >>>> retrieval pattern we need depends entirely on what we're solving for. >>>> >>>> I did my research and I decided to wrap everything up into 3 relevant >>>> use >>>> cases focusing on retrieval patterns, what we have, and what's missing. >>>> I also added a broader look at agentic workloads in section 2. >>>> >>>> Before I dive in, I want to put forward a framing that I think should >>>> guide >>>> us (coming out of our last discussion): >>>> >>>> """ >>>> Vector databases answer "what's relevant?" >>>> Fluss can answer "what's relevant right now, what just happened, and >>>> what's >>>> the current state?" — in one system. >>>> For workloads that need shared state, event-driven coordination, and >>>> fresh >>>> context alongside semantic retrieval (especially agentic workloads), >>>> Fluss is the real-time data backbone that vector DBs were never >>>> designed to >>>> be. >>>> """ >>>> >>>> With that in mind, let's go! >>>> >>>> ## 1. Use Cases and Their Retrieval Patterns >>>> >>>> As I see it, there are three distinct use cases where Fluss can play a >>>> role >>>> in AI workloads, and each implies a different retrieval pattern: >>>> >>>> ### 1a. Real-time feature serving for ML inference >>>> - Pattern: Point lookup by key (e.g. "give me the latest embedding for >>>> user_id 12345") >>>> - Fluss capability today: KV tables already support this. The value just >>>> happens to be a vector. >>>> - What's missing: A native VECTOR type to make the intent explicit. >>>> Otherwise this works today. >>>> >>>> ### 1b. Keeping external vector DBs fresh (Fluss as CDC source) >>>> - Pattern: Changelog subscription — downstream systems (Milvus, Qdrant, >>>> LanceDB) consume the log to keep their indexes up to date. >>>> - Fluss capability today: Log tables with CDC already support this. >>>> - What's missing: Again, the VECTOR type for schema clarity, and >>>> potentially optimized serialization for high-dimensional vectors. But >>>> functionally, this works. >>>> >>>> ### 1c. Real-time RAG / context engineering >>>> - Pattern: Similarity search — "find the 10 most relevant chunks to this >>>> query from the last N minutes + historical data." >>>> - Fluss capability today: This does not exist. Point lookup doesn't help >>>> because you don't know which keys are relevant — that's what similarity >>>> search tells you. >>>> - What's missing: This is the big gap. It requires vector search >>>> capability >>>> on the hot layer (brute-force for small datasets as Lorenzo suggested), >>>> union-read with the cold tier (Lance/Paimon), and a query interface to >>>> express nearest-neighbor queries. >>>> >>>> Use cases 1a and 1b are achievable with relatively small investments >>>> (VECTOR type + documentation). Use case 1c is the most compelling for >>>> the >>>> AI narrative but requires the most new capability. >>>> >>>> >>>> ## 2. Agentic Workloads — Where Fluss Can Truly Differentiate >>>> >>>> I've been thinking about where Fluss fits in the emerging agentic AI >>>> landscape, and I believe this is where our architecture gives us a >>>> genuine >>>> edge over pure vector databases. Agents have fundamentally different >>>> data >>>> patterns from traditional RAG — they're stateful, multi-step, often >>>> collaborative, and they produce data as much as they consume it. >>>> >>>> Here's how I see the agent data needs mapping to Fluss's capabilities: >>>> >>>> ### 2a. Shared context across agents >>>> When multiple agents collaborate (research agent + coding agent + review >>>> agent on the same task), they need shared, evolving context. This isn't >>>> a >>>> vector search problem — it's a real-time state synchronization problem. >>>> This maps directly to Fluss's KV table with changelog: each agent writes >>>> contributions keyed by task/step ID, others subscribe or do point >>>> lookups. >>>> This works today. >>>> >>>> ### 2b. Agent memory at multiple time horizons >>>> - Working memory (current task state, intermediate results): hot, >>>> mutable, >>>> keyed state → Fluss KV table. Works today. >>>> - Episodic memory (what worked in past similar tasks): semantic search >>>> over >>>> historical data → Log tiered to Lance/Paimon, needs similarity search on >>>> cold tier. Partially works (storage/tiering yes, retrieval no). >>>> - Organizational memory (policies, reference docs, past decisions): >>>> classic >>>> RAG corpus. Fluss keeps it fresh via CDC, but the heavy retrieval is >>>> delegated to the cold tier or an external system. >>>> >>>> The key insight: vector search is only needed for episodic and >>>> organizational memory, not for working memory. And working memory is the >>>> highest-frequency, most latency-sensitive pattern — exactly where >>>> Fluss's >>>> KV layer is strongest. >>>> >>>> ### 2c. Tool call and action logging >>>> Agents make tool calls and take actions that need durable, ordered >>>> logging >>>> for observability, debugging, and learning. This is what Fluss's >>>> append-only log was literally built for. Hot tier for real-time >>>> monitoring, >>>> cold tier (via tiering) for post-hoc analysis and fine-tuning data >>>> extraction. >>>> >>>> ### 2d. Event-driven agent coordination >>>> "When the research agent publishes findings, wake up the synthesis >>>> agent." >>>> This is event-driven orchestration via log subscription — native Fluss >>>> capability. Kafka can do this too, but Fluss adds KV serving and >>>> tiering in >>>> the same system. >>>> >>>> ### 2e. Context window management >>>> Agents need to decide what goes into the LLM context window at each >>>> step. >>>> This requires fast retrieval with both recency awareness and semantic >>>> relevance. A pure vector DB gives relevance but no recency guarantees. A >>>> pure stream gives recency but no semantic retrieval. Fluss with >>>> hot-layer >>>> brute-force + KV lookups + cold-tier indexed search could serve all >>>> three >>>> needs. >>>> >>>> To summarize the gap analysis for agents: >>>> >>>> - Shared mutable state: Vector DBs are not designed for this. Fluss KV >>>> tables are built for it. >>>> - Real-time event reactions: Vector DBs have no pub/sub. Fluss log >>>> subscription handles it. >>>> - Ordered action history: Vector DBs aren't logs. Fluss is. >>>> - Data freshness: Vector DBs depend on re-indexing latency. Fluss hot >>>> layer >>>> is immediately available. >>>> - Semantic retrieval over history: This is where vector DBs excel and >>>> where >>>> Fluss has a gap today. >>>> >>>> Fluss is strong where vector DBs are weak, and weak where vector DBs are >>>> strong. We don't need to beat Milvus at similarity search — we need to >>>> complement it, or provide "good enough" search for the hot tier while >>>> excelling at everything else. >>>> >>>> >>>> ## 3. On the "All-in-One" Framing >>>> >>>> I want to be honest about something: even with the embedding pipeline >>>> pattern (Fluss → Flink + embedding UDF → Fluss → Lance tiering), this is >>>> not truly an all-in-one system. The architecture still involves Fluss >>>> (storage), Flink (compute/orchestration), and an external embedding >>>> model >>>> (Claude, OpenAI, local model). That's three systems minimum. >>>> >>>> What Fluss actually offers is a simplification of the pipeline topology >>>> — >>>> from 4-5 systems (Kafka → embedding service → queue → vector DB → lake) >>>> down to 2-3 (Fluss → Flink w/ embedding → Fluss w/ tiering). Fewer >>>> moving >>>> parts, unified storage, one operational surface. That's genuinely >>>> valuable, >>>> but we should frame it as integration simplification rather than >>>> consolidation into a single system. I believe this framing is more >>>> credible >>>> and still compelling. >>>> >>>> >>>> ## 4. The VECTOR Type as Foundation >>>> >>>> Regardless of which use case we prioritize, the VECTOR type is the >>>> common >>>> prerequisite. Giannis correctly pointed to Paimon's PIP-40 as a >>>> reference. >>>> I think this should be our first concrete deliverable — it's low risk, >>>> high >>>> signal, and unblocks everything else. Without it, we're storing vectors >>>> as >>>> BYTES or ARRAY, which works mechanically but tells neither the system >>>> nor >>>> the user anything about what the data actually is. >>>> >>>> ## 5. On Post-Append Enrichment >>>> >>>> Keith, your idea about populating a vector column in-place rather than >>>> writing to a separate table is one of the most compelling >>>> differentiators >>>> I've heard in this discussion. A table that's immediately queryable for >>>> raw >>>> content and becomes vector-searchable shortly after — that's genuinely >>>> different from what Kafka or any current streaming system offers. >>>> >>>> That said, I think this deserves its own focused discussion and >>>> potentially >>>> its own FIP, as it touches fundamental assumptions about log >>>> immutability. >>>> The most pragmatic path might be leveraging KV table upsert semantics >>>> (write raw first, update with vector second), but there are more >>>> ambitious >>>> designs worth exploring. I'd suggest we scope this separately so it >>>> doesn't >>>> block the foundational work. >>>> >>>> >>>> ## 6. Suggested Priorities >>>> >>>> 1. VECTOR type (FIP, aligning with Paimon PIP-40) — common prerequisite, >>>> unblocks everything >>>> 2. Embedding pipeline quickstart (building on PR #2716 with batched >>>> calls) >>>> — proves the story with minimal effort >>>> 3. Use case gathering (Giannis's outreach to Yelp, Booking, etc.) — >>>> validate which retrieval patterns and which agentic patterns matter to >>>> real >>>> users >>>> 4. Agentic patterns documentation — KV tables for agent state, log >>>> tables >>>> for coordination. This works today and is an untold story. >>>> 5. Post-append enrichment design (separate FIP) — Keith's >>>> differentiating >>>> idea, scoped independently >>>> 6. Hot-layer similarity search + union-read — once use case 1c and the >>>> agentic context engineering need are validated >>>> >>>> I think collecting all of this in a shared document would be a great >>>> next >>>> step. Happy to contribute — especially on the agentic workload framing >>>> and >>>> the retrieval pattern analysis. >>>> >>>> Looking forward to the discussion. >>>> >>>> On Wed, Apr 8, 2026 at 1:20 PM Keith Lee <[email protected]> wrote: >>>> >>>> > Hello, >>>> > >>>> > I really like the "Fluss as embedding pipeline" direction. >>>> > >>>> > The core pattern of ingestion -> Flink batching + embedding -> >>>> writing back >>>> > to Fluss -> Lance Tiering is already possible today, albeit without >>>> > modality awareness. >>>> > This PR on Lance Quickstart demonstrates exactly that (minus >>>> batching): >>>> > Flink reads from Fluss log table A, calls an embedding model >>>> (streaming), >>>> > then writes into Fluss log table B. We can update the PR to use >>>> batched >>>> > calls. >>>> > >>>> > >>>> > >>>> https://github.com/apache/fluss/pull/2716/changes#diff-e549694d35816df1240d10f4597fa20c6df050bd845998f248ce5f883782d93dR391-R398 >>>> > >>>> > That said, I think there's room to architect Fluss more intentionally >>>> for >>>> > this use case. Reading from a log table, generating embeddings, and >>>> writing >>>> > to another log table is something Kafka can be used for - the Lance >>>> tiering >>>> > is the differentiator, but it's not enough on its own. >>>> > One concrete improvement worth exploring: allowing post-append >>>> enrichment >>>> > on log tables (e.g. populating a row's vector column in-place). This >>>> would >>>> > significantly reduce data movement between tables. >>>> > >>>> > +1 on collecting this in a document to help us brainstorm further - >>>> happy >>>> > to contribute to that. >>>> > >>>> > Best regards >>>> > Keith >>>> > >>>> > On Wed, Apr 8, 2026 at 6:55 AM Giannis Polyzos <[email protected] >>>> > >>>> > wrote: >>>> > >>>> > > Exciting indeed 😄 >>>> > > >>>> > > Actually, you are correct, that "multimodal ingestion" is currently >>>> a >>>> > > storage/transport capability (BYTES for content, ARRAY for >>>> pre-computed >>>> > > vectors) rather than type-system-level semantic awareness. A native >>>> > VECTOR >>>> > > type would make the design intent explicit. I see paimon has a >>>> proposal >>>> > for >>>> > > a Vector type, which might be something relevant to Fluss as well >>>> > > >>>> > > >>>> > >>>> https://cwiki.apache.org/confluence/display/PAIMON/PIP-40%3A+Introduce+a+new+Vector+data+type >>>> > > >>>> > > On freshness: Fluss reduces every bottleneck except embedding model >>>> > > latency, which it cannot eliminate but can mitigate via micro-batch >>>> > > embedding in a Flink job. The hot log layer also enables immediate >>>> > > raw-content retrieval before vector indexes are updated, which has a >>>> > > standalone value for hybrid retrieval in context engineering. >>>> > > >>>> > > The lowest-complexity path forward could be "Fluss as embedding >>>> pipeline" >>>> > > pattern: ingest -> Flink batching + embedding UDF or AI Functions -> >>>> > write >>>> > > vectors back to Fluss -> Lance (or even Paimon, assuming it goes >>>> down >>>> > that >>>> > > direction) tiering. This consolidates what currently requires Kafka >>>> + >>>> > > external embedding service + vector DB into a single system. The >>>> missing >>>> > > piece is making this pattern well-documented and ergonomic (ideally >>>> with >>>> > a >>>> > > VECTOR type). >>>> > > >>>> > > It might be worth collecting all this information in a document to >>>> better >>>> > > help us brainstorm, but I think this might be a good first approach. >>>> > > >>>> > > Best, >>>> > > Giannis >>>> > > >>>> > > On Wed, Apr 8, 2026 at 1:17 AM Keith Lee <[email protected]> wrote: >>>> > > >>>> > > > Hi Giannis, dev, >>>> > > > >>>> > > > Thank you for following up and your input. I agree in general on >>>> not >>>> > > > fixating on the details of technical implementation. Adding my >>>> > > observations >>>> > > > here. >>>> > > > >>>> > > > > Currently, Fluss supports the ingestion of multi-modal data and >>>> > > > tiering on the Lance format >>>> > > > > ingestion of multi-modal data >>>> > > > > fast serving so that it can be used for context engineering use >>>> case >>>> > > > >>>> > > > I am not certain that Fluss currently supports ingestion of >>>> multimodal >>>> > > > data. Or, at least, it is not aware of image and video on the type >>>> > > system / >>>> > > > metadata level. We do have to think about how ingestion and >>>> serving >>>> > will >>>> > > > look like here if we decide to defer vector processing e.g. will >>>> a hot >>>> > > > layer for multi modal data be useful for context engineering if >>>> there’s >>>> > > no >>>> > > > vector query capability? >>>> > > > >>>> > > > The discussion we had left me thinking around the aspect of using >>>> Fluss >>>> > > as >>>> > > > hot layer on top of existing format / infra used for multi-modal >>>> > context >>>> > > > engineering. Specifically, it’d be important for us to understand >>>> > > > 1. what is the de-facto average / worse case data freshness >>>> (vector >>>> > index >>>> > > > freshness?) achievable with existing format / tools? >>>> > > > 2. will adding Fluss on top of existing format / infra actually >>>> help >>>> > > > improve data freshness (vector index freshness)? I imagine that >>>> vector >>>> > > > embedding might be the bottleneck (Fluss will need to call a >>>> model to >>>> > get >>>> > > > vector embedding) >>>> > > > 3. Can Fluss help in a different way e.g. achieve similar data >>>> > freshness >>>> > > at >>>> > > > lower complexity / cost? E.g. Fluss performing vector embedding by >>>> > > batching >>>> > > > and calling an embedding model (locally or cloud) >>>> > > > >>>> > > > Exciting discussions! >>>> > > > >>>> > > > Best regards >>>> > > > Keith >>>> > > > >>>> > > > >>>> > > > >>>> > > > On Mon, 6 Apr 2026 at 08:52, Giannis Polyzos < >>>> [email protected]> >>>> > > > wrote: >>>> > > > >>>> > > > > Hi devs, >>>> > > > > >>>> > > > > Following up on our discussions and Fluss direction on Vector >>>> data >>>> > > > > support, i >>>> > > > > wanted to leave here my two cents. >>>> > > > > >>>> > > > > I wanna start by saying that im trying to follow-up with a few >>>> > > companies >>>> > > > > that work with vectors - like Yelp and Booking to understand >>>> their >>>> > use >>>> > > > > cases and ideally get some feedback from them to better help us >>>> shape >>>> > > > this >>>> > > > > direction. Currently, Fluss supports the ingestion of >>>> multi-modal >>>> > data >>>> > > > and >>>> > > > > tiering on the Lance format.. Seems like Paimon will also invest >>>> > > towards >>>> > > > > that direction. >>>> > > > > So I think a good first step for Fluss in that direction would >>>> be to >>>> > > act >>>> > > > as >>>> > > > > a streaming storage layer that can support: >>>> > > > > 1. The ingestion of multi-modal data >>>> > > > > 2. Fast serving of that data so it can be used for context >>>> > engineering >>>> > > > use >>>> > > > > cases >>>> > > > > 3. Continue its support and enhancement on paimon and Lance >>>> format - >>>> > > for >>>> > > > > example supporting the Primary Key table there. >>>> > > > > >>>> > > > > I think for now these would be some good first steps, >>>> considering >>>> > there >>>> > > > is >>>> > > > > already ground work there, the Lance format seems to be getting >>>> some >>>> > > good >>>> > > > > community adoption. >>>> > > > > So my suggestion would be to use the above as guideliness and >>>> not >>>> > spend >>>> > > > too >>>> > > > > much time now at processing vectors and defer that to >>>> integrations, >>>> > for >>>> > > > > example a LanceDB integration and then as we collect more >>>> feedback >>>> > > > > re-iterate. >>>> > > > > >>>> > > > > Another thing that may be good to think about is how users can >>>> > > integrate >>>> > > > > existing unstructured data --- think legal documents that >>>> already >>>> > live >>>> > > on >>>> > > > > S3 or other object storage -- and make fluss aware of them for >>>> > serving >>>> > > > them >>>> > > > > again as part of some context engineering jobs. >>>> > > > > https://fluss.apache.org/blog/fluss-for-ai/ >>>> > > > > I think that what we have in Fluss for AI is already a >>>> compelling >>>> > story >>>> > > > and >>>> > > > > allow fluss to act as a centralized data repository for all >>>> types of >>>> > > > data, >>>> > > > > so lets focus on that as a first step. >>>> > > > > >>>> > > > > Let me know your thoughts, and if there are more suggestions and >>>> > > > proposal I >>>> > > > > would be eager to hear your thoughts. >>>> > > > > >>>> > > > > Best, >>>> > > > > Giannis >>>> > > > > >>>> > > > > On Mon, Mar 9, 2026 at 5:20 PM Lorenzo Affetti < >>>> > > > > [email protected]> wrote: >>>> > > > > >>>> > > > > > Thanks guys for the valuable feedback. >>>> > > > > > >>>> > > > > > I will put this on the table with Wangcheng and Giannis >>>> Polyzos (I >>>> > > know >>>> > > > > he >>>> > > > > > has quite a vision for the future of Fluss for AI: >>>> > > > > > https://fluss.apache.org/blog/fluss-for-ai/). >>>> > > > > > So that we can come up with a roadmap and put that under the >>>> > > discussion >>>> > > > > > thread on Github. >>>> > > > > > >>>> > > > > > Thrilled! >>>> > > > > > >>>> > > > > > On Mon, Mar 2, 2026 at 1:35 PM ForwardXu <[email protected]> >>>> wrote: >>>> > > > > > >>>> > > > > >> Hi all, >>>> > > > > >> I think it makes perfect sense to create a dedicated roadmap >>>> for >>>> > > Lance >>>> > > > > >> support. This will help us clarify our priorities and ensure >>>> we >>>> > can >>>> > > > > deliver >>>> > > > > >> more comprehensive support, including advanced features like >>>> > complex >>>> > > > > data >>>> > > > > >> types and blob types, among others. >>>> > > > > >> Looking forward to discussing this further on Slack. >>>> > > > > >> >>>> > > > > >> Best, >>>> > > > > >> Forwardxu >>>> > > > > >> >>>> > > > > >> 原始邮件 >>>> > > > > >> ------------------------------ >>>> > > > > >> 发件人:Lorenzo Affetti via dev <[email protected]> >>>> > > > > >> 发件时间:2026年3月2日 18:48 >>>> > > > > >> 收件人:dev <[email protected]> >>>> > > > > >> 抄送:forwardxu <[email protected]>, Lorenzo Affetti < >>>> > > > > >> [email protected]> >>>> > > > > >> 主题:Re: Analysis of Lance storage format support >>>> > > > > >> >>>> > > > > >> Hello! Thanks for wrapping this up! >>>> > > > > >> >>>> > > > > >> I do understand both Cheng and Keith. >>>> > > > > >> For sure Lance support should be on par with other lake >>>> formats. >>>> > If >>>> > > > > >> something is not supported, there should be a concrete >>>> reason why >>>> > > > (apart >>>> > > > > >> from a lack of resources :) ). >>>> > > > > >> >>>> > > > > >> Still, input from the Lance community would be essential for >>>> > > > > >> understanding evolution areas of the support itself. >>>> > > > > >> >>>> > > > > >> For this item, I would take an approach similar to what >>>> Mehul did >>>> > > for >>>> > > > > >> Iceberg support. >>>> > > > > >> I think there is a lack of a roadmap for Lance support in >>>> 2026. >>>> > > > > >> >>>> > > > > >> Having a roadmap doesn't actually mean we will accomplish >>>> > > everything, >>>> > > > > but, >>>> > > > > >> it signals that we understand the problem space and have an >>>> idea >>>> > of >>>> > > > the >>>> > > > > >> sequence of actions to take. >>>> > > > > >> >>>> > > > > >> @cheng, I think you are the de-facto owner of the Lance >>>> module. >>>> > > > > >> Would it make sense to dedicate some of our resources to >>>> discuss >>>> > > this >>>> > > > > via >>>> > > > > >> Slack and start drafting a roadmap? >>>> > > > > >> >>>> > > > > >> On Sun, Mar 1, 2026 at 2:11 PM Keith Lee < >>>> > > [email protected] >>>> > > > > >>>> > > > > >> wrote: >>>> > > > > >> >>>> > > > > >> > Hello Cheng, >>>> > > > > >> > >>>> > > > > >> > Good call. I agree that gathering input from Lance >>>> community >>>> > will >>>> > > be >>>> > > > > >> >>>> > > > > >> > beneficial to inform integration of features such as vector >>>> > > search, >>>> > > > > vector >>>> > > > > >> > indexing and hybrid search. >>>> > > > > >> > >>>> > > > > >> >>>> > > > > >> > However, the issues I’ve outlined only meant to cover the >>>> scope >>>> > of >>>> > > > > bringing >>>> > > > > >> > current fluss lance integration up to parity to other >>>> lakehouses >>>> > > > like >>>> > > > > >> > paimon or iceberg e.g. batch or union read without lance >>>> feature >>>> > > > such >>>> > > > > as >>>> > > > > >> >>>> > > > > >> > vector search. As such, I believe these can be decoupled >>>> and we >>>> > > can >>>> > > > > have a >>>> > > > > >> >>>> > > > > >> > separate effort, gathering input from lance community and >>>> FIP >>>> > > > > proposal for >>>> > > > > >> > integrating vector search into feature such as union read. >>>> > > > > >> > >>>> > > > > >> > Let me know what your thoughts are on this. Thank you! >>>> > > > > >> > >>>> > > > > >> > Best regards >>>> > > > > >> > Keith Lee >>>> > > > > >> > >>>> > > > > >> > >>>> > > > > >> > On Sun, 1 Mar 2026 at 10:20, Cheng Wang <[email protected]> >>>> > wrote: >>>> > > > > >> > >>>> > > > > >> > > Hello Keith, >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> >>>> > > > > >> > > Regarding our plan to implement union read for Lance >>>> using >>>> > > Flink, >>>> > > > > might >>>> > > > > >> > it >>>> > > > > >> > > be beneficial to first gather input from the Lance >>>> community? >>>> > > > > >> > Understanding >>>> > > > > >> >>>> > > > > >> > > the primary scenarios where union read would help in the >>>> > machine >>>> > > > > learning >>>> > > > > >> > > scenario, along with the most popular execution engine in >>>> > Lance >>>> > > > > >> > ecosystem, >>>> > > > > >> > > could ensure we're building the right integration to >>>> maximize >>>> > > its >>>> > > > > >> > adoption. >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> > > Regards, >>>> > > > > >> > > Cheng Wang >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> > > ------------------ Original ------------------ >>>> > > > > >> > > From: >>>> > > > > >> > > "dev" >>>> > > > > >> > > >>>> > > < >>>> > > > > >> > > [email protected]>; >>>> > > > > >> > > Date: Sat, Feb 28, 2026 11:20 PM >>>> > > > > >> > > To: "dev"<[email protected]>; >>>> > > > > >> > > Cc: "Cheng Wang"<[email protected]>;"forwardxu"< >>>> > > > > >> > > [email protected]>; >>>> > > > > >> > > Subject: Re: Analysis of Lance storage format >>>> support >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> > > >>>> > > > > >> > > This is extremely helpful, thanks for putting this >>>> together. >>>> > > > > >> > > >>>> > > > > >> >>>> > > > > >> > > Maybe we can create an umbrella ticket on GitHub to keep >>>> track >>>> > > on >>>> > > > > these >>>> > > > > >> > and >>>> > > > > >> > > open individual tasks, for tracking. >>>> > > > > >> > > >>>> > > > > >> > > Best, >>>> > > > > >> > > Giannis >>>> > > > > >> > > >>>> > > > > >> > > On Sat, 28 Feb 2026 at 3:52 PM, Keith Lee < >>>> > > > > >> [email protected] >>>> > > > > >> > > >>>> > > > > >> > > wrote: >>>> > > > > >> > > >>>> > > > > >> > > > Hello, >>>> > > > > >> > > > >>>> > > > > >> >>>> > > > > >> > > > As discussed on community sync yesterday on >>>> analysing >>>> > where >>>> > > > we >>>> > > > > are >>>> > > > > >> > at >>>> > > > > >> > > the >>>> > > > > >> > > > moment in terms of Lance format support. >>>> > > > > >> > > > Here are my findings as part of working on Lance >>>> > QuickStart >>>> > > > > >> > > documentation >>>> > > > > >> > > > [1]. Lance lake tiering works in general, however >>>> there >>>> > are >>>> > > > > some >>>> > > > > >> > gaps >>>> > > > > >> > > that >>>> > > > > >> >>>> > > > > >> > > > to be addressed to bring Lance format support in >>>> parity >>>> > > with >>>> > > > > Paimon >>>> > > > > >> > / >>>> > > > > >> > > > Iceberg. >>>> > > > > >> > > > >>>> > > > > >> >>>> > > > > >> > > > - (Merged) Support for Arrow FixedSizeList to enable >>>> > > pylance >>>> > > > > native >>>> > > > > >> > > vector >>>> > > > > >> > > > search [2] >>>> > > > > >> > > > - (In progress) Support Flink SQL Union Read query >>>> > against >>>> > > > > Lance >>>> > > > > >> > > table [3] >>>> > > > > >> > > > - (Open) Support Flink SQL batch query against Lance >>>> > table >>>> > > > [4] >>>> > > > > >> > > > - (Blocked) Primary Key table support - I believe >>>> this is >>>> > > > still >>>> > > > > >> > > blocking on >>>> > > > > >> > > > Lance format support for delete API [5] >>>> > > > > >> > > > >>>> > > > > >> > > > Finally there is also a gap in the ability of >>>> performing >>>> > > > vector >>>> > > > > >> > > search on >>>> > > > > >> > > > hot data / via union read. After discussion with >>>> Mehul, >>>> > > > native >>>> > > > > >> > vector >>>> > > > > >> > > > indexing on hot data in Fluss would be a separate, >>>> bigger >>>> > > > > effort >>>> > > > > >> > that >>>> > > > > >> > > we >>>> > > > > >> > > > can evolve towards if there's demand for it. >>>> > > > > >> > > > >>>> > > > > >> > > > Appreciate feedback here from Cheng, Forward and >>>> anyone >>>> > > else >>>> > > > > with >>>> > > > > >> >>>> > > > > >> > > > familiarity around this area as I have only started >>>> > dipping >>>> > > > my >>>> > > > > toes >>>> > > > > >> > > into >>>> > > > > >> > > > Lance. >>>> > > > > >> > > > >>>> > > > > >> > > > *Additionally, if anyone wants to help contributing >>>> in >>>> > this >>>> > > > > area, >>>> > > > > >> > > please >>>> > > > > >> > > > reach out. * >>>> > > > > >> > > > >>>> > > > > >> > > > Best regards >>>> > > > > >> > > > Keith Lee >>>> > > > > >> > > > >>>> > > > > >> > > > Reference >>>> > > > > >> > > > [1] https://github.com/apache/fluss/pull/2716 >>>> > > > > >> > > > [2] https://github.com/apache/fluss/issues/2706 >>>> > > > > >> > > > [3] https://github.com/apache/fluss/issues/2715 >>>> > > > > >> > > > [4] https://github.com/apache/fluss/issues/2751 >>>> > > > > >> > > > [5] >>>> https://github.com/lance-format/lance/issues/3961 >>>> > > > > >> > > > >>>> > > > > >> > >>>> > > > > >> >>>> > > > > >> >>>> > > > > >> -- >>>> > > > > >> Lorenzo Affetti >>>> > > > > >> Senior Software Engineer @ Flink Team >>>> > > > > >> Ververica <http://www.ververica.com> >>>> > > > > >> >>>> > > > > >> >>>> > > > > >> >>>> > > > > > >>>> > > > > > -- >>>> > > > > > Lorenzo Affetti >>>> > > > > > Senior Software Engineer @ Flink Team >>>> > > > > > Ververica <http://www.ververica.com> >>>> > > > > > >>>> > > > > >>>> > > > >>>> > > >>>> > >>>> >>>> >>>> -- >>>> Lorenzo Affetti >>>> Senior Software Engineer @ Flink Team >>>> Ververica <http://www.ververica.com> >>>> >>> > > -- > Lorenzo Affetti > Senior Software Engineer @ Flink Team > Ververica <http://www.ververica.com> > -- Lorenzo Affetti Senior Software Engineer @ Flink Team Ververica <http://www.ververica.com>
