Hi , If someone got a chance to look at the poc, Please let me know if this is at a stage when I can create an fip for this.
Sagar. On Sat, 15 Aug 2026 at 9:54 PM, Sagar <[email protected]> wrote: > Thanks Mehul! > > I Wanted to share an update on the VECTOR data type for Fluss. This PR has > an e2e POC for the same: > > https://github.com/apache/fluss/pull/4004. > > It mainly contains an IT for Lance Vector Tiering > (LanceVectorTieringITCase) with 2 tests: testVectorTiering and > testNullVectorTiering. I am seeing some build errors on fluss-spark but the > IT seems to work locally. At a high level, it: > > - Writes records with vector embeddings (VECTOR(dim)) into a Fluss > table. > - A Flink background job reads those vector records from Fluss and > writes them into Lance storage. > - The test opens the resulting Lance dataset directly and checks that > every vector, float value, row count, and null field made it through > accurately without corruption or data loss. > > While this seems to work, there are a couple of things worth calling out: > > 1) Flink SQL Limitation: Flink SQL doesn't have a native VECTOR data type > (it treats embeddings as ARRAY<FLOAT>), so vector columns aren't exposed as > a distinct VECTOR type in SQL queries. Under the hood, Fluss models this as > a dedicated VECTOR(dim) type mapped directly to Arrow's > FixedSizeListVector<Float32>. This avoids the extra overhead of > variable-sized lists and aligns 1:1 with Lance's native vector storage > format. > > 2) Multi-Dimensional Array Limitation > Currently, we only support 1D fixed-size float vectors. Multi-dimensional > arrays (like matrices or nested arrays ARRAY<ARRAY<FLOAT>>) are not > supported for vector tiering yet. > > To support that, we would need more work across schema conversion to Lance > multi-level list format and even more things on Arrow converters. We can > take a look at it later i think. > > Please take a look and see if it is at a point where we can introduce a > FIP. I think it would be similar to this doc: > > > https://docs.google.com/document/d/1idmsgMLjYScgYj-l_ABD7rwvNbDvoaq8bbMgWcgZW20/edit?tab=t.0 > > Sagar. > > > On Sat, Jul 4, 2026 at 5:45 PM Mehul Batra <[email protected]> > wrote: > >> Hi Sagar, >> >> Thanks for the detailed findings and document, really appreciate the work >> here. >> >> I think this is the right direction as discussed on the community call, >> let’s keep phase one focused on just the vector data type. Excited to see >> the POC take shape, and a FIP to discuss the findings sounds like a great >> next step. >> >> A few things to keep in mind from our thread & discussions: >> >> • Fixed dimension + fixed element type at schema time (VECTOR(1536)) >> >> • Arrow-native layout (FixedSizeList<Float32>, zero-copy with >> Lance/Paimon) >> >> • Support FLOAT32, FLOAT16, INT8 for quantization from day one so we avoid >> a breaking change later. >> >> Best Regards, >> >> Mehul Batra >> On Sun, Jun 28, 2026 at 10:21 AM Sagar <[email protected]> wrote: >> >> > Hi, >> > >> > Following up again to see if there are any further comments or feedback. >> > >> > Sagar. >> > >> > On Tue, 16 Jun 2026 at 11:09 PM, Sagar <[email protected]> >> wrote: >> > >> > > Hi Jark >> > > >> > > Thanks for the detailed feedback! Please find my responses: >> > > >> > > >> > > *Point 1* >> > > >> > > *Property-to-API mapping and unexposed parameters* — Added a mapping >> > > table to Public Interfaces covering all three API calls. For >> parameters >> > > Fluss doesn't expose: replace is fixed True on both create_index() and >> > > create_fts_index() — any other value would break the idempotency the >> > > state machine relies on. use_tantivy is fixed False (see below). >> > num_bits, >> > > delete_unverified, and retrain are not exposed; LanceDB defaults >> apply. >> > > Since we route through JNI to lance-core (discussed below) rather than >> > the >> > > Python client, parameter semantics are identical to the Python >> > equivalents >> > > shown in the example. >> > > >> > > *FTS path* — Fluss targets the native FTS path (use_tantivy=False), >> now >> > > the LanceDB upstream default. The legacy Tantivy path is not >> supported; >> > it >> > > differs in both parameter surface and on-disk format. This is fixed at >> > the >> > > JNI layer. >> > > >> > > *Default divergences* — Checked against the LanceDB docs [1]: all FTS >> > > defaults in the FIP are identical to LanceDB's defaults, so no >> rationale >> > is >> > > needed there. The only divergences are on the vector side: >> > ef_construction >> > > (Fluss: 150, LanceDB: 300) and lance.index.m=(none), which maps to >> > > LanceDB's hardcoded 20. Both are now documented with rationale. >> > > >> > > *Side-by-side example* — Added below the existing full-configuration >> SQL >> > > block. >> > > >> > > *Point 2 — Execution model, LanceDB embedding, and horizontal scaling* >> > > >> > > *Where does the committer run?* >> > > >> > > Inside the Tiering Service worker, consistent with FIP-5. Modified the >> > FIP >> > > to update this >> > > >> > > *How is LanceDB embedded in the JVM?* >> > > >> > > com.lancedb:lance-core is a first-party JNI binding from the >> > > lance-format/lance monorepo — already a dependency in Fluss's existing >> > > Lance integration from FIP-5, not a new one introduced here. >> > > >> > > I verified the published 0.39.0 JAR directly. createIndex and >> listIndexes >> > > are present. The one gap is optimizeIndices — needed to fold newly >> > > written rows into existing indices after each tiering cycle. The JNI >> > > pattern is established in the codebase by nativeCreateIndex; >> contributing >> > > nativeOptimizeIndices is a single function addition in >> > > java/lance-jni/src/dataset.rs with a corresponding method pair in >> > > Dataset.java. This is a committed prerequisite of the FIP-44 >> > > implementation. No sidecar, no subprocess. FIP is updated with this >> > detail. >> > > >> > > Similarly, the APIs to add an FTS index also seem missing in the jni >> > > binding. We will need to add those as well. >> > > >> > > *How is index work distributed?* >> > > >> > > Per-table, scoped to the committer owning that table. Horizontal >> scaling >> > > is at table granularity. >> > > >> > > The single-table pinning concern is real but bounded: createIndex is >> > > non-blocking — the build runs async inside LanceDB's Tokio runtime, >> the >> > > committer thread is released immediately and polls listIndexes() on >> > > subsequent timer fires. The build itself is internally multi-threaded. >> > The >> > > constraint is cross-JVM-process parallelism, not single-threading. A >> > *Scaling >> > > Constraints* note will be added to the FIP. Coordinator-assigned index >> > > builds are a reasonable future extension but out of scope here. This >> is >> > > also added to the FIP. >> > > >> > > >> > > *3. Configuration drift after the table exists* >> > > >> > > For the initial scope of FIP-44, we will take the *'reject at DDL >> time'* >> > > approach. >> > > >> > > Index configurations will be treated as immutable once the index state >> > > enters IN_PROGRESS or COMPLETED. If a user attempts to modify >> properties >> > > like lance.index.type, metric, or num_partitions via ALTER TABLE, the >> DDL >> > > validator will reject it. >> > > >> > > *Rationale:* This keeps the FIP-44 state machine strictly linear >> (ABSENT >> > > → IN_PROGRESS → COMPLETED). It avoids the complexities of modeling >> > > PENDING_REBUILD states and protects Tiering workers from accidentally >> > > triggering massive background rebuilds due to a simple property tweak. >> > > Declarative background rebuilds for config drift can be tackled in a >> > future >> > > FIP. I will update the document to explicitly state this constraint >> > > >> > > Let me know what you think! >> > > >> > > >> > > Sagar. >> > > >> > > [1]: https://docs.lancedb.com/search/full-text-search#advanced-usage >> > > >> > > On Sun, May 31, 2026 at 10:34 AM Jark Wu <[email protected]> wrote: >> > > >> > >> Hi Sagar, >> > >> >> > >> Thanks for the detailed FIP. Three comments below. >> > >> >> > >> ## 1. Public-interface docs need a mapping and a worked example >> > >> >> > >> The `lance.*` properties currently stand alone in the FIP — to >> > >> understand any of them, a reader has to cross-reference the LanceDB >> > >> docs. I'd like the FIP to add three things to the public-interface >> > >> section: >> > >> >> > >> - An explicit table mapping each Fluss property to the LanceDB API >> > >> call and parameter it maps to (e.g. `lance.index.type` → >> > >> `Table.create_index(index_type=...)`). >> > >> - An explicit mapping of each Fluss default to the corresponding >> > >> LanceDB default, with rationale for any deliberate divergence. >> > >> Skimming the FIP, several `lance.fts.*` defaults look like they >> differ >> > >> from LanceDB upstream defaults (e.g. `stem`, `remove_stop_words`, >> > >> `ascii_folding`), and `lance.index.m`'s `(none)` effectively means >> > >> LanceDB's hardcoded `20`. The reasons aren't stated. >> > >> - A side-by-side example showing the same index expressed as (a) a >> > >> Fluss `CREATE TABLE ... WITH (...)` statement, and (b) the equivalent >> > >> LanceDB Python call. That makes the abstraction concrete for both >> > >> reviewers and future users. >> > >> >> > >> Two specific things worth pinning down while you're in there: >> > >> >> > >> - `create_fts_index` has a legacy Tantivy path and a newer native FTS >> > >> path (`use_tantivy=False`, now the upstream default). Which one is >> the >> > >> FIP targeting? The parameter surface and on-disk format both differ. >> > >> - `Table.create_index` and `Table.optimize` have additional >> parameters >> > >> (`replace`, `num_bits`, `delete_unverified`, `retrain`, …) that >> aren't >> > >> currently mapped. Either include them or explain why they're >> > >> deliberately hidden — `replace` in particular matters because the >> > >> state machine relies on `create_index` being idempotent, which is >> only >> > >> true with `replace=True`. >> > >> >> > >> ## 2. Who builds the index? Execution model and horizontal scaling >> > >> >> > >> The FIP assigns the index lifecycle to the `LanceLakeCommitter`, but >> > >> the deeper execution-model question is not yet answered: >> > >> >> > >> - Where does the committer (and therefore `create_index()` / >> > >> `optimize()`) physically run? My reading of FIP-5 is that the >> > >> committer lives inside the Tiering Service workers. Is that the >> intent >> > >> here? >> > >> >> > >> - If so, the Tiering Service now has to **embed LanceDB**. LanceDB is >> > >> a Rust core with Python and Node bindings — there is no first-party >> > >> Java client today. How is it embedded into the JVM-based tiering >> > >> worker? JNI over the Rust core? A sidecar subprocess? Something else? >> > >> This is a non-trivial dependency to take on and deserves explicit >> > >> discussion in the FIP. >> > >> >> > >> - How is index work **distributed** across Tiering Service workers? >> > >> Per-table affinity? Coordinator-assigned? With a single large table >> > >> whose one heavy index takes hours to build, does the work pin to one >> > >> worker, or can it be split? If the asynchronous build effectively >> runs >> > >> in-process inside the worker that initiated it, then horizontal >> > >> scaling is per-table at best. >> > >> >> > >> >> > >> ## 3. Configuration drift after the table exists >> > >> >> > >> What happens if a user changes `lance.index.type` (or `metric`, >> > >> `num_partitions`, …) on a table that already has a COMPLETED index? >> > >> The state machine only models `ABSENT → IN_PROGRESS → COMPLETED`, >> with >> > >> no "config changed, rebuild" transition. We need an explicit answer >> > >> here — silently keep the old index, force a rebuild, or reject the >> > >> property change at DDL time. Each option has different operational >> > >> implications and the FIP should commit to one. >> > >> >> > >> Looking forward to your thoughts. >> > >> >> > >> Best, >> > >> Jark >> > >> >> > >> On Fri, 29 May 2026 at 22:41, Sagar <[email protected]> >> wrote: >> > >> > >> > >> > Hi , >> > >> > >> > >> > Bumping this thread. Please take a look. >> > >> > >> > >> > Sagar. >> > >> > >> > >> > On Sat, 23 May 2026 at 9:53 AM, Sagar <[email protected]> >> > >> wrote: >> > >> > >> > >> > > Hi, >> > >> > > >> > >> > > I created FIP-44 >> > >> > > < >> > >> >> > >> https://cwiki.apache.org/confluence/pages/viewpage.action?pageId=429064608 >> > > >> > >> to >> > >> > > enhance the LanceDB integration with Fluss. >> > >> > > >> > >> > > Please review. >> > >> > > >> > >> > > Sagar. >> > >> > > >> > >> >> > > >> > >> >
