Hi Sagar, And thanks for tackling this.. Currently everyone is super busy with the release coming these days. In case no one has time to take a look within then next few days, allow me a week or two when I will be fully back to look into this, because itβs something Iβm excited about.
Apologies for the delays and do know your work is highly appreciated π We can also sync on the upcoming Fluss community call. Best, Giannis On Sat, 22 Aug 2026 at 4:16β―AM, Sagar <[email protected]> wrote: > Hi , > > If someone got a chance to look at the poc, Please let me know if this is > at a stage when I can create an fip for this. > > Sagar. > > On Sat, 15 Aug 2026 at 9:54β―PM, Sagar <[email protected]> wrote: > > > Thanks Mehul! > > > > I Wanted to share an update on the VECTOR data type for Fluss. This PR > has > > an e2e POC for the same: > > > > https://github.com/apache/fluss/pull/4004. > > > > It mainly contains an IT for Lance Vector Tiering > > (LanceVectorTieringITCase) with 2 tests: testVectorTiering and > > testNullVectorTiering. I am seeing some build errors on fluss-spark but > the > > IT seems to work locally. At a high level, it: > > > > - Writes records with vector embeddings (VECTOR(dim)) into a Fluss > > table. > > - A Flink background job reads those vector records from Fluss and > > writes them into Lance storage. > > - The test opens the resulting Lance dataset directly and checks that > > every vector, float value, row count, and null field made it through > > accurately without corruption or data loss. > > > > While this seems to work, there are a couple of things worth calling out: > > > > 1) Flink SQL Limitation: Flink SQL doesn't have a native VECTOR data type > > (it treats embeddings as ARRAY<FLOAT>), so vector columns aren't exposed > as > > a distinct VECTOR type in SQL queries. Under the hood, Fluss models this > as > > a dedicated VECTOR(dim) type mapped directly to Arrow's > > FixedSizeListVector<Float32>. This avoids the extra overhead of > > variable-sized lists and aligns 1:1 with Lance's native vector storage > > format. > > > > 2) Multi-Dimensional Array Limitation > > Currently, we only support 1D fixed-size float vectors. Multi-dimensional > > arrays (like matrices or nested arrays ARRAY<ARRAY<FLOAT>>) are not > > supported for vector tiering yet. > > > > To support that, we would need more work across schema conversion to > Lance > > multi-level list format and even more things on Arrow converters. We can > > take a look at it later i think. > > > > Please take a look and see if it is at a point where we can introduce a > > FIP. I think it would be similar to this doc: > > > > > > > https://docs.google.com/document/d/1idmsgMLjYScgYj-l_ABD7rwvNbDvoaq8bbMgWcgZW20/edit?tab=t.0 > > > > Sagar. > > > > > > On Sat, Jul 4, 2026 at 5:45β―PM Mehul Batra <[email protected]> > > wrote: > > > >> Hi Sagar, > >> > >> Thanks for the detailed findings and document, really appreciate the > work > >> here. > >> > >> I think this is the right direction as discussed on the community call, > >> letβs keep phase one focused on just the vector data type. Excited to > see > >> the POC take shape, and a FIP to discuss the findings sounds like a > great > >> next step. > >> > >> A few things to keep in mind from our thread & discussions: > >> > >> β’ Fixed dimension + fixed element type at schema time (VECTOR(1536)) > >> > >> β’ Arrow-native layout (FixedSizeList<Float32>, zero-copy with > >> Lance/Paimon) > >> > >> β’ Support FLOAT32, FLOAT16, INT8 for quantization from day one so we > avoid > >> a breaking change later. > >> > >> Best Regards, > >> > >> Mehul Batra > >> On Sun, Jun 28, 2026 at 10:21β―AM Sagar <[email protected]> > wrote: > >> > >> > Hi, > >> > > >> > Following up again to see if there are any further comments or > feedback. > >> > > >> > Sagar. > >> > > >> > On Tue, 16 Jun 2026 at 11:09β―PM, Sagar <[email protected]> > >> wrote: > >> > > >> > > Hi Jark > >> > > > >> > > Thanks for the detailed feedback! Please find my responses: > >> > > > >> > > > >> > > *Point 1* > >> > > > >> > > *Property-to-API mapping and unexposed parameters* β Added a mapping > >> > > table to Public Interfaces covering all three API calls. For > >> parameters > >> > > Fluss doesn't expose: replace is fixed True on both create_index() > and > >> > > create_fts_index() β any other value would break the idempotency the > >> > > state machine relies on. use_tantivy is fixed False (see below). > >> > num_bits, > >> > > delete_unverified, and retrain are not exposed; LanceDB defaults > >> apply. > >> > > Since we route through JNI to lance-core (discussed below) rather > than > >> > the > >> > > Python client, parameter semantics are identical to the Python > >> > equivalents > >> > > shown in the example. > >> > > > >> > > *FTS path* β Fluss targets the native FTS path (use_tantivy=False), > >> now > >> > > the LanceDB upstream default. The legacy Tantivy path is not > >> supported; > >> > it > >> > > differs in both parameter surface and on-disk format. This is fixed > at > >> > the > >> > > JNI layer. > >> > > > >> > > *Default divergences* β Checked against the LanceDB docs [1]: all > FTS > >> > > defaults in the FIP are identical to LanceDB's defaults, so no > >> rationale > >> > is > >> > > needed there. The only divergences are on the vector side: > >> > ef_construction > >> > > (Fluss: 150, LanceDB: 300) and lance.index.m=(none), which maps to > >> > > LanceDB's hardcoded 20. Both are now documented with rationale. > >> > > > >> > > *Side-by-side example* β Added below the existing full-configuration > >> SQL > >> > > block. > >> > > > >> > > *Point 2 β Execution model, LanceDB embedding, and horizontal > scaling* > >> > > > >> > > *Where does the committer run?* > >> > > > >> > > Inside the Tiering Service worker, consistent with FIP-5. Modified > the > >> > FIP > >> > > to update this > >> > > > >> > > *How is LanceDB embedded in the JVM?* > >> > > > >> > > com.lancedb:lance-core is a first-party JNI binding from the > >> > > lance-format/lance monorepo β already a dependency in Fluss's > existing > >> > > Lance integration from FIP-5, not a new one introduced here. > >> > > > >> > > I verified the published 0.39.0 JAR directly. createIndex and > >> listIndexes > >> > > are present. The one gap is optimizeIndices β needed to fold newly > >> > > written rows into existing indices after each tiering cycle. The JNI > >> > > pattern is established in the codebase by nativeCreateIndex; > >> contributing > >> > > nativeOptimizeIndices is a single function addition in > >> > > java/lance-jni/src/dataset.rs with a corresponding method pair in > >> > > Dataset.java. This is a committed prerequisite of the FIP-44 > >> > > implementation. No sidecar, no subprocess. FIP is updated with this > >> > detail. > >> > > > >> > > Similarly, the APIs to add an FTS index also seem missing in the jni > >> > > binding. We will need to add those as well. > >> > > > >> > > *How is index work distributed?* > >> > > > >> > > Per-table, scoped to the committer owning that table. Horizontal > >> scaling > >> > > is at table granularity. > >> > > > >> > > The single-table pinning concern is real but bounded: createIndex is > >> > > non-blocking β the build runs async inside LanceDB's Tokio runtime, > >> the > >> > > committer thread is released immediately and polls listIndexes() on > >> > > subsequent timer fires. The build itself is internally > multi-threaded. > >> > The > >> > > constraint is cross-JVM-process parallelism, not single-threading. A > >> > *Scaling > >> > > Constraints* note will be added to the FIP. Coordinator-assigned > index > >> > > builds are a reasonable future extension but out of scope here. This > >> is > >> > > also added to the FIP. > >> > > > >> > > > >> > > *3. Configuration drift after the table exists* > >> > > > >> > > For the initial scope of FIP-44, we will take the *'reject at DDL > >> time'* > >> > > approach. > >> > > > >> > > Index configurations will be treated as immutable once the index > state > >> > > enters IN_PROGRESS or COMPLETED. If a user attempts to modify > >> properties > >> > > like lance.index.type, metric, or num_partitions via ALTER TABLE, > the > >> DDL > >> > > validator will reject it. > >> > > > >> > > *Rationale:* This keeps the FIP-44 state machine strictly linear > >> (ABSENT > >> > > β IN_PROGRESS β COMPLETED). It avoids the complexities of modeling > >> > > PENDING_REBUILD states and protects Tiering workers from > accidentally > >> > > triggering massive background rebuilds due to a simple property > tweak. > >> > > Declarative background rebuilds for config drift can be tackled in a > >> > future > >> > > FIP. I will update the document to explicitly state this constraint > >> > > > >> > > Let me know what you think! > >> > > > >> > > > >> > > Sagar. > >> > > > >> > > [1]: > https://docs.lancedb.com/search/full-text-search#advanced-usage > >> > > > >> > > On Sun, May 31, 2026 at 10:34β―AM Jark Wu <[email protected]> wrote: > >> > > > >> > >> Hi Sagar, > >> > >> > >> > >> Thanks for the detailed FIP. Three comments below. > >> > >> > >> > >> ## 1. Public-interface docs need a mapping and a worked example > >> > >> > >> > >> The `lance.*` properties currently stand alone in the FIP β to > >> > >> understand any of them, a reader has to cross-reference the LanceDB > >> > >> docs. I'd like the FIP to add three things to the public-interface > >> > >> section: > >> > >> > >> > >> - An explicit table mapping each Fluss property to the LanceDB API > >> > >> call and parameter it maps to (e.g. `lance.index.type` β > >> > >> `Table.create_index(index_type=...)`). > >> > >> - An explicit mapping of each Fluss default to the corresponding > >> > >> LanceDB default, with rationale for any deliberate divergence. > >> > >> Skimming the FIP, several `lance.fts.*` defaults look like they > >> differ > >> > >> from LanceDB upstream defaults (e.g. `stem`, `remove_stop_words`, > >> > >> `ascii_folding`), and `lance.index.m`'s `(none)` effectively means > >> > >> LanceDB's hardcoded `20`. The reasons aren't stated. > >> > >> - A side-by-side example showing the same index expressed as (a) a > >> > >> Fluss `CREATE TABLE ... WITH (...)` statement, and (b) the > equivalent > >> > >> LanceDB Python call. That makes the abstraction concrete for both > >> > >> reviewers and future users. > >> > >> > >> > >> Two specific things worth pinning down while you're in there: > >> > >> > >> > >> - `create_fts_index` has a legacy Tantivy path and a newer native > FTS > >> > >> path (`use_tantivy=False`, now the upstream default). Which one is > >> the > >> > >> FIP targeting? The parameter surface and on-disk format both > differ. > >> > >> - `Table.create_index` and `Table.optimize` have additional > >> parameters > >> > >> (`replace`, `num_bits`, `delete_unverified`, `retrain`, β¦) that > >> aren't > >> > >> currently mapped. Either include them or explain why they're > >> > >> deliberately hidden β `replace` in particular matters because the > >> > >> state machine relies on `create_index` being idempotent, which is > >> only > >> > >> true with `replace=True`. > >> > >> > >> > >> ## 2. Who builds the index? Execution model and horizontal scaling > >> > >> > >> > >> The FIP assigns the index lifecycle to the `LanceLakeCommitter`, > but > >> > >> the deeper execution-model question is not yet answered: > >> > >> > >> > >> - Where does the committer (and therefore `create_index()` / > >> > >> `optimize()`) physically run? My reading of FIP-5 is that the > >> > >> committer lives inside the Tiering Service workers. Is that the > >> intent > >> > >> here? > >> > >> > >> > >> - If so, the Tiering Service now has to **embed LanceDB**. LanceDB > is > >> > >> a Rust core with Python and Node bindings β there is no first-party > >> > >> Java client today. How is it embedded into the JVM-based tiering > >> > >> worker? JNI over the Rust core? A sidecar subprocess? Something > else? > >> > >> This is a non-trivial dependency to take on and deserves explicit > >> > >> discussion in the FIP. > >> > >> > >> > >> - How is index work **distributed** across Tiering Service workers? > >> > >> Per-table affinity? Coordinator-assigned? With a single large table > >> > >> whose one heavy index takes hours to build, does the work pin to > one > >> > >> worker, or can it be split? If the asynchronous build effectively > >> runs > >> > >> in-process inside the worker that initiated it, then horizontal > >> > >> scaling is per-table at best. > >> > >> > >> > >> > >> > >> ## 3. Configuration drift after the table exists > >> > >> > >> > >> What happens if a user changes `lance.index.type` (or `metric`, > >> > >> `num_partitions`, β¦) on a table that already has a COMPLETED index? > >> > >> The state machine only models `ABSENT β IN_PROGRESS β COMPLETED`, > >> with > >> > >> no "config changed, rebuild" transition. We need an explicit answer > >> > >> here β silently keep the old index, force a rebuild, or reject the > >> > >> property change at DDL time. Each option has different operational > >> > >> implications and the FIP should commit to one. > >> > >> > >> > >> Looking forward to your thoughts. > >> > >> > >> > >> Best, > >> > >> Jark > >> > >> > >> > >> On Fri, 29 May 2026 at 22:41, Sagar <[email protected]> > >> wrote: > >> > >> > > >> > >> > Hi , > >> > >> > > >> > >> > Bumping this thread. Please take a look. > >> > >> > > >> > >> > Sagar. > >> > >> > > >> > >> > On Sat, 23 May 2026 at 9:53β―AM, Sagar <[email protected] > > > >> > >> wrote: > >> > >> > > >> > >> > > Hi, > >> > >> > > > >> > >> > > I created FIP-44 > >> > >> > > < > >> > >> > >> > > >> > https://cwiki.apache.org/confluence/pages/viewpage.action?pageId=429064608 > >> > > > >> > >> to > >> > >> > > enhance the LanceDB integration with Fluss. > >> > >> > > > >> > >> > > Please review. > >> > >> > > > >> > >> > > Sagar. > >> > >> > > > >> > >> > >> > > > >> > > >> > > >
