zy-kkk opened a new pull request, #68707:
URL: https://github.com/apache/doris/pull/68707

   ### What problem does this PR solve?
   
   Issue Number: close #68625
   
   Related PR: #68453 (time travel, tags, branches and managed versioning for 
Lance table scans)
   
   Problem Summary:
   
   `vector_search()` and `full_text_search()` always searched the latest 
version of `main`, and their `table` argument rejected `FOR VERSION AS OF`, 
`@tag` and `@branch`. A search could not be reproduced against a released tag, 
compared across versions, or run on an experimental branch.
   
   Both TVFs now take four optional properties, with the semantics of the table 
syntax from #68453:
   
   | Property | Value | Table syntax | Combines with |
   |-|-|-|-|
   | `version` | positive integer | `FOR VERSION AS OF N` | `branch` |
   | `timestamp` | `yyyy-MM-dd HH:mm:ss[.SSS]`, session time zone | `FOR TIME 
AS OF` | `branch` |
   | `tag` | tag name | `@tag(name)` | none |
   | `branch` | branch name, `main` is the main chain | `@branch(name)` | 
`version` or `timestamp` |
   
   ```sql
   SELECT id, _distance FROM vector_search(
       "table" = "lance_catalog.db.documents", "column" = "embedding",
       "query_vector" = "[0.1, 0.2, 0.3]", "version" = "3")
   ORDER BY _distance;
   
   SELECT id, _score FROM full_text_search(
       "table" = "lance_catalog.db.documents", "column" = "content", "query" = 
"database",
       "branch" = "dev", "timestamp" = "2026-09-20 12:00:00")
   ORDER BY _score DESC;
   ```
   
   How it works:
   
   - **FE.** The TVF parses the properties into a `LanceRefSelector` and loads 
its metadata through the same `LanceCatalogClient.readTableSnapshot` as table 
scans, so storage-versioned and namespace-managed tables follow the #68453 
rules unchanged. Schema, search column and field id, vector type and dimension, 
fragments and index segments all come from the selected snapshot, and planning 
never reopens the dataset. The splits carry the branch directory URI and the 
version, so no thrift change is needed.
   - **Each execution resolves again.** A bound TVF is never cached, prepared 
statements re-analyze on every `EXECUTE`, and TVF statements skip the SQL 
cache, so the latest version, a future time, a branch head and a moved tag are 
resolved per execution and stay fixed within its plan. `CREATE TABLE AS SELECT` 
plans its query twice (schema, then rows), so there a moving selector can 
resolve twice, as it already can for a table scan with an explicit selector; 
pinning `version` gives an exact copy. This is documented.
   - **No fallback.** A snapshot that cannot be selected, or whose index files 
are missing, is an error; nothing falls back to the latest version.
   - **EXPLAIN** of a search node now shows `lanceCatalogType`, 
`lanceManagedVersioning` and `lanceBranch` next to `lanceVersion`, as table 
scans do. Planner errors that name "dataset version N" also name the branch, 
since branch versions overlap main's.
   
   Decisions worth calling out:
   
   - `version` takes only a number. `FOR VERSION AS OF 'x'` names a tag because 
SQL has one clause for both; the TVF has a `tag` property, and Lance allows tag 
names made of digits, which `FOR VERSION AS OF` cannot reach.
   - An empty selector value is an error rather than "not set", so an empty 
client-side variable cannot silently search the latest version of `main`.
   - The selector properties must be constants. Accepting `?` would leave 
PREPARE without a snapshot to take the result schema from. `full_text_search` 
now rejects `?` in any property: until now a placeholder reached it as the 
literal text "?", so `'query'=?` searched for "?".
   - Each search relation resolves its own snapshot, as each explicit selector 
already does for table scans. `vector_search(t) JOIN t` can read two versions 
if a commit lands in between; pinning both with `version` / `FOR VERSION AS OF` 
avoids that. This is documented.
   - A managed table's branch is opened by a URI Doris joins 
(`<table>/tree/<branch>`), which Lance does not validate on that path, so Doris 
now applies Lance's branch-name rules there, with Rust's 
`char::is_alphanumeric` for letters and digits. A name such as `../other` or a 
percent-encoded dot segment is rejected before the branch is opened. Other 
tables check a branch out through Lance, which validates the name itself.
   
   BE changes:
   
   - A search split or a two-phase row fetch without a positive version is 
rejected instead of opening the latest version.
   - Errors from opening the dataset, reading batches, preparing the FTS query 
context and the two-phase row fetch name the snapshot: "... at Lance dataset 
version N at <uri>", with the query string and any password removed from the 
URI. Missing index files surface at execution because Lance loads indexes 
lazily.
   - Fragment ids from the FE must exist in the opened snapshot. lance-c 
silently skips unknown fragment ids, which would drop rows without an error.
   - The profile records `LanceDatasetVersion` and `LanceDatasetUri`, the 
snapshot an execution actually read. EXPLAIN cannot show this for moving 
selectors because it resolves again.
   - **Behavior change for every vector search:** a split without index 
segments now always runs flat (`use_index=false`). The FE plans such splits as 
flat search, and EXPLAIN and the profile count them as flat, but Lance used to 
search them with the first index on the column in manifest order, which need 
not be the index the FE selected (for example while an index is rebuilt under a 
new name). Indexes the FE cannot plan with (no fragment coverage metadata, no 
recorded metric, or a legacy index whose segments all lack index details, which 
the FE skips; `lanceVectorIndexStatus=UNKNOWN_COVERAGE` / `METRIC_MISMATCH` / 
`NO_MATCH`) are therefore no longer used implicitly, and such searches become 
exact flat searches. Results can change only where the unplanned index was 
approximate; latency can grow on such datasets.
   - A vector split that carries index segments while the request says 
`use_index=false` is rejected. The FE plans segments only when `use_index` 
allows them, so it never sends such a split; the BE now takes `use_index` from 
the segments and would otherwise have to override either the plan or the 
request.
   
   Known limitations (documented):
   
   - A branch of a branch, or a branch of a shallow clone, can fail to find the 
index files it inherited, because Lance writers before v13 record them under 
the intermediate branch (lance-format/lance#9176). The search fails with an 
error.
   - After a branch is deleted and recreated under the same name, or a dataset 
is replaced at the same URI, version numbers repeat. Since #68615 the BE's 
Lance keys the index list and the row-id index by the manifest e_tag 
(lance-format/lance#8904, #8930), but the BE's process-wide Lance session still 
keys two entries by version alone, until eviction or restart. On datasets with 
stable row ids, the live-row mask that vector and full-text search prefilter 
with can come from the old data, so rows deleted in the new data can reappear 
or live rows can be missing (lance-format/lance#9079, open). For a manifest 
without an e_tag (V1 naming, or a store that returns no ETag), the index list 
can come from the old data, on the BE and in the FE's Lance session until the 
catalog is refreshed. Neither is introduced by this PR.
   - `SHOW INDEX` and `lance_index_entries()` keep showing the latest version 
of `main`.
   
   ### Release note
   
   `vector_search()` and `full_text_search()` accept `version`, `timestamp`, 
`tag` and `branch` to search a historical version, the version at a time, a tag 
or a branch of a Lance table, including REST Namespace managed tables. Vector 
search no longer uses an index the planner did not select, and 
`full_text_search()` rejects prepared-statement placeholders.
   
   ### Check List (For Author)
   
   - Test: Regression test / Unit Test
       - Regression: `external_table_p0/lance/test_lance_search_snapshot` 
(filesystem catalog) and `test_lance_rest_search_snapshot` (REST, 
storage-versioned and managed) on a new fixture `search_snapshot.lance`. It has 
seven versions (appends, a vector index, an FTS index, an append no index 
covers, a delete, a vector index rebuild), a tag, and a branch `dev` whose 
versions and fragment ids overlap main's with different rows, plus a tag at a 
non-latest branch version. `search_snapshot_pruned.lance` lost a version to 
cleanup and the index files of a tagged version. 
`search_snapshot_evolved.lance` adds a column and renames the vector column 
after its index is built, then appends data the index does not cover. The 
suites check rows, distances, two-phase reads (on a branch, a tag, the columns 
of a version before and after a rename, and a managed table through REST), 
EXPLAIN, full-text search across a delete, `use_index=false` with a selector, 
and every selector error. `test_lance_rest_c
 atalog.out` and `test_lance_rest_time_travel.out` list the new REST tables. 
All 29 suites in `external_table_p0/lance` pass locally.
       - FE unit tests: `LanceSearchSelectorTest` (property rules, time zones, 
selector forwarding), `LancePreparedSearchTest` (both TVFs pass their selector 
to the table; each analysis resolves it exactly once, and a moved tag is 
resolved again on the next `EXECUTE`; search statements bypass the SQL cache; 
placeholders rejected), `LanceScanNodeTest` (EXPLAIN, errors naming the 
branch), `LanceCatalogClientTest` (branch names, including the Unicode letters 
and digits Lance accepts), and, with real datasets and locally only like the 
other JNI tests, `LanceSearchSnapshotTest` (historical index coverage, 
branches, tags, schema evolution, cleanup, a snapshot planned before later 
commits) and `LanceManagedVersioningTest` (managed versions, tags, branches, 
unfinalized heads through the search entry points).
       - BE unit tests in `lance_reader_test.cpp`: missing fragment ids on 
normal, vector and FTS splits, unpinned versions, error and profile contents, a 
split without index segments searching flat although an index covers it, and a 
split with index segments under `use_index=false` being rejected. The indexed 
variants of the existing multi-vector tests now pass the index segments the FE 
would plan.
   - Behavior changed: Yes. New TVF properties; vector splits without index 
segments always search flat; `full_text_search` rejects `?`; managed branch 
names are validated; the error for a time in removed history now reads "Lance 
cannot select the version at or before '...'" instead of "Lance cannot resolve 
FOR TIME AS OF '...'", because the TVF reaches it too; BE errors name the 
dataset version and URI.
   - Does this need documentation: Yes (apache/doris-website#4185)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to