This is an automated email from the ASF dual-hosted git repository.

Gabriel39 pushed a commit to branch master
in repository https://gitbox.apache.org/repos/asf/doris-website.git


The following commit(s) were added to refs/heads/master by this push:
     new 4024df5fd79 [docs] Add Lance hybrid search best practices and correct 
type mappings (#4182)
4024df5fd79 is described below

commit 4024df5fd795c1828853dd38d22a829cb8b3b193
Author: Gabriel <[email protected]>
AuthorDate: Wed Sep 30 11:46:17 2026 +0800

    [docs] Add Lance hybrid search best practices and correct type mappings 
(#4182)
    
    Filtered Lance retrieval can repeat scalar-index work when scalar
    segments cross search-split boundaries. Add an English/Chinese guide
    explaining filter placement, physical-segment split planning,
    fragment-coverage alignment, construction and maintenance workflows,
    shared-cache pressure, and profile diagnosis.
    
    The guide distinguishes fragment-scoped V2 scalar pruning from legacy
    reader paths, treats segment sizing as a workload experiment, and
    explains why logical indexes do not automatically align different
    columns. It includes filtered vector/FTS SQL and a clearly labeled SDK
    orchestration pattern with separate per-index commits.
    
    Also correct the Catalog type mapping to document Arrow/Lance JSON as
    Doris JSON, together with the already-supported Duration, nullable Null,
    and BFloat16 mappings. Preserve storage-type checks,
    unsupported-extension restrictions, and the distinction between JSON
    extensions and ordinary strings containing JSON text. Remove JSON from
    the unsupported-column exclusion example. Type behavior was checked
    against the current branch-4.1 converter and reader, including
    apache/doris#67325.
    
    Validation:
    - Docusaurus 3.6.3 MDX compilation passed for all six touched pages.
    - Front matter, Markdown structure, new-page SEO, targeted
    links/anchors, and diff checks passed.
    - Checked bilingual SQL/Python examples, Python syntax, metric names,
    heading anchors, sidebar JSON/reference, and sensitive-data absence in
    additions.
    - The i18n checker reports only missing `current` counterparts. These
    pages intentionally follow the existing 4.x-only Lance Catalog placement
    and retain its documented experimental/4.2 availability boundary.
    - No full-site build, dataset execution, or cluster benchmark was
    performed. No performance improvement is claimed as measured.
    
    Self-review:
    - Goal: covers filtered retrieval tuning and aligned construction;
    corrects the requested JSON support statement.
    - Scope: two new localized pages, focused Catalog corrections, entry
    links, and one sidebar item.
    - Information architecture: follows the existing 4.x Lakehouse Best
    Practices layout; locales are synchronized.
    - Links/navigation: existing URLs and headings preserved; new page
    linked from the sidebar, Catalog, and cache guide.
    - UI/configuration: no component, styling, routing configuration, or
    build-script changes; sidebar JSON remains valid.
    - Verification: targeted checks above cover document parsing and
    navigation; runtime validation limits are explicit.
    - Other issues: no additional correctness, accessibility, SEO, or
    maintainability issues found in self-review.
---
 .../lakehouse/best-practices/doris-lance.mdx       |   2 +
 .../best-practices/lance-hybrid-search.mdx         | 193 +++++++++++++++++++++
 .../lakehouse/catalogs/lance-catalog.mdx           |  17 +-
 .../lakehouse/best-practices/doris-lance.mdx       |   2 +
 .../best-practices/lance-hybrid-search.mdx         | 193 +++++++++++++++++++++
 .../lakehouse/catalogs/lance-catalog.mdx           |  17 +-
 versioned_sidebars/version-4.x-sidebars.json       |   1 +
 7 files changed, 417 insertions(+), 8 deletions(-)

diff --git 
a/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/best-practices/doris-lance.mdx
 
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/best-practices/doris-lance.mdx
index acf509f5c27..b3d8c061c00 100644
--- 
a/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/best-practices/doris-lance.mdx
+++ 
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/best-practices/doris-lance.mdx
@@ -12,6 +12,8 @@
 Lance Catalog 是实验性功能,从 Apache Doris 4.2 开始支持。请先通过 [Lance 
Catalog](../catalogs/lance-catalog.mdx) 配置访问,并检查其 [Reader 
兼容性](../catalogs/lance-catalog.mdx#lance-版本与兼容性)。Doris 读取已有的 Lance 
索引;构建和维护索引需要使用兼容的 Lance 工具。
 :::
 
+标量过滤与向量或全文检索组合时,请参阅 [Lance 
混合检索性能最佳实践](lance-hybrid-search.mdx),了解过滤条件位置、索引对齐构建和 Profile 诊断。
+
 ## 1. 根据场景选择调优路径 {#choose-a-scenario}
 
 | 查询场景 | 从哪里开始 | 验证重点 |
diff --git 
a/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/best-practices/lance-hybrid-search.mdx
 
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/best-practices/lance-hybrid-search.mdx
new file mode 100644
index 00000000000..2fb00d038b1
--- /dev/null
+++ 
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/best-practices/lance-hybrid-search.mdx
@@ -0,0 +1,193 @@
+---
+{
+    "title": "Lance 混合检索性能最佳实践",
+    "language": "zh-CN",
+    "description": "指导 Apache Doris 用户优化 Lance 标量过滤与向量、全文混合检索:对齐索引 Fragment 
覆盖范围、选择 Segment 粒度、配置共享缓存并分析查询 Profile。"
+}
+---
+
+当标量条件参与 Lance 向量或全文检索时,可以使用本指南确定过滤条件位置、组织索引 
Segment,并识别重复过滤开销。本文的混合检索指检索与标量过滤组合;分别调用向量和全文 TVF 不会自动融合两者的排名。
+
+:::note
+Lance Catalog 为实验性功能,从 Apache Doris 4.2 开始支持。请先按 [Lance 
Catalog](../catalogs/lance-catalog.mdx) 配置访问并确认读写版本兼容性。Doris 读取已有 Lance 
索引,索引构建和维护由兼容的 Lance 工具完成。最新 Lance SDK 或上游提案不代表当前 Doris 内置读取器已经支持对应能力。
+:::
+
+## 1. 将候选生成所需的过滤条件放在检索内部 {#candidate-filters}
+
+假设已配置的表包含整数列 `id`、字符串列 `category`、字符串列 `content` 和四维向量列 
`embedding`,对必须影响候选生成的条件,应使用 TVF 的 `filter` 参数。向量索引需要匹配维度和距离度量;文本列需要已有提交完成的 FTS 
索引。
+
+```sql
+SELECT id, _distance
+FROM vector_search(
+    "table" = "lance_catalog.default.items",
+    "column" = "embedding",
+    "query_vector" = "[0.1,0.2,0.3,0.4]",
+    "metric" = "l2",
+    "top_k" = "20",
+    "filter" = "category = 'book'",
+    "use_index" = "true"
+)
+ORDER BY _distance ASC, id;
+```
+
+这会在满足条件的行中检索最多 20 个近邻。这是近似检索,返回 20 行不能证明召回率。带过滤条件的全文检索示例如下:
+
+```sql
+SELECT id, _score
+FROM full_text_search(
+    "table" = "lance_catalog.default.items",
+    "column" = "content",
+    "query" = "storage engine",
+    "query_type" = "match",
+    "operator" = "and",
+    "top_k" = "20",
+    "filter" = "category = 'book'",
+    "coverage_mode" = "strict"
+)
+ORDER BY _score DESC, id;
+```
+
+它返回最多 20 篇匹配文档,按 BM25 排序。外层 `WHERE category = 'book'` 的语义不同:Doris 在 Lance 
已生成的候选上过滤,然后执行全局 TopN,被删除的候选不会自动补齐。即使 `EXPLAIN` 将外层条件显示在 Doris Scan 
算子内部,也应只把有意在检索后执行的条件放在 TVF 外。
+
+## 2. 理解并行工作的单位 {#execution-and-splits}
+
+```text
+查询快照
+  -> FE 按物理检索索引 Segment 规划 Split
+  -> 每个 Split:标量 Prefilter -> 向量或全文候选
+  -> Doris 残余条件过滤 -> 局部 TopN -> 全局 TopN
+  -> 可选的输出列延迟回表
+```
+
+| 查询路径 | Split 切分依据 | 对性能的影响 |
+|---|---|---|
+| 使用索引的向量检索 | 选中的物理向量索引 Segment;未覆盖 Fragment 使用 Flat Search Split | 标量索引 
Segment 不会独立决定向量查询的 Split 数。未索引数据可能主导延迟。 |
+| 全文检索 | 选中的物理 FTS 索引 Segment | 默认 `strict` 模式拒绝未覆盖 Fragment;`index_only` 
会忽略它们。不存在全文 Flat Search 回退。 |
+| 普通标量扫描 | 自身扫描规划,可能使用标量索引 Segment 或 Fragment 分组 | 不能从纯标量查询的计划推断带过滤向量查询的并行度。 |
+
+`lance_fragments_per_split` 控制普通 Fragment 扫描的分组,不控制物理向量索引 Segment 数。增加 
Pipeline 实例数或 Scanner 并发不能产生比规划结果更多的向量 Segment Split。Lance Scanner 
内部也可能并行执行,因此一个 Doris Scanner 不等于一个 CPU 线程。
+
+对 `S` 个检索 Split,每个最多返回 `top_k + offset` 个候选,因此残余过滤和全局 TopN 前的候选上界是 `S * (top_k 
+ offset)`。增加 Segment 有助于分发工作,也会增加每段初始化、过滤、候选合并和并发内存需求。
+
+## 3. 对齐 Fragment 覆盖范围,而不是 Segment 数量 {#fragment-alignment}
+
+Fragment 是数据单位;索引 Segment 是覆盖一组 Fragment 的独立索引;逻辑索引将物理 Segment 组织在同一个索引名下。IVF 
Partition、BTree Page、FTS Posting Block 都是索引内部单位,不是 Doris 检索 Split 的切分单位。
+
+以下示例包含八个数据 Fragment:
+
+| 布局 | 向量索引覆盖范围 | 标量索引覆盖范围 | 支持 Fragment 范围裁剪的标量读取路径中的影响 |
+|---|---|---|---|
+| 边界对齐 | `V0={0,1,2,3}`、`V1={4,5,6,7}` | `B0={0,1,2,3}`、`B1={4,5,6,7}` | 每个向量 
Split 只需要对应的标量 Segment。 |
+| 标量更细 | 同上 | 八个 Segment,每个覆盖一个 Fragment | 每个 Split 选择四个标量 Segment,没有跨越向量边界的 
Segment。 |
+| 标量更粗 | 同上 | `B0={0,1,2,3,4,5,6,7}` | 两个向量 Split 都可能查询同一个标量 Segment。 |
+| 数量相同但边界交叉 | 同上 | `B0={0,1,4,5}`、`B1={2,3,6,7}` | 两个标量 Segment 都与两个向量 Split 
相交。数量相等不代表对齐。 |
+
+按 Fragment 范围裁剪可以跳过完全不相交的标量 Segment。部分相交时仍会保留该 Segment,不会自动重建更小的 
BTree,也不会在计算谓词之前自动切小每个 Posting List。重复查询可以复用已缓存的索引内容,但仍可能重复搜索和构建 Row ID 集合。保留一个 
BTree Segment 不等于扫描它的所有 Page,实际成本取决于谓词和匹配结果。
+
+:::important
+收益依赖实际读取路径。在具备 Fragment 范围标量加载能力的 Lance V2 Filtered Read 路径中,可以在搜索前排除无关标量 
Segment。旧 V1 数据文件路径可能先在更大范围求值逻辑标量索引,再限制结果 Row ID。这里的 V1/V2 指**数据文件格式**,不是 Doris 
版本或向量索引类型。不要假设所有读取器版本或 FTS 路径都有相同裁剪能力,应分别验证各类查询的计划和 Profile。
+:::
+
+对经常组合查询的列,构建向量、标量和 FTS 索引时使用相同 Fragment 分组。标量 Segment 可以更细,但应完整包含在服务的检索 
Segment 内。如果向量和 FTS 需要不同粒度,可以从共同的基础分组分别合并,并确保标量 Segment 
不跨越两者的边界。这是减少潜在重叠的布局策略,不保证所有执行路径都会利用它。
+
+### 选择初始粒度
+
+不存在通用最优的每段行数或字节数。可以从参与查询的 BE 数出发,在固定召回率和目标并发下,比较每个 BE 一个均衡检索 Segment 与每个 BE 少量 
Segment 的布局。这些是实验起点,不是必须遵守的比例。
+
+| 场景 | 构建重点 | 需要衡量的取舍 |
+|---|---|---|
+| 低延迟的带过滤向量检索 | 向量分组均衡,标量覆盖范围包含在对应分组内 | 最慢 Split 与重复过滤、候选扇出的成本 |
+| 高并发检索 | 避免单查询产生过多 Split,保留可复用 Partition | 同时观察 QPS、P95、CPU 排队和峰值内存 |
+| 宽范围标量过滤或高频全文词项 | 分组时考虑匹配行数、文本长度和 Posting 大小 | 行数相等不代表过滤或搜索成本相等 |
+| 持续追加 | 对相同的新 Fragment 分组构建所有相关索引 | 新数据覆盖及时性与小 Segment 积累 |
+| 单个 Fragment 非常大 | 在构建索引之前规划数据文件粒度 | 基于 Fragment 的覆盖范围不能把一个 Fragment 
拆成多个互不相交的覆盖分组 |
+
+除行数外,还应衡量编码后向量与图的大小、标量值分布、文本和 Posting 大小。每个向量 Segment 应有足够的数据支持其 IVF/PQ 
训练配置。`nlist`(IVF Partition 数)与物理 Segment 数是两个独立选择,调整其中一个不会自动调整另一个。
+
+## 4. 按同一分组方案构建相关索引 {#build-aligned-indexes}
+
+Doris 没有自动让多列索引边界对齐的参数。应在兼容 SDK 支持的范围内使用 Lance 分布式索引构建 
API。以下展示编排方式,不是可直接运行的数据生成脚本或经过压测的参数配置:
+
+```python
+# ds is an existing LanceDataset at the chosen build snapshot.
+# groups contains non-empty, disjoint lists of actual fragment IDs.
+# vector_params and fts_params are validated for the schema and SDK version.
+# Keep writes and compaction paused for this simple coordinator example.
+visible = {fragment.fragment_id for fragment in ds.get_fragments()}
+assigned = [fragment_id for group in groups for fragment_id in group]
+assert groups and all(groups)
+assert len(assigned) == len(set(assigned))
+assert set(assigned) == visible
+
+specs = [
+    ("embedding_idx", "embedding", "IVF_PQ", vector_params),
+    ("category_idx", "category", "BTREE", {}),
+    ("content_idx", "content", "INVERTED", fts_params),
+]
+segments = {name: [] for name, _, _, _ in specs}
+for fragment_ids in groups:
+    for name, column, index_type, params in specs:
+        segment = ds.create_index_uncommitted(
+            column,
+            index_type,
+            name=name,
+            fragment_ids=fragment_ids,
+            **params,
+        )
+        segments[name].append(segment)
+
+# These are three separate commits, not an atomic multi-index publication.
+for name, column, _, _ in specs:
+    ds.commit_existing_index_segments(name, column, segments[name])
+```
+
+按以下步骤准备并验证输入:
+
+1. **固定构建方案。** 从同一快照枚举真实 Fragment ID,ID 
不一定连续。按行数和预估索引字节数均衡分组,所有相关列复用同一组分组结果。分布式 Worker 必须打开相同构建版本。
+2. **选择兼容配置。** 为 IVF_PQ 提供适合向量维度和样本量的 
metric、`num_partitions`、`num_sub_vectors`。同一逻辑索引内保持配置一致。FTS 
保持分析器和分词配置一致,需要短语检索时启用 `with_position`。
+3. **选择向量模型策略。** 支持的多 Segment 构建流程允许各段独立训练。如果后续物理合并要求共享模型,应给 Worker 提供相同 IVF 
Centroid 和 PQ Codebook。不要直接合并独立训练的 Segment,也不要假设 HNSW 图会在合并后保留,应确认 SDK 支持的流程。
+4. **先构建,再发布。** 每次构建返回具有独立 UUID 的物理 Segment。收集成功结果,校验覆盖范围和兼容性,再分别提交每个逻辑索引。不同列或 
Worker 输出不要共用一个物理 UUID。示例假设使用新的逻辑索引名,替换、重试和失败清理需要额外的协调策略。
+5. **验证最终 Manifest。** 对每个索引,将实际 Segment 到 Fragment 的覆盖关系与规划分组比较,检查遗漏和重叠。只比较 
Segment 数量不够,应使用 Lance 索引元数据。Doris 对 Filesystem Catalog 表的 `SHOW INDEX` 
有助于检查逻辑索引名、类型和字段,但不能代替覆盖关系检查。
+
+共享分组方案不要求把所有 Worker 输出合并为一个 Segment,否则会丢失刻意保留的边界。构建、提交 API 和支持的合并流程参见 [Lance 
分布式索引指南](https://lance.org/guide/distributed_indexing/)。
+
+### 在维护中保持对齐
+
+- 对相同的新 Fragment 分组追加相关向量、标量和 FTS Segment。每次发布后检查查询快照的索引覆盖率。从数据追加到索引发布完成之间,FTS 
`strict` 模式可能拒绝查询。
+- 将 Compaction、数据重写和索引优化视为布局变化。不同操作和读取器版本可能重映射或重建索引,应检查最终覆盖关系,不要假设旧边界始终保留。
+- 按同一布局方案合并小分组。如果混合检索性能依赖对齐,不要独立把各索引优化为互不相关的分组。
+- 在线维护需要具备快照、冲突处理和恢复能力的协调器。逐索引提交不提供跨列的原子切换。TVF 读取最新 `main` 快照,其 `table` 
参数不支持选择历史版本或 Tag。
+
+### 逻辑索引与上游规划
+
+Lance 的[索引格式](https://lance.org/format/index/)定义逻辑索引和物理 Segment,[分布式索引跟踪 
Issue](https://github.com/lance-format/lance/issues/6309) 跟踪各索引类型共享的 Segment 
生命周期。该模型**不要求不同列拥有相同 Fragment 边界**。跨列分组和维护仍由编排方负责,不能把上游路线图理解成 Doris 已经提供自动对齐功能。
+
+## 5. 同时调优过滤和内存 {#workload-tuning}
+
+| 现象 | 下一步实验 | 正确性或资源约束 |
+|---|---|---|
+| 过滤很强,向量结果不足或召回率低 | 检查过滤位置与索引覆盖率,对比 Probe 设置,并在受控数据集上建立精确基线 | 减少 Probe 
可能降低召回率,增加 Probe 也不能产生不满足过滤条件的行。 |
+| 宽范围过滤产生很大的 Prefilter 集合 | 对比对齐布局以及标量谓词实际选择性 | 使用索引不能消除大量匹配 Row ID 的物化成本。 |
+| 向量 Split 增多时标量工作重复 | 比较 Segment 覆盖范围交集与读取路径 | 即使缓存已热,增加 Split 也可能放大过滤开销。 |
+| 冷查询慢 | 使用有代表性的向量和词项比较冷、热 Profile | 一个查询的预热不能覆盖生产工作集。 |
+| 热缓存下混合流量变慢 | 在目标并发下观察共享缓存未命中和 CPU 排队 | 同一 BE 上的标量、向量、FTS 内容可能竞争同一 Session 
索引缓存预算。 |
+| `top_k` 大、`offset` 深或输出列宽 | 在语义允许时减少候选需求,检查延迟回表 | 候选堆、Row ID Mask、Refine 
和回表缓冲属于额外查询内存。 |
+
+按**合并后的**可复用索引工作集规划缓存,为并发查询内存保留空间。缓存容量不是进程内存上限,增加 Metadata Cache 也不会扩大 Index 
Cache。各索引载荷公式、HNSW 图开销、缓存作用域和 Doris BE 配置示例参见 [Lance 查询最佳实践](doris-lance.mdx)。
+
+## 6. 用计划和 Profile 验证 {#validation}
+
+使用相同查询集、数据快照、索引类型与配置、`top_k` 
和召回率目标进行比较,每次只修改一种布局或参数。覆盖强过滤与宽范围过滤、不同向量和词项、冷缓存与热缓存、目标并发,记录 
P50/P95、吞吐、结果数量、召回率和每个 BE 的峰值内存。
+
+| 阶段 | 证据 | 如何解释 |
+|---|---|---|
+| 规划与覆盖率 | `EXPLAIN` 中的 
`lanceSearchIndexSegments`、`lanceSearchUnindexedFragments`;Profile 中 Scanner 数 
| 确认 Split 数、回退工作量以及计划能否利用预期布局。 |
+| Prefilter | 
`LancePrefilterInputRows`、`LancePrefilterRowIds`、`LancePrefilterLoadTime`、`LancePrefilterBuildTime`
 | Load 包含构建过滤器的耗时,不应相加。没有远程 I/O 时,大 Row ID 集合仍可能主导开销。 |
+| 标量 Segment 路径 | 有统计值时的 
`LanceScalarIndexSegmentsRequested`、`LanceScalarIndexSegmentsSearched`、`LanceScalarIndexCandidateRows`
 | 这些计数描述已埋点的标量 Segment 路径,不一定覆盖 ANN/FTS 内部所有标量查询。零值不能证明没有标量过滤。 |
+| 索引加载与搜索 | 
`LanceIndexPartitionCacheMissLoads`、`LanceExecutionIOBytesRead`、`LanceIndexComparisons`,以及可用的详细索引计时器
 | 区分缓存重新加载、热缓存搜索和候选处理。I/O 计数为零不代表没有 CPU 工作。 |
+| 输出回表 | 存在时的 Materialization 算子和 Row ID Fetch 计时 | 将输出列读取与候选检索分开归因。 |
+| 端到端延迟 | 最慢 Scanner、依赖等待、FE 等待/取结果/写结果和客户端墙钟时间 | 不能将重叠算子或所有并行 Scanner 
的时间直接相加作为查询延迟。 |
+
+Profile 详细程度取决于 Doris 构建和内置 Lance 埋点。缺失指标不等于实测为零。`FileScannerV2` 包含 Scanner 
打开、读取和关闭工作;读取区间可能包括 Prefilter 加载、索引 I/O、运行时等待、ANN 搜索、Refine 
和转换。嵌套或并行计时不是可直接相加的分解。如果现有指标无法解释该区间,应收集 CPU/I/O Trace 
或使用带所需详细埋点的构建,再判断剩余开销是否来自存储。
+
+对无过滤条件的向量查询,如果读取器具有 Segment 范围优化,且选中的索引 Segment 完整位于扫描 Fragment 范围内,可以避免枚举所有 
Row ID。真实标量谓词仍需计算,删除和可见性处理也可能保留。对齐能减少可避免的工作,但不保证带过滤检索的 Prefilter 开销为零。
diff --git 
a/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
 
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
index a18f25b0ec6..66939e915ec 100644
--- 
a/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
+++ 
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
@@ -472,6 +472,8 @@ nested_index   nested_label_btree    1             
attributes.`child.with.dot`
 
 | Lance / Arrow 类型 | Doris 类型 | 说明 |
 |---|---|---|
+| `null`(可空字段) | `NULL` | 所有值均为 SQL NULL;非可空 Null 字段不支持 |
+| `duration(s/ms/us/ns)` | `BIGINT` | 保留字段声明单位的有符号计数,不统一转换为秒 |
 | `bool` | `BOOLEAN` | |
 | `int8` | `TINYINT` | |
 | `uint8` | `SMALLINT` | 无符号整数无损提升 |
@@ -487,6 +489,9 @@ nested_index   nested_label_btree    1             
attributes.`child.with.dot`
 | `decimal128(P,S)` | `DECIMAL(P,S)` | 最大精度为 38 |
 | `decimal256(P,S)` | `DECIMAL(P,S)` | 最大精度为 76 |
 | `utf8`、`large_utf8` | `TEXT` | |
+| `arrow.json` Extension | `JSON` | 底层类型为 `utf8` 或 `large_utf8` |
+| `lance.json` Extension | `JSON` | 底层类型为 `large_binary`;按 JSON 逻辑类型读取 |
+| `lance.bfloat16` Extension | `FLOAT` | 底层类型为 `fixed_size_binary(2)`;无损提升为 
Float32 |
 | `binary`、`large_binary` | `VARBINARY(2147483647)` | |
 | `fixed_size_binary(N)` | `VARBINARY(N)` | 保留固定字节宽度 |
 | `date32(day)`、`date64(ms)` | `DATE` | `date64` 应表示完整自然日 |
@@ -501,17 +506,21 @@ nested_index   nested_label_btree    1             
attributes.`child.with.dot`
 | `list`、`large_list`、`fixed_size_list` | `ARRAY` | 元素类型递归映射 |
 | `map` | `MAP` | Key 和 Value 类型递归映射 |
 
+Doris 根据 `ARROW:extension:name` 元数据识别上述扩展类型,并校验其底层存储类型。JSON 扩展映射为 Doris 
`JSON`(内部类型 `JSONB`),普通 `utf8`/`large_utf8` 列即使包含 JSON 文本,仍映射为 `TEXT`,不会自动变成 
JSON。支持 JSON 读取不代表任意 JSON 路径表达式都能下推;下推范围见[谓词下推](#谓词下推)。
+
+JSON、Duration、Null 和 BFloat16 读取需要 FE 与参与查询的 BE 均包含对应类型支持。旧构建可能仍显示不支持;升级时应同时确认 
Schema 识别和实际读取能力。
+
 当前不支持以下类型:
 
-- Arrow `null` 和 `duration`。
-- 带有 `ARROW:extension:name` 元数据的 Arrow/Lance Extension 类型,例如 Lance Blob 
v2、Arrow JSON Extension 和 Lance BFloat16 Extension。
+- 非可空的 Arrow `null` 字段。
+- 未列为支持的 Arrow/Lance Extension 类型,例如 Lance Blob v2;已知扩展类型的底层存储与上表不匹配时也不支持。
 - 无法递归映射其子类型的复杂类型。
 - 保留了 Dictionary 标记的 Arrow Dictionary 类型。
 
 对于不支持的顶层列,Catalog 表的 `DESC` 和 Lance 文件 TVF 的 `DESC FUNCTION` 都会保留该列,并显示 
`unknown type: 
UNSUPPORTED_TYPE`。如果复杂类型的任一子字段无法映射,则整个顶层复杂列会标记为不支持。查询只投影支持的列仍可正常执行;当 SQL 
投影不支持的列时,Doris 会在分析阶段报错。例如:
 
 ```sql
-SELECT * EXCEPT(blob_col, json_col)
+SELECT * EXCEPT(blob_col)
 FROM lance_catalog.default.all_types;
 ```
 
@@ -995,7 +1004,7 @@ lance_data_cache_read_block_size_bytes = 1048576
 
 配置容量时应同时考虑索引缓存、元数据缓存和查询工作内存。上述容量不是 Lance 查询总内存的上限,也不是 BE 
进程内存的上限。磁盘缓存还需要本地磁盘空间和文件描述符资源;增大读取块可以减少块数量,但小范围读取可能产生更多额外 
I/O。调整读取块大小或磁盘布局时使用新的缓存目录。
 
-按场景估算缓存、选择配置、分析共享缓存竞争以及确定 Index Segment 粒度,请参阅 [Lance 
查询最佳实践](../best-practices/doris-lance.mdx)。
+按场景估算缓存、选择配置、分析共享缓存竞争以及确定 Index Segment 粒度,请参阅 [Lance 
查询最佳实践](../best-practices/doris-lance.mdx)。过滤与检索组合的调优方法参见 [Lance 
混合检索性能最佳实践](../best-practices/lance-hybrid-search.mdx)。
 
 ### 查看缓存效果
 
diff --git 
a/versioned_docs/version-4.x/lakehouse/best-practices/doris-lance.mdx 
b/versioned_docs/version-4.x/lakehouse/best-practices/doris-lance.mdx
index 137642cdc95..4cbfac6a4b5 100644
--- a/versioned_docs/version-4.x/lakehouse/best-practices/doris-lance.mdx
+++ b/versioned_docs/version-4.x/lakehouse/best-practices/doris-lance.mdx
@@ -12,6 +12,8 @@ Use this guide to choose a per-BE cache budget and index 
segment layout for your
 Lance Catalog is experimental and supported starting from Apache Doris 4.2. 
First configure access using [Lance Catalog](../catalogs/lance-catalog.mdx) and 
check its [reader 
compatibility](../catalogs/lance-catalog.mdx#lance-version-and-compatibility). 
Doris reads existing Lance indexes; use compatible Lance tooling to build or 
maintain them.
 :::
 
+For scalar filters combined with vector or full-text retrieval, see [Lance 
Hybrid Search Performance Best Practices](lance-hybrid-search.mdx) for filter 
placement, aligned index construction, and profile diagnosis.
+
 ## 1. Choose a Tuning Path {#choose-a-scenario}
 
 | Your workload | Start with | What to validate |
diff --git 
a/versioned_docs/version-4.x/lakehouse/best-practices/lance-hybrid-search.mdx 
b/versioned_docs/version-4.x/lakehouse/best-practices/lance-hybrid-search.mdx
new file mode 100644
index 00000000000..26e07f480d9
--- /dev/null
+++ 
b/versioned_docs/version-4.x/lakehouse/best-practices/lance-hybrid-search.mdx
@@ -0,0 +1,193 @@
+---
+{
+    "title": "Lance Hybrid Search Performance Best Practices",
+    "language": "en",
+    "description": "Tune filtered vector and full-text search over Lance in 
Apache Doris: align index coverage, size segments, budget shared caches, and 
diagnose query profiles."
+}
+---
+
+Use this guide when scalar conditions participate in vector or full-text 
retrieval over Lance. It explains how to place filters, organize index 
segments, and identify repeated filtering work. Here, hybrid search means 
retrieval combined with scalar filtering; calling the vector and full-text TVFs 
separately does not automatically fuse their rankings.
+
+:::note
+Lance Catalog is experimental and supported starting from Apache Doris 4.2. 
Configure access and check the reader/writer compatibility requirements in 
[Lance Catalog](../catalogs/lance-catalog.mdx) first. Doris reads existing 
Lance indexes; build and maintain them with compatible Lance tooling. The 
latest Lance SDK or an upstream proposal does not guarantee support in the 
reader bundled with your Doris build.
+:::
+
+## 1. Put Candidate Filters Inside the Search {#candidate-filters}
+
+For a configured table containing an integer `id`, string `category`, string 
`content`, and a four-dimensional vector `embedding`, use the TVF's `filter` 
parameter for conditions that must affect candidate generation. The vector 
index must match the dimension and metric; the text column must have a 
committed FTS index.
+
+```sql
+SELECT id, _distance
+FROM vector_search(
+    "table" = "lance_catalog.default.items",
+    "column" = "embedding",
+    "query_vector" = "[0.1,0.2,0.3,0.4]",
+    "metric" = "l2",
+    "top_k" = "20",
+    "filter" = "category = 'book'",
+    "use_index" = "true"
+)
+ORDER BY _distance ASC, id;
+```
+
+This searches for up to 20 neighbors among matching rows. It is approximate 
search: returning 20 rows does not prove recall. For filtered full-text 
retrieval:
+
+```sql
+SELECT id, _score
+FROM full_text_search(
+    "table" = "lance_catalog.default.items",
+    "column" = "content",
+    "query" = "storage engine",
+    "query_type" = "match",
+    "operator" = "and",
+    "top_k" = "20",
+    "filter" = "category = 'book'",
+    "coverage_mode" = "strict"
+)
+ORDER BY _score DESC, id;
+```
+
+This returns up to 20 matching documents ranked by BM25. An outer `WHERE 
category = 'book'` has different semantics: Doris filters candidates that Lance 
has already generated, before global TopN. Removed candidates are not 
automatically replenished. Keep only intentional post-retrieval conditions 
outside the TVF, even if `EXPLAIN` places them inside a Doris Scan operator.
+
+## 2. Understand the Unit of Parallel Work {#execution-and-splits}
+
+```text
+Query snapshot
+  -> FE plans physical search-index segment splits
+  -> Each split: scalar Prefilter -> vector or full-text candidates
+  -> Doris residual filtering -> local TopN -> global TopN
+  -> Optional deferred fetch of output columns
+```
+
+| Query path | Split basis | Consequence |
+|---|---|---|
+| Indexed vector search | Selected physical vector index segments; uncovered 
fragments get Flat Search splits | Scalar index segments do not independently 
set vector-query split count. Unindexed data can dominate latency. |
+| Full-text search | Selected physical FTS index segments | Default `strict` 
coverage rejects uncovered fragments; `index_only` omits them. There is no 
full-text Flat Search fallback. |
+| Ordinary scalar scan | Its own scan planning, potentially using scalar-index 
segments or fragment groups | Do not infer filtered vector-query parallelism 
from a scalar-only query plan. |
+
+`lance_fragments_per_split` controls ordinary fragment scan grouping, not the 
number of physical vector index segments. Increasing pipeline instances or 
scanner concurrency cannot create more vector segment splits than the planner 
produces. A Lance scanner can also execute work internally in parallel, so one 
Doris scanner is not equivalent to one CPU thread.
+
+For `S` search splits, each can return at most `top_k + offset` candidates. 
The pre-merge candidate bound is therefore `S * (top_k + offset)`, before 
residual filtering and global TopN. More segments can improve distribution 
while increasing per-segment setup, filtering, candidate merging, and 
concurrent memory demand.
+
+## 3. Align Fragment Coverage, Not Segment Counts {#fragment-alignment}
+
+A fragment is a data unit. An index segment is a self-contained index covering 
a set of fragments. A logical index groups physical segments under one index 
name. An IVF partition, BTree page, or FTS posting block is an internal index 
unit; none of these defines a Doris search split.
+
+The following example uses eight data fragments:
+
+| Layout | Vector coverage | Scalar coverage | Filtering implication on a 
fragment-scoped scalar reader |
+|---|---|---|---|
+| Aligned | `V0={0,1,2,3}`, `V1={4,5,6,7}` | `B0={0,1,2,3}`, `B1={4,5,6,7}` | 
Each vector split needs only its corresponding scalar segment. |
+| Scalar is finer | Same as above | Eight segments, one per fragment | Each 
split selects four scalar segments; none crosses the vector boundary. |
+| Scalar is coarser | Same as above | `B0={0,1,2,3,4,5,6,7}` | Both vector 
splits can query the same scalar segment. |
+| Same count, crossing boundaries | Same as above | `B0={0,1,4,5}`, 
`B1={2,3,6,7}` | Both scalar segments overlap both vector splits. Equal counts 
do not provide alignment. |
+
+Fragment-scoped pruning skips scalar segments with no overlap. Partial overlap 
keeps the segment: it does not automatically rebuild a smaller BTree or slice 
every posting list before predicate evaluation. Repeated queries can reuse 
cached index content but still repeat searches and row-ID set construction. A 
BTree lookup does not necessarily scan all pages of a retained segment; cost 
depends on the predicate and matches.
+
+:::important
+This benefit depends on the actual reader path. In Lance's V2 filtered-read 
path with fragment-scoped scalar loading, unrelated scalar segments can be 
excluded before search. Legacy V1 data-file paths can evaluate the logical 
scalar index more broadly and restrict row IDs afterward. V1/V2 here refers to 
the **data-file format**, not the Doris release or vector index type. Do not 
assume identical pruning for every reader version or for the FTS path; verify 
the plan and profile of each que [...]
+:::
+
+For frequently combined columns, use the same fragment groups for vector, 
scalar, and FTS index construction. Scalar segments may be finer if each is 
contained within the search segment it serves. If vector and FTS segments need 
different sizes, consider coarsenings of common base groups and keep scalar 
segments inside both sets of boundaries. This is a layout policy to reduce 
potential overlap, not a guarantee that every execution path exploits it.
+
+### Choose a Starting Granularity
+
+There is no universal optimal row count or byte size per segment. Start from 
the BEs eligible for the workload, then compare one balanced search segment per 
BE with a few segments per BE at fixed recall and target concurrency. These are 
experiment points, not required ratios.
+
+| Workload | Construction priority | Main tradeoff to measure |
+|---|---|---|
+| Latency-sensitive filtered vector search | Balanced vector groups and scalar 
coverage contained within them | Slowest split versus repeated filtering and 
candidate fan-out |
+| High-concurrency retrieval | Avoid excessive per-query split fan-out; retain 
reusable partitions | QPS, P95 latency, CPU queues, and peak memory together |
+| Broad scalar filters or common FTS terms | Consider match counts, text 
length, and posting sizes when balancing groups | Equal row counts can still 
produce unequal filtering/search costs |
+| Frequent appends | Build all related indexes for the same new fragment 
groups | Fresh-data coverage versus accumulating tiny segments |
+| Very large individual fragments | Plan data-file sizing before index 
construction | Fragment-based coverage cannot split one fragment into multiple 
disjoint coverage groups |
+
+Measure encoded vector/graph size, scalar value distribution, and text/posting 
sizes as well as rows. Keep enough training data in each vector segment for its 
IVF/PQ configuration. `nlist` (IVF partition count) and physical segment count 
are separate choices; changing one does not automatically adjust the other.
+
+## 4. Build Related Indexes From One Grouping Plan {#build-aligned-indexes}
+
+Doris has no parameter that automatically aligns indexes across columns. Use 
the Lance distributed-build APIs where supported by your compatible SDK. The 
following is an orchestration pattern, not a ready-to-run dataset generator or 
a benchmark configuration:
+
+```python
+# ds is an existing LanceDataset at the chosen build snapshot.
+# groups contains non-empty, disjoint lists of actual fragment IDs.
+# vector_params and fts_params are validated for the schema and SDK version.
+# Keep writes and compaction paused for this simple coordinator example.
+visible = {fragment.fragment_id for fragment in ds.get_fragments()}
+assigned = [fragment_id for group in groups for fragment_id in group]
+assert groups and all(groups)
+assert len(assigned) == len(set(assigned))
+assert set(assigned) == visible
+
+specs = [
+    ("embedding_idx", "embedding", "IVF_PQ", vector_params),
+    ("category_idx", "category", "BTREE", {}),
+    ("content_idx", "content", "INVERTED", fts_params),
+]
+segments = {name: [] for name, _, _, _ in specs}
+for fragment_ids in groups:
+    for name, column, index_type, params in specs:
+        segment = ds.create_index_uncommitted(
+            column,
+            index_type,
+            name=name,
+            fragment_ids=fragment_ids,
+            **params,
+        )
+        segments[name].append(segment)
+
+# These are three separate commits, not an atomic multi-index publication.
+for name, column, _, _ in specs:
+    ds.commit_existing_index_segments(name, column, segments[name])
+```
+
+Prepare and validate the inputs as follows:
+
+1. **Freeze the build plan.** Enumerate actual fragment IDs from one snapshot; 
IDs need not be contiguous. Balance groups by rows and estimated index bytes. 
Reuse these exact groups for every related column. All distributed workers must 
open the same build version.
+2. **Choose compatible configurations.** For IVF_PQ, supply an appropriate 
metric, `num_partitions`, and `num_sub_vectors` for the vector dimension and 
sample size. Keep index configurations consistent across a logical index. For 
FTS, keep analyzer/tokenizer settings consistent and enable `with_position` 
when phrase search is needed.
+3. **Choose the vector model strategy.** Supported multi-segment builds can 
train independently per segment. If later physical merging requires shared 
models, supply the same IVF centroids and PQ codebook to workers. Do not 
blindly merge independently trained segments or assume HNSW graphs survive a 
merge; check the SDK's supported workflow.
+4. **Build first, then publish.** Each build returns a physical segment with 
its own UUID. Collect successful outputs and verify coverage and compatibility 
before committing each logical index. Do not assign one shared physical UUID to 
different columns or worker outputs. The example assumes new logical index 
names; replacement, retries, and failure cleanup require a separate coordinator 
policy.
+5. **Verify the resulting manifest.** For each index, compare the actual 
segment-to-fragment coverage with the planned groups and check for missing or 
overlapping coverage. Matching segment counts alone are insufficient. Use Lance 
index metadata; Doris `SHOW INDEX` on Filesystem Catalog tables helps check 
logical index names, types, and fields but does not replace this coverage audit.
+
+A shared grouping plan does not require merging all worker outputs into one 
segment. Such a merge would remove the boundaries you intended to preserve. 
Consult [Lance Distributed 
Indexing](https://lance.org/guide/distributed_indexing/) for the build/commit 
APIs and supported merge workflows.
+
+### Preserve Alignment During Maintenance
+
+- Append related vector, scalar, and FTS segments using the same new fragment 
groups. After each publication, check the query snapshot's coverage. FTS 
`strict` mode can reject queries between the data append and completion of 
index publication.
+- Treat compaction, rewrites, and index optimization as layout changes. 
Depending on the operation and reader version, indexes may be remapped or 
rebuilt; inspect the resulting coverage instead of assuming previous boundaries 
survive.
+- Coalesce small groups according to one layout plan. Do not independently 
optimize each index into unrelated groups if mixed-query performance depends on 
alignment.
+- For online maintenance, use a coordinator with snapshot/conflict handling 
and recovery. Separate per-index commits do not provide an atomic switch for 
all columns. The TVFs read the latest `main` snapshot; their `table` argument 
does not select a historical version or tag.
+
+### Logical Indexes and Upstream Planning
+
+Lance's [index format](https://lance.org/format/index/) defines logical 
indexes and physical segments. The [distributed-index tracking 
issue](https://github.com/lance-format/lance/issues/6309) tracks the shared 
segment lifecycle across index families. This model does **not** impose 
identical fragment boundaries on different columns. Cross-column grouping and 
maintenance remain orchestration responsibilities; do not treat an upstream 
roadmap as an automatic alignment feature in Doris.
+
+## 5. Tune Filters and Memory Together {#workload-tuning}
+
+| Observation | Next experiment | Correctness or resource constraint |
+|---|---|---|
+| Very selective filter, too few vector results or weak recall | Check filter 
placement and index coverage; compare probe settings and an exact baseline on a 
bounded dataset | Reducing probes for latency can reduce recall. More probes 
cannot create rows that fail the filter. |
+| Broad filter produces large Prefilter sets | Compare aligned layouts and the 
scalar predicate's actual selectivity | Index use does not eliminate the cost 
of materializing many matching row IDs. |
+| Repeated scalar work as vector split count rises | Compare segment coverage 
intersections and reader paths | Increasing splits can amplify filtering even 
with a warm cache. |
+| Cold queries are slow | Compare cold and warm profiles with representative 
vectors and terms | Warmup for one query does not cover the production working 
set. |
+| Warm mixed traffic becomes slow | Measure shared cache misses and CPU queues 
at target concurrency | Scalar, vector, and FTS index content can compete for 
the same session index-cache budget on a BE. |
+| Large `top_k`, deep `offset`, or wide output | Reduce requested candidates 
where semantics permit; inspect deferred fetching | Candidate heaps, row-ID 
masks, refinement, and fetch buffers are additional query memory. |
+
+Size the **combined** reusable index working set and leave headroom for 
concurrent query memory. Cache capacity is not a process-memory limit, and 
increasing metadata cache does not increase index cache. Use [Lance Query Best 
Practices](doris-lance.mdx) for per-index payload formulas, HNSW graph costs, 
cache scope, and Doris BE configuration examples.
+
+## 6. Verify With Plans and Profiles {#validation}
+
+Compare the same query set, dataset snapshot, index type/configuration, 
`top_k`, and recall target. Change one layout or parameter at a time. Include 
selective and broad filters, diverse vectors/terms, cold and warm runs, and the 
intended concurrency. Record P50/P95, throughput, result counts, recall, and 
per-BE peak memory.
+
+| Stage | Evidence | Interpretation |
+|---|---|---|
+| Planning and coverage | `EXPLAIN`: `lanceSearchIndexSegments`, 
`lanceSearchUnindexedFragments`; profile scanner counts | Confirm split count, 
fallback work, and whether the plan can use the intended layout. |
+| Prefilter | `LancePrefilterInputRows`, `LancePrefilterRowIds`, 
`LancePrefilterLoadTime`, `LancePrefilterBuildTime` | Load includes building 
the filter; do not add these timers. Large row-ID sets can dominate even 
without remote I/O. |
+| Scalar segment path | `LanceScalarIndexSegmentsRequested`, 
`LanceScalarIndexSegmentsSearched`, `LanceScalarIndexCandidateRows`, where 
populated | These counters describe the instrumented scalar-segment path, not 
necessarily every scalar lookup inside ANN/FTS. Zero is not proof that no 
scalar filtering happened. |
+| Index loading and search | `LanceIndexPartitionCacheMissLoads`, 
`LanceExecutionIOBytesRead`, `LanceIndexComparisons`; detailed index timers 
where available | Separate cache reloads from warm search and 
candidate-processing work. Zero I/O counters do not imply zero CPU work. |
+| Output fetch | Materialization operator and row-ID fetch timings, where 
present | Attribute output-column reads separately from candidate search. |
+| End-to-end latency | Slowest scanner, dependency waits, FE wait/fetch/write, 
and client wall time | Do not sum overlapping operators or all parallel scanner 
times into query latency. |
+
+Profile detail depends on the Doris build and bundled Lance instrumentation. 
Missing counters are not measured zeros. `FileScannerV2` includes scanner 
open/read/close work; its read interval can include Prefilter loading, index 
I/O, runtime waits, ANN search, refinement, and conversion. Nested or parallel 
counters are not an additive breakdown. If the exposed metrics cannot explain 
the interval, collect a CPU/I/O trace or use a build with the necessary 
detailed instrumentation before att [...]
+
+For an unfiltered vector query, a reader with the segment-scope optimization 
can avoid enumerating every row ID when the selected index segment is wholly 
inside the scan's fragment scope. Actual scalar predicates still need 
evaluation; deletion and visibility handling can also remain. Alignment reduces 
avoidable work but does not promise a zero-cost Prefilter for filtered 
retrieval.
diff --git a/versioned_docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx 
b/versioned_docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
index 7a0f9427e5c..d7eb0961850 100644
--- a/versioned_docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
+++ b/versioned_docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
@@ -472,6 +472,8 @@ nested_index   nested_label_btree    1             
attributes.`child.with.dot`
 
 | Lance / Arrow Type | Doris Type | Description |
 |---|---|---|
+| `null` (nullable field) | `NULL` | Every value is SQL NULL; non-nullable 
Null fields are unsupported |
+| `duration(s/ms/us/ns)` | `BIGINT` | Preserves the signed count in the 
declared unit; does not normalize to seconds |
 | `bool` | `BOOLEAN` | |
 | `int8` | `TINYINT` | |
 | `uint8` | `SMALLINT` | Losslessly widened unsigned integer |
@@ -487,6 +489,9 @@ nested_index   nested_label_btree    1             
attributes.`child.with.dot`
 | `decimal128(P,S)` | `DECIMAL(P,S)` | Maximum precision is 38 |
 | `decimal256(P,S)` | `DECIMAL(P,S)` | Maximum precision is 76 |
 | `utf8`, `large_utf8` | `TEXT` | |
+| `arrow.json` Extension | `JSON` | Requires `utf8` or `large_utf8` storage |
+| `lance.json` Extension | `JSON` | Requires `large_binary` storage; read as 
the logical JSON type |
+| `lance.bfloat16` Extension | `FLOAT` | Requires `fixed_size_binary(2)` 
storage; widened to Float32 without precision loss |
 | `binary`, `large_binary` | `VARBINARY(2147483647)` | |
 | `fixed_size_binary(N)` | `VARBINARY(N)` | Preserves the fixed byte width |
 | `date32(day)`, `date64(ms)` | `DATE` | A `date64` value must represent a 
complete calendar day |
@@ -501,17 +506,21 @@ nested_index   nested_label_btree    1             
attributes.`child.with.dot`
 | `list`, `large_list`, `fixed_size_list` | `ARRAY` | Element types are mapped 
recursively |
 | `map` | `MAP` | Key and value types are mapped recursively |
 
+Doris recognizes these extensions through `ARROW:extension:name` metadata and 
validates their storage types. JSON extensions map to Doris `JSON` (internally 
`JSONB`). An ordinary `utf8`/`large_utf8` column containing JSON text still 
maps to `TEXT`; its contents do not automatically change its type. JSON read 
support does not imply pushdown of arbitrary JSON path expressions; see 
[Predicate Pushdown](#predicate-pushdown) for the supported scope.
+
+Reading JSON, Duration, Null, and BFloat16 requires the corresponding type 
support in both the FE and the BEs serving the query. Older builds may still 
report these types as unsupported; verify schema discovery and actual reads 
when upgrading.
+
 The following types are not currently supported:
 
-- Arrow `null` and `duration`.
-- Arrow/Lance Extension types with `ARROW:extension:name` metadata, including 
Lance Blob v2, Arrow JSON Extension, and Lance BFloat16 Extension.
+- Non-nullable Arrow `null` fields.
+- Arrow/Lance Extension types not listed as supported, such as Lance Blob v2; 
known extensions with incompatible storage types are also unsupported.
 - Complex types whose child types cannot be mapped recursively.
 - Arrow Dictionary types that preserve the Dictionary marker.
 
 For an unsupported top-level column, `DESC` on a Catalog table and `DESC 
FUNCTION` on a Lance file TVF both preserve the column and display `unknown 
type: UNSUPPORTED_TYPE`. If any child of a complex type cannot be mapped, the 
whole top-level complex column is marked unsupported. Queries can still project 
only supported columns. Doris reports an error during analysis when SQL 
projects an unsupported column. For example:
 
 ```sql
-SELECT * EXCEPT(blob_col, json_col)
+SELECT * EXCEPT(blob_col)
 FROM lance_catalog.default.all_types;
 ```
 
@@ -995,7 +1004,7 @@ After the BE restarts, the shared cache initializes when 
the first Lance Dataset
 
 Budget for index and metadata caches alongside query working memory. These 
capacities do not limit total Lance query memory or BE process memory. The disk 
cache also consumes local disk space and file descriptors. Larger read blocks 
reduce the number of blocks but can increase extra I/O for small reads. Use a 
new cache directory when changing the read block size or disk layout.
 
-For workload-specific sizing formulas, configuration examples, shared-cache 
contention, and index segment granularity, see [Lance Query Best 
Practices](../best-practices/doris-lance.mdx).
+For workload-specific sizing formulas, configuration examples, shared-cache 
contention, and index segment granularity, see [Lance Query Best 
Practices](../best-practices/doris-lance.mdx). For filtering and retrieval 
together, see [Lance Hybrid Search Performance Best 
Practices](../best-practices/lance-hybrid-search.mdx).
 
 ### Checking Cache Effectiveness
 
diff --git a/versioned_sidebars/version-4.x-sidebars.json 
b/versioned_sidebars/version-4.x-sidebars.json
index 5ca7499aa21..9aed452fed9 100644
--- a/versioned_sidebars/version-4.x-sidebars.json
+++ b/versioned_sidebars/version-4.x-sidebars.json
@@ -890,6 +890,7 @@
               "items": [
                 "lakehouse/best-practices/optimization",
                 "lakehouse/best-practices/doris-lance",
+                "lakehouse/best-practices/lance-hybrid-search",
                 "lakehouse/best-practices/kerberos",
                 "lakehouse/best-practices/tpch",
                 "lakehouse/best-practices/tpcds"


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to