This is an automated email from the ASF dual-hosted git repository.

Gabriel39 pushed a commit to branch master
in repository https://gitbox.apache.org/repos/asf/doris-website.git


The following commit(s) were added to refs/heads/master by this push:
     new 4d1b13ca039 [docs] Add standalone Lance query best practices (#4181)
4d1b13ca039 is described below

commit 4d1b13ca0396e81376fff8b55da9ec7b65ada143
Author: Gabriel <[email protected]>
AuthorDate: Wed Sep 30 11:02:03 2026 +0800

    [docs] Add standalone Lance query best practices (#4181)
    
    Add a standalone Lance Query Best Practices guide so users can choose
    cache budgets and index segment granularity for their own workload. Keep
    the Lance Catalog page as the access, parameter, and metric reference,
    with an entry link to the new guide.
    
    The English and Chinese guides follow seven steps:
    1. Choose a tuning path for low-latency or concurrent vector search,
    scalar filtering, FTS, mixed indexes, or data-file reads.
    2. Collect workload and resource inputs and reserve memory for
    concurrent queries.
    3. Estimate vector, BTree, Bitmap/LabelList, NGram, FTS, ZoneMap,
    BloomFilter, and RTree payloads, with component-by-component tables,
    layout-specific formulas, and limits. Separate vector values/codes, row
    IDs, shared IVF/quantizer structures, HNSW graph buffers, and
    scalar/full-text structures; explicitly mark allocations requiring
    measurement.
    4. Budget for index types and tables competing within one BE-wide index
    cache; distinguish metadata and disk caches.
    5. Choose physical index segment granularity and evaluate candidate
    amplification, concurrency, and recall.
    6. Validate a representative mixed workload.
    7. Diagnose cache misses, poor BE utilization, high memory, and high
    latency before increasing resources.
    
    Register the new page under Lakehouse Best Practices in the 4.x sidebar.
    The examples and upstream references are kept in the guide, avoiding
    duplicate tuning content in the Catalog reference. HNSW has a dedicated
    graph breakdown and numeric examples for readers retaining or skipping
    edge distances. Subtotals exclude explicitly listed auxiliary
    allocations; shared allocations are counted once.
    
    Validation:
    - All four changed MDX pages compile with the repository's Docusaurus
    3.6.3 compiler.
    - Changed-file frontmatter and Markdown structure checks pass; new-page
    SEO checks pass.
    - New-page internal-link/anchor checks pass. External reference URLs
    were checked when introduced.
    - Bilingual example parity, seven-step structure, unique anchors,
    reciprocal entry/reference links, sidebar registration, original Catalog
    headings, and example arithmetic verified, including both HNSW layouts,
    table column counts, and bilingual inline-formula parity.
    - Full diff self-reviewed; disclosure checks and `git diff --check`
    pass. A full website build was not run.
    
    Scope and self-review:
    - Goal/scope: scenario-based, independently navigable guidance; two new
    localized pages, two Catalog entry links, and one sidebar entry.
    - Information architecture: the guide follows the existing 4.x Lance
    Catalog. The i18n checker flags missing `current` counterparts for both
    topics; neither a current-version Catalog nor its guide is added in this
    PR. English and Chinese 4.x content are synchronized.
    - Links/navigation: existing published Catalog headings, URLs, and
    frontmatter are preserved. Tuning sections introduced by this unmerged
    PR now live in the new guide with updated links.
    - UI/configuration: only documentation navigation changes; no React,
    styling, or build changes.
    - Other findings: no unresolved self-review issues. Formulas are
    loaded-payload estimates, and segment counts are experiment starting
    points, not benchmarked optima or process-memory caps.
---
 .../lakehouse/best-practices/doris-lance.mdx       | 326 +++++++++++++++++++++
 .../lakehouse/catalogs/lance-catalog.mdx           |   2 +
 .../lakehouse/best-practices/doris-lance.mdx       | 324 ++++++++++++++++++++
 .../lakehouse/catalogs/lance-catalog.mdx           |   2 +
 versioned_sidebars/version-4.x-sidebars.json       |   1 +
 5 files changed, 655 insertions(+)

diff --git 
a/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/best-practices/doris-lance.mdx
 
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/best-practices/doris-lance.mdx
new file mode 100644
index 00000000000..acf509f5c27
--- /dev/null
+++ 
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/best-practices/doris-lance.mdx
@@ -0,0 +1,326 @@
+---
+{
+    "title": "Lance 查询最佳实践",
+    "language": "zh-CN",
+    "description": "根据查询场景估算 Lance 索引和元数据缓存,选择 Index Segment 粒度,并调优 Apache 
Doris 向量、标量过滤和全文检索的性能与内存。"
+}
+---
+
+本文帮助你为 Lance 查询选择每个 BE 的缓存预算和 Index Segment 
布局。从工作负载出发,估算需要保留的缓存内容,再在目标并发下验证。示例是容量规划的起点,不是经过压测的最优值或 BE 内存上限。
+
+:::note
+Lance Catalog 是实验性功能,从 Apache Doris 4.2 开始支持。请先通过 [Lance 
Catalog](../catalogs/lance-catalog.mdx) 配置访问,并检查其 [Reader 
兼容性](../catalogs/lance-catalog.mdx#lance-版本与兼容性)。Doris 读取已有的 Lance 
索引;构建和维护索引需要使用兼容的 Lance 工具。
+:::
+
+## 1. 根据场景选择调优路径 {#choose-a-scenario}
+
+| 查询场景 | 从哪里开始 | 验证重点 |
+|---|---|---|
+| 单查询低延迟的向量搜索 | 保留经常探测的向量分区;按[粒度建议](#index-segment-granularity),相对于可用 BE 
数对比均衡的 Segment 数量 | 相近召回率下的 P95 延迟、最慢扫描器、峰值内存 |
+| 高并发向量搜索 | 为不同查询向量访问的热点分区并集做预算;为并发搜索保留内存和 CPU,而不是让每条查询都使用最多 Split | 
目标并发下的吞吐、P95 延迟、缓存重新加载和各 BE 内存 |
+| BTree 范围/点查或 Bitmap/LabelList 过滤 | 
估算[热点页面或位图条目](#scalar-index-cache-sizing),包含查找结构;确认谓词和索引实际使用情况 | 页面加载、选择率、临时 
row ID 集合及输出列读取 |
+| 全文或短语检索 | 估算词典、文档元信息、倒排,以及启用时的位置数据;预热时覆盖不同词项 | 高频词查询、短语查询内存、候选归并 |
+| 向量、标量和全文查询共享 BE | 将各自的[共享工作集](#shared-index-cache-budget)相加,并混合测试 | 
一类查询是否淘汰了另一类查询的缓存 |
+| 输出列回表或扫描占主要耗时 | 检查数据文件读取及[独立的磁盘数据缓存](../catalogs/lance-catalog.mdx#缓存类型和范围) 
| 有效读取量与读放大,解码或回表是否占主要耗时 |
+
+标量过滤也可能是向量查询 Prefilter 的一部分,因此一条查询可以同时需要标量和向量索引内容。应选择所有适用的场景,不必将它们视为互斥选项。
+
+## 2. 收集输入并确定内存预算 {#planning-inputs}
+
+为每个 BE 收集以下信息,或在真实调度条件下选择有代表性的 BE:
+
+- **工作负载:** 目标 QPS/并发、延迟目标、召回率要求、`top_k`、过滤条件、短语查询和输出列。
+- **资源:** 本次查询可用的 BE、CPU 能力,以及扣除其他 Doris 工作负载和进程/操作系统需求后可用的内存。
+- **索引:** 类型、版本、物理 Segment 数、覆盖情况、向量维度/编码、IVF 分区数,以及标量值宽度或字符串长度。
+- **复用情况:** 有代表性的查询时间窗口内访问的不同分区、页面、词项或位图条目。应使用跨查询的工作集,而不是一条查询返回的行数。
+- **查询峰值:** 目标并发下正在加载的分区、Prefilter/结果 row ID 集合、候选堆、解码、精排和二阶段回表内存。
+
+使用[索引公式](#index-cache-sizing)估算载荷,再检查总预算是否容得下:
+
+```text
+available_Lance_cache_budget_on_BE
+    = BE memory budget
+    - other BE memory needs
+    - peak concurrent query working memory
+    - safety headroom
+
+index_cache_target + metadata_cache_target
+    <= available_Lance_cache_budget_on_BE
+```
+
+即从 BE 内存预算中扣除其他 BE 
需求、并发查询峰值和安全余量,剩余部分才用于索引与元数据缓存。这是规划检查,不是强制的内存预留;缓存容量不限制查询内存或进程 
RSS。不要将可复用的缓存工作集直接乘以并发查询数,但要为这些查询额外的工作内存预留空间;并发查询访问不同数据时,还应计入缓存内容并集的增长。
+
+完整参数定义和默认值保留在 [BE 配置](../catalogs/lance-catalog.mdx#be-配置)中。将 `lance_*` BE 
参数写入各 BE 的 `be.conf`,并重启该 BE。它们不是 Catalog 属性或 SQL 会话变量。Index Segment 数量和 IVF 
构建参数需要通过 Lance 索引构建工具设置,不由这些缓存参数控制。
+
+## 3. 估算索引工作集 {#index-cache-sizing}
+
+应根据每个 BE 需要复用的索引分区,为 Index Cache 配置容量。存储中的完整索引大小、缓存工作集和查询峰值内存是三个不同的量。以下估算覆盖 
Doris 
支持的六种[向量索引类型](../catalogs/lance-catalog.mdx#支持的向量索引类型),用于容量规划,不代表精确的缓存计量或内存上限。
+
+### 各索引类型的向量载荷
+
+使用以下符号,所有公式的单位均为字节;`1 GiB = 1073741824 bytes`。
+
+| 符号 | 含义 |
+|---|---|
+| `N` | 待估算分区中的索引向量数。普通单向量列对应已索引行数;多向量列应统计已索引的子向量数量,并单独预留行映射和精排内存。 |
+| `D` | 向量维度。 |
+| `w` | 未量化向量中每个存储元素的字节数:Float16 为 2,Float32 为 4,Float64 为 8。对于二进制 `uint8` 
向量,以实际字节长度作为 `D`,并取 `w=1`。 |
+| `m`、`b` | PQ 子向量数量和每个子向量编码的位数。这是索引构建参数,不是查询参数。 |
+| `G` | 加载后的 HNSW 图内存,包括所有层级及辅助结构。 |
+
+**需要累加所有适用组件,不能只计算向量编码。** 第一张表拆分逐向量载荷,后续表格列出共享结构和 HNSW 图组件。公式估算加载后的缓冲区;标为 
**需测量** 的项目同样要预留预算,不代表零开销。共享分配只计算一次。
+
+| 索引类型 | 向量值 / 编码 | UInt64 row ID | HNSW 图 | 还需计入的预算 |
+|---|---|---|---|---|
+| `IVF_FLAT` | `N * D * w` | `8 * N` | 无 | 下表中的 IVF 结构及通用开销 |
+| `IVF_SQ` | SQ8 为 `N * D` | `8 * N` | 无 | IVF 结构、SQ 元信息及通用开销 |
+| `IVF_PQ` | `N * ceil(m * b / 8)` | `8 * N` | 无 | IVF 结构、PQ 码本及通用开销 |
+| `IVF_HNSW_FLAT` | `N * D * w` | `8 * N` | **加上 `G`** | IVF 结构及通用开销 |
+| `IVF_HNSW_SQ` | SQ8 为 `N * D` | `8 * N` | **加上 `G`** | IVF 结构、SQ 元信息及通用开销 |
+| `IVF_HNSW_PQ` | `N * ceil(m * b / 8)` | `8 * N` | **加上 `G`** | IVF 结构、PQ 
码本及通用开销 |
+
+PQ 的 `b=8` 时,每个向量的编码占 `m` 字节;`b=4` 时,在支持的偶数 `m` 配置下占 `m/2` 字节。例如,100 
万个向量、`m=96`、`b=8`,编码占 `96000000` 字节,row ID 占 `8000000` 字节,合计 `104000000` 字节,约 
99.2 MiB。压缩程度较高的索引尤其不能忽略 row ID 的开销。
+
+| 额外组件 | 适用类型 | 加载后载荷 / 预算方法 |
+|---|---|---|
+| IVF 聚类中心 | 全部六种类型 | 每套保留的 FP32 聚类中心为 `num_partitions * D * 4`;其他布局按实际元素宽度计算 |
+| IVF 分区目录 / 子索引元信息 | 全部六种类型 | **需测量**:保留的分区偏移、长度及 Reader 结构;独立于向量值和聚类中心 |
+| PQ 码本 | `IVF_PQ`、`IVF_HNSW_PQ` | 每份保留的标准 FP32 码本为 `2^b * D * 4` |
+| SQ 量化器元信息 | `IVF_SQ`、`IVF_HNSW_SQ` | **需测量**:保留的量化边界/参数;不能假设各版本采用固定布局 |
+| 范数、行覆盖结构及额外映射 | Reader 实际保留时 | **需测量**:实际数组/集合;多向量的行映射需单独预留预算 |
+| 索引对象、Arrow 有效性缓冲区、闲置容量及副本 | 全部六种类型 | **需测量**:逻辑载荷以外的分配;共享缓冲区只计算一次 |
+
+独立的 Index Segment 可能分别保留聚类中心和码本。不要将共享分配按每个分区重复计算,也不要假设所有副本都能共享。额外结构取决于嵌入的 
Lance 版本。
+
+#### HNSW 图的载荷拆分 {#hnsw-graph-payload}
+
+**图内存需要在向量编码和 row ID 之外额外计算。** 对于使用 UInt32 节点/邻居编号、可能保留 Float32 边距离的 Arrow 
布局,设 `E` 为已加载的所有层级中有向邻居条目总数,`V` 为这些层级的节点出现次数之和,`A` 为独立保留的图批次数。一个节点出现在多个层级时,需要在 
`V` 中计入多次。
+
+| 图组件 | 主要载荷,单位为字节 | 统计口径 |
+|---|---|---|
+| 邻居编号 | `4 * E` | 每个有向邻居条目一个 UInt32,统计所有层级 |
+| 保留的边距离 | `4 * E` | 每个邻居条目对应一个 Float32;**Reader 跳过距离列时为零** |
+| 图节点编号 | `4 * V` | 每次节点出现对应一个 UInt32;与数据集的 UInt64 row ID 不同 |
+| 邻居列表偏移 | `4 * (V + A)` | Int32 列表偏移,每个批次包含一个末尾偏移 |
+| 距离列表偏移 | `4 * (V + A)` | 第二组 Int32 列表偏移数组;**Reader 跳过距离列时为零** |
+| **保留距离时的小计** | **`8 * E + 12 * V + 8 * A`** | 大批次可近似为 `8 * E + 12 * 
V`,尚不是完整的 `G` |
+| **不保留距离时的小计** | **`4 * E + 8 * V + 4 * A`** | 仅保留节点编号和邻居列表时使用;两种小计选择其一,不累加 |
+| 层级偏移、入口点及节点查找结构 | **需测量** | 计入每个已加载图保留的元信息和查找结构分配 |
+| 有效性缓冲区、图对象及闲置分配容量 | **需测量** | 计入上述缓冲区之外的实际分配,不重复计算共享缓冲区 |
+| **完整图预算 `G`** | **主要小计 + 测得的图额外开销** | 与向量/row ID 载荷及其他索引结构相加 |
+
+这是特定布局的估算,不是每个节点固定占用的保证。HNSW 的 `m` 控制图连接数,与 PQ 的 `m` 无关。不要用 `N * m * 4` 代替 
`G`:实际度数、上层节点、保留的距离以及 Reader 版本都会影响结果。应加载有代表性的分区,测量其保留的分配以校准估算。
+
+例如,假设缓存的 `IVF_HNSW_PQ` 分区共包含 100 万个向量,PQ 
`m=96`、`b=8`,`E=32000000`、`V=1100000`、`A=1`。假设此 Reader 保留距离列。这里的图规模是示例输入,不是由 
HNSW 的 `m` 推导出的数值:
+
+| 组件 | 示例载荷 |
+|---|---|
+| PQ 编码 | `96000000` 字节 |
+| 数据集 row ID | `8000000` 字节 |
+| 主要图缓冲区 | `8 * 32000000 + 12 * 1100000 + 8 = 269200008` 字节,约 256.7 MiB |
+| **编码 + row ID + 主要图缓冲区** | **`373200008` 字节,约 355.9 MiB** |
+| 备选:跳过距离列时的主要图缓冲区 | `4 * 32000000 + 8 * 1100000 + 4 = 136800004` 字节,约 130.5 
MiB |
+| **备选合计:编码 + row ID + 不保留距离的图缓冲区** | **`240800004` 字节,约 229.6 MiB** |
+| 两种情况均需加上 | 上表中的图额外开销、IVF 聚类中心/目录、PQ 码本及通用开销 |
+
+按 Reader 实际保留的列选择对应示例。两种小计都不是完整的缓存容量建议。
+
+上游 [Lance Performance Guide](https://lance.org/guide/performance/) 
提供向量载荷和索引构建内存的估算,[Lance 向量索引格式](https://lance.org/format/index/vector/) 说明 IVF 
分区、子索引和量化存储的组成。上游最新文档可能包含 Doris 嵌入版本之后的功能,兼容性应以 Lance Catalog 页面的支持范围为准。
+
+### 标量与全文索引的载荷 {#scalar-index-cache-sizing}
+
+标量索引也可以估算,但不同索引类型需要的输入不同。应统计单个 BE 实际保留的条目,而不只是查询返回的行数。以下方法描述 Lance 的存储布局,不扩展 
Doris 的谓词下推或索引兼容性范围。应确认嵌入的 Lance 版本支持该索引,并确认查询实际使用了它。
+
+**BTree。** 加载后的叶子页包含字段值和 UInt64 row ID。设缓存叶子页共包含 `R` 条记录、`H` 个页面,定长值宽度为 
`w`,这些页面中字符串的平均字节长度为 `L`:
+
+| 叶子页组件 / 表示方式 | 值缓冲区 | 偏移数组 | UInt64 row ID |
+|---|---|---|---|
+| Int32 / Float32 | `4 * R` | 无 | `8 * R` |
+| Int64 / Float64 / 64 位时间戳 | `8 * R` | 无 | `8 * R` |
+| 其他定长值 | `w * R`;按位存储的 Boolean 使用实际位图缓冲区 | 无 | `8 * R` |
+| Arrow Utf8 / Binary | `R * L` | Int32 偏移为 `4 * (R + H)` | `8 * R` |
+| Arrow LargeUtf8 / LargeBinary | `R * L` | Int64 偏移为 `8 * (R + H)` | `8 * R` |
+
+**叶子页小计为后三列载荷之和。** 每个已打开的 Segment 还需计入下列组件;这里的 `P` 是该 Segment 的全部叶子页数,不只是缓存页数:
+
+| BTree 额外组件 | 主要载荷 / 预算方法 |
+|---|---|
+| Lookup 最小值 | 定长值为 `P * w`;变长值使用实际值/偏移缓冲区 |
+| Lookup 最大值 | 定长值为 `P * w`;变长值使用实际值/偏移缓冲区 |
+| Lookup NULL 计数 | UInt32 布局为 `4 * P` |
+| Lookup 页号 | UInt32 布局为 `4 * P` |
+| **定长 Lookup 缓冲区小计** | **`P * (2 * w + 8)`**,不含下面的额外结构 |
+| NULL 页面列表及查找对象 | **需测量**:保留的列表、对象及实现中的副本 |
+| 叶子页/Lookup 有效性缓冲区、Arrow 对象及闲置容量 | **需测量**:逻辑值/偏移/row ID 以外的分配 |
+
+字符串公式假设使用表中的 Arrow 布局;字典或 view 表示需要使用实际缓冲区。磁盘压缩后的大小不能代表解码后的页面大小。BTree 默认每页 
4096 **条记录**,不是字节;应使用索引的实际页大小。Lookup 属于索引工作集,不能因为它描述页面就划入 BE Metadata Cache 预算。
+
+例如,1 亿个 Int64 值的叶子页值和 row ID 载荷为 `100000000 * 16 = 1600000000` 字节,约 1.49 
GiB。如果缓存页面覆盖 1000 万条记录,则该载荷约为 152.6 MiB。按每页 4096 条记录计算,完整索引约有 24415 页;lookup 
batch 的 min/max/计数/页号载荷约 572 KiB,尚未包含其他结构。返回十条匹配记录不代表只缓存十条记录:Reader 加载的是页面。
+
+**Bitmap、LabelList 与 NGram。** 累加下表中适用的组件。成员关系位图不能直接按每个匹配行 8 字节计算。LabelList 
按不同标签统计成员关系,同一行可以出现在多个标签位图中。NGram 按不同 gram/行组合统计成员关系,同一行内重复出现同一个 gram 
不增加成员关系。这里的 NGram 索引与使用 n-gram 分词器的 FTS 索引不同。
+
+| 索引类型 | 组件 | 加载后载荷 / 预算方法 |
+|---|---|---|
+| Bitmap | 键查找结构 | **需测量**:保留的键值、条目引用及查找对象 |
+| LabelList | 标签查找结构 / 包装结构 | **需测量**:不同标签键、条目引用及包装对象 |
+| NGram | Gram 查找结构 | **需测量**:保留的 gram、条目引用及查找对象 |
+| 全部三种类型 | 成员关系位图:数组容器 | 每个覆盖 65536 个取值范围、含 `c` 个 UInt16 条目的常规 Roaring 数组容器约为 
`2 * c` |
+| 全部三种类型 | 成员关系位图:稠密容器 | 每个覆盖 65536 个取值范围的常规稠密容器为 `8192` 字节 |
+| 全部三种类型 | 成员关系位图:run / 完整 Fragment 表示 | **需测量**:实际表示;每个容器按实际选用的表示计算,不将各备选表示累加 
|
+| 全部三种类型 | 位图目录 / 容量 | **需测量**:Fragment/高位目录、容器对象及闲置容量 |
+| 实际存在时 | NULL / NULL 列表跟踪 | **需测量**:单独计入保留的 NULL 成员关系结构 |
+| **每种类型的合计** | **查找结构 + 已加载的成员关系表示 + 目录 + NULL 跟踪** | 
共享分配只计算一次;序列化位图大小不能保证驻留内存大小 |
+
+例如,256 个已加载的稠密容器有 2 MiB 的 bitset 载荷,**尚未包含**键、目录、NULL 跟踪和对象开销。
+
+**其他标量与全文索引。** 每行列出一个独立组件;小计行汇总之前的组件,不能再次累加。只统计该 BE 保留的内容。**需测量**表示需要使用 Reader 
的实际分配,不表示该组件可以忽略。
+
+| 索引类型 | 组件 | 加载后载荷 / 预算方法 |
+|---|---|---|
+| FTS / Inverted | 词典 | **需测量**:加载的词项、词典结构及查找分配 |
+| FTS / Inverted | 文档元信息 | **需测量**:保留的文档映射、长度及其他 Reader 元信息 |
+| FTS / Inverted | 未压缩倒排 row ID | `T` 个 UInt64 词项/文档条目为 `8 * T` |
+| FTS / Inverted | 未压缩倒排词频 | Float32 词频为 `4 * T`;ID 与词频合计 `12 * T` |
+| FTS / Inverted | 位置数据,启用时 | `O` 次存储的 Int32 出现位置为 `4 * O` |
+| FTS / Inverted | 位置偏移 | 实际保留的偏移数组字节数,独立于 `4 * O` |
+| FTS / Inverted | 压缩倒排布局 | 使用实际保留的压缩块**替代**对应的未压缩布局公式 |
+| FTS / Inverted | 评分 / 块元信息及分配开销 | **需测量**:保留的块描述、评分数据、对象及闲置容量 |
+| ZoneMap | 最小值 | `Z` 个加载的 zone、宽度为 `w` 的定长值,逻辑载荷为 `Z * w` |
+| ZoneMap | 最大值 | 逻辑载荷为 `Z * w`;min/max 合计 `2 * Z * w` |
+| ZoneMap | 计数及 zone 边界 | **需测量**:保留的计数和边界;按各 Fragment 统计 zone,包括不足一个 zone 的尾部 
|
+| ZoneMap | NULL 跟踪及装箱值/对象开销 | **需测量**:可选的 NULL 行集合及实际标量分配;可能大于 min/max 的逻辑字节数 
|
+| BloomFilter | 过滤器位数组 | 累加实际分配的过滤器字节数;`Z` 个等大的过滤器、每个 `F` 字节时为 `Z * F` |
+| BloomFilter | Zone 描述 / 过滤器对象 | **需测量**:保留的描述、过滤器参数及对象开销 |
+| BloomFilter | NULL 跟踪 | **需测量**:可选的 NULL 行集合 |
+| RTree | 包围盒坐标 | 每个二维条目有四个 Float64 坐标,为 `32 * Q`;`Q` 统计所有缓存树层级的条目 |
+| RTree | 行 / 子页面 ID | UInt64 ID 为 `8 * Q`;坐标与 ID 合计 `40 * Q` |
+| RTree | 树元信息及页面偏移 | **需测量**:保留的树描述及页面偏移缓冲区 |
+| RTree | NULL 跟踪、Arrow/对象及闲置容量 | **需测量**:实际额外分配 |
+| JSON-path 包装类型 | 底层索引 | 使用所选标量索引类型的完整组件预算 |
+| JSON-path 包装类型 | 路径及包装结构 | **需测量**:在底层索引之外保留的路径和包装分配 |
+
+BloomFilter 
应使用构建器实际按块取整后的分配;仅凭行数和误判率不能描述所有布局。对于不熟悉或实验性的布局,应测量加载后的分配,避免套用其他索引类型的公式。匹配行数很少,也可能仍需较大的词典、树或文档元信息。
+
+**查询工作内存与上述缓存表格分开预留:**临时交集/并集、结果 row ID 集合、近似过滤器复核、RTree 
精确复核读取的原始几何值,以及输出列读取都需要单独预算。
+
+底层结构可参考上游 [BTree 格式](https://lance.org/format/index/scalar/btree/)、[Bitmap 
格式](https://lance.org/format/index/scalar/bitmap/)和 [FTS 
格式](https://lance.org/format/index/scalar/fts/)。这里估算的是加载后的载荷,不是压缩文件大小或查询总 RSS。
+
+### 从索引大小推导每个 BE 的工作集
+
+对于 Segment `s`,令 `N_s` 为索引向量数,`P_s` 为 IVF 分区数,`H_s` 为某个 BE 需要保留的不同分区数。分区较均衡时:
+
+```text
+cached_vectors_on_BE ≈ sum over segments (N_s * H_s / P_s)
+index_cache_target ≈ cached vector/row-ID payload
+                   + cached HNSW graphs
+                   + centroids, codebooks and other index allocations
+                   + measured headroom
+```
+
+即:缓存目标容量约为缓存向量及 row ID 载荷、HNSW 图、聚类中心与码本等索引分配,加上经测量确定的余量。分区倾斜时,应使用实际分区行数。`H_s` 
是跨查询的工作集,不是某次查询的 `nprobes`;不同查询可能逐渐访问所有分区,重复或并发查询也可能共享同一分区缓存。
+
+按不同的 `(dataset, index segment, partition)` 统计缓存项,并计入该 BE 访问的所有表、索引和版本。一个 BE 
的缓存不会预热其他 BE,调度也不保证每个 Index Segment 始终在同一个 BE 
上执行。除非实际分配和复用情况支持这一假设,否则不要直接用完整索引大小除以 BE 数量。
+
+查询还会额外占用内存:正在加载的分区、被缓存淘汰但仍由查询持有的对象、Prefilter 
状态、距离表、候选堆、解码批次、精排和二阶段回表。`top_k`、`ef`、`refine_factor`、扫描并发和 `nprobes` 
会影响这些内存或访问的工作集,但不会改变每条向量的存储字节数。缓存配置不是 Lance 或 BE RSS 
的硬上限,增大缓存也不会消除距离计算和每次查询的过滤工作。
+
+### 配置示例与调优
+
+假设一个 Index Segment 中有 1 亿个 FP32、768 维向量,索引为 `IVF_PQ`,`m=96`、`b=8`,分为 4096 
个均衡分区。某个 BE 上的典型工作负载反复访问其中 1024 个不同分区:
+
+```text
+Full code/row-ID payload = 100000000 * (96 + 8) = 10400000000 bytes ≈ 9.69 GiB
+Cached payload          = 10400000000 * 1024 / 4096 = 2600000000 bytes ≈ 2.42 
GiB
+IVF centroids           = 4096 * 768 * 4 = 12582912 bytes = 12 MiB
+PQ codebook             = 256 * 768 * 4 = 786432 bytes = 0.75 MiB
+```
+
+如果这是该 BE 上唯一较大的索引工作集,且 BE 内存预算允许,可将 4 GiB 作为初始 Index Cache 
容量,在计算出的载荷之外为其他索引分配留出空间。该值需要验证,不代表通用的开销比例。如果工作负载需要完整索引常驻,4 GiB 就不够;即使默认的 10 
GiB,也已接近未计入其他分配的 9.69 GiB 载荷。
+
+若 BE 在预留查询峰值、其他 Doris 工作负载、进程和操作系统余量后,仍可分配 4 GiB 索引缓存和 1 GiB 元数据缓存,可配置:
+
+```properties
+lance_index_cache_size_bytes = 4294967296
+lance_metadata_cache_size_bytes = 1073741824
+```
+
+重启该 BE 后生效。应分别为每个 BE 做预算;这些是 BE 配置参数,不是 Catalog 属性或 SQL 会话变量。
+
+1. 使用有代表性的查询及并发预热,覆盖不同查询向量、过滤条件和表。仅重复同一个查询向量会低估多样化工作负载的工作集。
+2. 对比 `doris_be_lance_session_index_cache_usage_bytes` 与 
`doris_be_lance_session_index_cache_capacity_bytes`、时间窗口内 
`hits_total`/`misses_total` 的增量,以及每次查询的 
`LanceIndexPartitionCacheMissLoads`。使用量接近容量且持续未命中,可能意味着缓存反复淘汰;新分区、新索引版本或切换 BE 
也会造成未命中。只有确认存在复用需求且内存允许时才增大容量。
+3. BE 内存预算允许时,Metadata Cache 可从默认的 1 GiB 开始,根据其 `usage_bytes` 
和命中/未命中指标调整。文件/Fragment 数量、Schema/页面元信息、删除信息和 row ID 
映射决定其需求,不能按向量索引大小的固定比例配置。BE Metadata Cache 与 FE 表访问缓存相互独立。
+4. 增大任何一个缓存前,都应在目标并发下检查 BE 
进程内存和查询峰值。缓存命中率高但查询仍慢,可能是评分、Prefilter、解码或回表开销,增大缓存不一定有效。
+5. 单独为精排和输出列回表等重复数据文件读取配置 `lance_data_cache_disk_capacity_bytes`。增大它不会扩大 Index 
Cache,也不能弥补索引缓存未命中:索引文件绕过这个磁盘数据缓存。`lance_data_cache_read_block_size_bytes` 
建议先保持默认的 1 MiB,再根据有效读取量与读放大调整。
+
+## 4. 为共享缓存分配预算 {#shared-index-cache-budget}
+
+每个 BE 使用共享的 Lance Session 和一份 Index Cache 容量。向量分区、BTree 页面、Bitmap/LabelList 
条目和全文索引内容会竞争这份空间,同一 BE 访问的不同表、不同 Catalog 的索引也包括在内。缓存键用于区分条目身份,不代表预留容量。Lance BE 
缓存参数没有提供按表或索引类型设置配额、固定驻留的控制能力。
+
+只有加载后的内容占用缓存,仅创建多个索引不会自动将它们全部加载。查询开始使用这些索引后,加载一个索引可能淘汰另一个索引的条目。例如,大范围 BTree 
扫描或多样化的全文查询可能替换已经预热的向量分区,后续向量查询就需要重新加载。索引版本变化也可能在旧条目仍驻留时引入新条目。缓存淘汰不会立即释放仍被运行中查询引用的对象。
+
+应为混合工作集做预算,避免重复计算共享分配:
+
+```text
+index_cache_target_on_BE ≈ vector working set
+                        + BTree working set
+                        + Bitmap / LabelList / NGram working sets
+                        + FTS and other index working sets
+                        + measured headroom
+```
+
+即将各类索引工作集相加,再预留经测量确定的余量。例如,测量和载荷估算得出向量占 2.5 GiB、BTree 占 0.25 GiB、Bitmap 占 0.5 
GiB、FTS 占 0.75 GiB,且都已包含各自的索引结构,总工作集为 4 GiB。初始配置 5 GiB Index Cache 可为增长和估算误差留出 
1 GiB,但不保证适合所有工作负载。如果 BE 还容得下 1 GiB Metadata Cache,**以及**查询峰值内存、其他 Doris 
工作负载和进程/操作系统余量:
+
+```properties
+lance_index_cache_size_bytes = 5368709120
+lance_metadata_cache_size_bytes = 1073741824
+```
+
+Index Cache 和 Metadata Cache 容量独立:元数据条目不会直接淘汰索引条目,增大其中一个也不会扩大另一个。但两者仍消耗同一个 BE 
进程的内存。磁盘数据缓存是第三份预算,且不缓存索引文件。同一 BE 上拆分 Catalog 不能隔离 Index 
Cache;需要隔离的工作负载,应使用部署环境支持的独立 BE 资源和路由能力。
+
+排查相互抢占时,可先预热一类查询,记录延迟和缓存指标增量;再运行另一类查询,随后重复第一类查询。在索引版本和 BE 
分配具有可比性的前提下,如果接近容量时重新出现未命中或加载,说明可能存在淘汰压力。Session 命中/未命中指标汇总了多种工作负载,需要结合查询 
Profile 分析。预热所有索引本身也可能淘汰希望保留的工作集。应优先使用有代表性的混合工作负载预热,仅在有复用需求且内存允许时增大容量。
+
+## 5. 选择 Index Segment 粒度 {#index-segment-granularity}
+
+Index Segment、数据 Fragment、IVF Partition 和 BTree 叶子页是不同单位。当前 Doris 
向量搜索会将每个选中的物理 Index Segment 转换成一个扫描 Split,未被索引覆盖的 Fragment 增加 Flat Search 
Split。FTS 同样使用物理索引 Segment。单个向量 Segment 内的 IVF Partition 不会被拆成独立 Doris Split 
分配到多个 BE,但 Lance 内部仍可以并行执行。普通标量扫描是否按 Segment 
分组,取决于选中的索引和执行计划,不能假设所有标量索引都采用向量搜索的拆分策略。
+
+设 `B` 为本次查询实际可用的 BE 数量,`S` 为选中的向量 Segment 数量。如果 `S` 小于 `B`,这条查询就没有足够的索引扫描 
Split 分配到所有这些 BE。增加 `S` 会增加调度机会,但不保证所有 Split 同时执行;扫描器限制、Lance 
内部并发、CPU、I/O、缓存驻留情况及其他并发查询都会影响结果。仅增加 Doris Pipeline 实例不会将一个 Index Segment 继续拆小。
+
+对于较大数据集,可将 `B`、`2 * B`、`4 * B` 个相对均衡的向量 Segment 
作为**初始对比实验**,不是官方最优值,也不是小数据集必须满足的要求。以一个示例性的 1 亿向量数据集、8 个可用 BE 为例:
+
+| Segment 数量 | 平均每 Segment 向量数 | 评估重点 |
+|---|---|---|
+| 8 | 1250 万 | Split 足以覆盖 BE,每 Segment 的重复操作较少 |
+| 16 | 625 万 | 调度更灵活,单个大 Segment 对耗时的影响较小 |
+| 32 | 312.5 万 | 任务更多,但搜索、加载和合并开销可能增加 |
+
+应尽量均衡预计搜索工作量和加载后的索引字节数,而不只是 Fragment 个数。对于维度和编码相同的向量,行数均衡是合理起点;HNSW 图大小、IVF 
倾斜、标量过滤选择率和文档长度都可能改变负载。没有通用的每 Segment 
行数或文件大小阈值。高吞吐场景可能适合减少每条查询的任务数,为并发查询保留资源;单查询延迟优先的场景则可能受益于更多 Split。
+
+Segment 更多也意味着更多的逐 Segment 初始化、元数据/量化器对象和局部候选生成。每个向量/FTS Split 最多保留 `top_k + 
offset` 个候选,Doris 不会将这个上限除以 `S`。因此扫描侧候选上限为 `S * (top_k + 
offset)`,同时受实际匹配数量限制。例如,`top_k=1000`、无 offset、32 个 Split,在 Doris 归并前最多有 32000 
个候选。本地 TopN 可能减少网络传输,因此这不是网络行数的保证。`top_k` 较大时尤其要关注候选处理开销。
+
+Segment 数应与每个 Segment 的 IVF `num_partitions`、查询 `nprobes` 一起调优。分区均衡时,粗略探测向量数为 
`sum(N_s * nprobes_s / P_s)`,每个 Segment 最多为 `N_s`。这不是 HNSW 比较次数的公式。重建 Segment 
或改变分区数,即使 `nprobes` 相同也可能改变召回率;应在相近召回率下比较延迟。HNSW 的 `ef` 和精排配置也会影响对比。
+
+使用 Lance 工具构建和维护 Segment,再提交为目标逻辑索引。多个名字不同、分别覆盖全表的索引,不能替代一个包含多个物理 Segment 
的索引。优先按完整 Fragment 组成均衡分组,避免小批追加不断积累大量微小 Segment,并使用嵌入 Lance 版本支持的维护 API。构建后在 
Doris `EXPLAIN`/Profile 中确认物理 Segment 数和未覆盖 Fragment 数,不要从 `SHOW INDEX` 
的结果行数推断。将所有 Segment 合并成一个,可能降低 Doris 扫描并行度。构建及提交的概念可参考 [Lance 
分布式索引](https://lance.org/guide/distributed_indexing/)。
+
+对于 BTree 点查或高选择性的 Bitmap 查询,过多 Segment 可能增加查找和打开开销,而每个 Segment 减少的工作量很少。FTS 
会增加逐 Segment 的词典/倒排搜索和候选归并。应分别评估这些工作负载,不要为所有索引类型套用向量索引的 Segment 数。拆小 Segment 
也不能消除不必要的逐行 Prefilter 构建,应独立确认 Reader 的过滤行为。
+
+## 6. 验证混合工作负载 {#validate-workload}
+
+1. 确认索引选择、物理 Split 和索引覆盖情况,将回退扫描和二阶段回表纳入延迟分析。
+2. 估算每个索引保留的载荷,将共享工作集相加。查询内存单独预留,包括并行加载、row ID 集合、候选堆和解码。
+3. 使用真实并发对比冷缓存和预热后的混合工作负载,跟踪 P95 延迟、吞吐量、适用时的召回率、各 BE 峰值内存、缓存使用量与命中/未命中增量、索引加载 
I/O 以及扫描时间倾斜。
+4. 每次改变一个维度:缓存容量、Segment 数或搜索参数。若重建 Segment 改变了模型,应重新验证召回率。继续增加 Split 
或缓存不再改善吞吐、延迟或内存表现时,停止增加。
+5. 大量追加数据、重建索引、扩缩 BE 或改变查询组合后重新评估。针对一个查询向量或一张表校准的配置,不能作为所有工作负载的固定预算。
+
+## 7. 增加资源前先定位问题 {#troubleshooting}
+
+| 现象 | 优先检查 | 可评估的调整 |
+|---|---|---|
+| 缓存接近容量,索引持续未命中 | 索引/版本变化、BE 分配和总工作集 | 内存允许时增大 
`lance_index_cache_size_bytes`,减少无用预热,或用独立 BE 资源隔离工作负载 |
+| 标量或全文查询后,向量查询变慢 | 按[共享缓存预算](#shared-index-cache-budget)重复跨工作负载淘汰检查 | 
为两类查询一起做预算;同一 BE 上拆分 Catalog 不能隔离缓存 |
+| 大型已索引数据集只有少量 BE 扫描 | 物理 Segment Split 数、可用 BE 和扫描器调度 | 对比更多且均衡的 Segment;仅增大 
Pipeline 实例不能拆分 Segment |
+| Segment 增多后 CPU、内存或归并耗时上升 | 每 Split 候选数、`top_k + offset` 和查询并发 | 减少 
Segment,或避免请求超过应用需求的候选数量;搜索参数变化后重新检查召回率 |
+| 缓存命中率高但延迟仍高 | 评分、Prefilter 工作、解码和二阶段回表耗时 | 优化主要耗时阶段;增大 Index Cache 不能消除计算或回表 
|
+| 元数据反复加载 | Metadata Cache 使用量及命中/未命中增量、Fragment/文件数和版本变化 | 有复用需求且内存允许时,单独调整 
`lance_metadata_cache_size_bytes` |
+| 缓存未满但进程内存很高 | 并发查询分配、正在加载的数据、仍被引用的对象和其他 BE 工作负载 | 
根据测得的峰值降低并发或缓存预算;缓存容量不是进程内存上限 |
+| 数据文件 I/O 持续较高 | 磁盘缓存计数、复用情况、读取块大小和输出列 | 调整数据缓存预算或读取粒度;索引文件不使用这个磁盘数据缓存 |
+
+使用[查看缓存效果](../catalogs/lance-catalog.mdx#查看缓存效果)中的 Profile 计数器和 BE 
指标。仅凭汇总命中率无法判断哪个索引被淘汰,或哪个阶段占主要耗时。每次改变一个配置,记录每组对比的工作负载、索引版本、BE 分配、延迟、吞吐、召回率和峰值内存。
diff --git 
a/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
 
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
index 7e6328025a7..a18f25b0ec6 100644
--- 
a/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
+++ 
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
@@ -995,6 +995,8 @@ lance_data_cache_read_block_size_bytes = 1048576
 
 配置容量时应同时考虑索引缓存、元数据缓存和查询工作内存。上述容量不是 Lance 查询总内存的上限,也不是 BE 
进程内存的上限。磁盘缓存还需要本地磁盘空间和文件描述符资源;增大读取块可以减少块数量,但小范围读取可能产生更多额外 
I/O。调整读取块大小或磁盘布局时使用新的缓存目录。
 
+按场景估算缓存、选择配置、分析共享缓存竞争以及确定 Index Segment 粒度,请参阅 [Lance 
查询最佳实践](../best-practices/doris-lance.mdx)。
+
 ### 查看缓存效果
 
 在查询 Profile 的 `LanceReader` 下查看以下计数器:
diff --git 
a/versioned_docs/version-4.x/lakehouse/best-practices/doris-lance.mdx 
b/versioned_docs/version-4.x/lakehouse/best-practices/doris-lance.mdx
new file mode 100644
index 00000000000..137642cdc95
--- /dev/null
+++ b/versioned_docs/version-4.x/lakehouse/best-practices/doris-lance.mdx
@@ -0,0 +1,324 @@
+---
+{
+    "title": "Lance Query Best Practices",
+    "language": "en",
+    "description": "Size Lance index and metadata caches, choose index segment 
granularity, and tune Apache Doris vector, scalar, and full-text queries for 
your workload."
+}
+---
+
+Use this guide to choose a per-BE cache budget and index segment layout for 
your Lance queries. Start with your workload, estimate the content that needs 
to stay cached, and validate it under the intended concurrency. The examples 
are planning starting points, not benchmarked optima or BE memory limits.
+
+:::note
+Lance Catalog is experimental and supported starting from Apache Doris 4.2. 
First configure access using [Lance Catalog](../catalogs/lance-catalog.mdx) and 
check its [reader 
compatibility](../catalogs/lance-catalog.mdx#lance-version-and-compatibility). 
Doris reads existing Lance indexes; use compatible Lance tooling to build or 
maintain them.
+:::
+
+## 1. Choose a Tuning Path {#choose-a-scenario}
+
+| Your workload | Start with | What to validate |
+|---|---|---|
+| Vector search with low single-query latency | Retain frequently probed 
vector partitions; compare balanced segment counts relative to eligible BEs 
using the [segment guide](#index-segment-granularity) | P95 latency at 
comparable recall; slowest scanner; peak memory |
+| High-concurrency vector search | Budget for the union of hot partitions 
across different query vectors; leave memory and CPU for concurrent searches 
instead of maximizing splits for each query | Throughput and P95 latency at 
target concurrency, cache reloads, per-BE memory |
+| BTree range/point filters or Bitmap/LabelList filters | Estimate [hot pages 
or bitmap entries](#scalar-index-cache-sizing), including lookup structures; 
verify predicate/index use | Page loads, selectivity, temporary row-ID sets, 
and output-column reads |
+| Full-text or phrase search | Estimate dictionaries, document metadata, 
postings, and positions when enabled; include diverse terms in warmup | 
Common-term queries, phrase-query memory, and candidate merging |
+| Vector, scalar, and full-text queries sharing BEs | Add their [shared 
working sets](#shared-index-cache-budget) and test them together | Whether one 
query family evicts another's entries |
+| Output-column fetches or scans dominate | Inspect data-file reads and the 
[separate disk data cache](../catalogs/lance-catalog.mdx#cache-types-and-scope) 
| Useful bytes versus read amplification; whether decoding or fetches dominate |
+
+A scalar filter may also be part of a vector query's Prefilter, so a single 
query can need both scalar and vector index content. Choose all applicable rows 
rather than treating the scenarios as mutually exclusive.
+
+## 2. Collect Inputs and Set a Memory Budget {#planning-inputs}
+
+Collect these inputs for each BE or for a representative BE under realistic 
placement:
+
+- **Workload:** target QPS/concurrency, latency target, required recall, 
`top_k`, filters, phrase queries, and output columns.
+- **Resources:** BEs eligible for the queries, CPU capacity, and memory 
available after other Doris workloads and process/OS needs.
+- **Indexes:** types, versions, physical segment counts, coverage, vector 
dimensions/encoding, IVF partition counts, and scalar value widths or string 
lengths.
+- **Reuse:** distinct partitions, pages, terms, or bitmap entries accessed 
across a representative query window. Use the working set across queries, not 
one query's returned rows.
+- **Query peaks:** memory for active partition loads, Prefilter/result row-ID 
sets, candidate heaps, decoding, refinement, and second-phase fetches at the 
intended concurrency.
+
+Use the [index formulas](#index-cache-sizing) to estimate payload, then check 
that the combined budget fits:
+
+```text
+available_Lance_cache_budget_on_BE
+    = BE memory budget
+    - other BE memory needs
+    - peak concurrent query working memory
+    - safety headroom
+
+index_cache_target + metadata_cache_target
+    <= available_Lance_cache_budget_on_BE
+```
+
+This is a planning check, not an enforced memory reservation. Cache capacities 
do not cap query memory or process RSS. Do not multiply a reusable cached 
working set by the number of concurrent queries; do reserve for their 
additional query working memory, and account for a larger union of cached 
content when concurrent queries access different data.
+
+The full parameter definitions and defaults remain in [BE 
Configuration](../catalogs/lance-catalog.mdx#be-configuration). Set the 
`lance_*` BE parameters in each BE's `be.conf` and restart that BE. They are 
not catalog properties or SQL session variables. Index segment counts and IVF 
construction parameters are set through Lance index-building tooling, not these 
cache parameters.
+
+## 3. Estimate the Index Working Set {#index-cache-sizing}
+
+Size the index cache for the index partitions that each BE needs to reuse. The 
total index on storage, the cached working set, and peak query memory are 
different quantities. The estimates below cover the six [vector index types 
supported by 
Doris](../catalogs/lance-catalog.mdx#supported-vector-index-types); they are 
planning estimates, not exact cache accounting or memory limits.
+
+### Vector Payload by Index Type
+
+Use the following symbols. All formulas return bytes; `1 GiB = 1073741824 
bytes`.
+
+| Symbol | Meaning |
+|---|---|
+| `N` | Number of indexed vectors in the partitions being estimated. For an 
ordinary single-vector column this is the indexed row count. For multi-vector 
columns, count the indexed subvectors and allow separately for row mappings and 
refinement. |
+| `D` | Vector dimension. |
+| `w` | Bytes per stored, unquantized vector element: 2 for Float16, 4 for 
Float32, or 8 for Float64. For binary `uint8` vectors, use the actual byte 
length as `D` and `w=1`. |
+| `m`, `b` | PQ subvector count and bits per subvector code. These are 
index-build settings, not query parameters. |
+| `G` | Memory occupied by the loaded HNSW graph, including all levels and 
auxiliary structures. |
+
+**Add all applicable components, not just the vector codes.** The first table 
separates the per-vector payload; the following tables list shared structures 
and HNSW graph components. All formulas estimate loaded buffers. Rows marked 
**Measure** still need a budget; they are not zero-cost items. Count each 
shared allocation once.
+
+| Index type | Vector values / codes | UInt64 row IDs | HNSW graph | Other 
required budget |
+|---|---|---|---|---|
+| `IVF_FLAT` | `N * D * w` | `8 * N` | None | IVF structures + common overhead 
below |
+| `IVF_SQ` | `N * D` for SQ8 | `8 * N` | None | IVF structures + SQ metadata + 
common overhead |
+| `IVF_PQ` | `N * ceil(m * b / 8)` | `8 * N` | None | IVF structures + PQ 
codebook + common overhead |
+| `IVF_HNSW_FLAT` | `N * D * w` | `8 * N` | **Add `G`** | IVF structures + 
common overhead |
+| `IVF_HNSW_SQ` | `N * D` for SQ8 | `8 * N` | **Add `G`** | IVF structures + 
SQ metadata + common overhead |
+| `IVF_HNSW_PQ` | `N * ceil(m * b / 8)` | `8 * N` | **Add `G`** | IVF 
structures + PQ codebook + common overhead |
+
+For PQ with `b=8`, the code occupies `m` bytes per vector; with `b=4`, it 
occupies `m/2` bytes for supported, even `m`. For example, 1 million vectors 
with `m=96` and `b=8` have `96000000` bytes of codes and `8000000` bytes of row 
IDs: `104000000` bytes (about 99.2 MiB) combined. The row IDs must not be 
omitted when estimating strongly compressed indexes.
+
+| Additional component | Applies to | Loaded payload / how to budget |
+|---|---|---|
+| IVF centroids | All six types | FP32: `num_partitions * D * 4` per retained 
centroid set; use the actual element width for other layouts |
+| IVF partition directory / sub-index metadata | All six types | **Measure** 
retained partition offsets, lengths, and reader structures; distinct from 
vector values and centroids |
+| PQ codebook | `IVF_PQ`, `IVF_HNSW_PQ` | Standard FP32 codebook: `2^b * D * 
4` per retained copy |
+| SQ quantizer metadata | `IVF_SQ`, `IVF_HNSW_SQ` | **Measure** retained 
quantizer bounds/parameters; do not assume one fixed layout across versions |
+| Norms, row-coverage structures, additional mappings | Where retained by the 
reader | **Measure** actual arrays/sets; multi-vector row mappings need their 
own budget |
+| Index objects, Arrow validity buffers, spare capacity and copies | All six 
types | **Measure** allocations beyond the logical payload; count shared 
buffers once |
+
+Independent index segments can retain separate centroid sets and codebooks. Do 
not multiply shared allocations by every partition or assume all copies are 
shared. The extra structures depend on the embedded Lance version.
+
+#### HNSW Graph Components {#hnsw-graph-payload}
+
+**The graph is additional to vector codes and row IDs.** For an Arrow-backed 
Lance graph with UInt32 node/neighbor IDs and optional retained Float32 edge 
distances, let `E` be the stored directed neighbor entries across all loaded 
levels, `V` the node occurrences across those levels, and `A` the number of 
independently retained graph batches. A node present at several levels 
contributes several times to `V`.
+
+| Graph component | Main payload, in bytes | Counting rule |
+|---|---|---|
+| Neighbor IDs | `4 * E` | UInt32 per directed neighbor entry; count all 
levels |
+| Stored edge distances | `4 * E` | Float32 per neighbor entry; **zero when 
the reader skips the distance column** |
+| Graph node IDs | `4 * V` | UInt32 per node occurrence; these are distinct 
from dataset UInt64 row IDs |
+| Neighbor-list offsets | `4 * (V + A)` | Int32 list offsets, including one 
terminal offset per batch |
+| Distance-list offsets | `4 * (V + A)` | A second Int32 list-offset array; 
**zero when the reader skips the distance column** |
+| **Subtotal with distances** | **`8 * E + 12 * V + 8 * A`** | Approximately 
`8 * E + 12 * V` for large batches; not the complete `G` |
+| **Subtotal without distances** | **`4 * E + 8 * V + 4 * A`** | Use when only 
node IDs and neighbor lists are retained; select one subtotal, not both |
+| Level offsets, entry points and node lookup structures | **Measure** | 
Include retained metadata and lookup allocations for each loaded graph |
+| Validity buffers, graph objects and spare allocation capacity | **Measure** 
| Add actual allocations beyond the buffers above, without double-counting 
shared buffers |
+| **Complete graph budget `G`** | **Main subtotal + measured graph extras** | 
Add this to the vector/row-ID payload and other index structures |
+
+This is a layout-specific estimate, not a fixed bytes-per-node guarantee. 
HNSW's `m` controls graph connectivity and is unrelated to PQ's `m`. Do not 
substitute `N * m * 4` for `G`: actual degree, upper levels, retained 
distances, and the reader version matter. Calibrate by loading representative 
partitions and measuring their retained allocations.
+
+For example, suppose the cached `IVF_HNSW_PQ` partitions contain 1 million 
vectors, PQ `m=96`, `b=8`, `E=32000000`, `V=1100000`, and `A=1`. Assume this 
reader retains the distance column. These graph counts are illustrative inputs, 
not values derived from HNSW's `m`:
+
+| Component | Example payload |
+|---|---|
+| PQ codes | `96000000` bytes |
+| Dataset row IDs | `8000000` bytes |
+| Main graph buffers | `8 * 32000000 + 12 * 1100000 + 8 = 269200008` bytes, 
about 256.7 MiB |
+| **Codes + row IDs + main graph buffers** | **`373200008` bytes, about 355.9 
MiB** |
+| Alternative: main graph buffers when distances are skipped | `4 * 32000000 + 
8 * 1100000 + 4 = 136800004` bytes, about 130.5 MiB |
+| **Alternative total: codes + row IDs + graph buffers without distances** | 
**`240800004` bytes, about 229.6 MiB** |
+| Still to add in either case | Graph extras, IVF centroids/directory, PQ 
codebook and common overhead from the tables above |
+
+Select the example matching the reader's retained columns. Neither subtotal is 
a complete cache-capacity recommendation.
+
+The upstream [Lance Performance Guide](https://lance.org/guide/performance/) 
explains vector payload and index-build memory estimates. The [Lance vector 
index format](https://lance.org/format/index/vector/) describes the separation 
of IVF partitions, sub-indexes, and quantized storage. Their latest content may 
describe features newer than the Lance version embedded in Doris; use the Lance 
Catalog support matrix for compatibility.
+
+### Scalar and Full-Text Index Payloads {#scalar-index-cache-sizing}
+
+Scalar indexes can also be estimated, but the inputs differ by index family. 
Count the entries actually retained on one BE, not just the rows returned by a 
query. The following methods describe Lance storage layouts; they do not extend 
the Doris predicate-pushdown or index compatibility guarantees. Confirm that 
the embedded Lance version supports the index and that the query uses it.
+
+**BTree.** A loaded leaf page contains values and UInt64 row IDs. Let `R` be 
the number of records in the cached leaf pages, `H` the number of those pages, 
`w` the fixed-width value size, and `L` the average string length in bytes in 
those pages:
+
+| Leaf-page component / representation | Value buffer | Offsets | UInt64 row 
IDs |
+|---|---|---|---|
+| Int32 / Float32 | `4 * R` | None | `8 * R` |
+| Int64 / Float64 / 64-bit timestamp | `8 * R` | None | `8 * R` |
+| Other fixed-width values | `w * R`; use actual bitmap buffers for bit-packed 
Boolean | None | `8 * R` |
+| Arrow Utf8 / Binary | `R * L` | `4 * (R + H)` for Int32 offsets | `8 * R` |
+| Arrow LargeUtf8 / LargeBinary | `R * L` | `8 * (R + H)` for Int64 offsets | 
`8 * R` |
+
+**The leaf-page subtotal is the sum of the three payload columns.** Add these 
components for each opened segment; `P` below counts all leaf pages in that 
segment, not only cached pages:
+
+| Additional BTree component | Main payload / how to budget |
+|---|---|
+| Lookup minimum values | `P * w` for fixed-width values; actual value/offset 
buffers for variable-width values |
+| Lookup maximum values | `P * w` for fixed-width values; actual value/offset 
buffers for variable-width values |
+| Lookup NULL counts | `4 * P` for the UInt32 layout |
+| Lookup page numbers | `4 * P` for the UInt32 layout |
+| **Fixed-width lookup-buffer subtotal** | **`P * (2 * w + 8)`**; excludes the 
structures below |
+| NULL-page lists and lookup objects | **Measure** retained lists, objects and 
implementation-specific copies |
+| Leaf/lookup validity buffers, Arrow objects and spare capacity | **Measure** 
allocations beyond logical values/offsets/row IDs |
+
+String formulas assume the listed Arrow layouts; dictionary or view 
representations need their actual buffers. Disk compression does not determine 
decoded page size. The default BTree page size is 4096 **records**, not bytes; 
use the index's actual page size. The lookup belongs to the index working set, 
not the BE metadata-cache budget merely because it describes pages.
+
+For example, 100 million Int64 values have `100000000 * 16 = 1600000000` bytes 
(about 1.49 GiB) of leaf value/row-ID payload. If cached pages cover 10 million 
records, that payload is about 152.6 MiB. With 4096 records per page, the full 
index has about 24415 pages; the lookup batch's min/max/count/page-number 
payload is about 572 KiB before extra structures. Returning ten matching rows 
does not imply caching only ten records: the reader loads pages.
+
+**Bitmap, LabelList and NGram.** Sum the applicable component rows below. 
Membership bitmaps do not automatically cost eight bytes per matching row. 
LabelList counts membership per distinct label; a row can appear in multiple 
label bitmaps. NGram counts distinct gram/row memberships; repeated occurrences 
in one row do not add memberships. This NGram index differs from an FTS index 
using an n-gram tokenizer.
+
+| Index family | Component | Loaded payload / how to budget |
+|---|---|---|
+| Bitmap | Key lookup | **Measure** retained key values, entry references and 
lookup objects |
+| LabelList | Label lookup / wrapper | **Measure** distinct label keys, entry 
references and wrapper objects |
+| NGram | Gram lookup | **Measure** retained grams, entry references and 
lookup objects |
+| All three | Membership bitmap: array containers | About `2 * c` per 
conventional Roaring array container with `c` UInt16 entries in a 65536-value 
range |
+| All three | Membership bitmap: dense containers | `8192` bytes per 
conventional dense container covering a 65536-value range |
+| All three | Membership bitmap: run / full-fragment representations | 
**Measure** actual representation; use the representation selected for each 
container, not all alternatives added together |
+| All three | Bitmap directories / capacity | **Measure** fragment/high-bit 
directories, container objects and spare capacity |
+| Where present | NULL / null-list tracking | **Measure** retained NULL 
membership structures separately |
+| **Each family total** | **Lookup + its loaded membership representations + 
directories + NULL tracking** | Count shared allocations once; serialized 
bitmap size is not a resident-memory guarantee |
+
+For example, 256 loaded dense containers have 2 MiB of bitset payload, 
**before** keys, directories, NULL tracking and object overhead.
+
+**Other scalar and full-text indexes.** Each row identifies a separate 
component; subtotal rows summarize preceding components and must not be added 
again. Count only retained content on the BE. **Measure** means the reader's 
actual allocations are required, not that the component can be omitted.
+
+| Index family | Component | Loaded payload / how to budget |
+|---|---|---|
+| FTS / Inverted | Term dictionary | **Measure** loaded terms, dictionary 
structure and lookup allocations |
+| FTS / Inverted | Document metadata | **Measure** retained document mappings, 
lengths and other reader metadata |
+| FTS / Inverted | Plain posting row IDs | `8 * T` for `T` UInt64 
term/document entries |
+| FTS / Inverted | Plain posting frequencies | `4 * T` for Float32 
frequencies; IDs + frequencies total `12 * T` |
+| FTS / Inverted | Positions, when enabled | `4 * O` for `O` stored Int32 
occurrences |
+| FTS / Inverted | Position offsets | Actual retained offset-array bytes; 
separate from `4 * O` |
+| FTS / Inverted | Compressed posting layout | Use actual retained compressed 
blocks **instead of** the corresponding plain-layout formulas |
+| FTS / Inverted | Scoring / block metadata and allocation overhead | 
**Measure** retained block descriptors, scoring data, objects and spare 
capacity |
+| ZoneMap | Minimum values | `Z * w` logical bytes for `Z` loaded zones and 
fixed-width values of width `w` |
+| ZoneMap | Maximum values | `Z * w` logical bytes; min + max total `2 * Z * 
w` |
+| ZoneMap | Counts and zone boundaries | **Measure** retained counts and 
boundaries; count zones per Fragment, including partial zones |
+| ZoneMap | NULL tracking and boxed-value/object overhead | **Measure** 
optional NULL row sets and actual scalar allocations; these can dominate 
logical min/max bytes |
+| BloomFilter | Filter bit arrays | Sum actual allocated filter bytes; `Z * F` 
for `Z` equal-sized filters of `F` bytes each |
+| BloomFilter | Zone descriptors / filter objects | **Measure** retained 
descriptors, filter parameters and object overhead |
+| BloomFilter | NULL tracking | **Measure** optional NULL row sets |
+| RTree | Bounding-box coordinates | `32 * Q` for four Float64 coordinates per 
2D entry; count `Q` entries across all cached tree levels |
+| RTree | Row / child-page IDs | `8 * Q` for UInt64 IDs; coordinates + IDs 
total `40 * Q` |
+| RTree | Tree metadata and page offsets | **Measure** retained tree 
descriptors and page-offset buffers |
+| RTree | NULL tracking, Arrow/objects and spare capacity | **Measure** actual 
additional allocations |
+| JSON-path wrappers | Underlying index | Use the complete component budget of 
the selected scalar index family |
+| JSON-path wrappers | Path and wrapper structures | **Measure** retained 
paths and wrapper allocations in addition to the underlying index |
+
+Bloom filters require the builder's actual block-rounded allocations; row 
count and false-positive probability alone do not describe every layout. For 
unfamiliar or experimental layouts, measure loaded allocations instead of 
applying another family's formula. A small number of matching rows can still 
require substantial dictionary, tree, or document metadata.
+
+**Keep query working memory separate from all cache tables:** temporary 
intersections/unions, result row-ID sets, rechecks of approximate filters, 
original geometry reads for exact RTree checks, and output-column reads need 
their own budget.
+
+See the upstream [BTree format](https://lance.org/format/index/scalar/btree/), 
[Bitmap format](https://lance.org/format/index/scalar/bitmap/), and [FTS 
format](https://lance.org/format/index/scalar/fts/) for the underlying 
structures. The formulas above estimate loaded payload, not the compressed file 
size or total query RSS.
+
+### From Index Size to a Per-BE Working Set
+
+For segment `s`, let `N_s` be its indexed vector count, `P_s` its IVF 
partition count, and `H_s` the number of distinct partitions to retain on a 
particular BE. With reasonably balanced partitions:
+
+```text
+cached_vectors_on_BE ≈ sum over segments (N_s * H_s / P_s)
+index_cache_target ≈ cached vector/row-ID payload
+                   + cached HNSW graphs
+                   + centroids, codebooks and other index allocations
+                   + measured headroom
+```
+
+Use actual partition row counts when partitions are skewed. `H_s` describes 
the working set across queries, not just the `nprobes` of one query. Diverse 
queries can eventually touch every partition; repeated or concurrent queries 
can share the same cached partition. Count distinct `(dataset, index segment, 
partition)` entries and include all tables, indexes, and versions accessed by 
that BE. Cache residency on one BE does not warm another BE, and scheduling 
does not guarantee that every  [...]
+
+Query-time memory is additional: in-flight partition loads, objects still held 
by queries after cache eviction, Prefilter state, distance tables, candidate 
heaps, decoded batches, refinement, and second-phase fetches. `top_k`, `ef`, 
`refine_factor`, scan concurrency, and `nprobes` affect this memory or the 
accessed working set; they do not change the stored bytes per vector. A 
configured cache capacity is not a hard cap on Lance or BE RSS. Increasing the 
cache also does not remove distan [...]
+
+### Configuration Example and Tuning
+
+Assume an FP32, 768-dimensional `IVF_PQ` index with 100 million vectors, 
`m=96`, `b=8`, and 4096 balanced partitions in one index segment. A 
representative workload on one BE repeatedly accesses 1024 distinct partitions:
+
+```text
+Full code/row-ID payload = 100000000 * (96 + 8) = 10400000000 bytes ≈ 9.69 GiB
+Cached payload          = 10400000000 * 1024 / 4096 = 2600000000 bytes ≈ 2.42 
GiB
+IVF centroids           = 4096 * 768 * 4 = 12582912 bytes = 12 MiB
+PQ codebook             = 256 * 768 * 4 = 786432 bytes = 0.75 MiB
+```
+
+If this is the only substantial index working set on the BE and its memory 
budget permits, 4 GiB is a reasonable initial index-cache allocation for this 
example, leaving room above the calculated payload for additional index 
allocations. It is an initial value to validate, not a universal overhead 
ratio. If the workload instead needs the entire index resident, 4 GiB is 
insufficient, and even the default 10 GiB is close to the 9.69 GiB payload 
before other allocations.
+
+For a BE where 4 GiB of index cache and 1 GiB of metadata cache fit alongside 
peak query memory, other Doris workloads, and process/OS headroom, set:
+
+```properties
+lance_index_cache_size_bytes = 4294967296
+lance_metadata_cache_size_bytes = 1073741824
+```
+
+Restart that BE for these settings to take effect. Apply the budget separately 
to each BE; these are BE configuration parameters, not catalog properties or 
SQL session variables.
+
+1. Warm up with representative queries and concurrency, including different 
query vectors, filters, and tables. Repeating one vector alone understates a 
diverse workload's working set.
+2. Compare `doris_be_lance_session_index_cache_usage_bytes` with 
`doris_be_lance_session_index_cache_capacity_bytes`, the interval increases of 
`hits_total`/`misses_total`, and the per-query 
`LanceIndexPartitionCacheMissLoads`. Persistent misses with usage near capacity 
suggest cache churn; misses can also come from new partitions, index versions, 
or different BEs. Increase capacity only when reuse and available memory 
justify it.
+3. Start metadata-cache sizing from the default 1 GiB when the BE memory 
budget permits. Adjust using its `usage_bytes` and hit/miss metrics. 
File/Fragment counts, schema/page metadata, deletion information, and row-ID 
mappings determine demand; it is not a fixed percentage of vector-index size. 
The BE metadata cache is separate from the FE table-access cache.
+4. Check BE process memory and query peaks under the intended concurrency 
before increasing either cache. A high cache hit rate with high latency can 
indicate scoring, Prefilter, decoding, or fetch costs; additional cache is not 
necessarily useful.
+5. Size `lance_data_cache_disk_capacity_bytes` independently for repeated 
data-file reads such as refinement and output-column fetches. Increasing it 
does not enlarge the index cache or compensate for index-cache misses: index 
files bypass this disk data cache. Keep 
`lance_data_cache_read_block_size_bytes` at its 1 MiB default initially, then 
change it only after measuring useful reads versus read amplification.
+
+## 4. Budget for Shared Caches {#shared-index-cache-budget}
+
+On each BE, Lance uses a shared session and one index-cache capacity. Vector 
partitions, BTree pages, Bitmap/LabelList entries, and full-text index content 
can compete for this space, including indexes from different tables and 
catalogs accessed on the same BE. Cache keys separate their identities; they do 
not reserve capacity. The Lance BE cache settings provide no per-table or 
per-index-type quota or pinning control.
+
+Only loaded content consumes cache space; merely creating several indexes does 
not load them all. Once queries use them, loading one index can evict another 
index's entries. For example, a broad BTree range scan or diverse FTS queries 
can replace vector partitions that were warm. A later vector query then reloads 
them. Changes of index version can also introduce new entries while older 
entries remain resident. Eviction does not immediately free an object still 
referenced by a running query.
+
+Plan the mixed working set, without counting shared allocations twice:
+
+```text
+index_cache_target_on_BE ≈ vector working set
+                        + BTree working set
+                        + Bitmap / LabelList / NGram working sets
+                        + FTS and other index working sets
+                        + measured headroom
+```
+
+For example, suppose measurements and payload estimates give 2.5 GiB for 
vectors, 0.25 GiB for BTree, 0.5 GiB for Bitmap, and 0.75 GiB for FTS, 
including their index structures. Their combined working set is 4 GiB. A 5 GiB 
initial index cache leaves 1 GiB for growth and estimation error; it does not 
guarantee every workload will fit. If the BE also has room for 1 GiB of 
metadata cache **and** peak query memory, other Doris workloads, and process/OS 
headroom:
+
+```properties
+lance_index_cache_size_bytes = 5368709120
+lance_metadata_cache_size_bytes = 1073741824
+```
+
+The index and metadata caches have separate capacities: metadata entries do 
not directly evict index entries, and increasing one does not enlarge the 
other. Both still consume the same BE process memory. The disk data cache is a 
third budget and does not cache index files. Separate catalogs on the same BE 
do not provide index-cache isolation; workloads requiring isolation need 
separate BE resources and routing supported by the deployment.
+
+To diagnose contention, first warm one query family and record latency and 
cache-metric deltas. Run a second family, then repeat the first. Renewed misses 
or reloads with cache usage near capacity suggest eviction pressure, provided 
index versions and BE placement stayed comparable. Session hit/miss metrics 
aggregate workloads, so combine them with query profiles. Prewarming every 
index can itself evict the working set you wanted to retain. Prefer 
representative mixed-workload warmup, an [...]
+
+## 5. Choose Index Segment Granularity {#index-segment-granularity}
+
+An Index Segment, a data Fragment, an IVF partition, and a BTree leaf page are 
different units. In the current Doris vector-search path, each selected 
physical Index Segment becomes one scan Split; uncovered Fragments add Flat 
Search splits. FTS also uses physical index segments. One vector segment's IVF 
partitions are not separate Doris splits distributed across BEs. Lance can 
still execute work internally in parallel. For ordinary scalar scans, 
segment-based grouping depends on the sel [...]
+
+Let `B` be the number of BEs actually eligible for this query and `S` the 
number of selected vector segments. If `S` is smaller than `B`, there are too 
few indexed scan splits to use all those BEs for that query. Increasing `S` 
creates scheduling opportunities, but it does not guarantee that all splits run 
simultaneously. Scanner limits, Lance's internal concurrency, CPU, I/O, cache 
residency, and concurrent queries also matter. Increasing Doris pipeline 
instances alone does not subdivid [...]
+
+For a large dataset, compare `B`, `2 * B`, and `4 * B` reasonably balanced 
vector segments as an **initial experiment**, not an official optimum or a 
requirement for small datasets. For an illustrative 100-million-vector dataset 
and eight eligible BEs:
+
+| Segments | Average vectors per segment | What to evaluate |
+|---|---|---|
+| 8 | 12.5 million | Enough splits to cover the BEs, with fewer per-segment 
operations |
+| 16 | 6.25 million | More scheduling flexibility and less sensitivity to one 
large segment |
+| 32 | 3.125 million | More tasks, but potentially greater search, loading, 
and merge overhead |
+
+Balance estimated search work and loaded index bytes, not just Fragment 
counts. Equal rows are a useful starting point for vectors with the same 
dimension and encoding; HNSW graph size, IVF skew, scalar selectivity, and 
document length can change the balance. There is no universal rows-per-segment 
or file-size threshold. High-throughput workloads may favor fewer tasks per 
query to preserve capacity for concurrent queries; single-query latency may 
benefit from more splits.
+
+More segments also mean more per-segment setup, metadata/quantizer objects, 
and local candidate generation. Each vector/FTS split retains up to `top_k + 
offset` candidates; Doris does not divide this bound by `S`. Thus the scan-side 
candidate bound is `S * (top_k + offset)`, subject to available matches. With 
`top_k=1000`, no offset, and 32 splits, that is up to 32000 candidates before 
Doris reductions. Local TopN can reduce network traffic, so this is not a 
network-row guarantee. Candid [...]
+
+Tune segment count together with each segment's IVF `num_partitions` and query 
`nprobes`. For balanced partitions, a rough probed-vector count is `sum(N_s * 
nprobes_s / P_s)`, capped at `N_s` for each segment. This is not an HNSW 
comparison-count formula. Rebuilding segments or changing partition counts can 
change recall even at the same `nprobes`; compare latency at comparable recall. 
HNSW `ef` and refinement settings also affect the comparison.
+
+Build and maintain segments with Lance tooling, then commit them as the 
intended logical index. Several separately named full-table indexes are not a 
substitute for one index with multiple physical segments. Prefer balanced 
groups of whole Fragments, avoid accumulating many tiny segments from small 
appends, and use the embedded Lance version's supported maintenance APIs. 
Verify the resulting physical segment count and uncovered-fragment count in 
Doris `EXPLAIN`/Profile; do not infer it f [...]
+
+For BTree point lookups or selective Bitmap queries, splitting into many 
segments can add lookups and opens while little work is saved per segment. FTS 
adds per-segment dictionary/posting searches and candidate merging. Evaluate 
these workloads separately instead of applying a vector segment count to all 
index families. Splitting also does not remove unnecessary per-row Prefilter 
construction; verify the reader's filtering behavior independently.
+
+## 6. Validate the Mixed Workload {#validate-workload}
+
+1. Confirm index selection, physical splits, and index coverage. Include 
fallback scans and second-phase fetches in the latency analysis.
+2. Estimate each index's retained payload and add the shared working sets. 
Reserve query memory separately, including parallel loads, row-ID sets, 
candidate heaps, and decoding.
+3. Compare cold-cache and warmed mixed workloads with realistic concurrency. 
Track P95 latency, throughput, recall where applicable, per-BE peak memory, 
cache usage and hit/miss deltas, index-load I/O, and scan-time imbalance.
+4. Change one dimension at a time: cache capacity, segment count, or search 
parameters. If segment rebuilding changes the model, recheck recall. Stop 
increasing splits or cache when throughput, latency, or memory ceases to 
improve.
+5. Repeat after substantial appends, index rebuilds, BE scaling, or changes in 
query mix. A configuration calibrated for one vector or table is not a stable 
budget for all workloads.
+
+## 7. Diagnose Before Increasing Resources {#troubleshooting}
+
+| Observation | Check first | Action to evaluate |
+|---|---|---|
+| Repeated index misses with cache usage near capacity | Index/version 
changes, BE placement, and the combined working set | Increase 
`lance_index_cache_size_bytes` if memory permits, reduce unnecessary warmup, or 
isolate workloads on separate BE resources |
+| Vector queries slow down after scalar or FTS queries | Repeat the 
cross-workload eviction check in [shared-cache 
budgeting](#shared-index-cache-budget) | Budget for both query families 
together; separate catalogs on the same BE do not isolate the cache |
+| Few BEs scan a large indexed dataset | Physical segment Split count, 
eligible BEs, and scanner scheduling | Compare more balanced segments; 
increasing pipeline instances alone cannot split a segment |
+| More segments increase CPU, memory, or merge time | Per-split candidate 
count, `top_k + offset`, and query concurrency | Reduce segment count or avoid 
requesting more candidates than the application needs; recheck recall when 
search settings change |
+| Cache hits are high but latency remains high | Scoring, Prefilter work, 
decoding, and second-phase fetch time | Optimize the dominant stage; a larger 
index cache does not remove computation or fetches |
+| Metadata reloads persist | Metadata-cache usage and hit/miss deltas; 
Fragment/file counts and version churn | Adjust 
`lance_metadata_cache_size_bytes` independently if reuse and memory permit |
+| Process memory is high despite cache usage below capacity | Concurrent query 
allocations, active loads, retained objects, and other BE workloads | Lower 
concurrency or cache budgets based on measured peaks; cache capacity is not a 
process-memory cap |
+| Data-file I/O stays high | Disk-cache counters, reuse, read block size, and 
query output columns | Adjust the data-cache budget or read granularity; index 
files do not use this disk data cache |
+
+Use the profile counters and BE metrics in [Checking Cache 
Effectiveness](../catalogs/lance-catalog.mdx#checking-cache-effectiveness). An 
aggregate hit rate alone does not identify which index was evicted or which 
stage dominates latency. Change one setting at a time and keep a record of the 
workload, index version, BE placement, latency, throughput, recall, and peak 
memory for each comparison.
diff --git a/versioned_docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx 
b/versioned_docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
index bf1eb3474ec..7a0f9427e5c 100644
--- a/versioned_docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
+++ b/versioned_docs/version-4.x/lakehouse/catalogs/lance-catalog.mdx
@@ -995,6 +995,8 @@ After the BE restarts, the shared cache initializes when 
the first Lance Dataset
 
 Budget for index and metadata caches alongside query working memory. These 
capacities do not limit total Lance query memory or BE process memory. The disk 
cache also consumes local disk space and file descriptors. Larger read blocks 
reduce the number of blocks but can increase extra I/O for small reads. Use a 
new cache directory when changing the read block size or disk layout.
 
+For workload-specific sizing formulas, configuration examples, shared-cache 
contention, and index segment granularity, see [Lance Query Best 
Practices](../best-practices/doris-lance.mdx).
+
 ### Checking Cache Effectiveness
 
 Inspect the following counters under `LanceReader` in the query Profile:
diff --git a/versioned_sidebars/version-4.x-sidebars.json 
b/versioned_sidebars/version-4.x-sidebars.json
index 68d980fe0f5..5ca7499aa21 100644
--- a/versioned_sidebars/version-4.x-sidebars.json
+++ b/versioned_sidebars/version-4.x-sidebars.json
@@ -889,6 +889,7 @@
               "label": "Lakehouse Best Practices",
               "items": [
                 "lakehouse/best-practices/optimization",
+                "lakehouse/best-practices/doris-lance",
                 "lakehouse/best-practices/kerberos",
                 "lakehouse/best-practices/tpch",
                 "lakehouse/best-practices/tpcds"


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to