leaves12138 commented on code in PR #9721:
URL: https://github.com/apache/paimon/pull/9721#discussion_r3978755224


##########
docs/docs/multimodal-table/blob.mdx:
##########
@@ -276,72 +162,14 @@ schema = Schema.from_pyarrow_schema(
         'data-evolution.enabled': 'true',
     }
 )
+catalog.create_table('my_db.image_table', schema, False)

Review Comment:
   [P2] Keep the add-column examples in the newly created database
   
   This new setup creates `my_db.image_table`, but the Python example under 
**Adding a Blob Column** still calls 
`catalog.alter_table('default.image_table', ...)`. Following the walkthrough on 
a fresh catalog therefore raises `TableNotExistException: Table 
default.image_table does not exist`. The same mismatch occurs in `vector.mdx`: 
its setup creates `my_db.vector_table`, while **Adding a Vector Column** alters 
`default.vector_table`. I executed both create/alter sequences and verified 
that changing the ALTER targets to their respective `my_db` identifiers 
succeeds. Please align the identifiers across both walkthroughs.



##########
docs/docs/multimodal-table/global-index.mdx:
##########
@@ -27,425 +26,203 @@ under the License.
 
 # Global Index
 
-## Overview
+<span id="overview" />
 
-Global Index is a powerful indexing mechanism for Data Evolution (append) 
tables. It enables efficient row-level lookups and filtering
-without full-table scans. Paimon supports multiple global index types:
+Global indexes map predicates or search requests to row IDs, which Paimon uses 
to
+read matching table rows. Choose scalar indexes for filtering, vector indexes 
for
+similarity search, and full-text indexes for scored text retrieval. Hybrid 
search
+combines scored routes; it is a query API rather than another stored index 
type.
 
-- **[BTree Index](./global-index/btree)**: A B-tree based index for scalar 
column lookups. Supports equality, IN, range predicates, and can be combined 
across multiple columns with AND/OR logic.
-- **[Bitmap Index](./global-index/bitmap)**: A bitmap based index for 
enum-like scalar dimensions and tag columns. Supports equality, IN, prefix 
match on string columns, complement predicates, and null checks with compressed 
row-id bitmaps.
-- **[Multivalue Index](./global-index/multivalue)**: A bitmap-backed index for 
element-membership predicates on `ARRAY` columns.
-- **[FM Index](./global-index/fm)**: An exact partitioned substring index for 
`CONTAINS` predicates on character columns.
-- **[Vector Index](./global-index/vector)**: An approximate nearest neighbor 
(ANN) index powered by Paimon's vector index library for vector similarity 
search.
-- **[Full-Text Index](./global-index/full-text)**: A full-text search index 
backed by the native full-text engine for text retrieval. Supports term 
matching and relevance scoring.
-- **[Hybrid Search](./global-index/hybrid-search)**: A multi-route search API 
that combines results from multiple vector routes, multiple full-text routes, 
or both before reading table rows.
+This guide covers Data Evolution append tables. Primary-key tables have a 
separate
+[index configuration and lifecycle](../primary-key-table/global-index).
+
+## Choose an Index
 
 | Index Type | Best For | Notes |
 |---|---|---|
-| BTree | Scalar filters on numeric, string, date, and timestamp columns | 
Best when predicates are selective, such as equality, IN, range, and null 
checks. |
-| Bitmap | Enum-like dimensions and tag columns | Best for equality, IN, 
string prefix match, complement predicates, and null checks over compressed 
row-id bitmaps. |
-| Multivalue | Membership tests on arrays of supported scalar elements | Best 
for `ARRAY_CONTAINS`, `ARRAYS_OVERLAP`, and `ARRAY_CONTAINS_ALL`. |
-| FM | Exact substring filters on character columns | Supports needles of any 
byte length and partitions the indexed text for bounded-memory construction and 
demand-loaded reads. |
-| Vector | Top-K similarity search on embeddings | Uses ANN algorithms. Tune 
build-time and search-time options to balance recall, latency, and index size. |
-| Full-Text | Keyword search over text columns | Uses full-text scoring and 
tokenizer configuration stored with each index file. |
-| Hybrid Search | Combining multiple vector routes, multiple full-text routes, 
or vector and full-text retrieval together | Runs multiple scored routes and 
merges them with a ranker before reading rows. |
+| [BTree](./global-index/btree) | Scalar filters on numeric, string, date, and 
timestamp columns | Best when predicates are selective, such as equality, IN, 
range, and null checks. |
+| [Bitmap](./global-index/bitmap) | Enum-like dimensions and tag columns | 
Best for equality, IN, string prefix match, complement predicates, and null 
checks over compressed row-id bitmaps. |
+| [Multivalue](./global-index/multivalue) | Membership tests on arrays of 
supported scalar elements | Best for `ARRAY_CONTAINS`, `ARRAYS_OVERLAP`, and 
`ARRAY_CONTAINS_ALL`. |
+| [FM](./global-index/fm) | Exact substring filters on character columns | 
Supports needles of any byte length and partitions the indexed text for 
bounded-memory construction and demand-loaded reads. |
+| [Vector](./global-index/vector) | Top-K similarity search on embeddings | 
Uses ANN algorithms. Tune build-time and search-time options to balance recall, 
latency, and index size. |
+| [Full-Text](./global-index/full-text) | Keyword search over text columns | 
Uses full-text scoring and tokenizer configuration stored with each index file. 
|
+| [Hybrid Search](./global-index/hybrid-search) | Combining multiple vector 
routes, multiple full-text routes, or vector and full-text retrieval together | 
Runs multiple scored routes and merges them with a ranker before reading rows. |
 
-Global indexes work on top of Data Evolution tables. To use global indexes, 
your table **must** have:
+:::caution Index coverage affects results
 
-- `'row-tracking.enabled' = 'true'`
-- `'data-evolution.enabled' = 'true'`
+The default search mode is `fast`. If an index covers only part of the table,
+matching rows outside that coverage can be omitted. Creating a table or 
appending
+data does not build a global index. Read [Coverage and 
Freshness](./global-index/manage-indexes#coverage-and-freshness)
+before using indexed queries on changing data.
 
-> Global index queries may not be exact when the index only covers part of the 
table data. If a query predicate matches the index, Paimon returns only the 
results from the indexed portion. Matching records in data that has not been 
indexed yet will not be returned.
+:::
 
 ## Prerequisites
 
-Create a table with the required properties:
+Start with a [configured Paimon catalog](./quick-start). Data Evolution 
requires
+an append table with row tracking enabled. The following examples assume the 
`db`
+database exists and that SQL uses that database. Load data before building 
indexes.
+
+`global-index.enabled` defaults to `true`; setting it does not create index 
files.
+
+The example table below is shared by the BTree, Bitmap, Multivalue, Vector, and
+Full-Text guides. It contains three-dimensional sample vectors and is 
partitioned
+by `dt`, so the partition-scoped build and drop examples also apply. Create and
+populate it once using one of the following alternatives.
 
 <Tabs groupId="global-index-create-table">
 
-<TabItem value="sql" label="SQL">
+<TabItem value="sql" label="Spark SQL">
 
 ```sql
 CREATE TABLE my_table (
     id INT,
     name STRING,
+    category STRING,
+    tag STRING,
     tags ARRAY<STRING>,
     embedding ARRAY<FLOAT>,
-    content STRING
-) TBLPROPERTIES (
+    content STRING,
+    dt STRING
+) PARTITIONED BY (dt) TBLPROPERTIES (
     'row-tracking.enabled' = 'true',
-    'data-evolution.enabled' = 'true',
-    'global-index.enabled' = 'true'
+    'data-evolution.enabled' = 'true'
 );
-```
-
-</TabItem>
 
-<TabItem value="python-sdk" label="Python SDK">
-
-```python
-import pyarrow as pa
-
-from pypaimon import Schema
-
-schema = Schema.from_pyarrow_schema(
-    pa.schema([
-        pa.field("id", pa.int32()),
-        pa.field("name", pa.string()),
-        pa.field("tags", pa.list_(pa.string())),
-        pa.field("embedding", pa.list_(pa.float32())),
-        pa.field("content", pa.string()),
-    ]),
-    options={
-        "row-tracking.enabled": "true",
-        "data-evolution.enabled": "true",
-        "global-index.enabled": "true",
-    },
-)
-
-catalog.create_table("db.my_table", schema, ignore_if_exists=False)
+INSERT INTO my_table VALUES
+    (101, 'a200', 'electronics', 'vip', array('blue', 'green'),
+     array(1.0f, 0.0f, 0.0f), 'paimon lake format and search', '2026-06-18'),
+    (102, 'a300', 'electronics', 'trial', array('green'),
+     array(0.9f, 0.1f, 0.0f), 'paimon full text search', '2026-06-18'),
+    (103, 'a400', 'books', 'standard', array('red'),
+     array(0.0f, 1.0f, 0.0f), 'vector search tutorial', '2026-06-19');
 ```
 
 </TabItem>
 
-</Tabs>
-
-## Lifecycle
-
-Create global indexes for all partitions or only selected partitions:
-
-<Tabs groupId="global-index-build">
-
-<TabItem value="sql" label="SQL">
+<TabItem value="flink-sql" label="Flink SQL">
 
 ```sql
-CALL sys.create_global_index(
-    table => 'db.my_table',
-    index_column => 'name',
-    index_type => 'btree'
-);
-
-CALL sys.create_global_index(
-    table => 'db.my_table',
-    index_column => 'name',
-    index_type => 'btree',
-    partitions => 'dt=2026-06-18;dt=2026-06-19'
-);
-```
-
-</TabItem>
-
-<TabItem value="python-sdk" label="Python SDK">
-
-```python
-table = catalog.get_table("db.my_table")
-
-added_files = table.create_global_index("name")
-print(added_files)
-```
-
-The API returns the number of committed index files. You can pass build 
options and restrict the
-build to selected partitions:
-
-```python
-added_files = table.create_global_index(
-    "name",
-    index_type="btree",
-    partitions=[{"dt": "2026-06-18"}, {"dt": "2026-06-19"}],
-    options={"sorted-index.records-per-range": "10000000"},
-)
-```
-
-</TabItem>
-
-</Tabs>
+SET 'execution.runtime-mode' = 'batch';
 
-PyPaimon global index build currently supports single-column BTree indexes,
-single-column Bitmap indexes, single-column paimon-vindex vector indexes,
-and single-column full-text indexes on tables with row tracking enabled.
-
-Drop index files:
-
-<Tabs groupId="global-index-drop">
-
-<TabItem value="sql" label="SQL">
-
-```sql
-CALL sys.drop_global_index(
-    table => 'db.my_table',
-    index_column => 'name',
-    index_type => 'btree'
+CREATE TABLE my_table (
+    id INT,
+    name STRING,
+    category STRING,
+    tag STRING,
+    tags ARRAY<STRING>,
+    embedding ARRAY<FLOAT>,
+    content STRING,
+    dt STRING
+) PARTITIONED BY (dt) WITH (
+    'row-tracking.enabled' = 'true',
+    'data-evolution.enabled' = 'true'
 );
-```
-
-</TabItem>
 
-<TabItem value="python-sdk" label="Python SDK">
-
-```python
-table = catalog.get_table("db.my_table")
-
-dropped_files = table.drop_global_index("name", index_type="btree")
-print(dropped_files)
-```
-
-You can also restrict the drop to selected partitions, or count matched files 
without committing:
-
-```python
-matched_files = table.drop_global_index(
-    "name",
-    index_type="btree",
-    partitions=[{"dt": "2026-06-18"}, {"dt": "2026-06-19"}],
-    dry_run=True,
-)
+INSERT INTO my_table VALUES
+    (101, 'a200', 'electronics', 'vip', ARRAY['blue', 'green'],
+     ARRAY[CAST(1.0 AS FLOAT), CAST(0.0 AS FLOAT), CAST(0.0 AS FLOAT)],
+     'paimon lake format and search', '2026-06-18'),
+    (102, 'a300', 'electronics', 'trial', ARRAY['green'],
+     ARRAY[CAST(0.9 AS FLOAT), CAST(0.1 AS FLOAT), CAST(0.0 AS FLOAT)],
+     'paimon full text search', '2026-06-18'),
+    (103, 'a400', 'books', 'standard', ARRAY['red'],
+     ARRAY[CAST(0.0 AS FLOAT), CAST(1.0 AS FLOAT), CAST(0.0 AS FLOAT)],
+     'vector search tutorial', '2026-06-19');
 ```
 
-</TabItem>
-
-</Tabs>
-
-Global indexes are stored in index files and recorded in table metadata. To 
inspect index files and
-their row-id coverage, query the `table_indexes` system table:
-
-<Tabs groupId="global-index-table-indexes">
-
-<TabItem value="sql" label="SQL">
-
-```sql
-SELECT index_type, index_field_name, row_range_start, row_range_end
-FROM my_table$table_indexes
-WHERE index_field_name IS NOT NULL;
-```
+Wait for the insert job to finish before building an index.
 
 </TabItem>
 
 <TabItem value="python-sdk" label="Python SDK">
 
 ```python
-import pyarrow.compute as pc
-
-table_indexes = catalog.get_table("db.my_table$table_indexes")
-read_builder = table_indexes.new_read_builder().with_projection([
-    "index_type",
-    "index_field_name",
-    "row_range_start",
-    "row_range_end",
-])
-
-pa_table = read_builder.new_read().to_arrow(
-    read_builder.new_scan().plan().splits()
-)
-pa_table = pa_table.filter(pc.is_valid(pa_table["index_field_name"]))
-print(pa_table)
-```
-
-</TabItem>
-
-</Tabs>
-
-You can also query `file_key_ranges` to inspect data file row-id ranges and 
diagnose coverage:
-
-<Tabs groupId="global-index-file-key-ranges">
-
-<TabItem value="sql" label="SQL">
-
-```sql
-SELECT file_path, first_row_id, record_count
-FROM my_table$file_key_ranges;
-```
-
-</TabItem>
-
-<TabItem value="python-sdk" label="Python SDK">
+import pyarrow as pa
+from pypaimon import Schema
 
-```python
-file_key_ranges = catalog.get_table("db.my_table$file_key_ranges")
-read_builder = file_key_ranges.new_read_builder().with_projection([
-    "file_path",
-    "first_row_id",
-    "record_count",
+# `catalog` is the configured Paimon catalog; database `db` already exists.
+pa_schema = pa.schema([
+    ('id', pa.int32()),
+    ('name', pa.string()),
+    ('category', pa.string()),
+    ('tag', pa.string()),
+    ('tags', pa.list_(pa.string())),
+    ('embedding', pa.list_(pa.float32())),
+    ('content', pa.string()),
+    ('dt', pa.string()),
 ])
-
-pa_table = read_builder.new_read().to_arrow(
-    read_builder.new_scan().plan().splits()
+schema = Schema.from_pyarrow_schema(
+    pa_schema,
+    partition_keys=['dt'],

Review Comment:
   [P2] Make the shared fixture work with the partition-scoped Python builds
   
   After executing this new setup, the build/rebuild workflow advertised in 
Manage Global Indexes and the vector guide fails:
   
   ```python
   table.create_global_index('name', index_type='btree')
   table.create_global_index(
       'name', index_type='btree',
       partitions=[{'dt': '2026-06-18'}, {'dt': '2026-06-19'}],
   )
   ```
   
   The second call raises `IndexError: Position 7 is out of bounds for row 
arity 1`. `dt` is column 7 in this fixture: 
`CreateGlobalIndexBuilder._resolve_partition_filter` constructs a full-row 
predicate, but `build_plan.indexed_row_ranges` applies it directly to the 
one-field index-manifest partition row. I also reproduced this with the 
documented IVF-flat calls. Moving `dt` to the first schema position in a 
control fixture makes the BTree rebuild return `0` as expected. The underlying 
SDK bug predates this PR, but the new shared walkthrough encounters it. Please 
provide a working fixture/workaround or explicitly avoid this unsupported 
sequence until the predicate projection is fixed.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to