leaves12138 commented on code in PR #9721:
URL: https://github.com/apache/paimon/pull/9721#discussion_r3978755224
##########
docs/docs/multimodal-table/blob.mdx:
##########
@@ -276,72 +162,14 @@ schema = Schema.from_pyarrow_schema(
'data-evolution.enabled': 'true',
}
)
+catalog.create_table('my_db.image_table', schema, False)
Review Comment:
[P2] Keep the add-column examples in the newly created database
This new setup creates `my_db.image_table`, but the Python example under
**Adding a Blob Column** still calls
`catalog.alter_table('default.image_table', ...)`. Following the walkthrough on
a fresh catalog therefore raises `TableNotExistException: Table
default.image_table does not exist`. The same mismatch occurs in `vector.mdx`:
its setup creates `my_db.vector_table`, while **Adding a Vector Column** alters
`default.vector_table`. I executed both create/alter sequences and verified
that changing the ALTER targets to their respective `my_db` identifiers
succeeds. Please align the identifiers across both walkthroughs.
##########
docs/docs/multimodal-table/global-index.mdx:
##########
@@ -27,425 +26,203 @@ under the License.
# Global Index
-## Overview
+<span id="overview" />
-Global Index is a powerful indexing mechanism for Data Evolution (append)
tables. It enables efficient row-level lookups and filtering
-without full-table scans. Paimon supports multiple global index types:
+Global indexes map predicates or search requests to row IDs, which Paimon uses
to
+read matching table rows. Choose scalar indexes for filtering, vector indexes
for
+similarity search, and full-text indexes for scored text retrieval. Hybrid
search
+combines scored routes; it is a query API rather than another stored index
type.
-- **[BTree Index](./global-index/btree)**: A B-tree based index for scalar
column lookups. Supports equality, IN, range predicates, and can be combined
across multiple columns with AND/OR logic.
-- **[Bitmap Index](./global-index/bitmap)**: A bitmap based index for
enum-like scalar dimensions and tag columns. Supports equality, IN, prefix
match on string columns, complement predicates, and null checks with compressed
row-id bitmaps.
-- **[Multivalue Index](./global-index/multivalue)**: A bitmap-backed index for
element-membership predicates on `ARRAY` columns.
-- **[FM Index](./global-index/fm)**: An exact partitioned substring index for
`CONTAINS` predicates on character columns.
-- **[Vector Index](./global-index/vector)**: An approximate nearest neighbor
(ANN) index powered by Paimon's vector index library for vector similarity
search.
-- **[Full-Text Index](./global-index/full-text)**: A full-text search index
backed by the native full-text engine for text retrieval. Supports term
matching and relevance scoring.
-- **[Hybrid Search](./global-index/hybrid-search)**: A multi-route search API
that combines results from multiple vector routes, multiple full-text routes,
or both before reading table rows.
+This guide covers Data Evolution append tables. Primary-key tables have a
separate
+[index configuration and lifecycle](../primary-key-table/global-index).
+
+## Choose an Index
| Index Type | Best For | Notes |
|---|---|---|
-| BTree | Scalar filters on numeric, string, date, and timestamp columns |
Best when predicates are selective, such as equality, IN, range, and null
checks. |
-| Bitmap | Enum-like dimensions and tag columns | Best for equality, IN,
string prefix match, complement predicates, and null checks over compressed
row-id bitmaps. |
-| Multivalue | Membership tests on arrays of supported scalar elements | Best
for `ARRAY_CONTAINS`, `ARRAYS_OVERLAP`, and `ARRAY_CONTAINS_ALL`. |
-| FM | Exact substring filters on character columns | Supports needles of any
byte length and partitions the indexed text for bounded-memory construction and
demand-loaded reads. |
-| Vector | Top-K similarity search on embeddings | Uses ANN algorithms. Tune
build-time and search-time options to balance recall, latency, and index size. |
-| Full-Text | Keyword search over text columns | Uses full-text scoring and
tokenizer configuration stored with each index file. |
-| Hybrid Search | Combining multiple vector routes, multiple full-text routes,
or vector and full-text retrieval together | Runs multiple scored routes and
merges them with a ranker before reading rows. |
+| [BTree](./global-index/btree) | Scalar filters on numeric, string, date, and
timestamp columns | Best when predicates are selective, such as equality, IN,
range, and null checks. |
+| [Bitmap](./global-index/bitmap) | Enum-like dimensions and tag columns |
Best for equality, IN, string prefix match, complement predicates, and null
checks over compressed row-id bitmaps. |
+| [Multivalue](./global-index/multivalue) | Membership tests on arrays of
supported scalar elements | Best for `ARRAY_CONTAINS`, `ARRAYS_OVERLAP`, and
`ARRAY_CONTAINS_ALL`. |
+| [FM](./global-index/fm) | Exact substring filters on character columns |
Supports needles of any byte length and partitions the indexed text for
bounded-memory construction and demand-loaded reads. |
+| [Vector](./global-index/vector) | Top-K similarity search on embeddings |
Uses ANN algorithms. Tune build-time and search-time options to balance recall,
latency, and index size. |
+| [Full-Text](./global-index/full-text) | Keyword search over text columns |
Uses full-text scoring and tokenizer configuration stored with each index file.
|
+| [Hybrid Search](./global-index/hybrid-search) | Combining multiple vector
routes, multiple full-text routes, or vector and full-text retrieval together |
Runs multiple scored routes and merges them with a ranker before reading rows. |
-Global indexes work on top of Data Evolution tables. To use global indexes,
your table **must** have:
+:::caution Index coverage affects results
-- `'row-tracking.enabled' = 'true'`
-- `'data-evolution.enabled' = 'true'`
+The default search mode is `fast`. If an index covers only part of the table,
+matching rows outside that coverage can be omitted. Creating a table or
appending
+data does not build a global index. Read [Coverage and
Freshness](./global-index/manage-indexes#coverage-and-freshness)
+before using indexed queries on changing data.
-> Global index queries may not be exact when the index only covers part of the
table data. If a query predicate matches the index, Paimon returns only the
results from the indexed portion. Matching records in data that has not been
indexed yet will not be returned.
+:::
## Prerequisites
-Create a table with the required properties:
+Start with a [configured Paimon catalog](./quick-start). Data Evolution
requires
+an append table with row tracking enabled. The following examples assume the
`db`
+database exists and that SQL uses that database. Load data before building
indexes.
+
+`global-index.enabled` defaults to `true`; setting it does not create index
files.
+
+The example table below is shared by the BTree, Bitmap, Multivalue, Vector, and
+Full-Text guides. It contains three-dimensional sample vectors and is
partitioned
+by `dt`, so the partition-scoped build and drop examples also apply. Create and
+populate it once using one of the following alternatives.
<Tabs groupId="global-index-create-table">
-<TabItem value="sql" label="SQL">
+<TabItem value="sql" label="Spark SQL">
```sql
CREATE TABLE my_table (
id INT,
name STRING,
+ category STRING,
+ tag STRING,
tags ARRAY<STRING>,
embedding ARRAY<FLOAT>,
- content STRING
-) TBLPROPERTIES (
+ content STRING,
+ dt STRING
+) PARTITIONED BY (dt) TBLPROPERTIES (
'row-tracking.enabled' = 'true',
- 'data-evolution.enabled' = 'true',
- 'global-index.enabled' = 'true'
+ 'data-evolution.enabled' = 'true'
);
-```
-
-</TabItem>
-<TabItem value="python-sdk" label="Python SDK">
-
-```python
-import pyarrow as pa
-
-from pypaimon import Schema
-
-schema = Schema.from_pyarrow_schema(
- pa.schema([
- pa.field("id", pa.int32()),
- pa.field("name", pa.string()),
- pa.field("tags", pa.list_(pa.string())),
- pa.field("embedding", pa.list_(pa.float32())),
- pa.field("content", pa.string()),
- ]),
- options={
- "row-tracking.enabled": "true",
- "data-evolution.enabled": "true",
- "global-index.enabled": "true",
- },
-)
-
-catalog.create_table("db.my_table", schema, ignore_if_exists=False)
+INSERT INTO my_table VALUES
+ (101, 'a200', 'electronics', 'vip', array('blue', 'green'),
+ array(1.0f, 0.0f, 0.0f), 'paimon lake format and search', '2026-06-18'),
+ (102, 'a300', 'electronics', 'trial', array('green'),
+ array(0.9f, 0.1f, 0.0f), 'paimon full text search', '2026-06-18'),
+ (103, 'a400', 'books', 'standard', array('red'),
+ array(0.0f, 1.0f, 0.0f), 'vector search tutorial', '2026-06-19');
```
</TabItem>
-</Tabs>
-
-## Lifecycle
-
-Create global indexes for all partitions or only selected partitions:
-
-<Tabs groupId="global-index-build">
-
-<TabItem value="sql" label="SQL">
+<TabItem value="flink-sql" label="Flink SQL">
```sql
-CALL sys.create_global_index(
- table => 'db.my_table',
- index_column => 'name',
- index_type => 'btree'
-);
-
-CALL sys.create_global_index(
- table => 'db.my_table',
- index_column => 'name',
- index_type => 'btree',
- partitions => 'dt=2026-06-18;dt=2026-06-19'
-);
-```
-
-</TabItem>
-
-<TabItem value="python-sdk" label="Python SDK">
-
-```python
-table = catalog.get_table("db.my_table")
-
-added_files = table.create_global_index("name")
-print(added_files)
-```
-
-The API returns the number of committed index files. You can pass build
options and restrict the
-build to selected partitions:
-
-```python
-added_files = table.create_global_index(
- "name",
- index_type="btree",
- partitions=[{"dt": "2026-06-18"}, {"dt": "2026-06-19"}],
- options={"sorted-index.records-per-range": "10000000"},
-)
-```
-
-</TabItem>
-
-</Tabs>
+SET 'execution.runtime-mode' = 'batch';
-PyPaimon global index build currently supports single-column BTree indexes,
-single-column Bitmap indexes, single-column paimon-vindex vector indexes,
-and single-column full-text indexes on tables with row tracking enabled.
-
-Drop index files:
-
-<Tabs groupId="global-index-drop">
-
-<TabItem value="sql" label="SQL">
-
-```sql
-CALL sys.drop_global_index(
- table => 'db.my_table',
- index_column => 'name',
- index_type => 'btree'
+CREATE TABLE my_table (
+ id INT,
+ name STRING,
+ category STRING,
+ tag STRING,
+ tags ARRAY<STRING>,
+ embedding ARRAY<FLOAT>,
+ content STRING,
+ dt STRING
+) PARTITIONED BY (dt) WITH (
+ 'row-tracking.enabled' = 'true',
+ 'data-evolution.enabled' = 'true'
);
-```
-
-</TabItem>
-<TabItem value="python-sdk" label="Python SDK">
-
-```python
-table = catalog.get_table("db.my_table")
-
-dropped_files = table.drop_global_index("name", index_type="btree")
-print(dropped_files)
-```
-
-You can also restrict the drop to selected partitions, or count matched files
without committing:
-
-```python
-matched_files = table.drop_global_index(
- "name",
- index_type="btree",
- partitions=[{"dt": "2026-06-18"}, {"dt": "2026-06-19"}],
- dry_run=True,
-)
+INSERT INTO my_table VALUES
+ (101, 'a200', 'electronics', 'vip', ARRAY['blue', 'green'],
+ ARRAY[CAST(1.0 AS FLOAT), CAST(0.0 AS FLOAT), CAST(0.0 AS FLOAT)],
+ 'paimon lake format and search', '2026-06-18'),
+ (102, 'a300', 'electronics', 'trial', ARRAY['green'],
+ ARRAY[CAST(0.9 AS FLOAT), CAST(0.1 AS FLOAT), CAST(0.0 AS FLOAT)],
+ 'paimon full text search', '2026-06-18'),
+ (103, 'a400', 'books', 'standard', ARRAY['red'],
+ ARRAY[CAST(0.0 AS FLOAT), CAST(1.0 AS FLOAT), CAST(0.0 AS FLOAT)],
+ 'vector search tutorial', '2026-06-19');
```
-</TabItem>
-
-</Tabs>
-
-Global indexes are stored in index files and recorded in table metadata. To
inspect index files and
-their row-id coverage, query the `table_indexes` system table:
-
-<Tabs groupId="global-index-table-indexes">
-
-<TabItem value="sql" label="SQL">
-
-```sql
-SELECT index_type, index_field_name, row_range_start, row_range_end
-FROM my_table$table_indexes
-WHERE index_field_name IS NOT NULL;
-```
+Wait for the insert job to finish before building an index.
</TabItem>
<TabItem value="python-sdk" label="Python SDK">
```python
-import pyarrow.compute as pc
-
-table_indexes = catalog.get_table("db.my_table$table_indexes")
-read_builder = table_indexes.new_read_builder().with_projection([
- "index_type",
- "index_field_name",
- "row_range_start",
- "row_range_end",
-])
-
-pa_table = read_builder.new_read().to_arrow(
- read_builder.new_scan().plan().splits()
-)
-pa_table = pa_table.filter(pc.is_valid(pa_table["index_field_name"]))
-print(pa_table)
-```
-
-</TabItem>
-
-</Tabs>
-
-You can also query `file_key_ranges` to inspect data file row-id ranges and
diagnose coverage:
-
-<Tabs groupId="global-index-file-key-ranges">
-
-<TabItem value="sql" label="SQL">
-
-```sql
-SELECT file_path, first_row_id, record_count
-FROM my_table$file_key_ranges;
-```
-
-</TabItem>
-
-<TabItem value="python-sdk" label="Python SDK">
+import pyarrow as pa
+from pypaimon import Schema
-```python
-file_key_ranges = catalog.get_table("db.my_table$file_key_ranges")
-read_builder = file_key_ranges.new_read_builder().with_projection([
- "file_path",
- "first_row_id",
- "record_count",
+# `catalog` is the configured Paimon catalog; database `db` already exists.
+pa_schema = pa.schema([
+ ('id', pa.int32()),
+ ('name', pa.string()),
+ ('category', pa.string()),
+ ('tag', pa.string()),
+ ('tags', pa.list_(pa.string())),
+ ('embedding', pa.list_(pa.float32())),
+ ('content', pa.string()),
+ ('dt', pa.string()),
])
-
-pa_table = read_builder.new_read().to_arrow(
- read_builder.new_scan().plan().splits()
+schema = Schema.from_pyarrow_schema(
+ pa_schema,
+ partition_keys=['dt'],
Review Comment:
[P2] Make the shared fixture work with the partition-scoped Python builds
After executing this new setup, the build/rebuild workflow advertised in
Manage Global Indexes and the vector guide fails:
```python
table.create_global_index('name', index_type='btree')
table.create_global_index(
'name', index_type='btree',
partitions=[{'dt': '2026-06-18'}, {'dt': '2026-06-19'}],
)
```
The second call raises `IndexError: Position 7 is out of bounds for row
arity 1`. `dt` is column 7 in this fixture:
`CreateGlobalIndexBuilder._resolve_partition_filter` constructs a full-row
predicate, but `build_plan.indexed_row_ranges` applies it directly to the
one-field index-manifest partition row. I also reproduced this with the
documented IVF-flat calls. Moving `dt` to the first schema position in a
control fixture makes the BTree rebuild return `0` as expected. The underlying
SDK bug predates this PR, but the new shared walkthrough encounters it. Please
provide a working fixture/workaround or explicitly avoid this unsupported
sequence until the predicate projection is fixed.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]