pvary commented on code in PR #16961: URL: https://github.com/apache/iceberg/pull/16961#discussion_r4151400392
########## format/index-spec.md: ########## @@ -0,0 +1,748 @@ +--- +title: "Index Spec" +--- +<!-- + - Licensed to the Apache Software Foundation (ASF) under one or more + - contributor license agreements. See the NOTICE file distributed with + - this work for additional information regarding copyright ownership. + - The ASF licenses this file to You under the Apache License, Version 2.0 + - (the "License"); you may not use this file except in compliance with + - the License. You may obtain a copy of the License at + - + - http://www.apache.org/licenses/LICENSE-2.0 + - + - Unless required by applicable law or agreed to in writing, software + - distributed under the License is distributed on an "AS IS" BASIS, + - WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + - See the License for the specific language governing permissions and + - limitations under the License. + --> +# Iceberg Index Specification + +## Background and Motivation + +An index is a secondary store of data from a table that is structured to accelerate specific access patterns. +An index is derived from the rows of a source table and is stored separately from table data, so it can be built, +refreshed, or dropped without rewriting the table. + +An index is most valuable when it is a property of the table rather than of the engine that built it. This +specification defines a common format for index metadata and a common storage architecture for index data, so that any +engine can build an index, maintain it, and use it to plan queries against the table. + +## Goals + +* **Independence** -- An index is committed as a separate object, without modifying its source table. +* **Consistency** -- An index reflects exactly the live rows of a single state of its source table. +* **Read-optimized** -- An index is structured for fast reads, at the cost of extra work when writing. +* **Incrementality** -- An index is refreshed by writing only the index data affected by the table's changes. +* **Scalability** -- An index supports any table size that the table spec supports. +* **Portability** -- An index is readable and maintainable by any engine, not only the one that wrote it. +* **Extensibility** -- An index type or ordering strategy can be added without disrupting existing indexes or + engines that do not implement it. + +## Overview + +Index state is maintained in index metadata files. All changes to index state create a new metadata file and replace +the old metadata with an atomic swap, as defined in [Commits and Concurrency](#commits-and-concurrency). An index +metadata file tracks the index definition, index properties, and extracts of a single index. A table may have +any number of indexes, including multiple indexes of the same type. + +An extract represents the state of an index for one snapshot of its source table and is used to access the +complete set of index data files for that state. The index data of an extract is organized as a +[tracking file](#tracking-file) that lists a set of [region files](#region-files). Index data files are immutable and +may be referenced by more than one extract. + +## Specification + +### Terms + +* **Index** -- A structure that accelerates retrieval of rows from a source table. +* **Extract** -- The state of an index for a single snapshot of the source table. +* **Index entry** -- The values produced by the index fields for one indexed row of the source table. +* **Ordering key** -- The tuple of values that determines the position of an index entry within an extract. +* **Tracking file** -- A file that lists the region files of an extract; one per extract. +* **Region file** -- A file that stores the index entries for a range of ordering keys; a subset of an extract. + +### Locations in Metadata + +Location strings stored in index metadata are classified and resolved as defined by +[paths in metadata](spec.md#paths-in-metadata) in the table specification. Relative locations are resolved against the +index `location`, which must be an absolute location. + +### Index Definition + +An index is defined by a source table, an index type, identity fields, materialized fields, non-materialized fields, and +an ordering key. The definition is fixed when the index is created and must not change for the lifetime of the index, +so region files remain readable through every extract that references them. A different definition requires a +new index. + +Index properties configure how an index is written and maintained, such as the target size of region files. Any commit +may change properties, and readers must not depend on them. + +#### Index Type + +The index type defines the logical category of an index and the class of queries it accelerates. + +| Type | Description | +|----------|------------------------------------------------------------------------------------------------------------------------| +| `scalar` | Accelerates point lookups on ordering key fields, and range filters when the ordering expressions are order preserving | + +This specification defines a single index type, `scalar`. Future specifications may define additional types. + +Writers must write `index-type` in lower case. Readers must match it case-insensitively. + +#### Index Fields + +An index field defines one value of an index entry, produced for an indexed row of the source table. An index declares +three lists of index fields: + +* [Identity fields](#identity-fields) store a source table field in region files as is, so a reader can return the + indexed values and distinguish entries that share an ordering key. +* [Materialized fields](#materialized-fields) store a value computed from an indexed row in region files. A value is + materialized when a reader cannot recompute it from the stored fields, such as the file and position, or when + recomputing it would cost more than storing it, such as a bucket or Hilbert value. +* [Non-materialized fields](#non-materialized-fields) keep only statistics in + [tracking file entries](#tracking-file-entry), for a value a reader can recompute from the stored fields, so it can + take part in ordering and pruning without being stored for every entry. + +Every index field has a field ID that must be unique across the three lists. + +Every source table field that an index field references must be present in the source schema of an [extract](#extract). +When a referenced field has been dropped, no new extract can be created, but existing extracts remain readable using +their own source schema. + +##### Identity Fields + +`identity-fields` is a non-empty list of unique source table field IDs. Each entry must reference a data field. +[Metadata columns](spec.md#reserved-field-ids) are not allowed. Each listed field is stored in the +[region files](#region-files) under its own field ID and takes its type from the source schema of an +[extract](#extract). + +Every source table field referenced by an expression field in the [ordering key](#ordering-key) must be an identity +field. + +##### Expression Fields + +Both [materialized fields](#materialized-fields) and [non-materialized fields](#non-materialized-fields) are expression +fields; they differ only in where their values are kept. + +The value of an expression field is produced by evaluating an +[Iceberg value expression](expressions-spec.md#value-expressions) for an indexed row of the source table. +An expression field has the following fields: + +| Requirement | Field name | Type | Description | +|-------------|---------------|-------------------|--------------------------------------------------------------| +| _required_ | `field-id` | `int` | ID that uniquely identifies the index field | +| _required_ | `type` | `expr-value` | Expression field representation | +| _required_ | `data-type` | Iceberg type | Type produced by the expression | +| _required_ | `expr` | JSON expression | Value expression that produces the field, serialized as JSON | + +Each expression field must satisfy the following requirements: + +- `expr` must contain only ID references to source table fields or + [metadata columns](spec.md#reserved-field-ids). Named references must not be used. The `_deleted`, `_change_type`, + `_change_ordinal`, and `_commit_snapshot_id` metadata columns must not be referenced, and neither must the + `file_path`, `pos`, and `row` columns of delete files. +- `expr` must be deterministic and must produce the declared `data-type`. +- `field-id` must not be a [reserved field ID](spec.md#reserved-field-ids). +- `data-type` must not change. A source table schema change that makes `expr` incompatible with `data-type` requires a + new index definition and prevents new extracts from being created. + +Expressions are serialized using the [JSON serialization](expressions-spec.md#appendix-b-json-serialization) defined by +the expressions specification. Types are serialized using the [type serialization](spec.md#schemas) defined by the table +specification. + +###### Materialized Fields + +`materialized-fields` is a list of expression fields whose values are stored in the [region files](#region-files). +Evaluating the identity fields and the materialized fields for one indexed row produces one region file row. + +###### Non-Materialized Fields + +`non-materialized-fields` is a list of expression fields whose row values are not stored in region files. Only their +field statistics are stored, in [tracking file entries](#tracking-file-entry). + +#### Ordering Key + +`ordering-key` is a list of field IDs from `identity-fields`, `materialized-fields`, and `non-materialized-fields`. +The values of the referenced fields, in list order, form the ordering key of an indexed row and determine the row's +position in the index, as defined in [Ordering](#ordering). The list must not be empty. +Every referenced field must have a primitive type. + +### Index Metadata + +The index metadata file stores the index definition and extract history. It is encoded as JSON. + +#### Index Metadata File + +The index metadata file has the following fields: + +| Requirement | Field name | Type | Description | +|-------------|---------------------------|----------------------------|----------------------------------------------------------------------------------------------------| +| _required_ | `format-version` | `int` | Index format version; must be `1` | +| _required_ | `index-uuid` | `string` | Stable UUID assigned at creation | +| _required_ | `table-uuid` | `string` | UUID of the indexed table | +| _required_ | `location` | `string` | Index root location | +| _required_ | `last-updated-ms` | `long` | Timestamp when the index was last updated (ms from epoch) [1] | +| _required_ | `index-type` | `string` | Logical index type | +| _required_ | `identity-fields` | `list<int>` | Source table fields stored in region files, see [Identity Fields](#identity-fields) | +| _optional_ | `materialized-fields` | `list<expression-field>` | Expression fields stored in region files, see [Materialized Fields](#materialized-fields) | +| _optional_ | `non-materialized-fields` | `list<expression-field>` | Fields stored only in tracking statistics, see [Non-Materialized Fields](#non-materialized-fields) | +| _required_ | `ordering-key` | `list<int>` | Field IDs that form the ordering key, see [Ordering Key](#ordering-key) | +| _optional_ | `properties` | `map<string, string>` | Index properties applicable for every extract | +| _optional_ | `extracts` | `list<extract>` | Extracts [2] | +| _optional_ | `metadata-log` | `list<metadata-log-entry>` | Previous index metadata files, see [Metadata Log](#metadata-log) | +| _optional_ | `encryption-keys` | `list<encryption-key>` | Encryption keys used by the index, see [Encryption Keys](#encryption-keys) | + +A missing optional list must be read as an empty list. + +Notes: + +1. Each index metadata file should update `last-updated-ms` just before writing. +2. An index that has not been built yet has no extracts. +3. Index names are not stored in index metadata. It is the catalog's responsibility to map index names to metadata file + locations. +4. How the indexes of a table are discovered is out of scope for this specification and is defined by the catalog + specification. + +#### Extract + +An extract is an immutable version of the index data generated from a specific source table snapshot. It +references a complete set of index files through the location of a single [tracking file](#tracking-file). + +An extract must index exactly the live rows of the referenced table snapshot. + +The referenced snapshot must have a `schema-id`. The schema it identifies is the **source schema** of the extract, the +schema that index fields resolve source table fields against. Review Comment: The schema-id is only optional, so very old snapshots could be reused. We fill this for a very long time, and most likely any relevant snapshot will have it. If the snapshot is removed, the schema could be removed as well. If the snapshot is removed, likely the index extract (snapshot) should be removed as well. I would prefer not to duplicate this information. If we want to duplicate something, then we should store the range file schema instead, but I prefer not to. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
