Hi everyone,

I opened RFC-4 as a pull request for review.
https://github.com/apache/incubator-xtable/pull/933

The RFC follows the earlier indexing thread [1] and the feature request
#887 [2].

What the RFC proposes:

- A canonical index model for XTable: a record-level (primary) index and a
secondary index first, with expression and vector indexes left for later
work.
- The Hudi metadata table as the initial storage for these indexes, so
Iceberg, Delta Lake, Paimon and Parquet tables get indexes today. Iceberg
has a secondary index design in progress but no index support in the
specification yet. When Iceberg defines an index format, XTable can convert
indexes between formats the way XTable converts table metadata today.
- A small lookup API that returns the file and the row position of each
matching row, so an engine can accelerate joins and merges instead of
scanning the target table.
- Index builds reuse the existing sync loop. Record-level index builds need
a distributed engine, so they run behind an execution provider interface,
and xtable-core stays runnable without Spark.

The RFC has three Mermaid flowcharts, and the plain diff shows only the
source of each one. Two ways to see the rendered diagrams:

- Open the rendered file directly:
https://github.com/vinishjail97/onetable/blob/rfc-4-index-support/rfc/rfc-4/rfc-4.md

- Or open the Files changed tab on the pull request and click the  "Display
the rich diff" button on the file header.

I would like feedback from the community on the index model, the lookup
API, the phasing and anything the RFC misses. Please comment on the pull
request or reply on this thread.

Thanks,
Vinish

[1] https://lists.apache.org/thread/kb3p85rgy0qkvjgdgn3580h36y1qov8c
[2] https://github.com/apache/incubator-xtable/issues/887

Reply via email to