chrevanthreddy opened a new pull request, #20107:
URL: https://github.com/apache/hudi/pull/20107

   ## Summary
   Port the proven RFC-109 IVF/RaBitQ vector-index bootstrap and Spark search 
implementation onto a clean current-master branch, then mechanically separate 
the posting scanner, RLI arbitration, and Spark read-path components. This 
draft is a **review/benchmark staging PR**, not a claim of completed 10M 
qualification.
   
   Related: RFC-109 #19309; prior incremental draft #19802 (historical 
implementation; please review this refactor against it rather than treating 
both as separate features). Partitioned-RLI integration/coverage is tracked 
separately in #20106.
   
   ## Included
   - Vector MDT Avro/payload, RaBitQ factor contract, bootstrap posting writer, 
generation manifest/cache, Spark SQL CREATE INDEX wiring, and option validation.
   - Historical packed posting scanner, delta overlay/tombstones, RaBitQ 
scoring, finalist RLI arbitration, and Spark approximate/exact search with 
positional Parquet/key-fallback and log-resident fetch paths.
   - Brute-force TVF filter/max-distance backward compatibility, with IVF 
runtime tuning options.
   - Mechanical read-path extractions into `VectorSearchAlgorithms`, 
`VectorExactSearchPlanner`, `VectorIndexRliArbitrator`, posting serde and 
candidate helpers. The preserved read-path was ported first, then split.
   
   ## Scope boundaries / review warnings
   - **Not yet 10M benchmarked on this branch.** Historical BigANN 
SparkApplication YAML targets Spark 4.0 / Scala 2.13 and shared writable GCS 
paths; this branch builds Spark 3.5 / Scala 2.12. Do not apply old manifests 
unchanged or attribute historical benchmark results to this branch. Current 
workstation GCS access requires interactive reauthentication; 10M prerequisites 
and isolated outputs remain unverified.
   - Partitioned RLI detection is intentionally hardcoded `false` in 
`IvfRaBitQMdtSearchAlgorithm.isPartitionedRecordIndex` for this initial 
baseline. The existing partition-aware arbiter path exists, but must not be 
claimed correct for partitioned-RLI tables until #20106 is addressed.
   - HFile cache telemetry/performance experiment, range prefetch, and 
unrelated generic core fixes are excluded. Existing defaults and 
SI/RLI/bootstrap behavior should not be changed by optional caching work.
   - This is 18 commits / 87 files from current merge base because it includes 
schema/DDL/bootstrap foundations as well as read path. Earlier RFC-109 PRs 
overlap; reviewers may prefer splitting/rebasing after benchmark validation.
   
   ## Validation done (Spark 3.5 / Scala 2.12)
   - Reactor `test-compile` passed (16 modules; RAT, Checkstyle, Scalastyle).
   - Focused Spark JUnit: planner 8/8, Parquet locator 1/1, TVF argument 
compatibility 5/5 (14/14 total).
   - Existing execution-level TVF tests: single- and batch-query combined 
filter/max-distance and invalid filter handling: 3/3 passed. Maven continued 
into unrelated ScalaTest suites; it was stopped after selected tests passed, so 
this is **not** a full test-suite pass.
   - `TestVectorIndexOptions`: 10/10; posting/RaBitQ/arbiter focused tests 
passed in earlier refactor slices.
   - Spark 3.5/Scala 2.12 shaded bundle packaged successfully at 
`packaging/hudi-spark-bundle/target/hudi-spark3.5-bundle_2.12-1.3.0-SNAPSHOT.jar`
 (local artifact; not deployed).
   
   ## Remaining review gates
   - Arrange GCS reauthentication and compatible Spark 3.5/Scala 2.12 cluster 
runtime; stage the bundle and adapted BigANN app under an isolated, immutable 
run ID; require source `_SUCCESS` markers.
   - Run 10M COW bootstrap/index/query with exact and approximate modes, 
collect latency/recall and inspect Spark logs plus metrics `_SUCCESS` before 
claiming success. Add real numbers and artifacts here.
   - Confirm non-partitioned RLI lifecycle and broaden regression coverage; 
#20106 separately tracks partitioned RLI.
   
   Please keep this PR **draft** until the 10M gate and reviewer feedback are 
resolved.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to