SteNicholas opened a new issue, #409:
URL: https://github.com/apache/paimon-cpp/issues/409

   ### Search before asking
   
   - [x] I searched in the 
[issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.
   
   ### Motivation
   
   Sub-issue of #399 (step 5: primary-key full-text index, write side).
   
   `BucketedPrimaryKeyIndexMaintainer` keeps only BTree definitions 
(`src/paimon/core/index/pk/bucketed_primary_key_index_maintainer.cpp:128-132`), 
and restore keeps only BTree payloads (`IsPrimaryKeyBTreePayload`). This has 
two effects:
   
   - Paimon C++ writers never build primary-key full-text archives.
   - When C++ compacts a table that already has Java-written archives, those 
archives are neither rebuilt nor retired. They keep pointing at data files that 
compaction has removed, and the new compacted files are not indexed.
   
   Java model (apache/paimon#8649, apache/paimon#8651, apache/paimon#8672, 
apache/paimon#8992):
   
   - **One archive per level.** Each (partition, bucket, non-zero data level) 
has exactly one immutable archive. It covers all eligible files of that level, 
sorted by file name.
   - **Eligible files.** A file is eligible when `fileSource == COMPACT && 
level > 0` (`PrimaryKeyIndexSourcePolicy`).
   - **Row ids.** An archive's row ids are the concatenated physical row 
positions of its source files, including rows deleted by deletion vectors and 
null rows.
   - **`PkFullTextIndexFile`** builds one archive:
     - index type `full-text`;
     - every source has the same level > 0 and a positive row count;
     - it calls `GlobalIndexSingleColumnWriter#write(text, sourceOffset + 
rowPos)` for each row;
     - `finish()` must return exactly one entry whose row count equals the 
total source row count.
     - The resulting `IndexFileMeta` has `GlobalIndexMeta(0, total - 1, 
fieldId, null, indexMeta, sourceMeta)`:
       - `indexMeta` is the flat JSON of the prefix-stripped options;
       - `sourceMeta` is `PrimaryKeyIndexSourceMeta(level, sourceFiles)`.
     - The file is named `index-{uuid}-{N}` under the index directory, or in 
the bucket directory when `index-file-in-data-file-dir` is set.
   - **`PkFullTextDataFileReader`** reads one text value per physical row, with 
no deletion-vector filtering. **`PkFullTextIndexBuilder`** builds an archive 
for one file or a list of files.
   - **`PkFullTextBucketIndexState#fromActiveDataFiles`** classifies payloads:
     - A payload is current only if its (level, ordered source files) exactly 
matches the level's eligible active files and its row counts match.
     - A payload that fails this is stale. So is every payload of a level that 
has more than one match, and any payload whose metadata cannot be parsed.
   - **Level planning** (`PrimaryKeyIndexLevels`, shared with the sorted 
indexes) picks the lowest level whose payload is missing or out of date. A plan 
with no source files removes the payload.
   - **`BucketedFullTextIndexMaintainer`**:
     - **Restore:** stale payloads are retired and emitted as deletions in the 
next commit.
     - **`prepareCommit(append, compact, waitCompaction)`:**
       1. Applies the data transition: removes `compactBefore` files and adds 
eligible `compactAfter` files.
       2. Finishes or starts a single background build.
       3. Atomically replaces the level's archive when the build is still valid 
(`canAccept`); otherwise deletes the generated file.
       4. Routes the result to the compact increment if there was a compact 
transition, and to the append increment otherwise.
       5. Rolls back and deletes generated files on failure. The returned 
commit carries an `abort` hook.
     - A failed build is not retried within the same call. The next 
`prepareCommit` plans again.
   - **Wiring:**
     - `BucketedPrimaryKeyIndexMaintainer.Factory` creates the full-text 
maintainer for fixed-bucket writes and for postpone-bucket compaction 
(apache/paimon#8992).
     - `IndexFileHandler#pkFullTextIndex(partition, bucket)` provides the index 
file.
     - Writers are not closed while a build is pending.
   
   ### Solution
   
   - Port the classes above.
   - Reuse the existing C++ `PrimaryKeyIndexSourceMeta`, 
`PrimaryKeyIndexSourcePolicy` and `PrimaryKeyIndexSourceFile`. Extract the 
level planning currently embedded in `src/paimon/core/index/pksorted/` so it 
can be shared.
   - Wire the full-text maintainer into `BucketedPrimaryKeyIndexMaintainer`: 
restore, `prepareCommit`, abort, and merging the increments. Create it through 
the `full-text` indexer from {{S1}}, using the options resolved in {{S9}}.
   - Add tests aligned with Java `PkFullTextIndexFileTest`, 
`PkFullTextDataFileReaderTest`, `PkFullTextBucketIndexStateTest` and 
`BucketedFullTextIndexMaintainerTest`, plus:
     - C++ compaction of a table with Java-written primary-key full-text 
archives; stale archives are retired and new ones built;
     - Java reading archives written by C++.
   
   ### Anything else?
   
   - Depends on {{S1}} and {{S9}}.
   - The realtime path still rejects primary-key global indexes 
(`src/paimon/core/utils/primary_key_table_utils.cpp:136-141`). That is out of 
scope here.
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to