SteNicholas opened a new issue, #409: URL: https://github.com/apache/paimon-cpp/issues/409
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar. ### Motivation Sub-issue of #399 (step 5: primary-key full-text index, write side). `BucketedPrimaryKeyIndexMaintainer` keeps only BTree definitions (`src/paimon/core/index/pk/bucketed_primary_key_index_maintainer.cpp:128-132`), and restore keeps only BTree payloads (`IsPrimaryKeyBTreePayload`). This has two effects: - Paimon C++ writers never build primary-key full-text archives. - When C++ compacts a table that already has Java-written archives, those archives are neither rebuilt nor retired. They keep pointing at data files that compaction has removed, and the new compacted files are not indexed. Java model (apache/paimon#8649, apache/paimon#8651, apache/paimon#8672, apache/paimon#8992): - **One archive per level.** Each (partition, bucket, non-zero data level) has exactly one immutable archive. It covers all eligible files of that level, sorted by file name. - **Eligible files.** A file is eligible when `fileSource == COMPACT && level > 0` (`PrimaryKeyIndexSourcePolicy`). - **Row ids.** An archive's row ids are the concatenated physical row positions of its source files, including rows deleted by deletion vectors and null rows. - **`PkFullTextIndexFile`** builds one archive: - index type `full-text`; - every source has the same level > 0 and a positive row count; - it calls `GlobalIndexSingleColumnWriter#write(text, sourceOffset + rowPos)` for each row; - `finish()` must return exactly one entry whose row count equals the total source row count. - The resulting `IndexFileMeta` has `GlobalIndexMeta(0, total - 1, fieldId, null, indexMeta, sourceMeta)`: - `indexMeta` is the flat JSON of the prefix-stripped options; - `sourceMeta` is `PrimaryKeyIndexSourceMeta(level, sourceFiles)`. - The file is named `index-{uuid}-{N}` under the index directory, or in the bucket directory when `index-file-in-data-file-dir` is set. - **`PkFullTextDataFileReader`** reads one text value per physical row, with no deletion-vector filtering. **`PkFullTextIndexBuilder`** builds an archive for one file or a list of files. - **`PkFullTextBucketIndexState#fromActiveDataFiles`** classifies payloads: - A payload is current only if its (level, ordered source files) exactly matches the level's eligible active files and its row counts match. - A payload that fails this is stale. So is every payload of a level that has more than one match, and any payload whose metadata cannot be parsed. - **Level planning** (`PrimaryKeyIndexLevels`, shared with the sorted indexes) picks the lowest level whose payload is missing or out of date. A plan with no source files removes the payload. - **`BucketedFullTextIndexMaintainer`**: - **Restore:** stale payloads are retired and emitted as deletions in the next commit. - **`prepareCommit(append, compact, waitCompaction)`:** 1. Applies the data transition: removes `compactBefore` files and adds eligible `compactAfter` files. 2. Finishes or starts a single background build. 3. Atomically replaces the level's archive when the build is still valid (`canAccept`); otherwise deletes the generated file. 4. Routes the result to the compact increment if there was a compact transition, and to the append increment otherwise. 5. Rolls back and deletes generated files on failure. The returned commit carries an `abort` hook. - A failed build is not retried within the same call. The next `prepareCommit` plans again. - **Wiring:** - `BucketedPrimaryKeyIndexMaintainer.Factory` creates the full-text maintainer for fixed-bucket writes and for postpone-bucket compaction (apache/paimon#8992). - `IndexFileHandler#pkFullTextIndex(partition, bucket)` provides the index file. - Writers are not closed while a build is pending. ### Solution - Port the classes above. - Reuse the existing C++ `PrimaryKeyIndexSourceMeta`, `PrimaryKeyIndexSourcePolicy` and `PrimaryKeyIndexSourceFile`. Extract the level planning currently embedded in `src/paimon/core/index/pksorted/` so it can be shared. - Wire the full-text maintainer into `BucketedPrimaryKeyIndexMaintainer`: restore, `prepareCommit`, abort, and merging the increments. Create it through the `full-text` indexer from {{S1}}, using the options resolved in {{S9}}. - Add tests aligned with Java `PkFullTextIndexFileTest`, `PkFullTextDataFileReaderTest`, `PkFullTextBucketIndexStateTest` and `BucketedFullTextIndexMaintainerTest`, plus: - C++ compaction of a table with Java-written primary-key full-text archives; stale archives are retired and new ones built; - Java reading archives written by C++. ### Anything else? - Depends on {{S1}} and {{S9}}. - The realtime path still rejects primary-key global indexes (`src/paimon/core/utils/primary_key_table_utils.cpp:136-141`). That is out of scope here. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
