leaves12138 commented on code in PR #9718: URL: https://github.com/apache/paimon/pull/9718#discussion_r3977559550
########## docs/docs/primary-key-table/compaction.md: ########## @@ -24,172 +24,122 @@ under the License. # Compaction -When more and more records are written into the LSM tree, the number of sorted runs will increase. Because querying an -LSM tree requires all sorted runs to be combined, too many sorted runs will result in a poor query performance, or even -out of memory. +Compaction combines sorted runs and applies the [merge engine](./merge-engine/) to records with +the same key. It reduces file and merge overhead for reads, while consuming CPU and storage I/O +on the write side. Paimon's default strategy selects runs using a universal compaction policy. -To limit the number of sorted runs, we have to merge several sorted runs into one big sorted run once in a while. This -procedure is called compaction. + -However, compaction is a resource intensive procedure which consumes a certain amount of CPU time and disk IO, so too -frequent compaction may in turn result in slower writes. It is a trade-off between query and write performance. Paimon -currently adopts a compaction strategy similar to Rocksdb's [universal compaction](https://github.com/facebook/rocksdb/wiki/Universal-Compaction). +Depending on the table options, compaction also generates a +[changelog](./changelog-producer), maintains [deletion vectors](./table-mode#merge-on-write) and +[indexes](./global-index#maintenance-and-coverage), or applies record-level expiration. +Compaction itself does not mean that old files are immediately deleted: snapshot, tag, and +partition retention have separate lifecycles. See [Maintenance](../maintenance/). -Compaction solves: +## Choose a Compaction Strategy -1. Reduce Level 0 files to avoid poor query performance. -2. Produce changelog via [changelog-producer](./changelog-producer). -3. Produce deletion vectors for [MOW mode](./table-mode#merge-on-write). -4. Snapshot Expiration, Tag Expiration, Partitions Expiration. - -Limitation: - -- There can only be one job working on the same partition's compaction, otherwise it will cause conflicts and one side will throw an exception failure. - -Writing performance is almost always affected by compaction, so its tuning is crucial. +| Need | Configuration or operation | Trade-off | +| --- | --- | --- | +| Control overlapping runs | Default background compaction and sorted-run thresholds | Lower thresholds spend more write resources to reduce read work | +| Fresh lookup changelogs or MOW rows | Lookup compaction with the default wait behavior | Commit latency includes required compaction work | +| Favor write throughput | [Asynchronous Compaction](#asynchronous-compaction) | More pending files and potentially older query results | +| Isolate compaction resources or coordinate writers | [Dedicated compaction job](#dedicated-compaction-job) | Requires a separately operated job | +| Refresh a read-optimized view | `compaction.optimization-interval` | Freshness follows completed full compactions | +| Fully merge every N delta commits | `full-compaction.delta-commits` | Synchronous full compaction increases write amplification | ## Asynchronous Compaction -Compaction is inherently asynchronous, but if you want it to be completely asynchronous without blocking writes, -expecting a mode for maximum writing throughput, the compaction can be done slowly and not in a hurry. -You can use the following strategies for your table: +Writers normally run compaction in background threads. They can still wait when too many sorted +runs accumulate, or when lookup compaction is needed before a commit. -```shell +The following table options illustrate a configuration that relaxes those waits: + +```properties num-sorted-run.stop-trigger = 2147483647 sort-spill-threshold = 10 lookup-wait = false ``` -This configuration will generate more files during peak write periods and gradually merge them for optimal read -performance during low write periods. +This effectively removes the sorted-run write-stall limit and allows pending work to accumulate. +It is a throughput-oriented example, not a default recommendation. Size the spill storage and +monitor file counts, compaction backlog, and read latency before using it. + +By default, MOW and `first-row` batch reads exclude pending Level-0 data until lookup compaction +publishes it. MOW batch readers can opt into merging pending data with +`deletion-vectors.merge-on-read`; see [MOW visibility](./table-mode#merge-on-write). +For `changelog-producer = lookup`, generated changelogs are also delayed. A compactor that cannot +keep up with sustained input will keep falling behind; relaxing waits does not add capacity. ## Dedicated compaction job -In general, if you expect multiple jobs to be written to the same table, you need to separate the compaction. You can -use [dedicated compaction job](../maintenance/dedicated-compaction#dedicated-compaction-job). +Set `write-only = true` on ingest writers and run a +[dedicated compaction job](../maintenance/dedicated-compaction#dedicated-compaction-job) when +compaction needs separate resources or multiple writers need a single compaction owner. +Avoid overlapping compaction jobs for the same partition, which can cause commit conflicts. -## Record-Level expire +This does not lift [dynamic-bucket](./data-distribution#dynamic-bucket) restrictions on concurrent +writers to the same partition. -In compaction, you can configure record-Level expire time to expire records, you should configure: +## Record-Level expire -1. `'record-level.expire-time'`: time retain for records. -2. `'record-level.time-field'`: time field for record level expire. +Configure `record-level.expire-time` for the retention duration and `record-level.time-field` for +the field used to evaluate each record's age. -Expiration happens in compaction, and there is no strong guarantee to expire records in time. -You can trigger a full compaction manually to expire records which were not expired in time. +Expiration happens when compaction processes the records, so it has no strict wall-clock deadline. +A manual full compaction can process records that ordinary compaction has not reached. This is +separate from expiring snapshots or entire partitions. ## Full Compaction -Paimon Compaction uses [Universal-Compaction](https://github.com/facebook/rocksdb/wiki/Universal-Compaction). -By default, when there is too much incremental data, Full Compaction will be automatically performed. You don't usually -have to worry about it. +Full compaction merges all runs of a bucket into its highest level. The default strategy can +select a full compaction as data accumulates. To request one regularly: -Paimon also provides a configuration that allows for regular execution of Full Compaction. +| Option | Scheduling | Use | +| --- | --- | --- | +| `compaction.optimization-interval` | Time-based optimization compaction | Keep the [read-optimized table](../concepts/system-tables#read-optimized-table) reasonably fresh | +| `full-compaction.delta-commits` | Synchronous compaction after a number of delta commits | COW behavior or periodic full-compaction changelogs | -1. 'compaction.optimization-interval': Implying how often to perform an optimization full compaction, this - configuration is used to ensure the query timeliness of the read-optimized system table. -2. 'full-compaction.delta-commits': Full compaction will be constantly triggered after delta commits. Its disadvantage - is that it can only perform compaction synchronously, which will affect writing efficiency. +`full-compaction.delta-commits` is incompatible with `changelog-producer = lookup`. See +[Full Compaction Changelogs](./changelog-producer#full-compaction) for producer-specific defaults. ## Lookup Compaction -When primary key table is configured with `lookup` [changelog producer](./changelog-producer) -or `first-row` [merge-engine](./merge-engine/) -or has enabled `deletion vectors` for [MOW mode](./table-mode#merge-on-write), Paimon will -use a radical compaction strategy to force compacting level 0 files to higher levels for every compaction trigger. +Paimon uses lookup compaction for the `lookup` changelog producer, the `first-row` merge engine, +and tables with deletion vectors. It reconciles Level-0 records with existing rows. -Paimon also provides configurations to optimize the frequency of this compaction. +| Option | Behavior | +| --- | --- | +| `lookup-compact = radical` (default) | Force new Level-0 files into higher levels at compaction triggers | +| `lookup-compact = gentle` | Use the universal strategy with a configurable forced-compaction interval | +| `lookup-compact.max-interval` | Number of compaction-selection attempts without universal compaction work before forcing Level-0 compaction; only used in `gentle` mode | -1. 'lookup-compact': compact mode used for lookup compaction. Possible values: `radical`, will use - `ForceUpLevel0Compaction` strategy to radically compact new files; `gentle`, will use `UniversalCompaction` strategy - to gently compact new files; -2. 'lookup-compact.max-interval': The max interval for a forced L0 lookup compaction to be triggered in `gentle` mode. - This option is only valid when `lookup-compact` mode is `gentle`. +To defer forced Level-0 compaction, use `gentle` together with an explicit +`lookup-compact.max-interval`. The interval has no default value; leaving it unset still forces +Level-0 compaction immediately when the universal strategy selects no work. It counts selection Review Comment: **[P2] Document the effective default for gentle lookup compaction** Setting only `lookup-compact = gentle` already defers forced Level-0 compaction. Although the config declaration uses `noDefaultValue()`, `CoreOptions.lookupCompactMaxInterval()` computes `2 * num-sorted-run.compaction-trigger` when it is unset (10 with the default trigger), and `MergeTreeCompactManagerFactory.createCompactStrategy()` passes that non-null interval to `ForceUpLevel0Compaction`. The existing `KeyValueFileStoreWriteTest.testGentleLookupCompactStrategy` explicitly asserts 10 without configuring a max interval. Immediate fallback when the universal strategy selects no work is the **radical** path, which passes a null interval. The current wording therefore understates the potential data/changelog visibility delay when users select gentle mode alone. Please describe the computed runtime default instead; an explicitly configured interval is also clamped to at least `num-sorted-run.compaction-trigger`. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
