leaves12138 commented on code in PR #9719:
URL: https://github.com/apache/paimon/pull/9719#discussion_r3978428846


##########
docs/docs/append-table/incremental-clustering.mdx:
##########
@@ -27,268 +27,257 @@ under the License.
 
 # Incremental Clustering
 
-Paimon currently supports ordering append tables using SFC (Space-Filling 
Curve)(see [sort compact](../maintenance/dedicated-compaction#sort-compact) for 
more info). 
-The resulting data layout typically delivers better performance for queries 
that target clustering keys. 
-However, with the current SortCompaction, even when neither the data nor the 
clustering keys have changed, 
-each run still rewrites the entire dataset, which is extremely costly. 
+Incremental clustering improves the data layout of append tables by sorting 
selected files on frequently filtered
+columns. Compared with repeatedly sorting an entire partition, it can reduce 
the amount of data rewritten while
+improving [file-statistics 
pruning](./query-performance#file-statistics-and-clustering). A run may select 
no files when
+its compaction criteria are not met. Full mode considers all runs in the 
selected scope, but can skip work that is
+already clustered; see [file selection](#implement).
 
-To address this, Paimon introduced a more flexible, incremental clustering 
mechanism—Incremental Clustering. 
-On each run, it selects only a specific subset of files to cluster, avoiding a 
full rewrite. This enables low-cost, 
-sort-based optimization of the data layout and improves query performance. In 
addition, with Incremental Clustering, 
-you can adjust clustering keys without rewriting existing data, the layout 
evolves dynamically as cluster runs and 
-gradually converges to an optimal state, significantly reducing the 
decision-making complexity around data layout.
+Clustering also merges small files, respecting `target-file-size`. It changes 
the physical layout, not the rows returned
+by a query, and does not replace SQL `ORDER BY`.
 
+## Requirements
 
-Incremental Clustering supports:
-- Support incremental clustering; minimizing write amplification as possible.
-- Support small-file compaction; during rewrites, respect target-file-size.
-- Support changing clustering keys; newly ingested data is clustered according 
to the latest clustering keys.
-- Provide a full mode; when selected, the entire dataset will be reclustered.
+| Requirement | Unaware-bucket append (`bucket = -1`) | Bucketed append 
(`bucket > 0`) |
+| --- | --- | --- |
+| Primary key | Must not be defined. | Must not be defined. |
+| Enable clustering | `clustering.incremental = true` and nonempty 
`clustering.columns`. | Same. |
+| Append ordering | No bucket-order guarantee. | Must set 
`bucket-append-ordered = false`. |
+| Deletion vectors | Supported. | Must remain disabled. |
+| Compaction execution | Schedule explicit clustering jobs; the Flink sink's 
normal background compaction is disabled. | Writer compaction and dedicated 
compact jobs use the bucket clustering path. |
+| Global/local sort mode | Configurable for batch clustering jobs. | 
Clustering is performed within each partition and bucket; the global/local 
option does not select this path. |
+| Historical-partition auto-clustering | Supported. | 
`clustering.history-partition.*` does not apply. |
 
-Incremental Clustering is supported for append tables in both unaware-bucket 
mode (`bucket = -1`) and
-bucketed mode (`bucket > 0`). For bucketed append tables, additional 
requirements apply because
-clustering gives up the ordered append guarantee within buckets.
+Data Evolution tables cannot enable incremental clustering. If streaming 
consumers require ordered append reads from a
+bucketed table, keep that ordering and do not enable clustering.
 
 ## Enable Incremental Clustering
 
-To enable Incremental Clustering, the following configuration needs to be set 
for the table:
-<table className="table table-bordered">
-    <thead>
-    <tr>
-      <th className="text-left" style={{width: "20%"}}>Option</th>
-      <th className="text-left" style={{width: "10%"}}>Value</th>
-      <th className="text-left" style={{width: "5%"}}>Required</th>
-      <th className="text-left" style={{width: "10%"}}>Type</th>
-      <th className="text-left" style={{width: "55%"}}>Description</th>
-    </tr>
-    </thead>
-    <tbody>
-    <tr>
-      <td><h5>clustering.incremental</h5></td>
-      <td>true</td>
-      <td style={{wordWrap: "break-word"}}>Yes</td>
-      <td>Boolean</td>
-      <td>Must be set to true to enable incremental clustering. Default is 
false.</td>
-    </tr>
-    <tr>
-      <td><h5>clustering.columns</h5></td>
-      <td>'clustering-columns'</td>
-      <td style={{wordWrap: "break-word"}}>Yes</td>
-      <td>String</td>
-      <td>The clustering columns, in the format 'columnName1,columnName2'. It 
is not recommended to use partition keys as clustering keys.</td>
-    </tr>
-    <tr>
-      <td><h5>clustering.strategy</h5></td>
-      <td>'zorder' or 'hilbert' or 'order'</td>
-      <td style={{wordWrap: "break-word"}}>No</td>
-      <td>String</td>
-      <td>The ordering algorithm used for clustering. If not set, It'll 
decided from the number of clustering columns. 'order' is used for 1 column, 
'zorder' for less than 5 columns, and 'hilbert' for 5 or more columns.</td>
-    </tr>
-    <tr>
-      <td><h5>clustering.incremental.mode</h5></td>
-      <td>'global-sort' or 'local-sort'</td>
-      <td style={{wordWrap: "break-word"}}>No</td>
-      <td>Enum</td>
-      <td>The sort mode for incremental clustering compaction. Default is 
<code>global-sort</code>. <code>global-sort</code> performs a global range 
shuffle across tasks before local sorting, output files are globally ordered by 
the clustering columns at the cost of network shuffling. 
<code>local-sort</code> skips the global shuffle and sorts rows only within 
each compaction task independently, each output file is internally ordered but 
there is no global ordering across files, this mode is cheaper and sufficient 
for per-file Parquet lookup optimizations.</td>
-    </tr>
-    </tbody>
-
-</table>
-
-For bucketed append tables (`bucket > 0`), you must also set the following 
option:
-
-<table className="table table-bordered">
-    <thead>
-    <tr>
-      <th className="text-left" style={{width: "20%"}}>Option</th>
-      <th className="text-left" style={{width: "10%"}}>Value</th>
-      <th className="text-left" style={{width: "5%"}}>Required</th>
-      <th className="text-left" style={{width: "10%"}}>Type</th>
-      <th className="text-left" style={{width: "55%"}}>Description</th>
-    </tr>
-    </thead>
-    <tbody>
-    <tr>
-      <td><h5>bucket-append-ordered</h5></td>
-      <td>false</td>
-      <td style={{wordWrap: "break-word"}}>Yes</td>
-      <td>Boolean</td>
-      <td>Must be set to false for bucketed append tables with incremental 
clustering.</td>
-    </tr>
-    </tbody>
-
-</table>
-
-Bucketed append tables with Incremental Clustering do not support 
`deletion-vectors.enabled = true`.
-
-Example:
+Set the clustering keys on the table using the DDL for its layout. Choose the 
tab for your engine.
+
+### Unaware-Bucket Table
+
+For `my_table` from the [overview](./), enable clustering and specify the 
columns:
+
+<Tabs groupId="engine">
+<TabItem value="flink" label="Flink">
+
+```sql
+ALTER TABLE my_table SET (
+    'clustering.incremental' = 'true',
+    'clustering.columns' = 'product_id,price'
+);
+```
+
+</TabItem>
+<TabItem value="spark" label="Spark">
 
 ```sql
-ALTER TABLE T SET (
+ALTER TABLE my_table SET TBLPROPERTIES (
+    'clustering.incremental' = 'true',
+    'clustering.columns' = 'product_id,price'
+);
+```
+
+</TabItem>
+</Tabs>
+
+### Bucketed Table
+
+For `bucketed_table` from [Bucketed 
append](./bucketed#create-a-bucketed-table), disable append ordering in the 
**same**
+statement that enables clustering. Keep deletion vectors disabled. This 
explicitly opts out of
+[ordered streaming reads](./bucketed#bucketed-streaming).
+
+<Tabs groupId="engine">
+<TabItem value="flink" label="Flink">
+
+```sql
+ALTER TABLE bucketed_table SET (
     'bucket-append-ordered' = 'false',
     'clustering.incremental' = 'true',
-    'clustering.columns' = 'event_time,user_id',
-    'clustering.strategy' = 'zorder'
+    'clustering.columns' = 'product_id,price'
 );
 ```
 
-Once Incremental Clustering for a table is enabled, you can run Incremental 
Clustering in batch mode periodically 
-to continuously optimizes data layout of the table and deliver better query 
performance.
+</TabItem>
+<TabItem value="spark" label="Spark">
 
-**Note**: Since common compaction also rewrites files, it may disrupt the 
ordered data layout built by Incremental Clustering. 
-Therefore, when Incremental Clustering is enabled, the table no longer 
supports write-time compaction or dedicated compaction; 
-clustering and small-file merging must be performed exclusively via 
Incremental Clustering runs.
+```sql
+ALTER TABLE bucketed_table SET TBLPROPERTIES (
+    'bucket-append-ordered' = 'false',
+    'clustering.incremental' = 'true',
+    'clustering.columns' = 'product_id,price'
+);
+```
 
-## Run Incremental Clustering
-:::info
+</TabItem>
+</Tabs>
+
+### Clustering Options
 
-The following examples submit batch compact jobs. They are the recommended way 
to run Incremental Clustering explicitly.
+| Option | Default | How to use it |
+| --- | --- | --- |
+| `clustering.incremental` | `false` | Set to `true` to enable incremental 
clustering. |
+| `clustering.columns` | Not set | Comma-separated columns, such as 
`product_id,price`. Prefer frequently filtered data columns over partition 
columns. |
+| `clustering.strategy` | `auto` | `order`, `zorder`, or `hilbert`. Automatic 
selection uses `order` for one column, `zorder` for two to four, and `hilbert` 
for five or more. |
+| `clustering.incremental.mode` | `global-sort` | Sort execution mode for 
unaware-bucket batch clustering; see below. |
 
-:::
+## Choose a Sort Mode
 
-To run a Incremental Clustering job, follow these instructions. 
+For unaware-bucket tables, the mode controls how the **selected files in each 
partition** are sorted:
 
-You don't need to specify any clustering-related parameters when running 
Incremental Clustering,
-these options are already defined as table options. If you need to change 
clustering settings, please update the corresponding table options.
+| Mode | Execution | Tradeoff |
+| --- | --- | --- |
+| `global-sort` | Range-shuffles rows across tasks, then sorts within tasks 
using the configured clustering strategy. | Coordinates the layout across the 
selected output files, at the cost of a network shuffle. |
+| `local-sort` | Sorts rows independently within each compaction task, without 
the global range shuffle. | Less shuffle work; ranges in files produced by 
different tasks can overlap. Useful when ordering within files is sufficient, 
such as for Parquet lookup optimizations. |
 
-<Tabs groupId="incremental-clustering">
+Here, “global” refers to the selected clustering work within a partition. It 
does not imply that all existing files or
+all table partitions become globally ordered after an incremental run.
+
+## Run Incremental Clustering
 
-<TabItem value="spark-sql" label="Spark SQL">
+Run explicit compact jobs in batch mode. Table options supply the clustering 
columns and strategy. The examples below
+show routine incremental selection (`minor`) and full clustering (`full`) of 
the selected table or partition scope.
 
-Run the following sql:
+<Tabs groupId="engine">
+<TabItem value="spark" label="Spark SQL">
 
 ```sql
---set the write parallelism, if too big, may generate a large number of small 
files.
-SET spark.sql.shuffle.partitions=10;
+-- Choose parallelism for the workload; too many tasks can produce small files.
+SET spark.sql.shuffle.partitions = 10;
 
--- run incremental clustering
-CALL sys.compact(table => 'T')
+-- Select files using the incremental compaction strategy.
+CALL sys.compact(table => 'my_table', compact_strategy => 'minor');
 
--- run incremental clustering with full mode, this will recluster all data
-CALL sys.compact(table => 'T', compact_strategy => 'full')
+-- Alternatively, request full clustering; already-clustered runs can be 
skipped.
+CALL sys.compact(table => 'my_table', compact_strategy => 'full');
+```
 
--- run incremental clustering with global-sort mode (default)
--- performs a global range shuffle across tasks, output files are globally 
ordered
-CALL sys.compact(table => 'T', options => 
'clustering.incremental.mode=global-sort')
+To limit the work to one partition of `my_table`, use `partitions`:

Review Comment:
   **[P2] Qualify partition limits when historical auto-clustering is enabled**
   
   When `clustering.history-partition.idle-to-full-sort` is set, as in the 
later auto-clustering example, this call can also rewrite partitions 
**outside** `dt=2026-09-10`. `IncrementalClusterManager.createCompactUnits()` 
adds the output of `HistoryPartitionCluster.pickForHistoryPartitions()`, whose 
partition filter explicitly excludes the requested partitions and selects 
additional historical ones. Thus neither `partitions` nor `where` is a hard 
bound on the entire job in this configuration; following the new scope guidance 
can unexpectedly trigger large historical rewrites. Please qualify the 
Spark/Flink scoped examples and explain this expansion in the auto-clustering 
section. That section should also state that the auto-full path requires an 
explicit partition predicate: `HistoryPartitionCluster.create()` returns null 
without one, so setting the idle duration alone on the unscoped `minor` example 
does not activate it. 
`HistoryPartitionClusterTest.testHistoryPartitionAutoClusterin
 g` covers both the extra-partition selection and the no-predicate case.
   



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to