hudi-agent commented on code in PR #19799: URL: https://github.com/apache/hudi/pull/19799#discussion_r3900947906
########## website/docs/cli.md: ########## @@ -785,6 +873,431 @@ table change-table-type COW ╚════════════════════════════════════════════════╧══════════════════════════════════════╧══════════════════════════════════════╝ ``` +### Changing the Meta Fields Mode + +`hoodie.meta.fields.mode` decides which of Hudi's meta columns are physically written into base files. It is a +storage-layout decision baked in at write time, so a write can never change it on an existing table. +`table set-meta-fields-mode` is the sanctioned way to change it. + +```java +hudi:trips->table set-meta-fields-mode --target-mode COMMIT_TIME_ONLY +``` + +`--target-mode` accepts `ALL`, `NONE`, `COMMIT_TIME_ONLY`, `FILE_NAME_ONLY` or `COMMIT_TIME_AND_FILE_NAME`. The value +is resolved case-insensitively and trimmed, so `commit_time_only` is accepted too. Setting the mode the table is +already in is a no-op and reports as much. + +On a table that already has commits, two guards apply, because this command changes the table property without +rewriting a single existing file: + +- **Widening is refused outright**, and `--force` does not override it. Widening means the target mode populates a + meta column the current mode does not. Since earlier files are not rewritten, the table would advertise a column + that is null for every row written so far, and incremental queries and file-name lookups silently skip exactly + those rows. To widen, recreate the table. The CLI uses the same predicate as the write path + (`BaseHoodieWriteClient#validateAgainstTableProperties`), so the two cannot disagree about which transitions are + legal. +- **Narrowing needs `--force`** (default `false`). It leaves mixed-mode files: old commits keep the old layout, new + commits use the new one, and incremental and file-pruning semantics differ between the two sets. Passing `--force` + logs a warning recording the transition and the commit count. + +Neither guard applies to a table with no commits, where the mode can be set freely. + +### Inspecting the Timeline + +`commits show` lists completed commits. The timeline commands show every instant regardless of action and state, which +is what you want when diagnosing a stuck table: a compaction sitting in `REQUESTED`, or a rollback that never +completed, never appears in `commits show`. + +```java +hudi:trips->timeline show active --limit 10 +hudi:trips->timeline show incomplete +``` + +Both print `Instant`, `Action`, `State`, and the `Requested` / `Inflight` / `Completed` file modification times. +`timeline show incomplete` restricts the listing to instants that are not yet completed. + +| Option | Default | Applies to | Description | +| --- | --- | --- | --- | +| `--limit` | `10` | both | Number of rows to display. | +| `--sortBy` | unset | both | Field to sort by. | +| `--desc` | `false` | both | Reverse the ordering. | +| `--headeronly` | `false` | both | Print the header only. | +| `--show-rollback-info` | `false` | both | For rollback instants, also show the instant being rolled back. | +| `--show-time-seconds` | `false` | both | Include seconds in the instant file modification times. | +| `--with-metadata-table` | `false` | `timeline show active` only | Show the metadata table timeline alongside the data table, adding `MT Action`, `MT State` and the three matching MT time columns. | + +The metadata table has its own timeline, and the two can disagree when a metadata commit fails. To read it directly: + +```java +hudi:trips->metadata timeline show active --limit 10 +hudi:trips->metadata timeline show incomplete +``` + +These accept `--limit`, `--sortBy`, `--desc`, `--headeronly` and `--show-time-seconds`. They have no +`--show-rollback-info` and no `--with-metadata-table`, since they are already scoped to the metadata table. + +### Diffing a File or Partition + +`diff file` and `diff partition` replay the timeline and show every commit that touched a given file group or +partition, which is the quickest way to answer "what has been writing to this file". Both report the standard commit +columns plus the write statistics for the matching entries only. + +```java +hudi:trips->diff file --fileId 5f8a1e0b-1b4b-4a3f-9b1a-2c7d6e5f4a3b-0 --limit 10 +hudi:trips->diff partition --partitionPath 2026/08/26 --includeArchivedTimeline true +``` + +`diff file` takes `--fileId` and `diff partition` takes `--partitionPath` as a partition path relative to the table +base path. `diff partition` is only meaningful on a partitioned table. Both then share these options: + +| Option | Default | Description | +| --- | --- | --- | +| `--includeArchivedTimeline` | `false` | Also scan archived instants, not just the active timeline. | +| `--startTs` | unset, meaning now minus 10 days | Start of the instant range. Only applied when `--includeArchivedTimeline` is `true`. | +| `--endTs` | unset, meaning now minus 1 day | End of the instant range. Only applied when `--includeArchivedTimeline` is `true`. | +| `--limit` | `-1`, meaning no limit | Number of rows to display. | +| `--sortBy` | unset | Field to sort by. | +| `--desc` | `false` | Reverse the ordering. | +| `--headeronly` | `false` | Print the header only. | + +Note the interaction between the three range options, which is easy to get wrong. `--startTs` and `--endTs` are used +to select archived instants only. With the default `--includeArchivedTimeline false` the whole active timeline is +scanned and both bounds are ignored, so passing a narrow range does not restrict the output. Set +`--includeArchivedTimeline true` for the bounds to take effect, and note that the archived range is then merged with +the full active timeline rather than replacing it. + +### Repairing a Table + +`repair show empty commit metadata` scans completed instants on the active timeline and reports the ones whose +metadata file is empty, which is what a commit interrupted between file creation and metadata write leaves behind. + +```java +hudi:trips->repair show empty commit metadata +``` + +Note that this command writes its findings to the CLI log at `WARN` level rather than returning a table, so with the +default logging configuration you will see the `Empty Commit: ...` lines in the console log rather than in a rendered +result. It only reports; it does not modify the timeline. + +`rename partition` rewrites the data under one partition value into another, as a Spark job, and deletes the old +partition on success. + +```java +hudi:trips->set --conf SPARK_HOME=<SPARK_HOME> +hudi:trips->rename partition --oldPartition 2026/08/26 --newPartition 2026-08-26 --sparkMaster local[2] +``` + +`repair deprecated partition` is the special case of that rename for tables written before Hudi settled on its +placeholder for the null partition value: it rewrites data from the deprecated `default` partition into +`__HIVE_DEFAULT_PARTITION__`. + +```java +hudi:trips->repair deprecated partition --sparkMaster local[2] +``` + +Both take `--sparkProperties` (a Spark properties file path, empty by default), `--sparkMaster` (empty by default) and +`--sparkMemory` (`4G` by default). Both read the old partition, rewrite those records under the new partition value, +and then issue a `delete_partition` write against the old one, so the change goes through the timeline rather than +behind it. Both are a no-op when the old partition holds no records. + +They differ in one respect worth knowing: `rename partition` additionally removes the old partition directory from +storage after the delete write, logging a warning if that removal fails, whereas `repair deprecated partition` leaves +the emptied `default` directory in place. Either way these rewrite data, so take a savepoint first if the table +matters. + +### Auditing Storage Locks + +When a table uses a storage-based lock provider, the lock provider can record every lock transition to a set of JSONL +files, so that a suspected concurrency violation can be reconstructed after the fact. The audit is off by default and +is controlled by a config file next to the locks themselves, at +`<basePath>/.hoodie/.locks/audit_enabled.json`; the audit records land in `<basePath>/.hoodie/.locks/audit/`. + +```java +hudi:trips->locks audit enable +Lock audit enabled successfully. +Audit config written to: /user/hive/warehouse/table1/.hoodie/.locks/audit_enabled.json +Audit files will be stored at: /user/hive/warehouse/table1/.hoodie/.locks/audit +``` + +`locks audit status` reports whether auditing is on, and where both the config and the records live. A table that has +never had auditing enabled reports `DISABLED` with the config file marked `(not found)`. + +```java +hudi:trips->locks audit status +Lock Audit Status: ENABLED +Table: /user/hive/warehouse/table1 +Config file: /user/hive/warehouse/table1/.hoodie/.locks/audit_enabled.json +Audit files location: /user/hive/warehouse/table1/.hoodie/.locks/audit +``` + +`locks audit validate` is the reason to collect the records. It parses every `.jsonl` file in the audit folder into +transaction windows and checks them against each other. Overlapping windows are reported as errors, since two writers +holding the lock at once is exactly the violation the lock provider exists to prevent. A transaction that never +released its lock is reported as a warning, which usually means a driver OOM or a non-graceful shutdown rather than a +correctness problem. The verdict is `PASSED`, `WARNING` when only warnings were found, or `FAILED` when any error was. + +```java +hudi:trips->locks audit validate +Validation Result: PASSED +Audit Files: 12 total, 12 parsed successfully, 0 failed to parse +Transactions Validated: 12 +Issues Found: 0 +Details: All audit lock transactions validated successfully +``` + +With no audit folder or no audit files the command reports `PASSED` with zero transactions validated, so a `PASSED` +verdict on its own does not prove that auditing was ever on. Check `locks audit status` first. + +`locks audit cleanup` prunes old records. `--ageDays` defaults to `7` and `--dryRun` defaults to `false`, so run it +with `--dryRun true` first to see what it would remove. + +```java +hudi:trips->locks audit cleanup --dryRun true --ageDays 30 +``` + +`locks audit disable` turns auditing off. It keeps the existing records by default; pass `--keepAuditFiles false` to +delete them at the same time, which internally runs the same cleanup with no age threshold. + +```java +hudi:trips->locks audit disable --keepAuditFiles true +``` + +All five commands require a table to be connected, and report `No Hudi table loaded. Please connect to a table first.` +otherwise. + +## Command reference + +Every command `hudi-cli` exposes, grouped by area, with its options and their defaults. An option marked +`(required)` has no default and must be supplied; a value in backticks after an option is its default. +The sections above cover the commonly used ones in more depth. + +Some entries are aliases of the same command rather than distinct ones. `refresh`, `metadata refresh`, +`commits refresh`, `cleans refresh` and `savepoints refresh` are five names for one method that reloads the table +metadata, and `temp query` / `temp_query`, `temp delete` / `temp_delete` and `temps show` / `temps_show` are +underscore and space spellings of the same three commands. + +### Table and session + +- **`cleans refresh`** Refresh table metadata. +- **`commits refresh`** Refresh table metadata. +- **`connect`** Connect to a hoodie table. + <br />Options: `--path` (required), `--eventuallyConsistent` (`false`), `--initialCheckIntervalMs` (`2000`), `--maxWaitIntervalMs` (`300000`), `--maxCheckIntervalMs` (`7`), `--timeGeneratorType` (`WAIT_TO_ADJUST_SKEW`), `--maxExpectedClockSkewMs` (`200`), `--useDefaultLockProvider` (`false`) +- **`create`** Create a hoodie table if not present. + <br />Options: `--path` (required), `--tableName` (required), `--tableType` (`COPY_ON_WRITE`), `--archiveLogFolder`, `--tableVersion`, `--payloadClass` (`org.apache.hudi.common.model.HoodieAvroPayload`) +- **`desc`** Describe Hoodie Table properties. +- **`fetch table schema`** Fetches latest table schema. + <br />Options: `--outputFilePath` +- **`kerberos kdestroy`** Destroy Kerberos authentication. + <br />Options: `--krb5conf` (`/etc/krb5.conf`) +- **`kerberos kinit`** Perform Kerberos authentication. + <br />Options: `--krb5conf` (`/etc/krb5.conf`), `--principal` (required), `--keytab` (required) +- **`metadata refresh`** Refresh table metadata. +- **`refresh`** Refresh table metadata. +- **`savepoints refresh`** Refresh table metadata. +- **`set`** Set spark launcher env to cli. + <br />Options: `--conf` (required) +- **`show env`** Show spark launcher env by key. + <br />Options: `--key` (required) +- **`show envs all`** Show spark launcher envs. +- **`table change-table-type`** Change hudi table type to target type: COW or MOR. + <br />Options: `--target-type` (required), `--enable-compaction` (`true`), `--parallelism` (`3`), `--sparkMaster` (`local`), `--sparkMemory` (`4G`), `--retry` (`1`), `--propsFilePath`, `--hoodieConfigs` +- **`table delete-configs`** Delete the supplied table configs from the table. + <br />Options: `--comma-separated-configs` (required) +- **`table recover-configs`** Recover table configs, from update/delete that failed midway. +- **`table set-meta-fields-mode`** Set hoodie.meta.fields.mode on an existing table. This is the sanctioned way to change. + <br />Options: `--target-mode` (required), `--force` (`false`) +- **`table update-configs`** Update the table configs with configs with provided file. + <br />Options: `--props-file` (required) +- **`utils loadClass`** Load a class. + <br />Options: `--class` (required) + +### Commits and the timeline + +- **`commit show_write_stats`** Show write stats of a commit. + <br />Options: `--createView`, `--commit` (required), `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`), `--includeArchivedTimeline` (`false`) +- **`commit showfiles`** Show file level details of a commit. + <br />Options: `--createView`, `--commit` (required), `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`), `--includeArchivedTimeline` (`false`) +- **`commit showpartitions`** Show partition level details of a commit. + <br />Options: `--createView`, `--commit` (required), `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`), `--includeArchivedTimeline` (`false`) +- **`commits compare`** Compare commits with another Hoodie table. + <br />Options: `--path` (required) +- **`commits show`** Show the commits. + <br />Options: `--includeExtraMetadata` (`false`), `--createView`, `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`), `--partition`, `--includeArchivedTimeline` (`false`) +- **`commits show_inflights`** Show inflight instants that are left longer than a certain duration. + <br />Options: `--lookbackInMins` (`0`) +- **`commits showarchived`** Show the archived commits. + <br />Options: `--includeExtraMetadata` (`false`), `--createView`, `--startTs`, `--endTs`, `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`), `--partition` +- **`commits sync`** Sync commits with another Hoodie table. + <br />Options: `--path` (required) +- **`diff file`** Check how file differs across range of commits. + <br />Options: `--fileId` (required), `--startTs`, `--endTs`, `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`), `--includeArchivedTimeline` (`false`) +- **`diff partition`** Check how file differs across range of commits. It is meant to be used only for partitioned tables. + <br />Options: `--partitionPath` (required), `--startTs`, `--endTs`, `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`), `--includeArchivedTimeline` (`false`) +- **`metadata timeline show active`** List all instants in active timeline of metadata table. + <br />Options: `--limit` (`10`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`), `--show-time-seconds` (`false`) +- **`metadata timeline show incomplete`** List all incomplete instants in active timeline of metadata table. + <br />Options: `--limit` (`10`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`), `--show-time-seconds` (`false`) +- **`show archived commit stats`** Read commits from archived files and show file group details. + <br />Options: `--archiveFolderPattern`, `--limit` (`10`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`) +- **`show archived commits`** Read commits from archived files and show details. + <br />Options: `--skipMetadata` (`true`), `--limit` (`10`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`) +- **`timeline show active`** List all instants in active timeline. + <br />Options: `--limit` (`10`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`), `--with-metadata-table` (`false`), `--show-rollback-info` (`false`), `--show-time-seconds` (`false`) +- **`timeline show incomplete`** List all incomplete instants in active timeline. + <br />Options: `--limit` (`10`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`), `--show-rollback-info` (`false`), `--show-time-seconds` (`false`) +- **`trigger archival`** Trigger archival. + <br />Options: `--minCommits` (`20`), `--maxCommits` (`30`), `--commitsRetainedByCleaner` (`10`), `--enableMetadata` (`true`), `--sparkMemory` (`1G`), `--sparkMaster` (`local`) + +### Files, stats and log files + +- **`show fsview all`** Show entire file-system view. + <br />Options: `--pathRegex` (`*`), `--baseFileOnly` (`false`), `--maxInstant`, `--includeMax` (`false`), `--includeInflight` (`false`), `--excludeCompaction` (`false`), `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`) +- **`show fsview latest`** Show latest file-system view. + <br />Options: `--partitionPath`, `--baseFileOnly` (`false`), `--maxInstant`, `--merge` (`true`), `--includeMax` (`false`), `--includeInflight` (`false`), `--excludeCompaction` (`false`), `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`) +- **`show logfile metadata`** Read commit metadata from log files. + <br />Options: `--logFilePathPattern` (required), `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`) +- **`show logfile records`** Read records from log files. + <br />Options: `--limit` (`10`), `--logFilePathPattern` (required), `--mergeRecords` (`false`) +- **`stats filesizes`** File Sizes. Display summary stats on sizes of files. + <br />Options: `--partitionPath` (`*/*/*`), `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`) +- **`stats wa`** Write Amplification. Ratio of how many records were upserted to how many. + <br />Options: `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`) + +### Table services + +- **`clean showpartitions`** Show partition level details of a clean. + <br />Options: `--clean` (required), `--limit` (`-1`), `--sortBy`, `--desc` (`false`), `--headeronly` (`false`) +- **`cleans run`** Run clean. + <br />Options: `--sparkMemory` (`4G`), `--propsFilePath`, `--hoodieConfigs`, `--sparkMaster` +- **`cleans show`** Show the cleans. + <br />Options: `--limit` (`-1`), `--sortBy`, `--startTs`, `--endTs`, `--includeArchivedTimeline` (`false`), `--desc` (`false`), `--headeronly` (`false`) +- **`clustering run`** Run Clustering. + <br />Options: `--sparkMaster` (`SparkUtil.DEFAULT_SPARK_MASTER`), `--sparkMemory` (`4g`), `--parallelism` (`1`), `--retry` (`1`), `--clusteringInstant`, `--propsFilePath`, `--hoodieConfigs` Review Comment: 🤖 This entry lists the `--sparkMaster` default as the raw Java constant `SparkUtil.DEFAULT_SPARK_MASTER` rather than its value. That constant resolves to `yarn` (`SparkUtil.DEFAULT_SPARK_MASTER = "yarn"`), so it would help to substitute the concrete value here — the way the other Spark commands in this reference already do (e.g. compaction/marker show `local`). The same unresolved constant appears on the `clustering schedule` / `clustering scheduleAndExecute` entries (lines 1177, 1179) and on `metadata create` / `metadata init` / `metadata list-partitions` (lines 1246, 1252, 1256). It might also be worth a one-line note that this default is `yarn`, since that differs from the `local` default used by the compaction and other Spark commands and could surprise a reader running these on a non-YARN setup. <sub><i>⚠️ AI-generated; verify before applying. React 👍/👎 to flag quality.</i></sub> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
