764276020 opened a new issue, #68538: URL: https://github.com/apache/doris/issues/68538
### Environment - Doris **3.1.4**, **cloud / compute-storage-separation** mode (FDB-backed MetaService) - Large log table: 54+ TB, 184 daily partitions, ~23-39 buckets per partition (5754 tablets) - Operation: `ALTER TABLE coremail.<t> ADD INDEX ... USING INVERTED ...` (heavy schema-change path, `SchemaChangeJobV2` / shadow index) ### Symptom The schema change fails with a MetaService transaction-size error and the whole job is then CANCELLED: ``` schema change tasks failed, error reason: task type: ALTER, status_code: INVALID_ARGUMENT, status_message: [(doris-dwcloud-be-w4sata-20...)[INVALID_ARGUMENT]failed to commit tablet job: failed to commit job kv, err=Transaction exceeds byte limit], backendId: Backend [id=1743164044730, host=doris-dwcloud-be-w4sata-20...] ``` Then `SHOW ALTER TABLE COLUMN` shows the job as `CANCELLED` with that message. ### Root cause (code) `CloudSchemaChangeJob`'s COMMIT is executed by MetaService `finish_tablet_job` and everything is packed into **one FDB transaction**: - `cloud/src/meta-service/meta_service_job.cpp:1787` – "move rowsets [2-alter_version] to recycle" (one `RecycleRowsetPB` with a full `RowsetMetaCloudPB` per rowset) - `cloud/src/meta-service/meta_service_job.cpp:1963` – `for (size_t i = 0; i < schema_change.txn_ids().size(); ++i)` converts **every** tmp rowset to a formal rowset in the same txn - commit failure surfaces as `failed to commit job kv` (`:2204` in master, `:1641` in 3.1.4) - `cloud/src/meta-store/txn_kv_error.h:55` – `Transaction exceeds byte limit` (FDB 10 MB hard limit) Transaction size therefore scales with the number of rowsets of the tablet. Schema change writes `tmp rowset -> formal rowset` **1:1**, so a tablet with many small rowsets (frequent small loads + compaction lag) deterministically blows the 10 MB limit. There is **no batching, no `approximate_bytes` pre-check and no resumable split** in this path (searched for `approximate_bytes` / `max_txn_commit_byte` / batch in that function: none). For contrast, the **load** path already handles exactly this problem: - `cloud/src/common/config.h:356` `enable_cloud_txn_lazy_commit = true` - `cloud/src/common/config.h:358` `txn_lazy_commit_rowsets_thresold = 1000` - `cloud/src/common/config.h:362` `txn_lazy_max_rowsets_per_batch = 1000` - `cloud/src/meta-service/txn_lazy_committer.cpp` commits in batches of <= 1000 rowsets So "one txn must not hold too many rowsets" is known to the project; the schema-change commit path just was not covered. ### Impact - The failing tablet makes the job **permanently un-runnable** (deterministic, not transient): re-running the ALTER reproduces it. - No retry at the FE level: `AlterJobV2.getRetryTimes()` (`fe/fe-core/.../alter/AlterJobV2.java:267-275`) only retries `DELETE_BITMAP_LOCK_ERROR` / `NETWORK_ERROR`, and replica_num = 1 in cloud, so the first failed tablet task cancels the whole job (hours of conversion work lost, plus the rewritten objects in object storage). - On our 54 TB / 86 TB tables this makes ADD INDEX practically impossible without pre-compaction. ### Suggested fix Batch the COMMIT like the load path does, e.g. convert tmp->formal rowset and write the recycle records in chunks of N rowsets per transaction, making the operation resumable/idempotent (so the BE can continue instead of failing the task); or at least split when the txn size approaches the limit. ### Workaround we use Run `ADMIN COMPACT TABLE <db>.<tbl> PARTITION (<p>) WHERE type='BASE';` for the affected partitions first, to reduce the rowset count per tablet, then re-run the ALTER (one table at a time). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
