764276020 opened a new issue, #68537:
URL: https://github.com/apache/doris/issues/68537
### Environment
- Doris **3.1.4**, **cloud / compute-storage-separation** mode
- `ALTER TABLE coremail.udtrans_dellog ADD INDEX ... USING INVERTED ...`
(heavy SC path; job id `1789639373133`, shadow index id `1789639373134`)
- Minute-level log table, 369 daily partitions (3838 shadow tablets)
### Symptom
A long-running schema change was cancelled although nobody touched it, with:
```
[INTERNAL_ERROR]failed to commit rowset: failed to get tablet_idx,
err=KeyNotFound
tablet_id=1789639379947 key=01106d657461...tablet_index...a0aed1cbeb
```
and `SHOW ALTER TABLE COLUMN` shows the job CANCELLED with that message.
### Timeline (from MetaService / recycler / FE / BE logs)
- `11:58` job created; FE calls `prepareMaterializedIndex()` so MetaService
writes `RecycleIndexPB{state=PREPARED, expiration=jobCreateTime+timeout}` for
the shadow index (`CloudSchemaChangeJobV2.java:191,198-199`;
`meta_service_partition.cpp:263,326`). This PREPARED record is removed by
`commit_index()` on success (`meta_service_partition.cpp:395,532` "remove
recycle index") — i.e. **PREPARED means "index creation in progress"**.
- `13:57` BE starts converting shadow tablet `1789639379947`; writes
`.dat`/`.idx` to S3 normally for ~22 minutes.
- **`14:17:46` recycler: `begin to recycle index, instance_id=1638888
table_id=207246110 index_id=1789639373134 state=PREPARED`** → state set to
RECYCLING → all tablets of the shadow index are deleted (tablet meta +
`tablet_idx` + rowsets + objects).
- `14:19:29` BE: `failed to commit rowset: failed to get tablet_idx,
err=KeyNotFound tablet_id=1789639379947`
- `14:19:30` FE: `runRunningJob():619 schema change task failed,
failedTimes: 1, maxFailedTimes: 0` → `cancelImpl():841` → **job CANCELLED**;
afterwards hundreds of `failed to commit rowset: stale perpare rowset request`
from other BEs (job already gone).
### Root cause (code)
`calculate_index_expired_time()` returns 0 **unconditionally** when
`force_immediate_recycle` is on, without looking at the record state:
```cpp
// cloud/src/recycler/recycler.cpp:1903-1921 (master; same code at 3.1.4
:959-971)
int64_t calculate_index_expired_time(...) {
if (config::force_immediate_recycle) {
return 0L; // <- PREPARED records become "expired"
too
}
...
if (index_meta_pb.state() == RecycleIndexPB::DROPPED) { ... }
}
```
and the index recycle path then deletes the **whole index**:
- `recycler.cpp:2976` skips records that are not expired; `:2987` logs
`begin to recycle index ... state=...`
- `:3602` `tablet_belongs = "index"` → `recycle_tablets(table_id, index_id,
...)` deletes every tablet of the index (all partitions), including the
`tablet_idx` KV that the in-flight schema change needs at commit time.
Note the asymmetry: **other object types DO special-case PREPARED** —
partitions: `recycler.cpp:3259` "Partitions with PREPARED state MUST have no
data"; restore jobs: `:6137` "PREPARED or COMMITTED, change state to DROPPED
and return". The index path has no such handling.
We do understand `force_immediate_recycle` is documented as "**just for
TEST**" (`cloud/src/common/config.h:148-150`). We still report it, because:
1. it is a normal (non-hidden) mutable config that operators enable when
they need to reclaim object-storage space quickly;
2. when it is on, **in-progress DDL (index creation) is destroyed** and no
API/error message tells the user that this is the reason (the surfaced error is
`failed to get tablet_idx ... KeyNotFound`, which points at the cancelled job
instead of the recycler);
3. we verified the same code path exists both in 3.1.4 and in current
master, so it is not fixed.
### Suggested fix
- Exempt `state == PREPARED` from `force_immediate_recycle` in
`calculate_index_expired_time()` (or force `expiration + retention_seconds` for
PREPARED), and/or
- before recycling an index, skip it if it is still referenced by a running
schema-change job (`job_tablet_key` / `TabletJobInfoPB` are available in
MetaService).
### Workaround
Keep `force_immediate_recycle = false` permanently; reclaim space with
`ADMIN CLEAN TRASH`, retention tuning, or by waiting for the normal retention
windows (compacted 3h / dropped index 3h / tmp rowset 72h).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]