Gabriel39 commented on issue #66497:
URL: https://github.com/apache/doris/issues/66497#issuecomment-6053021847

   After reconsidering the implementation against the execution model of 
Iceberg `rewrite_data_files`, please simplify the design and implementation for 
this issue. This changes the previously accepted v5.1 direction: a durable 
asynchronous Lance index job framework should no longer be a requirement for 
the initial implementation.
   
   ### 1. Use synchronous SQL with bounded worker execution
   
   The SQL statement should validate the request, dispatch resource-isolated 
work to BE, wait for completion, refresh metadata after a confirmed commit, and 
return the result. Internal parallel execution does not require an asynchronous 
SQL/job interface.
   
   Keep heavy Lance work outside FE and the BE main process, with hard worker 
resource limits, bounded concurrency/queueing, deadlines, and reliable child 
cleanup. These protections are independent of durable job management and should 
remain.
   
   ### 2. Revert the job infrastructure already introduced
   
   Please prepare an explicit cleanup/revert on `branch-4.1`, rather than 
leaving the old framework disabled or maintaining two execution paths:
   
   - Revert the durable Lance index job state machine, manager, persistent 
same-name fences, unresolved-job quotas, and job replay infrastructure 
introduced by #67235.
   - Remove the job-specific parts of #67630: durable job creation and JobId 
results, `SHOW LANCE INDEX JOB(S)`, and catalog guards tied to unresolved jobs. 
Preserve and adapt the useful privilege, schema, parameter, snapshot, and 
IF-condition validation for synchronous execution.
   - Remove the associated unused configuration, serialization/RPC fields, 
tests, and documentation as appropriate. Review journal/image compatibility 
before removing persisted readers or operation codes; do not make existing 
metadata unreadable or reuse old codes. Any necessary compatibility handling 
should be minimal and should not preserve an active job framework.
   - Stop pursuing #67978 and #67754 in their current form. Rework 
#68668/#68669 around the synchronous worker lifecycle and its failure 
semantics, retaining useful isolation and query-consumption tests.
   
   Keep the read-only `SHOW INDEX` and physical-index inspection functionality, 
and the reusable target-aware DDL validation. Please update the issue design 
and PR roadmap to reflect this scope change.
   
   ### 3. State the simpler failure contract explicitly
   
   The initial version does not promise persistent job history, FE failover 
recovery, resumable builds, or durable same-name serialization across failures.
   
   A timeout, disconnect, worker loss, or lost response after dispatch must not 
be presented as proof that nothing committed. Return an explicit 
indeterminate-outcome error when applicable, do not automatically retry a 
potentially committed mutation, and let the user inspect authoritative index 
metadata. Observing a same-name index does not prove which request created it.
   
   Cancellation cannot undo an existing commit. Worker capacity must remain 
accounted for until termination is confirmed. A confirmed commit followed by 
metadata-refresh failure must be reported distinctly from a failed build.
   
   ### 4. Keep a path to distributed construction
   
   Synchronous SQL does not preclude distributed index construction. The 
desired extension is:
   
   `pin snapshot -> partition fragments -> parallel BE builds of uncommitted 
segments -> validate/collect outputs -> one coordinated final commit -> refresh 
-> return`
   
   Please keep build and commit responsibilities separable. Verify the APIs 
available in the pinned Lance bindings and add the necessary adapter support 
before implementing this distributed path; do not run several complete, 
independently committing `create_index` calls against the same logical index 
name.
   
   The initial implementation may use one worker. Multi-BE construction can 
follow without introducing durable jobs; persistent recovery should be a 
separate, justified follow-up requirement rather than a prerequisite for index 
lifecycle support.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to