Gabriel39 commented on issue #66497: URL: https://github.com/apache/doris/issues/66497#issuecomment-6053021847
After reconsidering the implementation against the execution model of Iceberg `rewrite_data_files`, please simplify the design and implementation for this issue. This changes the previously accepted v5.1 direction: a durable asynchronous Lance index job framework should no longer be a requirement for the initial implementation. ### 1. Use synchronous SQL with bounded worker execution The SQL statement should validate the request, dispatch resource-isolated work to BE, wait for completion, refresh metadata after a confirmed commit, and return the result. Internal parallel execution does not require an asynchronous SQL/job interface. Keep heavy Lance work outside FE and the BE main process, with hard worker resource limits, bounded concurrency/queueing, deadlines, and reliable child cleanup. These protections are independent of durable job management and should remain. ### 2. Revert the job infrastructure already introduced Please prepare an explicit cleanup/revert on `branch-4.1`, rather than leaving the old framework disabled or maintaining two execution paths: - Revert the durable Lance index job state machine, manager, persistent same-name fences, unresolved-job quotas, and job replay infrastructure introduced by #67235. - Remove the job-specific parts of #67630: durable job creation and JobId results, `SHOW LANCE INDEX JOB(S)`, and catalog guards tied to unresolved jobs. Preserve and adapt the useful privilege, schema, parameter, snapshot, and IF-condition validation for synchronous execution. - Remove the associated unused configuration, serialization/RPC fields, tests, and documentation as appropriate. Review journal/image compatibility before removing persisted readers or operation codes; do not make existing metadata unreadable or reuse old codes. Any necessary compatibility handling should be minimal and should not preserve an active job framework. - Stop pursuing #67978 and #67754 in their current form. Rework #68668/#68669 around the synchronous worker lifecycle and its failure semantics, retaining useful isolation and query-consumption tests. Keep the read-only `SHOW INDEX` and physical-index inspection functionality, and the reusable target-aware DDL validation. Please update the issue design and PR roadmap to reflect this scope change. ### 3. State the simpler failure contract explicitly The initial version does not promise persistent job history, FE failover recovery, resumable builds, or durable same-name serialization across failures. A timeout, disconnect, worker loss, or lost response after dispatch must not be presented as proof that nothing committed. Return an explicit indeterminate-outcome error when applicable, do not automatically retry a potentially committed mutation, and let the user inspect authoritative index metadata. Observing a same-name index does not prove which request created it. Cancellation cannot undo an existing commit. Worker capacity must remain accounted for until termination is confirmed. A confirmed commit followed by metadata-refresh failure must be reported distinctly from a failed build. ### 4. Keep a path to distributed construction Synchronous SQL does not preclude distributed index construction. The desired extension is: `pin snapshot -> partition fragments -> parallel BE builds of uncommitted segments -> validate/collect outputs -> one coordinated final commit -> refresh -> return` Please keep build and commit responsibilities separable. Verify the APIs available in the pinned Lance bindings and add the necessary adapter support before implementing this distributed path; do not run several complete, independently committing `create_index` calls against the same logical index name. The initial implementation may use one worker. Multi-BE construction can follow without introducing durable jobs; persistent recovery should be a separate, justified follow-up requirement rather than a prerequisite for index lifecycle support. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
