AlankritVerma01 commented on issue #34384:
URL: https://github.com/apache/superset/issues/34384#issuecomment-5999373595

   Hey @rusackas,
   
   Thanks. I reviewed the existing estimator and Beto's earlier draft. I 
propose adding optional estimate warnings to SQL Lab's Run workflow.
   
   These would be advisory warnings. Users could edit, cancel or run anyway, 
with database permissions and native limits still applying.
   
   I'm considering a separate feature flag for automatic warnings, disabled by 
default. It would give operators control over rolling out this new behavior. 
However, existing users would already stay unaffected because each connection's 
warning setting would also start disabled.
   
   The warnings would depend on estimation being available. If the existing 
estimation flag is off, the warning settings would be inactive with an 
explanation. I'm not settled on whether the separate flag is needed, or whether 
the existing estimation flag plus connection-level opt-in is sufficient.
   
   The flow would be:
   
   - An administrator opens Settings > Database Connections > Edit connection > 
Advanced > SQL Lab.
   - Below "Enable query cost estimation", add "Enable pre-run warnings" and a 
threshold field.
   - When someone clicks Run, estimate the query and compare it against that 
connection's threshold.
   - Within the threshold, proceed normally. Above it, show the estimate, 
threshold and limitations, with options to edit, cancel or run anyway.
   - SQL Lab would show whether checking is enabled and what it checks.
   
   The main use case is an ad hoc query where someone realizes, "I didn't 
expect this query to process that much data." I'd start with deliberate SQL Lab 
Run, shortcuts and rerun for single read-only queries. Dashboard queries can 
also be expensive, but dashboards can refresh or run several queries without a 
single user decision point. I'd leave that interaction for later and consider 
expanding support after validating the SQL Lab workflow.
   
   I'd start implementation with BigQuery and test the shared design with 
PostgreSQL early. BigQuery would use processed bytes; PostgreSQL would use 
relative planner cost. Each engine would retain its own units and limitations, 
with thresholds chosen for the connection's workload.
   
   The build would cover:
   
   - Add the connection settings and save the warning rule in Database.extra.
   - Get the estimate and compare it against the threshold before rounding 
numbers for display. Keep the existing manual estimate button and API backward 
compatible.
   - Check the selected query with the settings it will run under. Changed 
inputs require a new check.
   - Wait for the estimate and any user decision before submitting execution. 
Cancel submits no execution request.
   - Test the warning, continuation, cancellation and failure paths. Measure 
added delay and whether warnings help users revise queries.
   
   If an enabled check cannot provide a usable estimate, I'd explain why and 
offer an explicit choice to run without one. Identified input, permission and 
native-limit errors would remain errors. Unexpected failures would show an 
error rather than silently execute. Unsupported connections would run normally 
with checking shown as unavailable.
   
   Does the SQL Lab-first scope make sense, and would you prefer a separate 
warning flag or connection-level opt-in under the existing estimation flag?


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to