rusackas opened a new pull request, #44627: URL: https://github.com/apache/superset/pull/44627
### SUMMARY #42895 restored `MEDIAN`, `STDDEV_SAMP`, and `VAR_SAMP` as first-class metric aggregates (usable anywhere a metric aggregate is chosen — every chart type, SQL Lab, MCP), but only enabled them for Postgres, MySQL, DuckDB, and Redshift via each engine spec's `_extended_aggregations` opt-in dict. This PR adds Databricks and Snowflake to that list, as part of ongoing work to bring the Pivot Table's aggregation methods back to feature parity with pre-SIP-216 (#41184) behavior for customers on those engines (tracking doc: `docs/sip/pivot-table-aggregation-parity.md`). Both vendors document native `median()`/`stddev_samp()`/`var_samp()` aggregate functions with the correct *sample* (not population) semantics these names denote, and no dialect-specific SQL construct is needed — same simple pattern as DuckDB/Redshift, unlike Postgres's `MEDIAN`, which needs `percentile_cont(0.5) WITHIN GROUP (...)` since Postgres has no native `MEDIAN`. Confirmed against each vendor's own SQL function reference documentation, not yet against a live instance (same verification basis `SnowflakeEngineSpec`'s existing test already used for what it had before this PR). - **Databricks**: added to `DatabricksBaseEngineSpec`, which covers the Native, Python Connector, and ODBC connector variants (all three `DatabricksXxxEngineSpec` classes that build on it). Also added separately to `DatabricksHiveEngineSpec` (the "Interactive Cluster" connector), since that class inherits from `HiveEngineSpec`/`PrestoEngineSpec` instead and would otherwise silently fall back to the unimplemented default — Interactive Clusters run Spark SQL too, so the same functions apply. - **Snowflake**: `MEDIAN` is overridden to use Snowflake's own simpler native form rather than inheriting `PostgresBaseEngineSpec`'s `percentile_cont`/`WITHIN GROUP` workaround (Snowflake supports that form too, but doesn't need it). `STDDEV_SAMP`/`VAR_SAMP` are reused directly from the inherited Postgres dict since the spelling is identical on both engines. - Dropped `SnowflakeEngineSpec` from `test_extended_aggregations_unverified.py`'s negative "must reject" list now that it opts in explicitly. `DatabricksBaseEngineSpec`/`DatabricksHiveEngineSpec` were never members of that list — they aren't Postgres/MySQL-dialect-family specs, so there was never a risk of them silently inheriting the dict. ### TESTING INSTRUCTIONS ``` pytest tests/unit_tests/db_engine_specs/test_databricks.py tests/unit_tests/db_engine_specs/test_snowflake.py tests/unit_tests/db_engine_specs/test_extended_aggregations_unverified.py -v ``` New tests: `test_extended_aggregation_func_compiles_expected_sql` (parametrized over both Databricks spec classes × all 3 aggregates, asserting the exact compiled SQL), `test_databricks_hive_spec_shares_extended_aggregations`. `mypy`/`ruff`/`pylint` clean via `pre-commit`. ### ADDITIONAL INFORMATION - [ ] Has associated issue: - [ ] Required feature flags: - [ ] Changes UI - [ ] Includes DB Migration (follow approval process in [SIP-59](https://github.com/apache/superset/issues/13351)) - [ ] Migration is atomic, supports rollback & is backwards-compatible - [ ] Confirm DB migration upgrade and downgrade tested - [ ] Runtime estimates and downtime expectations provided - [ ] Introduces new feature or API - [ ] Removes existing feature or API 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
