rusackas commented on code in PR #44627:
URL: https://github.com/apache/superset/pull/44627#discussion_r4099940451
##########
superset/db_engine_specs/databricks.py:
##########
@@ -246,6 +248,17 @@ class DatabricksBaseEngineSpec(BaseEngineSpec):
identifier_quote_start: str = "`"
identifier_quote_end: str = "`"
+ # Databricks SQL documents native MEDIAN/STDDEV_SAMP/VAR_SAMP aggregate
+ # functions (docs.databricks.com/aws/en/sql/language-manual/functions/
+ # {median,stddev_samp,var_samp}), confirmed against that reference, not a
+ # live Databricks instance. All three are plain function calls, same
+ # spelling as Postgres/DuckDB, so no dialect-specific rewriting is needed.
+ _extended_aggregations: dict[str, Callable[[ColumnElement],
ColumnElement]] = {
+ "MEDIAN": sa.func.median,
+ "STDDEV_SAMP": sa.func.stddev_samp,
+ "VAR_SAMP": sa.func.var_samp,
+ }
Review Comment:
Fair, and worth being explicit about rather than just implying it in the
docstring: this ships doc-verified, not live-verified. Same basis Redshift's
own MEDIAN override already ships on in this file (see its comment above
`RedshiftEngineSpec._extended_aggregations`), so it's consistent with an
accepted pattern here, not a new risk class. Tracking live verification as a
follow-up rather than blocking on it.
##########
superset/db_engine_specs/databricks.py:
##########
@@ -986,6 +999,16 @@ class DatabricksHiveEngineSpec(HiveEngineSpec):
_time_grain_expressions = time_grain_expressions
+ # Interactive Clusters run Spark SQL, same as the primary Databricks
+ # connector above; same native MEDIAN/STDDEV_SAMP/VAR_SAMP functions
+ # apply here rather than the inherited (unimplemented) HiveEngineSpec/
+ # PrestoEngineSpec default.
+ _extended_aggregations: dict[str, Callable[[ColumnElement],
ColumnElement]] = {
+ "MEDIAN": sa.func.median,
+ "STDDEV_SAMP": sa.func.stddev_samp,
+ "VAR_SAMP": sa.func.var_samp,
+ }
Review Comment:
Same caveat as the primary connector above, and same answer: doc-verified
only, not against a live Interactive Cluster. Worth flagging that this spec
would otherwise silently fall back to the inherited, unimplemented Hive/Presto
default (a hard "not supported" error) rather than quietly misbehaving, so the
failure mode if the docs are ever wrong is a clear error at execution time, not
silent wrong data.
##########
superset/db_engine_specs/snowflake.py:
##########
@@ -142,10 +143,19 @@ class SnowflakeEngineSpec(PostgresBaseEngineSpec):
force_column_alias_quotes = True
max_column_name_length = 256
- # `PostgresBaseEngineSpec._extended_aggregations`
(MEDIAN/STDDEV_SAMP/VAR_SAMP)
- # is verified against real Postgres behavior, not Snowflake's; disable it
here
- # until someone confirms the same expressions against a live Snowflake
instance.
- _extended_aggregations: dict[str, Callable[[ColumnElement],
ColumnElement]] = {}
+ # Snowflake documents native MEDIAN/STDDEV_SAMP/VAR_SAMP aggregate
functions
+ #
(docs.snowflake.com/en/sql-reference/functions/{median,stddev_samp,var_samp}),
+ # confirmed against that reference, not a live Snowflake instance.
STDDEV_SAMP/
+ # VAR_SAMP are plain function calls, same spelling as inherited from
+ # `PostgresBaseEngineSpec`, so those are reused directly. MEDIAN is
overridden:
+ # Snowflake's own `MEDIAN(x)` is a plain function call, unlike Postgres's
+ # `percentile_cont(0.5) WITHIN GROUP (ORDER BY x)` (Postgres has no native
+ # MEDIAN), so there is no reason to compile the more complex inherited
form.
+ _extended_aggregations: dict[str, Callable[[ColumnElement],
ColumnElement]] = {
+ "MEDIAN": sa.func.median,
+ "STDDEV_SAMP":
PostgresBaseEngineSpec._extended_aggregations["STDDEV_SAMP"],
+ "VAR_SAMP": PostgresBaseEngineSpec._extended_aggregations["VAR_SAMP"],
+ }
Review Comment:
Same as the Databricks reply above: doc-verified against Snowflake's own
function reference, not a live instance. `SnowflakeEngineSpec` was actually
disabled outright before this PR for exactly this reason (see the diff), so
this is a deliberate, considered opt-in rather than an oversight, on the same
basis Redshift already ships on elsewhere in this file family.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]