sadpandajoe commented on code in PR #44816:
URL: https://github.com/apache/superset/pull/44816#discussion_r4152914591


##########
superset/models/helpers.py:
##########
@@ -4288,36 +4295,153 @@ def dttm_sql_literal(self, dttm: datetime, col: 
"TableColumn") -> str:
 
         return f"""'{dttm.strftime("%Y-%m-%d %H:%M:%S.%f")}'"""
 
-    def get_time_filter(  # pylint: disable=too-many-arguments  # noqa: C901
+    def _collect_partition_mirror_range(
         self,
-        time_col: "TableColumn",
+        mapping: Optional["PartitionMapping"],
+        sink: list[tuple[utils.FilterOperator, Any]],
+        column_name: str,
         start_dttm: Optional[sa.DateTime],
         end_dttm: Optional[sa.DateTime],
-        time_grain: Optional[str] = None,
-        label: Optional[str] = "__time",
-        template_processor: Optional[BaseTemplateProcessor] = None,
-    ) -> Optional[ColumnElement]:
-        col = (
-            time_col.get_timestamp_expression(
-                time_grain=time_grain,
-                label=label,
-                template_processor=template_processor,
-            )
-            if time_grain
-            else self.convert_tbl_column_to_sqla_col(
-                time_col, label=label, template_processor=template_processor
-            )
+        widen_bounds_by: Optional[relativedelta] = None,
+    ) -> None:
+        """
+        Record a time range for mirroring onto the partition column.
+
+        The bounds are adjusted first, by the same helper `get_time_filter` 
uses:
+        the mirrored predicate has to describe the same instants as the
+        timestamp predicate it stands in for, or the pruning is wrong by 
exactly
+        the dataset's timezone offset -- silently.
+
+        Either bound may be `None` for an open-ended range, in which case only
+        the bound that exists is mirrored.
+
+        `widen_bounds_by` is one grain bucket width, passed when the filter 
this
+        stands in for compares a *truncated* column. The real predicate then
+        keeps rows the raw bounds exclude, and widening both ends by one bucket
+        is the smallest range guaranteed to contain all of them -- see the 
"Time
+        grains" section of `superset.connectors.sqla.partition_mapping`. Note
+        the widened upper bound is `<=` rather than `<`: `ts < until + width`
+        only gives `T(ts) <= T(until + width)` for a transform that is 
monotonic
+        but not strictly so, such as `unix_timestamp` on a sub-second column.
+
+        A widened range can also be wide enough to fail `_bounds_are_ordered`
+        against a tighter filter on the same column. That fails open -- no
+        mirroring, no pruning, no dropped rows.
+        """
+        if mapping is None or column_name != mapping.mapped_column:
+            return
+        if not mapping.mirrors(utils.FilterOperator.TEMPORAL_RANGE):
+            return
+
+        # Resolve the column so the offset adjustment sees the same type
+        # `get_time_filter` does -- an offset the column's type cannot 
represent
+        # is dropped, and the mirrored bounds have to agree with the originals.
+        mapped_col = next(
+            (col for col in self.columns if col.column_name == column_name), 
None
         )
+        start_dttm, end_dttm = self.adjust_time_bounds(start_dttm, end_dttm, 
mapped_col)
+
+        # Adjust first, then widen. `adjust_time_bounds` moves the naive UI
+        # bounds into the frame the column is *stored* in, and the grain
+        # truncation in the real predicate happens in that same frame, so the
+        # bucket width has to be added there too. The two commute for the
+        # hour-offset branch but not for the `ZoneInfo` one, where widening
+        # across a DST boundary first lands an hour out.
+        upper_operator = utils.FilterOperator.LESS_THAN

Review Comment:
   A monotonic transform can collapse distinct timestamps: with `p = 
date_format(ts, 'yyyy-MM-dd')`, a 10:00 row satisfies `ts < '2026-02-01 12:00'` 
but fails the added `p < '2026-02-01'`. Can strict bounds become inclusive on 
the partition column, both here and for structured `<`/`>` filters?



##########
superset/connectors/sqla/partition_mapping.py:
##########
@@ -0,0 +1,977 @@
+# Licensed to the Apache Software Foundation (ASF) under one
+# or more contributor license agreements.  See the NOTICE file
+# distributed with this work for additional information
+# regarding copyright ownership.  The ASF licenses this file
+# to you under the Apache License, Version 2.0 (the
+# "License"); you may not use this file except in compliance
+# with the License.  You may obtain a copy of the License at
+#
+#   http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing,
+# software distributed under the License is distributed on an
+# "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+# KIND, either express or implied.  See the License for the
+# specific language governing permissions and limitations
+# under the License.
+"""
+Partition filter mapping.
+
+Datasets on Hadoop-family engines are commonly partitioned on a *technical*
+column -- an epoch integer, a lowercased region key -- that no analyst would
+filter on. Unless a query carries a predicate on that column the engine scans
+every partition.
+
+A dataset owner names one partition column ``p``, one business column that
+filters are mirrored from, and a value transform ``T`` (a SQL expression
+containing a ``:value`` placeholder). Superset then appends an equivalent
+predicate on ``p`` to every query, so chart authors change nothing and queries
+prune.
+
+The load-bearing assumption
+---------------------------
+Everything here reasons about ``T(col) op T(v)``, but what is emitted is
+``p op T(v)`` -- a predicate on a *physically different column*. The step from
+one to the other is::
+
+    p = T(mapped_col)   for every row in the table
+
+Superset cannot verify that; it is a property of whatever ETL populates the
+partition column. If that job lags, backfills with different logic, or writes
+the partition key in a different timezone, mirrored predicates silently drop
+real rows. The mapping is only as trustworthy as the pipeline behind it.
+
+Operator safety
+---------------
+A mirrored predicate ``P2`` may only be ``AND``-ed onto a query when the
+original predicate ``P1`` *implies* it:
+
+===========================================  =============================
+Original                                     Safe when
+===========================================  =============================
+``col = v``, ``col IN (...)``                always -- ``T`` is a function
+``col >=|>|<|<= v``, ``TEMPORAL_RANGE``      only if ``T`` is monotonic
+``col != v``, ``NOT IN``, ``LIKE``, ...      never
+===========================================  =============================
+
+Negations are never safe because ``T`` need not be injective:
+``lower(:value)`` with ``country != 'US'`` mirrors to ``region_key != 'us'``,
+which wrongly excludes rows whose ``country`` is already lowercase ``'us'`` --
+rows the original filter *keeps*.
+
+Monotonicity is a property of the transform, not of the column's data type:
+``hour(:value)``, ``date_format(:value, 'dd')`` and ``dayofweek(:value)`` are
+all reasonable transforms on a ``TIMESTAMP`` column and none of them preserve
+ordering. It is therefore declared by the owner, not inferred.
+
+Time grains
+-----------
+A grained filter -- drill-to-detail, mostly -- compares the *truncated* column,
+``trunc(col) op v``, so the raw bounds it carries do not describe the rows it
+keeps. A row in the final partial bucket satisfies ``trunc(col) < until`` while
+``col < until`` excludes it.
+
+Which direction a grain rounds is not knowable from the duration alone, and the
+obvious guess is wrong: ``WEEK_ENDING_SATURDAY`` rounds *forward* on Hive and
+Presto, and Ocient's grains are ``ROUND``, i.e. to nearest. What every grain
+does satisfy is a bound on the displacement::
+
+    |trunc(ts) - ts| < width(grain)
+
+so widening *both* bounds by one bucket width is no narrower than the real
+predicate whichever way the grain rounds -- and it needs no monotonicity of
+``trunc`` itself, which is what rescues Hive's oddly-anchored ``P1W``. Width is
+a property of the grain, not of the engine, which is what makes this
+maintainable; see `grain_bucket_width`. A grain whose SQL an operator supplied
+has no known width, and does not mirror at all.
+"""
+
+from __future__ import annotations
+
+import hashlib
+import logging
+import re
+from dataclasses import dataclass
+from functools import lru_cache
+from typing import Any, cast, TYPE_CHECKING
+
+import sqlalchemy as sa
+from dateutil.relativedelta import relativedelta
+from flask import current_app as app
+from flask_babel import lazy_gettext as _
+from sqlalchemy.engine.interfaces import Dialect
+from sqlalchemy.sql.elements import ColumnElement
+
+from superset.connectors.sqla.partition_mapping_storage import (
+    ColumnTransform,
+    load_partition_mapping,
+    StoredPartitionMapping,
+)
+from superset.constants import LRU_CACHE_MAX_SIZE, TimeGrain
+from superset.exceptions import SupersetParseError
+from superset.extensions import cache_manager, feature_flag_manager
+from superset.sql.parse import SQLStatement
+from superset.utils import json
+from superset.utils.core import FilterOperator
+
+if TYPE_CHECKING:
+    from superset.connectors.sqla.models import SqlaTable, TableColumn
+    from superset.models.core import Database
+
+logger = logging.getLogger(__name__)
+
+FEATURE_FLAG = "PARTITION_FILTER_MAPPING"
+
+#: Placeholder the owner writes in the transform, e.g. 
``unix_timestamp(:value)``.
+#: Matched with word boundaries so ``:values`` is not mistaken for it.
+VALUE_PLACEHOLDER_RE = re.compile(r":value\b")
+
+#: Balanced Jinja blocks. The probe would render these in a different context
+#: at a different time from the chart query, so they are rejected at save time.
+JINJA_BLOCK_RE = re.compile(r"\{\{.*?\}\}|\{%.*?%\}|\{#.*?#\}", re.DOTALL)
+
+#: Substituted for ``:value`` before parsing -- sqlglot rejects a bare 
``:value``
+#: on most dialects. Mirrors the ``_JINJA_BLOCK_RE`` -> ``NULL`` trick used by
+#: ``validate_stored_expression``.
+_PARSE_STANDIN = "NULL"
+
+#: Functions whose value depends on wall-clock time or randomness. The probe
+#: runs in a different session at a different moment from the chart query and
+#: its result is then cached, so any of these freezes a snapshot of probe time
+#: into the emitted predicate.
+NON_DETERMINISTIC_FUNCTIONS = {
+    "CURRENT_DATE",
+    "CURRENT_TIME",
+    "CURRENT_TIMESTAMP",
+    "NOW",
+    "RAND",
+    "RANDOM",
+    "UUID",
+}
+
+#: Functions that mean "now" only in their zero-argument form. On Hive and
+#: Impala ``unix_timestamp()`` is the current time while ``unix_timestamp(x)``
+#: -- the canonical transform for this feature -- is pure.
+NON_DETERMINISTIC_WHEN_NILADIC = {"UNIX_TIMESTAMP"}
+
+#: Safe for any function ``T``.
+MIRRORABLE_ALWAYS = {FilterOperator.EQUALS, FilterOperator.IN}
+
+#: Safe only when ``T`` preserves ordering.
+MIRRORABLE_IF_MONOTONIC = {
+    FilterOperator.GREATER_THAN,
+    FilterOperator.GREATER_THAN_OR_EQUALS,
+    FilterOperator.LESS_THAN,
+    FilterOperator.LESS_THAN_OR_EQUALS,
+    FilterOperator.TEMPORAL_RANGE,
+}
+
+
+def mirrorable_operators(is_monotonic: bool) -> set[FilterOperator]:
+    """
+    The operators whose predicates may be mirrored onto the partition column.
+
+    :param is_monotonic: whether the owner declared the transform
+        order-preserving
+    """
+    if is_monotonic:
+        return MIRRORABLE_ALWAYS | MIRRORABLE_IF_MONOTONIC
+    return set(MIRRORABLE_ALWAYS)
+
+
+#: How wide one bucket of each built-in time grain is.
+#:
+#: Keyed on the ISO duration a filter carries in its ``grain``. Written out
+#: rather than parsed: four of the week grains are ISO *intervals* with an
+#: anchor (``P1W/1970-01-03T00:00:00Z``) that ``isodate.parse_duration``
+#: rejects outright, and ``PT0.5H`` / ``P0.25Y`` are fractional. A literal
+#: table is also the thing a reviewer can check a line at a time.
+#:
+#: ``relativedelta`` rather than ``timedelta`` so the calendar grains stay
+#: calendar arithmetic: a month is not 30 days.
+GRAIN_BUCKET_WIDTHS: dict[str, relativedelta] = {
+    TimeGrain.SECOND: relativedelta(seconds=1),
+    TimeGrain.FIVE_SECONDS: relativedelta(seconds=5),
+    TimeGrain.THIRTY_SECONDS: relativedelta(seconds=30),
+    TimeGrain.MINUTE: relativedelta(minutes=1),
+    TimeGrain.FIVE_MINUTES: relativedelta(minutes=5),
+    TimeGrain.TEN_MINUTES: relativedelta(minutes=10),
+    TimeGrain.FIFTEEN_MINUTES: relativedelta(minutes=15),
+    TimeGrain.THIRTY_MINUTES: relativedelta(minutes=30),
+    TimeGrain.HALF_HOUR: relativedelta(minutes=30),
+    TimeGrain.HOUR: relativedelta(hours=1),
+    TimeGrain.SIX_HOURS: relativedelta(hours=6),
+    TimeGrain.DAY: relativedelta(days=1),
+    TimeGrain.WEEK: relativedelta(days=7),
+    TimeGrain.WEEK_STARTING_SUNDAY: relativedelta(days=7),
+    TimeGrain.WEEK_STARTING_MONDAY: relativedelta(days=7),
+    TimeGrain.WEEK_ENDING_SATURDAY: relativedelta(days=7),
+    TimeGrain.WEEK_ENDING_SUNDAY: relativedelta(days=7),
+    TimeGrain.MONTH: relativedelta(months=1),
+    TimeGrain.QUARTER: relativedelta(months=3),
+    TimeGrain.QUARTER_YEAR: relativedelta(months=3),
+    TimeGrain.YEAR: relativedelta(years=1),
+}
+
+
+def grain_bucket_width(grain: str | None, engine: str) -> relativedelta | None:
+    """
+    How far a grain's truncation can move a timestamp, or ``None`` if unknown.
+
+    A grained filter compares the *truncated* column, so the raw bounds it
+    carries do not describe the rows it keeps. Widening both bounds by one
+    bucket recovers a predicate that is no narrower than the real one -- see
+    `_collect_partition_mirror_range` for the argument. That only works for a
+    grain whose bucket width Superset knows, which excludes anything an
+    operator supplied.
+
+    :param grain: the ISO duration from the filter's ``grain``
+    :param engine: the engine spec's ``engine``, to check per-engine overrides
+    """
+    if not grain:
+        return None
+
+    # An operator-declared grain carries whatever duration string they typed,
+    # and `TIME_GRAIN_ADDON_EXPRESSIONS` can redefine a *built-in* grain's SQL
+    # per engine -- `P1D` could be anything at all. Neither has a width we can
+    # claim to know, so both fall back to not mirroring.
+    if grain in app.config["TIME_GRAIN_ADDONS"]:
+        return None
+    if grain in app.config["TIME_GRAIN_ADDON_EXPRESSIONS"].get(engine, {}):
+        return None
+
+    return GRAIN_BUCKET_WIDTHS.get(grain)
+
+
+@dataclass(frozen=True)
+class PartitionMapping:
+    """A resolved, usable partition filter mapping."""
+
+    partition_column: str
+    mapped_column: str
+    value_transform: str
+    is_monotonic: bool
+
+    def mirrors(self, operator: FilterOperator) -> bool:
+        return operator in mirrorable_operators(self.is_monotonic)
+
+
+def contains_value_placeholder(transform: str | None) -> bool:
+    """Whether the transform contains the ``:value`` placeholder."""
+    return bool(transform) and VALUE_PLACEHOLDER_RE.search(transform or "") is 
not None
+
+
+def contains_jinja(transform: str | None) -> bool:
+    """Whether the transform contains a balanced Jinja block."""
+    return bool(transform) and JINJA_BLOCK_RE.search(transform or "") is not 
None
+
+
+def parse_skeleton(transform: str) -> str:
+    """
+    The transform with ``:value`` substituted out, ready for a SQL parser.
+
+    ``sanitize_clause`` / sqlglot choke on a bare ``:value`` on most dialects,
+    so the placeholder is swapped for a benign literal first -- the same trick
+    ``validate_stored_expression`` uses for Jinja blocks.
+    """
+    return VALUE_PLACEHOLDER_RE.sub(_PARSE_STANDIN, transform)
+
+
+def _parse_skeleton(transform: str, engine: str) -> SQLStatement | None:
+    """
+    Parse ``SELECT <transform>`` with the placeholder substituted out.
+
+    Returns ``None`` when the transform does not parse.
+    """
+    try:
+        return SQLStatement(f"SELECT {parse_skeleton(transform)}", engine)
+    except SupersetParseError:
+        return None
+
+
+def is_parseable(transform: str | None, engine: str) -> bool:
+    """Whether the transform parses as a single select expression."""
+    if not transform or not transform.strip():
+        return False
+    return _parse_skeleton(transform, engine) is not None
+
+
+def find_non_deterministic_functions(transform: str, engine: str) -> set[str]:
+    """
+    Names of non-deterministic functions the transform calls.
+
+    ``UNIX_TIMESTAMP`` is only reported in its zero-argument form, which means
+    "now" on Hive and Impala; the one-argument form is the canonical temporal
+    transform and stays allowed.
+    """
+    statement = _parse_skeleton(transform, engine)
+    if statement is None:
+        return set()
+
+    found = {
+        name
+        for name in NON_DETERMINISTIC_FUNCTIONS
+        if statement.check_functions_present({name})
+    }
+    return found | _find_niladic_calls(statement)
+
+
+def _find_niladic_calls(statement: SQLStatement) -> set[str]:
+    """
+    Names from ``NON_DETERMINISTIC_WHEN_NILADIC`` called with no arguments.
+
+    Note some dialects resolve the zero-argument form themselves -- Hive parses
+    ``unix_timestamp()`` straight to ``CURRENT_TIMESTAMP`` -- in which case the
+    name-based check above has already caught it. This is the backstop for the
+    dialects that do not.
+    """
+    return NON_DETERMINISTIC_WHEN_NILADIC & statement.get_niladic_functions()
+
+
+def resolve_partition_mapping(datasource: SqlaTable) -> PartitionMapping | 
None:
+    """
+    Resolve the dataset's mapping, or ``None`` when nothing may be mirrored.
+
+    Every bail-out here is defensive as well as functional: save-time 
validation
+    rejects most of these, but rows predating the validation can still violate
+    the invariants, and a column sync can invalidate a mapping that was fine
+    when it was written.
+    """
+    if not feature_flag_manager.is_feature_enabled(FEATURE_FLAG):
+        return None
+
+    stored = load_partition_mapping(datasource)
+    partition_column = stored.partition_column
+    if not partition_column:
+        return None
+
+    columns_by_name = {column.column_name: column for column in 
datasource.columns}
+    if partition_column not in columns_by_name:
+        # The partition column was dropped by a column sync or at the source.
+        return None
+
+    mapped_column_name = 
stored.effective_mapped_column(datasource.main_dttm_col)
+    if not mapped_column_name or mapped_column_name not in columns_by_name:
+        return None
+
+    if mapped_column_name == partition_column:
+        # Self-mapping: the mirrored predicate would duplicate the original.
+        return None
+
+    transform_spec = stored.transform_for(mapped_column_name)
+    transform = transform_spec.value_transform
+    # `is_transform_active` rather than a query-time restatement of the same
+    # rules: it *is* `validate_transform` with the messages discarded, so the 
two
+    # cannot drift. The difference that matters is the non-deterministic
+    # denylist, which a Tier-2-only check omits -- and `CreateDatasetCommand`,
+    # import, and the editor's free-text Extra box all reach this storage 
without
+    # passing the save-time validation, so a transform calling `now()` would
+    # otherwise be mirrored and its result cached.
+    if not is_transform_active(transform, datasource.database.backend):
+        return None
+
+    if _has_active_advanced_data_type(columns_by_name[mapped_column_name]):
+        # `translate_filter` builds its own predicate shape from *translated*
+        # values, so the `(operator, value)` pair the operator matrix reasons
+        # about does not exist and mirroring would apply the wrong values.
+        return None
+
+    return PartitionMapping(
+        partition_column=str(partition_column),
+        mapped_column=str(mapped_column_name),
+        value_transform=cast(str, transform),
+        is_monotonic=transform_spec.is_monotonic,
+    )
+
+
+def _has_active_advanced_data_type(column: TableColumn) -> bool:
+    advanced_data_type = getattr(column, "advanced_data_type", None)
+    if not advanced_data_type:
+        return False
+    if not 
feature_flag_manager.is_feature_enabled("ENABLE_ADVANCED_DATA_TYPES"):
+        return False
+    return advanced_data_type in app.config.get("ADVANCED_DATA_TYPES", {})
+
+
+def build_probe_sql(
+    transform: str,
+    values: list[Any],
+    dialect: Dialect | None = None,
+) -> str:
+    """
+    Compile a single ``SELECT`` that evaluates the transform at every value.
+
+    Values are attacker-controlled (a Gamma user picks filter values), so they
+    are bound as parameters and rendered by the dialect's own literal processor
+    rather than interpolated into the SQL text.
+
+    Note this deliberately does *not* go through ``BaseEngineSpec``'s text
+    helper, which escapes ``:`` on every engine but Athena and would destroy 
the
+    ``:value`` placeholder before it can be bound.
+    """
+    selections = []
+    for index, value in enumerate(values):
+        clause = sa.text(transform).bindparams(sa.bindparam("value", 
value=value))

Review Comment:
   For timestamp filters supplied as epoch milliseconds, 
`filter_values_handler` returns a SQLAlchemy `literal_column` such as Hive's 
`CAST(... AS TIMESTAMP)`, which cannot be rendered as this bound literal. The 
probe catches the resulting `CompileError` and silently skips pruning for 
equality, IN, and comparison filters; can these values be normalized to actual 
temporal values before binding?



##########
superset/connectors/sqla/partition_mapping.py:
##########
@@ -0,0 +1,977 @@
+# Licensed to the Apache Software Foundation (ASF) under one
+# or more contributor license agreements.  See the NOTICE file
+# distributed with this work for additional information
+# regarding copyright ownership.  The ASF licenses this file
+# to you under the Apache License, Version 2.0 (the
+# "License"); you may not use this file except in compliance
+# with the License.  You may obtain a copy of the License at
+#
+#   http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing,
+# software distributed under the License is distributed on an
+# "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+# KIND, either express or implied.  See the License for the
+# specific language governing permissions and limitations
+# under the License.
+"""
+Partition filter mapping.
+
+Datasets on Hadoop-family engines are commonly partitioned on a *technical*
+column -- an epoch integer, a lowercased region key -- that no analyst would
+filter on. Unless a query carries a predicate on that column the engine scans
+every partition.
+
+A dataset owner names one partition column ``p``, one business column that
+filters are mirrored from, and a value transform ``T`` (a SQL expression
+containing a ``:value`` placeholder). Superset then appends an equivalent
+predicate on ``p`` to every query, so chart authors change nothing and queries
+prune.
+
+The load-bearing assumption
+---------------------------
+Everything here reasons about ``T(col) op T(v)``, but what is emitted is
+``p op T(v)`` -- a predicate on a *physically different column*. The step from
+one to the other is::
+
+    p = T(mapped_col)   for every row in the table
+
+Superset cannot verify that; it is a property of whatever ETL populates the
+partition column. If that job lags, backfills with different logic, or writes
+the partition key in a different timezone, mirrored predicates silently drop
+real rows. The mapping is only as trustworthy as the pipeline behind it.
+
+Operator safety
+---------------
+A mirrored predicate ``P2`` may only be ``AND``-ed onto a query when the
+original predicate ``P1`` *implies* it:
+
+===========================================  =============================
+Original                                     Safe when
+===========================================  =============================
+``col = v``, ``col IN (...)``                always -- ``T`` is a function
+``col >=|>|<|<= v``, ``TEMPORAL_RANGE``      only if ``T`` is monotonic
+``col != v``, ``NOT IN``, ``LIKE``, ...      never
+===========================================  =============================
+
+Negations are never safe because ``T`` need not be injective:
+``lower(:value)`` with ``country != 'US'`` mirrors to ``region_key != 'us'``,
+which wrongly excludes rows whose ``country`` is already lowercase ``'us'`` --
+rows the original filter *keeps*.
+
+Monotonicity is a property of the transform, not of the column's data type:
+``hour(:value)``, ``date_format(:value, 'dd')`` and ``dayofweek(:value)`` are
+all reasonable transforms on a ``TIMESTAMP`` column and none of them preserve
+ordering. It is therefore declared by the owner, not inferred.
+
+Time grains
+-----------
+A grained filter -- drill-to-detail, mostly -- compares the *truncated* column,
+``trunc(col) op v``, so the raw bounds it carries do not describe the rows it
+keeps. A row in the final partial bucket satisfies ``trunc(col) < until`` while
+``col < until`` excludes it.
+
+Which direction a grain rounds is not knowable from the duration alone, and the
+obvious guess is wrong: ``WEEK_ENDING_SATURDAY`` rounds *forward* on Hive and
+Presto, and Ocient's grains are ``ROUND``, i.e. to nearest. What every grain
+does satisfy is a bound on the displacement::
+
+    |trunc(ts) - ts| < width(grain)
+
+so widening *both* bounds by one bucket width is no narrower than the real
+predicate whichever way the grain rounds -- and it needs no monotonicity of
+``trunc`` itself, which is what rescues Hive's oddly-anchored ``P1W``. Width is
+a property of the grain, not of the engine, which is what makes this
+maintainable; see `grain_bucket_width`. A grain whose SQL an operator supplied
+has no known width, and does not mirror at all.
+"""
+
+from __future__ import annotations
+
+import hashlib
+import logging
+import re
+from dataclasses import dataclass
+from functools import lru_cache
+from typing import Any, cast, TYPE_CHECKING
+
+import sqlalchemy as sa
+from dateutil.relativedelta import relativedelta
+from flask import current_app as app
+from flask_babel import lazy_gettext as _
+from sqlalchemy.engine.interfaces import Dialect
+from sqlalchemy.sql.elements import ColumnElement
+
+from superset.connectors.sqla.partition_mapping_storage import (
+    ColumnTransform,
+    load_partition_mapping,
+    StoredPartitionMapping,
+)
+from superset.constants import LRU_CACHE_MAX_SIZE, TimeGrain
+from superset.exceptions import SupersetParseError
+from superset.extensions import cache_manager, feature_flag_manager
+from superset.sql.parse import SQLStatement
+from superset.utils import json
+from superset.utils.core import FilterOperator
+
+if TYPE_CHECKING:
+    from superset.connectors.sqla.models import SqlaTable, TableColumn
+    from superset.models.core import Database
+
+logger = logging.getLogger(__name__)
+
+FEATURE_FLAG = "PARTITION_FILTER_MAPPING"
+
+#: Placeholder the owner writes in the transform, e.g. 
``unix_timestamp(:value)``.
+#: Matched with word boundaries so ``:values`` is not mistaken for it.
+VALUE_PLACEHOLDER_RE = re.compile(r":value\b")
+
+#: Balanced Jinja blocks. The probe would render these in a different context
+#: at a different time from the chart query, so they are rejected at save time.
+JINJA_BLOCK_RE = re.compile(r"\{\{.*?\}\}|\{%.*?%\}|\{#.*?#\}", re.DOTALL)
+
+#: Substituted for ``:value`` before parsing -- sqlglot rejects a bare 
``:value``
+#: on most dialects. Mirrors the ``_JINJA_BLOCK_RE`` -> ``NULL`` trick used by
+#: ``validate_stored_expression``.
+_PARSE_STANDIN = "NULL"
+
+#: Functions whose value depends on wall-clock time or randomness. The probe
+#: runs in a different session at a different moment from the chart query and
+#: its result is then cached, so any of these freezes a snapshot of probe time
+#: into the emitted predicate.
+NON_DETERMINISTIC_FUNCTIONS = {
+    "CURRENT_DATE",
+    "CURRENT_TIME",
+    "CURRENT_TIMESTAMP",
+    "NOW",
+    "RAND",
+    "RANDOM",
+    "UUID",
+}
+
+#: Functions that mean "now" only in their zero-argument form. On Hive and
+#: Impala ``unix_timestamp()`` is the current time while ``unix_timestamp(x)``
+#: -- the canonical transform for this feature -- is pure.
+NON_DETERMINISTIC_WHEN_NILADIC = {"UNIX_TIMESTAMP"}
+
+#: Safe for any function ``T``.
+MIRRORABLE_ALWAYS = {FilterOperator.EQUALS, FilterOperator.IN}
+
+#: Safe only when ``T`` preserves ordering.
+MIRRORABLE_IF_MONOTONIC = {
+    FilterOperator.GREATER_THAN,
+    FilterOperator.GREATER_THAN_OR_EQUALS,
+    FilterOperator.LESS_THAN,
+    FilterOperator.LESS_THAN_OR_EQUALS,
+    FilterOperator.TEMPORAL_RANGE,
+}
+
+
+def mirrorable_operators(is_monotonic: bool) -> set[FilterOperator]:
+    """
+    The operators whose predicates may be mirrored onto the partition column.
+
+    :param is_monotonic: whether the owner declared the transform
+        order-preserving
+    """
+    if is_monotonic:
+        return MIRRORABLE_ALWAYS | MIRRORABLE_IF_MONOTONIC
+    return set(MIRRORABLE_ALWAYS)
+
+
+#: How wide one bucket of each built-in time grain is.
+#:
+#: Keyed on the ISO duration a filter carries in its ``grain``. Written out
+#: rather than parsed: four of the week grains are ISO *intervals* with an
+#: anchor (``P1W/1970-01-03T00:00:00Z``) that ``isodate.parse_duration``
+#: rejects outright, and ``PT0.5H`` / ``P0.25Y`` are fractional. A literal
+#: table is also the thing a reviewer can check a line at a time.
+#:
+#: ``relativedelta`` rather than ``timedelta`` so the calendar grains stay
+#: calendar arithmetic: a month is not 30 days.
+GRAIN_BUCKET_WIDTHS: dict[str, relativedelta] = {
+    TimeGrain.SECOND: relativedelta(seconds=1),
+    TimeGrain.FIVE_SECONDS: relativedelta(seconds=5),
+    TimeGrain.THIRTY_SECONDS: relativedelta(seconds=30),
+    TimeGrain.MINUTE: relativedelta(minutes=1),
+    TimeGrain.FIVE_MINUTES: relativedelta(minutes=5),
+    TimeGrain.TEN_MINUTES: relativedelta(minutes=10),
+    TimeGrain.FIFTEEN_MINUTES: relativedelta(minutes=15),
+    TimeGrain.THIRTY_MINUTES: relativedelta(minutes=30),
+    TimeGrain.HALF_HOUR: relativedelta(minutes=30),
+    TimeGrain.HOUR: relativedelta(hours=1),
+    TimeGrain.SIX_HOURS: relativedelta(hours=6),
+    TimeGrain.DAY: relativedelta(days=1),
+    TimeGrain.WEEK: relativedelta(days=7),
+    TimeGrain.WEEK_STARTING_SUNDAY: relativedelta(days=7),
+    TimeGrain.WEEK_STARTING_MONDAY: relativedelta(days=7),
+    TimeGrain.WEEK_ENDING_SATURDAY: relativedelta(days=7),
+    TimeGrain.WEEK_ENDING_SUNDAY: relativedelta(days=7),
+    TimeGrain.MONTH: relativedelta(months=1),
+    TimeGrain.QUARTER: relativedelta(months=3),
+    TimeGrain.QUARTER_YEAR: relativedelta(months=3),
+    TimeGrain.YEAR: relativedelta(years=1),
+}
+
+
+def grain_bucket_width(grain: str | None, engine: str) -> relativedelta | None:
+    """
+    How far a grain's truncation can move a timestamp, or ``None`` if unknown.
+
+    A grained filter compares the *truncated* column, so the raw bounds it
+    carries do not describe the rows it keeps. Widening both bounds by one
+    bucket recovers a predicate that is no narrower than the real one -- see
+    `_collect_partition_mirror_range` for the argument. That only works for a
+    grain whose bucket width Superset knows, which excludes anything an
+    operator supplied.
+
+    :param grain: the ISO duration from the filter's ``grain``
+    :param engine: the engine spec's ``engine``, to check per-engine overrides
+    """
+    if not grain:
+        return None
+
+    # An operator-declared grain carries whatever duration string they typed,
+    # and `TIME_GRAIN_ADDON_EXPRESSIONS` can redefine a *built-in* grain's SQL
+    # per engine -- `P1D` could be anything at all. Neither has a width we can
+    # claim to know, so both fall back to not mirroring.
+    if grain in app.config["TIME_GRAIN_ADDONS"]:
+        return None
+    if grain in app.config["TIME_GRAIN_ADDON_EXPRESSIONS"].get(engine, {}):
+        return None
+
+    return GRAIN_BUCKET_WIDTHS.get(grain)
+
+
+@dataclass(frozen=True)
+class PartitionMapping:
+    """A resolved, usable partition filter mapping."""
+
+    partition_column: str
+    mapped_column: str
+    value_transform: str
+    is_monotonic: bool
+
+    def mirrors(self, operator: FilterOperator) -> bool:
+        return operator in mirrorable_operators(self.is_monotonic)
+
+
+def contains_value_placeholder(transform: str | None) -> bool:
+    """Whether the transform contains the ``:value`` placeholder."""
+    return bool(transform) and VALUE_PLACEHOLDER_RE.search(transform or "") is 
not None
+
+
+def contains_jinja(transform: str | None) -> bool:
+    """Whether the transform contains a balanced Jinja block."""
+    return bool(transform) and JINJA_BLOCK_RE.search(transform or "") is not 
None
+
+
+def parse_skeleton(transform: str) -> str:
+    """
+    The transform with ``:value`` substituted out, ready for a SQL parser.
+
+    ``sanitize_clause`` / sqlglot choke on a bare ``:value`` on most dialects,
+    so the placeholder is swapped for a benign literal first -- the same trick
+    ``validate_stored_expression`` uses for Jinja blocks.
+    """
+    return VALUE_PLACEHOLDER_RE.sub(_PARSE_STANDIN, transform)
+
+
+def _parse_skeleton(transform: str, engine: str) -> SQLStatement | None:
+    """
+    Parse ``SELECT <transform>`` with the placeholder substituted out.
+
+    Returns ``None`` when the transform does not parse.
+    """
+    try:
+        return SQLStatement(f"SELECT {parse_skeleton(transform)}", engine)
+    except SupersetParseError:
+        return None
+
+
+def is_parseable(transform: str | None, engine: str) -> bool:
+    """Whether the transform parses as a single select expression."""
+    if not transform or not transform.strip():
+        return False
+    return _parse_skeleton(transform, engine) is not None
+
+
+def find_non_deterministic_functions(transform: str, engine: str) -> set[str]:
+    """
+    Names of non-deterministic functions the transform calls.
+
+    ``UNIX_TIMESTAMP`` is only reported in its zero-argument form, which means
+    "now" on Hive and Impala; the one-argument form is the canonical temporal
+    transform and stays allowed.
+    """
+    statement = _parse_skeleton(transform, engine)
+    if statement is None:
+        return set()
+
+    found = {
+        name
+        for name in NON_DETERMINISTIC_FUNCTIONS
+        if statement.check_functions_present({name})
+    }
+    return found | _find_niladic_calls(statement)
+
+
+def _find_niladic_calls(statement: SQLStatement) -> set[str]:
+    """
+    Names from ``NON_DETERMINISTIC_WHEN_NILADIC`` called with no arguments.
+
+    Note some dialects resolve the zero-argument form themselves -- Hive parses
+    ``unix_timestamp()`` straight to ``CURRENT_TIMESTAMP`` -- in which case the
+    name-based check above has already caught it. This is the backstop for the
+    dialects that do not.
+    """
+    return NON_DETERMINISTIC_WHEN_NILADIC & statement.get_niladic_functions()
+
+
+def resolve_partition_mapping(datasource: SqlaTable) -> PartitionMapping | 
None:
+    """
+    Resolve the dataset's mapping, or ``None`` when nothing may be mirrored.
+
+    Every bail-out here is defensive as well as functional: save-time 
validation
+    rejects most of these, but rows predating the validation can still violate
+    the invariants, and a column sync can invalidate a mapping that was fine
+    when it was written.
+    """
+    if not feature_flag_manager.is_feature_enabled(FEATURE_FLAG):
+        return None
+
+    stored = load_partition_mapping(datasource)
+    partition_column = stored.partition_column
+    if not partition_column:
+        return None
+
+    columns_by_name = {column.column_name: column for column in 
datasource.columns}
+    if partition_column not in columns_by_name:
+        # The partition column was dropped by a column sync or at the source.
+        return None
+
+    mapped_column_name = 
stored.effective_mapped_column(datasource.main_dttm_col)
+    if not mapped_column_name or mapped_column_name not in columns_by_name:
+        return None
+
+    if mapped_column_name == partition_column:
+        # Self-mapping: the mirrored predicate would duplicate the original.
+        return None
+
+    transform_spec = stored.transform_for(mapped_column_name)
+    transform = transform_spec.value_transform
+    # `is_transform_active` rather than a query-time restatement of the same
+    # rules: it *is* `validate_transform` with the messages discarded, so the 
two
+    # cannot drift. The difference that matters is the non-deterministic
+    # denylist, which a Tier-2-only check omits -- and `CreateDatasetCommand`,
+    # import, and the editor's free-text Extra box all reach this storage 
without
+    # passing the save-time validation, so a transform calling `now()` would
+    # otherwise be mirrored and its result cached.
+    if not is_transform_active(transform, datasource.database.backend):
+        return None
+
+    if _has_active_advanced_data_type(columns_by_name[mapped_column_name]):
+        # `translate_filter` builds its own predicate shape from *translated*
+        # values, so the `(operator, value)` pair the operator matrix reasons
+        # about does not exist and mirroring would apply the wrong values.
+        return None
+
+    return PartitionMapping(
+        partition_column=str(partition_column),
+        mapped_column=str(mapped_column_name),
+        value_transform=cast(str, transform),
+        is_monotonic=transform_spec.is_monotonic,
+    )
+
+
+def _has_active_advanced_data_type(column: TableColumn) -> bool:
+    advanced_data_type = getattr(column, "advanced_data_type", None)
+    if not advanced_data_type:
+        return False
+    if not 
feature_flag_manager.is_feature_enabled("ENABLE_ADVANCED_DATA_TYPES"):
+        return False
+    return advanced_data_type in app.config.get("ADVANCED_DATA_TYPES", {})
+
+
+def build_probe_sql(
+    transform: str,
+    values: list[Any],
+    dialect: Dialect | None = None,
+) -> str:
+    """
+    Compile a single ``SELECT`` that evaluates the transform at every value.
+
+    Values are attacker-controlled (a Gamma user picks filter values), so they
+    are bound as parameters and rendered by the dialect's own literal processor
+    rather than interpolated into the SQL text.
+
+    Note this deliberately does *not* go through ``BaseEngineSpec``'s text
+    helper, which escapes ``:`` on every engine but Athena and would destroy 
the
+    ``:value`` placeholder before it can be bound.
+    """
+    selections = []
+    for index, value in enumerate(values):
+        clause = sa.text(transform).bindparams(sa.bindparam("value", 
value=value))
+        compiled = clause.compile(
+            dialect=dialect,
+            compile_kwargs={"literal_binds": True},
+        )
+        selections.append(f"{compiled} AS v{index}")
+    return "SELECT " + ", ".join(selections)
+
+
+def evaluate_transform(
+    database: Database,
+    catalog: str | None,
+    schema: str | None,
+    transform: str,
+    values: list[Any],
+) -> list[Any] | None:
+    """
+    Evaluate ``transform`` against the engine once per distinct value.
+
+    Returns one result per input value, positionally aligned with ``values``, 
or
+    ``None`` if anything at all goes wrong. Failing open costs pruning, never
+    correctness: the chart query still runs, it just scans more partitions.
+
+    The probe is pinned to the dataset's catalog and schema so session settings
+    match the chart query as closely as the connection pool allows. It still
+    runs in a *different* session, which is why transforms calling
+    session-dependent functions are rejected at save time.
+    """
+    if not values:
+        return None
+
+    # Dedupe so a 200-value `IN` list costs one column, not 200.
+    distinct: list[Any] = []
+    seen: set[Any] = set()
+    for value in values:
+        key = _hashable(value)
+        if key not in seen:
+            seen.add(key)
+            distinct.append(value)
+
+    cache_key = _probe_cache_key(database, catalog, schema, transform, 
distinct)
+    cached = _cache_get(cache_key)
+    if cached is None:
+        cached = _run_probe(database, catalog, schema, transform, distinct)
+        if cached is None:
+            # Deliberately not cached: a transient engine blip would otherwise
+            # keep the dataset pruning-free for the whole cache timeout.
+            return None
+        _cache_set(cache_key, cached)
+
+    evaluated = dict(
+        zip((_hashable(value) for value in distinct), cached, strict=False)
+    )
+    return [evaluated[_hashable(value)] for value in values]
+
+
+def _run_probe(
+    database: Database,
+    catalog: str | None,
+    schema: str | None,
+    transform: str,
+    distinct: list[Any],
+) -> list[Any] | None:
+    try:
+        sql = build_probe_sql(transform, distinct, _dialect_for(database))

Review Comment:
   This probe bypasses the stored-expression subquery policy and RLS injection: 
an Alpha dataset owner without `sql_lab` can submit a scalar `SELECT` through 
preview and receive its result in `emitted_predicate`, and imported mappings 
reach the same sink. Can the probe enforce the same subquery/access/RLS checks 
as stored expressions before executing SQL?



##########
superset/connectors/sqla/partition_mapping.py:
##########
@@ -0,0 +1,977 @@
+# Licensed to the Apache Software Foundation (ASF) under one
+# or more contributor license agreements.  See the NOTICE file
+# distributed with this work for additional information
+# regarding copyright ownership.  The ASF licenses this file
+# to you under the Apache License, Version 2.0 (the
+# "License"); you may not use this file except in compliance
+# with the License.  You may obtain a copy of the License at
+#
+#   http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing,
+# software distributed under the License is distributed on an
+# "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+# KIND, either express or implied.  See the License for the
+# specific language governing permissions and limitations
+# under the License.
+"""
+Partition filter mapping.
+
+Datasets on Hadoop-family engines are commonly partitioned on a *technical*
+column -- an epoch integer, a lowercased region key -- that no analyst would
+filter on. Unless a query carries a predicate on that column the engine scans
+every partition.
+
+A dataset owner names one partition column ``p``, one business column that
+filters are mirrored from, and a value transform ``T`` (a SQL expression
+containing a ``:value`` placeholder). Superset then appends an equivalent
+predicate on ``p`` to every query, so chart authors change nothing and queries
+prune.
+
+The load-bearing assumption
+---------------------------
+Everything here reasons about ``T(col) op T(v)``, but what is emitted is
+``p op T(v)`` -- a predicate on a *physically different column*. The step from
+one to the other is::
+
+    p = T(mapped_col)   for every row in the table
+
+Superset cannot verify that; it is a property of whatever ETL populates the
+partition column. If that job lags, backfills with different logic, or writes
+the partition key in a different timezone, mirrored predicates silently drop
+real rows. The mapping is only as trustworthy as the pipeline behind it.
+
+Operator safety
+---------------
+A mirrored predicate ``P2`` may only be ``AND``-ed onto a query when the
+original predicate ``P1`` *implies* it:
+
+===========================================  =============================
+Original                                     Safe when
+===========================================  =============================
+``col = v``, ``col IN (...)``                always -- ``T`` is a function
+``col >=|>|<|<= v``, ``TEMPORAL_RANGE``      only if ``T`` is monotonic
+``col != v``, ``NOT IN``, ``LIKE``, ...      never
+===========================================  =============================
+
+Negations are never safe because ``T`` need not be injective:
+``lower(:value)`` with ``country != 'US'`` mirrors to ``region_key != 'us'``,
+which wrongly excludes rows whose ``country`` is already lowercase ``'us'`` --
+rows the original filter *keeps*.
+
+Monotonicity is a property of the transform, not of the column's data type:
+``hour(:value)``, ``date_format(:value, 'dd')`` and ``dayofweek(:value)`` are
+all reasonable transforms on a ``TIMESTAMP`` column and none of them preserve
+ordering. It is therefore declared by the owner, not inferred.
+
+Time grains
+-----------
+A grained filter -- drill-to-detail, mostly -- compares the *truncated* column,
+``trunc(col) op v``, so the raw bounds it carries do not describe the rows it
+keeps. A row in the final partial bucket satisfies ``trunc(col) < until`` while
+``col < until`` excludes it.
+
+Which direction a grain rounds is not knowable from the duration alone, and the
+obvious guess is wrong: ``WEEK_ENDING_SATURDAY`` rounds *forward* on Hive and
+Presto, and Ocient's grains are ``ROUND``, i.e. to nearest. What every grain
+does satisfy is a bound on the displacement::
+
+    |trunc(ts) - ts| < width(grain)
+
+so widening *both* bounds by one bucket width is no narrower than the real
+predicate whichever way the grain rounds -- and it needs no monotonicity of
+``trunc`` itself, which is what rescues Hive's oddly-anchored ``P1W``. Width is
+a property of the grain, not of the engine, which is what makes this
+maintainable; see `grain_bucket_width`. A grain whose SQL an operator supplied
+has no known width, and does not mirror at all.
+"""
+
+from __future__ import annotations
+
+import hashlib
+import logging
+import re
+from dataclasses import dataclass
+from functools import lru_cache
+from typing import Any, cast, TYPE_CHECKING
+
+import sqlalchemy as sa
+from dateutil.relativedelta import relativedelta
+from flask import current_app as app
+from flask_babel import lazy_gettext as _
+from sqlalchemy.engine.interfaces import Dialect
+from sqlalchemy.sql.elements import ColumnElement
+
+from superset.connectors.sqla.partition_mapping_storage import (
+    ColumnTransform,
+    load_partition_mapping,
+    StoredPartitionMapping,
+)
+from superset.constants import LRU_CACHE_MAX_SIZE, TimeGrain
+from superset.exceptions import SupersetParseError
+from superset.extensions import cache_manager, feature_flag_manager
+from superset.sql.parse import SQLStatement
+from superset.utils import json
+from superset.utils.core import FilterOperator
+
+if TYPE_CHECKING:
+    from superset.connectors.sqla.models import SqlaTable, TableColumn
+    from superset.models.core import Database
+
+logger = logging.getLogger(__name__)
+
+FEATURE_FLAG = "PARTITION_FILTER_MAPPING"
+
+#: Placeholder the owner writes in the transform, e.g. 
``unix_timestamp(:value)``.
+#: Matched with word boundaries so ``:values`` is not mistaken for it.
+VALUE_PLACEHOLDER_RE = re.compile(r":value\b")
+
+#: Balanced Jinja blocks. The probe would render these in a different context
+#: at a different time from the chart query, so they are rejected at save time.
+JINJA_BLOCK_RE = re.compile(r"\{\{.*?\}\}|\{%.*?%\}|\{#.*?#\}", re.DOTALL)
+
+#: Substituted for ``:value`` before parsing -- sqlglot rejects a bare 
``:value``
+#: on most dialects. Mirrors the ``_JINJA_BLOCK_RE`` -> ``NULL`` trick used by
+#: ``validate_stored_expression``.
+_PARSE_STANDIN = "NULL"
+
+#: Functions whose value depends on wall-clock time or randomness. The probe
+#: runs in a different session at a different moment from the chart query and
+#: its result is then cached, so any of these freezes a snapshot of probe time
+#: into the emitted predicate.
+NON_DETERMINISTIC_FUNCTIONS = {
+    "CURRENT_DATE",
+    "CURRENT_TIME",
+    "CURRENT_TIMESTAMP",
+    "NOW",
+    "RAND",
+    "RANDOM",
+    "UUID",
+}
+
+#: Functions that mean "now" only in their zero-argument form. On Hive and
+#: Impala ``unix_timestamp()`` is the current time while ``unix_timestamp(x)``
+#: -- the canonical transform for this feature -- is pure.
+NON_DETERMINISTIC_WHEN_NILADIC = {"UNIX_TIMESTAMP"}
+
+#: Safe for any function ``T``.
+MIRRORABLE_ALWAYS = {FilterOperator.EQUALS, FilterOperator.IN}
+
+#: Safe only when ``T`` preserves ordering.
+MIRRORABLE_IF_MONOTONIC = {
+    FilterOperator.GREATER_THAN,
+    FilterOperator.GREATER_THAN_OR_EQUALS,
+    FilterOperator.LESS_THAN,
+    FilterOperator.LESS_THAN_OR_EQUALS,
+    FilterOperator.TEMPORAL_RANGE,
+}
+
+
+def mirrorable_operators(is_monotonic: bool) -> set[FilterOperator]:
+    """
+    The operators whose predicates may be mirrored onto the partition column.
+
+    :param is_monotonic: whether the owner declared the transform
+        order-preserving
+    """
+    if is_monotonic:
+        return MIRRORABLE_ALWAYS | MIRRORABLE_IF_MONOTONIC
+    return set(MIRRORABLE_ALWAYS)
+
+
+#: How wide one bucket of each built-in time grain is.
+#:
+#: Keyed on the ISO duration a filter carries in its ``grain``. Written out
+#: rather than parsed: four of the week grains are ISO *intervals* with an
+#: anchor (``P1W/1970-01-03T00:00:00Z``) that ``isodate.parse_duration``
+#: rejects outright, and ``PT0.5H`` / ``P0.25Y`` are fractional. A literal
+#: table is also the thing a reviewer can check a line at a time.
+#:
+#: ``relativedelta`` rather than ``timedelta`` so the calendar grains stay
+#: calendar arithmetic: a month is not 30 days.
+GRAIN_BUCKET_WIDTHS: dict[str, relativedelta] = {
+    TimeGrain.SECOND: relativedelta(seconds=1),
+    TimeGrain.FIVE_SECONDS: relativedelta(seconds=5),
+    TimeGrain.THIRTY_SECONDS: relativedelta(seconds=30),
+    TimeGrain.MINUTE: relativedelta(minutes=1),
+    TimeGrain.FIVE_MINUTES: relativedelta(minutes=5),
+    TimeGrain.TEN_MINUTES: relativedelta(minutes=10),
+    TimeGrain.FIFTEEN_MINUTES: relativedelta(minutes=15),
+    TimeGrain.THIRTY_MINUTES: relativedelta(minutes=30),
+    TimeGrain.HALF_HOUR: relativedelta(minutes=30),
+    TimeGrain.HOUR: relativedelta(hours=1),
+    TimeGrain.SIX_HOURS: relativedelta(hours=6),
+    TimeGrain.DAY: relativedelta(days=1),
+    TimeGrain.WEEK: relativedelta(days=7),
+    TimeGrain.WEEK_STARTING_SUNDAY: relativedelta(days=7),
+    TimeGrain.WEEK_STARTING_MONDAY: relativedelta(days=7),
+    TimeGrain.WEEK_ENDING_SATURDAY: relativedelta(days=7),
+    TimeGrain.WEEK_ENDING_SUNDAY: relativedelta(days=7),
+    TimeGrain.MONTH: relativedelta(months=1),
+    TimeGrain.QUARTER: relativedelta(months=3),
+    TimeGrain.QUARTER_YEAR: relativedelta(months=3),
+    TimeGrain.YEAR: relativedelta(years=1),
+}
+
+
+def grain_bucket_width(grain: str | None, engine: str) -> relativedelta | None:
+    """
+    How far a grain's truncation can move a timestamp, or ``None`` if unknown.
+
+    A grained filter compares the *truncated* column, so the raw bounds it
+    carries do not describe the rows it keeps. Widening both bounds by one
+    bucket recovers a predicate that is no narrower than the real one -- see
+    `_collect_partition_mirror_range` for the argument. That only works for a
+    grain whose bucket width Superset knows, which excludes anything an
+    operator supplied.
+
+    :param grain: the ISO duration from the filter's ``grain``
+    :param engine: the engine spec's ``engine``, to check per-engine overrides
+    """
+    if not grain:
+        return None
+
+    # An operator-declared grain carries whatever duration string they typed,
+    # and `TIME_GRAIN_ADDON_EXPRESSIONS` can redefine a *built-in* grain's SQL
+    # per engine -- `P1D` could be anything at all. Neither has a width we can
+    # claim to know, so both fall back to not mirroring.
+    if grain in app.config["TIME_GRAIN_ADDONS"]:
+        return None
+    if grain in app.config["TIME_GRAIN_ADDON_EXPRESSIONS"].get(engine, {}):
+        return None
+
+    return GRAIN_BUCKET_WIDTHS.get(grain)
+
+
+@dataclass(frozen=True)
+class PartitionMapping:
+    """A resolved, usable partition filter mapping."""
+
+    partition_column: str
+    mapped_column: str
+    value_transform: str
+    is_monotonic: bool
+
+    def mirrors(self, operator: FilterOperator) -> bool:
+        return operator in mirrorable_operators(self.is_monotonic)
+
+
+def contains_value_placeholder(transform: str | None) -> bool:
+    """Whether the transform contains the ``:value`` placeholder."""
+    return bool(transform) and VALUE_PLACEHOLDER_RE.search(transform or "") is 
not None
+
+
+def contains_jinja(transform: str | None) -> bool:
+    """Whether the transform contains a balanced Jinja block."""
+    return bool(transform) and JINJA_BLOCK_RE.search(transform or "") is not 
None
+
+
+def parse_skeleton(transform: str) -> str:
+    """
+    The transform with ``:value`` substituted out, ready for a SQL parser.
+
+    ``sanitize_clause`` / sqlglot choke on a bare ``:value`` on most dialects,
+    so the placeholder is swapped for a benign literal first -- the same trick
+    ``validate_stored_expression`` uses for Jinja blocks.
+    """
+    return VALUE_PLACEHOLDER_RE.sub(_PARSE_STANDIN, transform)
+
+
+def _parse_skeleton(transform: str, engine: str) -> SQLStatement | None:
+    """
+    Parse ``SELECT <transform>`` with the placeholder substituted out.
+
+    Returns ``None`` when the transform does not parse.
+    """
+    try:
+        return SQLStatement(f"SELECT {parse_skeleton(transform)}", engine)
+    except SupersetParseError:
+        return None
+
+
+def is_parseable(transform: str | None, engine: str) -> bool:
+    """Whether the transform parses as a single select expression."""
+    if not transform or not transform.strip():
+        return False
+    return _parse_skeleton(transform, engine) is not None
+
+
+def find_non_deterministic_functions(transform: str, engine: str) -> set[str]:
+    """
+    Names of non-deterministic functions the transform calls.
+
+    ``UNIX_TIMESTAMP`` is only reported in its zero-argument form, which means
+    "now" on Hive and Impala; the one-argument form is the canonical temporal
+    transform and stays allowed.
+    """
+    statement = _parse_skeleton(transform, engine)
+    if statement is None:
+        return set()
+
+    found = {
+        name
+        for name in NON_DETERMINISTIC_FUNCTIONS
+        if statement.check_functions_present({name})
+    }
+    return found | _find_niladic_calls(statement)
+
+
+def _find_niladic_calls(statement: SQLStatement) -> set[str]:
+    """
+    Names from ``NON_DETERMINISTIC_WHEN_NILADIC`` called with no arguments.
+
+    Note some dialects resolve the zero-argument form themselves -- Hive parses
+    ``unix_timestamp()`` straight to ``CURRENT_TIMESTAMP`` -- in which case the
+    name-based check above has already caught it. This is the backstop for the
+    dialects that do not.
+    """
+    return NON_DETERMINISTIC_WHEN_NILADIC & statement.get_niladic_functions()
+
+
+def resolve_partition_mapping(datasource: SqlaTable) -> PartitionMapping | 
None:
+    """
+    Resolve the dataset's mapping, or ``None`` when nothing may be mirrored.
+
+    Every bail-out here is defensive as well as functional: save-time 
validation
+    rejects most of these, but rows predating the validation can still violate
+    the invariants, and a column sync can invalidate a mapping that was fine
+    when it was written.
+    """
+    if not feature_flag_manager.is_feature_enabled(FEATURE_FLAG):
+        return None
+
+    stored = load_partition_mapping(datasource)
+    partition_column = stored.partition_column
+    if not partition_column:
+        return None
+
+    columns_by_name = {column.column_name: column for column in 
datasource.columns}
+    if partition_column not in columns_by_name:
+        # The partition column was dropped by a column sync or at the source.
+        return None
+
+    mapped_column_name = 
stored.effective_mapped_column(datasource.main_dttm_col)
+    if not mapped_column_name or mapped_column_name not in columns_by_name:
+        return None
+
+    if mapped_column_name == partition_column:
+        # Self-mapping: the mirrored predicate would duplicate the original.
+        return None
+
+    transform_spec = stored.transform_for(mapped_column_name)
+    transform = transform_spec.value_transform
+    # `is_transform_active` rather than a query-time restatement of the same
+    # rules: it *is* `validate_transform` with the messages discarded, so the 
two
+    # cannot drift. The difference that matters is the non-deterministic
+    # denylist, which a Tier-2-only check omits -- and `CreateDatasetCommand`,
+    # import, and the editor's free-text Extra box all reach this storage 
without
+    # passing the save-time validation, so a transform calling `now()` would
+    # otherwise be mirrored and its result cached.
+    if not is_transform_active(transform, datasource.database.backend):
+        return None
+
+    if _has_active_advanced_data_type(columns_by_name[mapped_column_name]):
+        # `translate_filter` builds its own predicate shape from *translated*
+        # values, so the `(operator, value)` pair the operator matrix reasons
+        # about does not exist and mirroring would apply the wrong values.
+        return None
+
+    return PartitionMapping(
+        partition_column=str(partition_column),
+        mapped_column=str(mapped_column_name),
+        value_transform=cast(str, transform),
+        is_monotonic=transform_spec.is_monotonic,
+    )
+
+
+def _has_active_advanced_data_type(column: TableColumn) -> bool:
+    advanced_data_type = getattr(column, "advanced_data_type", None)
+    if not advanced_data_type:
+        return False
+    if not 
feature_flag_manager.is_feature_enabled("ENABLE_ADVANCED_DATA_TYPES"):
+        return False
+    return advanced_data_type in app.config.get("ADVANCED_DATA_TYPES", {})
+
+
+def build_probe_sql(
+    transform: str,
+    values: list[Any],
+    dialect: Dialect | None = None,
+) -> str:
+    """
+    Compile a single ``SELECT`` that evaluates the transform at every value.
+
+    Values are attacker-controlled (a Gamma user picks filter values), so they
+    are bound as parameters and rendered by the dialect's own literal processor
+    rather than interpolated into the SQL text.
+
+    Note this deliberately does *not* go through ``BaseEngineSpec``'s text
+    helper, which escapes ``:`` on every engine but Athena and would destroy 
the
+    ``:value`` placeholder before it can be bound.
+    """
+    selections = []
+    for index, value in enumerate(values):
+        clause = sa.text(transform).bindparams(sa.bindparam("value", 
value=value))
+        compiled = clause.compile(
+            dialect=dialect,
+            compile_kwargs={"literal_binds": True},
+        )
+        selections.append(f"{compiled} AS v{index}")
+    return "SELECT " + ", ".join(selections)
+
+
+def evaluate_transform(
+    database: Database,
+    catalog: str | None,
+    schema: str | None,
+    transform: str,
+    values: list[Any],
+) -> list[Any] | None:
+    """
+    Evaluate ``transform`` against the engine once per distinct value.
+
+    Returns one result per input value, positionally aligned with ``values``, 
or
+    ``None`` if anything at all goes wrong. Failing open costs pruning, never
+    correctness: the chart query still runs, it just scans more partitions.
+
+    The probe is pinned to the dataset's catalog and schema so session settings
+    match the chart query as closely as the connection pool allows. It still
+    runs in a *different* session, which is why transforms calling
+    session-dependent functions are rejected at save time.
+    """
+    if not values:
+        return None
+
+    # Dedupe so a 200-value `IN` list costs one column, not 200.
+    distinct: list[Any] = []
+    seen: set[Any] = set()
+    for value in values:
+        key = _hashable(value)
+        if key not in seen:
+            seen.add(key)
+            distinct.append(value)
+
+    cache_key = _probe_cache_key(database, catalog, schema, transform, 
distinct)
+    cached = _cache_get(cache_key)
+    if cached is None:
+        cached = _run_probe(database, catalog, schema, transform, distinct)
+        if cached is None:
+            # Deliberately not cached: a transient engine blip would otherwise
+            # keep the dataset pruning-free for the whole cache timeout.
+            return None
+        _cache_set(cache_key, cached)
+
+    evaluated = dict(
+        zip((_hashable(value) for value in distinct), cached, strict=False)
+    )
+    return [evaluated[_hashable(value)] for value in values]
+
+
+def _run_probe(
+    database: Database,
+    catalog: str | None,
+    schema: str | None,
+    transform: str,
+    distinct: list[Any],
+) -> list[Any] | None:
+    try:
+        sql = build_probe_sql(transform, distinct, _dialect_for(database))
+        frame = database.get_df(sql=sql, catalog=catalog, schema=schema)
+        if frame is None or frame.empty:
+            logger.warning(
+                "Partition transform probe returned no rows; skipping 
mirroring"
+            )
+            return None
+        row = frame.iloc[0]

Review Comment:
   `lower(:value), 'x'` passes the current parse check, but probing `IN ('US', 
'CA')` returns `('us', 'x', 'ca', 'x')`; taking the first two columns produces 
a partition filter that drops the matching CA rows. Can validation require 
exactly one expression and the probe reject any result whose column count 
differs from the number of inputs?



##########
superset/mcp_service/common/schema_discovery.py:
##########
@@ -365,6 +365,25 @@ def get_columns_from_model(
 # UUID search separately and converts it to an exact ``uuid`` filter.
 DATASET_SEARCH_COLUMNS = ["table_name", "description", "schema", "sql"]
 DATASET_EXTRA_COLUMNS: dict[str, ColumnMetadata] = {
+    # Partition filter mapping. Listed here rather than in
+    # `_COLUMN_DESCRIPTIONS` because `get_columns_from_model` walks
+    # `mapper.columns`, and these are read-only properties over the mapping
+    # store, not columns -- a description there would never be reached.
+    "partition_column": ColumnMetadata(
+        name="partition_column",
+        description="Physical column the engine partitions on",
+        type="str",
+        is_default=False,
+    ),

Review Comment:
   Agreed—the new selectable partition fields are absent from `DatasetInfo`, so 
`get_dataset_info` drops them when validating its output. Can the response 
schema carry both fields as well?



##########
superset/datasets/schemas.py:
##########
@@ -101,6 +192,22 @@ class DatasetColumnsPutSchema(Schema):
     datetime_format = fields.String(
         allow_none=True, validate=[Length(1, 100), validate_python_date_format]
     )
+    partition_value_transform = fields.String(
+        allow_none=True,
+        metadata={
+            "description": (
+                "SQL expression containing a :value placeholder. Filters on "
+                "this column are mirrored onto the dataset's partition column "
+                "with the value passed through this transform."
+            )
+        },
+    )

Review Comment:
   Agreed—a typed transform longer than 4096 characters can be accepted and 
stored, then `_as_transform` reads it back as unset and pruning silently stops. 
Can the typed field enforce the same length bound as the Extra schema and 
storage reader?



##########
superset-frontend/src/components/Datasource/DatasourceModal/index.tsx:
##########
@@ -183,6 +183,8 @@ const DatasourceModal: 
FunctionComponent<DatasourceModalProps> = ({
       currency_code_column: datasource.currency_code_column ?? null,
       normalize_columns: datasource.normalize_columns,
       always_filter_main_dttm: datasource.always_filter_main_dttm,
+      partition_column: datasource.partition_column ?? null,
+      partition_mapped_column: datasource.partition_mapped_column ?? null,

Review Comment:
   Agreed—on a dataset with no existing mapping, adding one in the documented 
Extra box still sends null table fields and null column transforms, and the DAO 
patch clears the newly entered mapping on a successful save. Can this editor 
avoid resending stale typed mapping values when Extra is the source of the edit?



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to