yujun777 commented on code in PR #68787:
URL: https://github.com/apache/doris/pull/68787#discussion_r4225691973


##########
fe/fe-core/src/main/java/org/apache/doris/nereids/rules/analysis/IvmNormalizeMTMV.java:
##########
@@ -1069,6 +1071,202 @@ private Map<String, Slot> collectIvmHiddenSlots(Plan 
normalizedChild) {
                 .collect(Collectors.toMap(Slot::getName, slot -> slot, (left, 
right) -> left, LinkedHashMap::new));
     }
 
+    /**
+     * Materializes the aggregate state columns this layer drops, so an 
incremental refresh can still
+     * read the old aggregate state from the MV.
+     *
+     * <p>Every state slot the apply stage reads must be a persisted MV 
column, because apply resolves
+     * the old state by column name from the MV physical table. Hidden state 
columns are hidden-named
+     * and therefore propagate through every layer, but the aggregate 
functions whose own value is
+     * their mergeable state (SUM, COUNT(expr), MIN/MAX, 
COLLECT_LIST/ARRAY_AGG, BITMAP_UNION) keep
+     * that state in their visible column, which disappears as soon as an 
upper layer consumes it
+     * inside an expression without projecting it, as in {@code SELECT SUM(v) 
* 100}. Such a slot is
+     * materialized here as a bare pass-through hidden alias, and every target 
reading it is rebound
+     * to that alias.
+     *
+     * <p>Rebinding matters as much as materializing: the column pool lets a 
target reuse a visible
+     * aggregate column as its own hidden state (AVG reusing a visible SUM 
column), so a reusing target
+     * must follow the column's owner onto the materialized carrier instead of 
reading a column that no
+     * longer reaches the MV.
+     *
+     * <p>The materialized name is the name the delta sub-plan already 
generates for that state
+     * ({@link IvmUtil#ivmAggHiddenColumnName}, keyed by the owning target's 
ordinal and kind), so delta
+     * aggregate outputs and delta slot lookups are unaffected and both sides 
of the merge agree on the
+     * column name. A slot an earlier layer already materialized, or that 
another target already
+     * materialized for the same aggregate state, is reused instead of 
materializing a duplicate.
+     */
+    private List<NamedExpression> 
materializeDroppedAggState(List<NamedExpression> outputs) {
+        IvmAggMeta aggMeta = rewriteResult.getAggMeta();
+        if (aggMeta == null) {
+            // Below the aggregate no target is known yet, so no aggregate 
state can be dropped here.
+            return outputs;
+        }
+        List<NamedExpression> extendedOutputs = new ArrayList<>(outputs);
+        List<IvmAggTarget> reboundTargets = new 
ArrayList<>(aggMeta.getAggTargets().size());
+        boolean rebound = false;
+        // While the layout is being created there is no MV yet (CREATE 
MATERIALIZED VIEW keeps only its
+        // name), so normalize decides the layout itself; a refresh is bound 
by the MV's own schema.
+        MTMV mtmv = statementContext.getIvmRewriteContext().get().getMtmv();
+        for (IvmAggTarget target : aggMeta.getAggTargets()) {
+            Slot valueStateSlot = valueStateSlotFor(target, mtmv, 
extendedOutputs, aggMeta);
+            ImmutableMap.Builder<IvmAggStateKey, Slot> hiddenStateSlots = 
ImmutableMap.builder();
+            for (Map.Entry<IvmAggStateKey, Slot> hiddenStateSlot : 
target.getHiddenStateSlots().entrySet()) {
+                hiddenStateSlots.put(hiddenStateSlot.getKey(),
+                        materializeAggStateSlot(extendedOutputs, 
hiddenStateSlot.getValue(), aggMeta));
+            }
+            IvmAggTarget reboundTarget = target.withStateSlots(valueStateSlot, 
hiddenStateSlots.build());
+            rebound |= reboundTarget != target;
+            reboundTargets.add(reboundTarget);
+        }
+        if (rebound) {
+            // Keep the rebinding visible to the layers above: a layer that 
passes a state column through
+            // under a different slot (the refresh sink rebinds the normalized 
hidden columns to the MV's
+            // own slots) changes which slot carries the state without adding 
any column.
+            rewriteResult.setAggMeta(aggMeta.withAggTargets(reboundTargets));
+        }
+        return extendedOutputs;
+    }
+
+    /**
+     * Returns the slot this target's own value state is read from, 
materializing a carrier column when the
+     * state needs a column the layout does not have yet.
+     *
+     * <p>When {@code mtmv} is null the layout is being created (CREATE 
MATERIALIZED VIEW knows only its
+     * name): the state is the visible aggregate column when that column 
survives into the MV, and a
+     * materialized carrier otherwise. On a refresh {@code mtmv} exists and 
its schema owns the layout: the
+     * state is the carrier only when the MV really has that column, and 
otherwise the MV's visible column
+     * carries it. That keeps the refresh path independent of the binder's 
insert projects, which rename a
+     * state column locally and rebuild it as the MV column under its own 
name, so a user column that
+     * happens to be named like the aggregate's generated alias never becomes 
the state.
+     */
+    private Slot valueStateSlotFor(IvmAggTarget target, MTMV mtmv, 
List<NamedExpression> outputs,
+            IvmAggMeta aggMeta) {
+        if (!aggFunctionRegistry.visibleColumnHoldsValueState(target)) {
+            // AVG, BITMAP_UNION_COUNT and COUNT(*) merge hidden state or the 
group count instead.
+            return target.getValueStateSlot();
+        }
+        // The column currently carrying this target's value: the carrier 
materialized by a lower layer if
+        // there is one, otherwise the visible column the aggregate still 
produces.
+        Slot stateSlot = target.getValueStateSlot() != null
+                ? target.getValueStateSlot() : target.getVisibleSlot();
+        if (mtmv == null) {
+            // Otherwise the projection decides: the state keeps the column 
the plan projects it as (the
+            // aggregate output itself or a rename of it), and a state the 
projection drops gets a carrier.
+            return materializeAggStateSlot(outputs, stateSlot, aggMeta);
+        }
+        String carrierName = 
IvmUtil.ivmAggHiddenColumnName(target.getOrdinal(),
+                target.getFunctionKind().name());
+        if (mtmv.getColumn(carrierName) != null) {
+            return materializeAggStateSlot(outputs, stateSlot, aggMeta);
+        }
+        // Without a carrier the MV keeps the state in the column the plan 
projected it as. That name is
+        // only usable when the MV really has it: the binder renames a state 
column to the MV's own column
+        // name for an unnamed aggregate (COUNT(*) becomes __count_0), while 
for a clamped key column it
+        // renames the state to a project-local name and rebuilds the MV 
column by coercion, in which case
+        // the MV's visible column carries the state.
+        NamedExpression projected = findProjectedKey(outputs, stateSlot);
+        return projected != null && mtmv.getColumn(projected.getName()) != null
+                ? projected.toSlot() : null;
+    }
+
+    /**
+     * Returns the slot that carries {@code stateSlot} above this layer, 
appending a hidden pass-through
+     * alias when this layer drops it.
+     *
+     * <p>A layer keeps the state alive when it projects the state slot itself 
or a pure rename of it, which
+     * is how the refresh sink maps the normalized hidden columns onto the 
MV's own slots. Only when the
+     * state slot is really gone from the layer's output does it need a 
carrier.
+     */
+    private Slot materializeAggStateSlot(List<NamedExpression> outputs, Slot 
stateSlot, IvmAggMeta aggMeta) {
+        NamedExpression projected = findProjectedKey(outputs, stateSlot);
+        if (projected != null) {
+            // The projecting output's own slot is what carries the value: the 
column itself, a column it
+            // projects through, or a carrier another target needed for this 
same state.
+            return projected.toSlot();
+        }
+        return materializeStateCarrier(outputs, stateSlot, aggMeta);
+    }
+
+    /**
+     * Appends a hidden carrier column for {@code stateSlot} and returns its 
slot.
+     *
+     * <p>The carrier is named after the owning target's ordinal and kind, 
which is the name the delta
+     * sub-plan already generates for the same state, so no delta-side lookup 
changes.
+     */
+    private Slot materializeStateCarrier(List<NamedExpression> outputs, Slot 
stateSlot, IvmAggMeta aggMeta) {
+        IvmAggTarget owner = aggTargetOwningVisibleSlot(stateSlot, aggMeta);
+        Alias carrier = new Alias(stateSlot,
+                IvmUtil.ivmAggHiddenColumnName(owner.getOrdinal(), 
owner.getFunctionKind().name()));

Review Comment:
   The mechanism you describe is real code behaviour (`rewriteIvmHiddenOutput` 
replaces the sink's typed hidden placeholder with the child's bare slot), but 
on this branch a carrier's plan type cannot diverge from its MV column type, 
and I could not reproduce the failure. What I ran against this head, each 
created and refreshed end to end:
   
   1. A legacy `DATE` base column is not constructible here: `CREATE TABLE ... 
d DATE` is rejected with "Disable to create table with `DATE` type columns, 
please use `DATEV2`". IVM base tables additionally have to be binlog/stream 
enabled, which is not released yet, so only new tables are in scope.
   2. Forcing a legacy state slot with an explicit cast, `SELECT k, 
YEAR(MIN(CAST(d AS DATEV1))) AS y FROM dt_base GROUP BY k`: `DESC` shows the 
carrier column as `date`, the incremental refresh succeeds, and the value 
matches the source query (`2024, 2024`, and `2023, 2024` after inserting an 
earlier date).
   3. A legacy temporal column, `dt DATETIMEV1` with `SELECT k, YEAR(MIN(dt)) 
AS y ... GROUP BY k`: the carrier column type is `datetime`, the incremental 
refresh succeeds, and the values match the source query.
   4. The nested case, `SELECT k, ARRAY_SORT(ARRAY_AGG(d)) AS lst ... GROUP BY 
k` over a `DATE` column: the carrier is `array<date>`, the incremental refresh 
succeeds, and the recorded sorted elements match the source query 
(`["2023-05-05", "2024-01-01", "2024-02-01"]`, `["2024-03-01"]`).
   
   So with `enable_date_conversion` at its default the plan slot type, the 
generated MV column type and the sink target agree for these carriers. If you 
have a concrete statement and plan ordering that makes them differ on this 
branch, please share it and I will add the coercion for it.
   



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to