David Anderson created FLINK-40368:
--------------------------------------

             Summary: Deduplicate output should only be insert-only when 
ordered by time attribute
                 Key: FLINK-40368
                 URL: https://issues.apache.org/jira/browse/FLINK-40368
             Project: Flink
          Issue Type: Bug
          Components: Table SQL / Planner
    Affects Versions: 2.1.3, 2.2.1, 2.3.0
            Reporter: David Anderson


FLINK-37005 went too far, and the planner now incorrectly treats deduplication 
with ordering ASC on an ordering column that not a time attribute as producing 
an insert-only result. 

Claude suggests this approach for fixing this:

Make the inference guard match the translation guard — add the 
{{sortOnTimeAttributeOnly}} condition at 
FlinkChangelogModeInferenceProgram.scala:264 so non-time top-1 ranks fall into 
the existing {{ALL_CHANGES}} branch at line 286. One wrinkle: you can't simply 
call {{canConvertToDeduplicate}} there, because it calls 
{{{}ChangelogPlanUtils.inputInsertOnly{}}}, which reads a trait that isn't set 
yet during inference — the same trap documented at 
FlinkRelMdModifiedMonotonicity.scala:252. So expose 
{{RankUtil.sortOnTimeAttributeOnly}} (currently private) and combine it with 
the locally computed {{insertOnly}} from {{{}children{}}}.

Test gap worth closing at the same time: {{DeduplicateTest}} only covers 
{{proctime/rowtime/PROCTIME()}} ordering — there is no case with a plain 
column. Relatedly, {{StreamPhysicalRank.getDeduplicateDescription}} hardcodes 
{{{}order=[PROCTIME|ROWTIME]{}}}, which is another sign the dedup path assumes 
a time attribute throughout.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to