David Anderson created FLINK-40368:
--------------------------------------
Summary: Deduplicate output should only be insert-only when
ordered by time attribute
Key: FLINK-40368
URL: https://issues.apache.org/jira/browse/FLINK-40368
Project: Flink
Issue Type: Bug
Components: Table SQL / Planner
Affects Versions: 2.1.3, 2.2.1, 2.3.0
Reporter: David Anderson
FLINK-37005 went too far, and the planner now incorrectly treats deduplication
with ordering ASC on an ordering column that not a time attribute as producing
an insert-only result.
Claude suggests this approach for fixing this:
Make the inference guard match the translation guard — add the
{{sortOnTimeAttributeOnly}} condition at
FlinkChangelogModeInferenceProgram.scala:264 so non-time top-1 ranks fall into
the existing {{ALL_CHANGES}} branch at line 286. One wrinkle: you can't simply
call {{canConvertToDeduplicate}} there, because it calls
{{{}ChangelogPlanUtils.inputInsertOnly{}}}, which reads a trait that isn't set
yet during inference — the same trap documented at
FlinkRelMdModifiedMonotonicity.scala:252. So expose
{{RankUtil.sortOnTimeAttributeOnly}} (currently private) and combine it with
the locally computed {{insertOnly}} from {{{}children{}}}.
Test gap worth closing at the same time: {{DeduplicateTest}} only covers
{{proctime/rowtime/PROCTIME()}} ordering — there is no case with a plain
column. Relatedly, {{StreamPhysicalRank.getDeduplicateDescription}} hardcodes
{{{}order=[PROCTIME|ROWTIME]{}}}, which is another sign the dedup path assumes
a time attribute throughout.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)