MartijnVisser commented on code in PR #29373:
URL: https://github.com/apache/flink/pull/29373#discussion_r4193916286


##########
flink-table/flink-table-runtime/src/main/java/org/apache/flink/table/runtime/operators/over/NonTimeRowsUnboundedPrecedingFunction.java:
##########
@@ -273,21 +273,25 @@ protected void processRemainingElements(
             // to comply with the sql rows syntax
             for (int j = 0; j < ids.size(); j++) {
                 RowData value = valueMapState.get(ids.get(j));
+                RowData prevAcc = 
accMapState.get(GenericRowData.of(ids.get(j)));
                 aggFuncs.accumulate(value);
-                RowData accData = 
accMapState.get(GenericRowData.of(ids.get(j)));
+                RowData newAcc = aggFuncs.getAccumulators();
                 // Logic to early out
                 // TODO: Move comparison to function i.e. canEarlyOut(prev, 
curr)
-                if (aggFuncs.getValue().equals(accData)) {
+                if (newAcc.equals(prevAcc)) {

Review Comment:
   On heap, `ARRAY_AGG`, `COLLECT` and `PERCENTILE` now fail here with the 
FLINK-40735 error instead of the cast error. That one is left to FLINK-40735, 
which will need to change this comparison too.



##########
flink-table/flink-table-runtime/src/test/java/org/apache/flink/table/runtime/operators/over/NonTimeRowsUnboundedPrecedingFunctionTest.java:
##########
@@ -96,6 +96,58 @@ void testInsertOnlyRecordsWithCustomSortKey() throws 
Exception {
         validateRows(actualRows, expectedRows);
     }
 
+    @Test
+    void testInsertWithDuplicateSortKeyAndLastValueAgg() throws Exception {
+        KeyedProcessOperator<RowData, RowData, RowData> operator =
+                new KeyedProcessOperator<>(
+                        new NonTimeRowsUnboundedPrecedingFunction<RowData>(
+                                0L,
+                                lastValueAggsHandleFunction,
+                                GENERATED_ROW_VALUE_EQUALISER,
+                                GENERATED_SORT_KEY_EQUALISER,
+                                GENERATED_SORT_KEY_COMPARATOR_ASC,
+                                lastValueAccTypes,
+                                inputFieldTypes,
+                                SORT_KEY_TYPES,
+                                SORT_KEY_SELECTOR) {});
+
+        OneInputStreamOperatorTestHarness<RowData, RowData> testHarness =
+                createTestHarness(operator);
+        testHarness.open();
+
+        testHarness.processElement(insertRecord("key1", 1L, 100L));
+        testHarness.processElement(insertRecord("key1", 2L, 200L));
+        testHarness.processElement(insertRecord("key1", 5L, 500L));
+        testHarness.processElement(insertRecord("key1", 6L, 600L));
+        testHarness.processElement(insertRecord("key1", 4L, 400L));
+        testHarness.processElement(insertRecord("key1", 5L, 503L));
+        testHarness.processElement(updateBeforeRecord("key1", 5L, 500L));

Review Comment:
   This only runs on heap, where the delete-path change goes unnoticed. With 
`EmbeddedRocksDBStateBackend` it fails at the `put` for both methods, once 
`TestRowValueEqualiser` handles `BinaryRowData`. Can you add a RocksDB run?



##########
flink-table/flink-table-runtime/src/main/java/org/apache/flink/table/runtime/operators/over/NonTimeRowsUnboundedPrecedingFunction.java:
##########
@@ -402,19 +406,27 @@ private RowData getPreviousAccumulator(
      */
     private void reAccumulateIdsAndEmitUpdates(
             List<Long> ids, int removeIndex, Collector<RowData> out) throws 
Exception {
+        RowData baseAcc = aggFuncs.getAccumulators();
         for (int j = removeIndex; j < ids.size(); j++) {
             RowData value = valueMapState.get(ids.get(j));
             if (j == removeIndex) {
-                collectDelete(out, value, 
accMapState.get(GenericRowData.of(ids.get(j))));
+                RowData deletedAcc = 
accMapState.get(GenericRowData.of(ids.get(j)));
+                collectDelete(out, value, 
setAccumulatorAndGetValue(deletedAcc));
+                aggFuncs.setAccumulators(baseAcc);
             } else {
+                RowData prevAcc = 
accMapState.get(GenericRowData.of(ids.get(j)));
                 aggFuncs.accumulate(value);
+                RowData newAcc = aggFuncs.getAccumulators();
                 // Logic to early out
-                if 
(aggFuncs.getValue().equals(accMapState.get(GenericRowData.of(ids.get(j))))) {
+                if (newAcc.equals(prevAcc)) {
                     break;

Review Comment:
   Pre-existing, but after this `break` the following sort keys restart from 
this id's accumulator. With `SUM(ts)` on heap over rows (1,0), (1,5), (1,7), 
(2,1), deleting the first one turns 13 into 6. Fix here or in a follow-up?



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to