[ 
https://issues.apache.org/jira/browse/FLINK-20654?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17340190#comment-17340190
 ] 

Matthias commented on FLINK-20654:
----------------------------------

I didn't have to analyze the UCITCase. I looked into another issue and was 
initially puzzled by the size of the Maven logs (only one of the two). I'm just 
concerned that the big file size for {{mvn*.log}} would make analyzing any 
issue a bit annoying. Just grepping for a String in the Maven logs takes 
already much longer.

I totally understand the necessity of why we set the log level to trace. I was 
just wondering whether we have some plan to disable it again. I could imagine 
the case where we just don't run into issues with the UC anymore. At what point 
in time do we decide to set the log level back to {{INFO}}? Alternatively, we 
could move those logs into it's own file? Not sure how easy it is with our 
current AzureCI setup.

> Unaligned checkpoint recovery may lead to corrupted data stream
> ---------------------------------------------------------------
>
>                 Key: FLINK-20654
>                 URL: https://issues.apache.org/jira/browse/FLINK-20654
>             Project: Flink
>          Issue Type: Bug
>          Components: Runtime / Checkpointing
>    Affects Versions: 1.12.0, 1.12.1
>            Reporter: Arvid Heise
>            Assignee: Piotr Nowojski
>            Priority: Critical
>              Labels: pull-request-available, test-stability
>             Fix For: 1.13.0, 1.12.3
>
>
> Fix of FLINK-20433 shows potential corruption after recovery for all 
> variations of UnalignedCheckpointITCase.
> To reproduce, run UCITCase a couple hundreds times. The issue showed for me 
> in:
> - execute [Parallel union, p = 5]
> - execute [Parallel union, p = 10]
> - execute [Parallel cogroup, p = 5]
> - execute [parallel pipeline with remote channels, p = 5]
> with decreasing frequency.
> The issue manifests as one of the following issues:
> - stream corrupted exception
> - EOF exception
> - assertion failure in NUM_LOST or NUM_OUT_OF_ORDER
> - (for union) ArithmeticException overflow (because the number that should be 
> [0;100000] has been mis-deserialized)



--
This message was sent by Atlassian Jira
(v8.3.4#803005)

Reply via email to