[
https://issues.apache.org/jira/browse/FLINK-20672?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Flink Jira Bot updated FLINK-20672:
-----------------------------------
Labels: auto-deprioritized-major stale-minor (was:
auto-deprioritized-major)
I am the [Flink Jira Bot|https://github.com/apache/flink-jira-bot/] and I help
the community manage its development. I see this issues has been marked as
Minor but is unassigned and neither itself nor its Sub-Tasks have been updated
for 180 days. I have gone ahead and marked it "stale-minor". If this ticket is
still Minor, please either assign yourself or give an update. Afterwards,
please remove the label or in 7 days the issue will be deprioritized.
> notifyCheckpointAborted RPC failure can fail JM
> -----------------------------------------------
>
> Key: FLINK-20672
> URL: https://issues.apache.org/jira/browse/FLINK-20672
> Project: Flink
> Issue Type: Bug
> Components: Runtime / Checkpointing
> Affects Versions: 1.11.3, 1.12.0
> Reporter: Roman Khachatryan
> Priority: Minor
> Labels: auto-deprioritized-major, stale-minor
>
> Introduced in FLINK-8871, aborted RPC notifications are done asynchonously:
>
> {code}
> private void sendAbortedMessages(long checkpointId, long timeStamp) {
> // send notification of aborted checkpoints asynchronously.
> executor.execute(() -> {
> // send the "abort checkpoint" messages to necessary
> vertices.
> // ..
> });
> }
> {code}
> However, the executor that eventually executes this request is created as
> follows
> {code}
> final ScheduledExecutorService futureExecutor =
> Executors.newScheduledThreadPool(
> Hardware.getNumberCPUCores(),
> new ExecutorThreadFactory("jobmanager-future"));
> {code}
> ExecutorThreadFactory uses UncaughtExceptionHandler that exits JVM on error.
> cc: [~yunta]
--
This message was sent by Atlassian Jira
(v8.20.1#820001)