[
https://issues.apache.org/jira/browse/NIFI-16248?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Sevag Keosseian updated NIFI-16248:
-----------------------------------
Description:
Connector FlowFile purge is an asynchronous operation replicated across all
nodes in a NiFi cluster. The Connector purge endpoints currently have no
dedicated cluster response merger.
As a result, the cluster coordinator can return an arbitrary node's purge
response instead of a cluster-wide response. If the selected node finishes
before another node, the client can observe the purge as finished and attempt
to delete the Connector while another node still has queued FlowFiles.
This causes intermittent Connector deletion failures such as:
Cannot delete Connector because its Process Group has FlowFiles queued; all
FlowFiles must be removed before it can be deleted.
The issue was exposed by
ClusteredConnectorBacklogIT.testClusteredReadQueueAttributesLoadBalancedConnection
during test teardown, but the underlying problem applies to any clustered
client using the Connector purge endpoints.
Sample failed run:
https://github.com/apache/nifi/actions/runs/32513358371/job/96902345347
was:
Connector FlowFile purge is an asynchronous operation replicated across all
nodes in a NiFi cluster. The Connector purge endpoints currently have no
dedicated cluster response merger.
As a result, the cluster coordinator can return an arbitrary node's purge
response instead of a cluster-wide response. If the selected node finishes
before another node, the client can observe the purge as finished and attempt
to delete the Connector while another node still has queued FlowFiles.
This causes intermittent Connector deletion failures such as:
Cannot delete Connector because its Process Group has FlowFiles queued; all
FlowFiles must be removed before it can be deleted.
The issue was exposed by
ClusteredConnectorBacklogIT.testClusteredReadQueueAttributesLoadBalancedConnection
during test teardown, but the underlying problem applies to any clustered
client using the Connector purge endpoints.
> Clustered Connector purge requests can report completion before all nodes
> finish
> --------------------------------------------------------------------------------
>
> Key: NIFI-16248
> URL: https://issues.apache.org/jira/browse/NIFI-16248
> Project: Apache NiFi
> Issue Type: Bug
> Components: Core Framework
> Reporter: Sevag Keosseian
> Assignee: Sevag Keosseian
> Priority: Minor
> Labels: Cluster, Connector, system-test-failure
>
> Connector FlowFile purge is an asynchronous operation replicated across all
> nodes in a NiFi cluster. The Connector purge endpoints currently have no
> dedicated cluster response merger.
> As a result, the cluster coordinator can return an arbitrary node's purge
> response instead of a cluster-wide response. If the selected node finishes
> before another node, the client can observe the purge as finished and attempt
> to delete the Connector while another node still has queued FlowFiles.
> This causes intermittent Connector deletion failures such as:
> Cannot delete Connector because its Process Group has FlowFiles queued; all
> FlowFiles must be removed before it can be deleted.
> The issue was exposed by
> ClusteredConnectorBacklogIT.testClusteredReadQueueAttributesLoadBalancedConnection
> during test teardown, but the underlying problem applies to any clustered
> client using the Connector purge endpoints.
> Sample failed run:
> https://github.com/apache/nifi/actions/runs/32513358371/job/96902345347
--
This message was sent by Atlassian Jira
(v8.20.10#820010)