[
https://issues.apache.org/jira/browse/IGNITE-23551?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17896353#comment-17896353
]
Jeremy McMillan commented on IGNITE-23551:
------------------------------------------
{quote}This causes the NPE to be triggered as the cache predicate for this
cache is null on the restarted node.
{quote}
Test code indicates this should never be null. Maybe tracing this back will
bear fruit.
https://github.com/apache/ignite/blob/7b1a4a1fe72110df4bf6623d87dd764b57302085/modules/core/src/main/java/org/apache/ignite/internal/managers/discovery/GridDiscoveryManager.java#L3422
> Restarted node fails with NullPointerException
> ----------------------------------------------
>
> Key: IGNITE-23551
> URL: https://issues.apache.org/jira/browse/IGNITE-23551
> Project: Ignite
> Issue Type: Bug
> Components: cache
> Affects Versions: 2.15, 2.16
> Environment: OS: Ubuntu/debian
> Java: Openjdk version "17.0.13" 2024-10-15
> Reporter: Ashu Pachauri
> Priority: Major
> Attachments: ignite-config.xml, ignite.log
>
>
> We are using Ignite as a persistant caching system primarily to write KVs
> using the redis interface; we define redis caches statically in the xml
> config.
> We have been plagued by an issue where restarting a node in an existing
> stable cluster does not work and the node fails every time trying to join the
> cluster giving a NullPointerException. This happens with any and every node
> in the cluster and persists no matter how many times the node is started up.
> After a full cluster restart the issue goes away.
>
> Following is the stacktrace we see in the logs of the failed node:
> {code:java}
> [10:51:34,769][SEVERE][tcp-disco-msg-worker-[fa915882
> 10.132.0.114:47500]-#2-#57][TcpDiscoverySpi] TcpDiscoverSpi's message worker
> thread failed abnormally. S
> topping the node in order to prevent cluster wide instability.
> java.lang.NullPointerException: Cannot invoke
> "org.apache.ignite.internal.managers.discovery.GridDiscoveryManager$CachePredicate.addClientNode(java.util.UUID,
> boolean)" because "p" is null
> at
> org.apache.ignite.internal.managers.discovery.GridDiscoveryManager.addClientNode(GridDiscoveryManager.java:428)
> at
> org.apache.ignite.internal.processors.cache.ClusterCachesInfo.addReceivedClientNodesToDiscovery(ClusterCachesInfo.java:1600)
> at
> org.apache.ignite.internal.processors.cache.ClusterCachesInfo.onGridDataReceived(ClusterCachesInfo.java:1519)
> at
> org.apache.ignite.internal.processors.cache.GridCacheProcessor.onGridDataReceived(GridCacheProcessor.java:3137)
> at
> org.apache.ignite.internal.managers.discovery.GridDiscoveryManager$4.onExchange(GridDiscoveryManager.java:1019)
> at
> org.apache.ignite.spi.discovery.tcp.TcpDiscoverySpi.onExchange(TcpDiscoverySpi.java:2197)
> at
> org.apache.ignite.spi.discovery.tcp.ServerImpl$RingMessageWorker.processNodeAddFinishedMessage(ServerImpl.java:5359)
> at
> org.apache.ignite.spi.discovery.tcp.ServerImpl$RingMessageWorker.processMessage(ServerImpl.java:3242)
> at
> org.apache.ignite.spi.discovery.tcp.ServerImpl$RingMessageWorker.processMessage(ServerImpl.java:2918)
> at
> org.apache.ignite.spi.discovery.tcp.ServerImpl$MessageWorker.body(ServerImpl.java:8048)
> at
> org.apache.ignite.spi.discovery.tcp.ServerImpl$RingMessageWorker.body(ServerImpl.java:3089)
> at
> org.apache.ignite.internal.util.worker.GridWorker.run(GridWorker.java:125)
> at
> org.apache.ignite.spi.discovery.tcp.ServerImpl$MessageWorkerThread.body(ServerImpl.java:7979)
> at org.apache.ignite.spi.IgniteSpiThread.run(IgniteSpiThread.java:58)
> {code}
> Attaching the config and logs for a test cluster for reference.
> [^ignite-config.xml]
> [^ignite.log]
--
This message was sent by Atlassian Jira
(v8.20.10#820010)