[ 
https://issues.apache.org/jira/browse/IGNITE-23551?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17896199#comment-17896199
 ] 

Ashu Pachauri commented on IGNITE-23551:
----------------------------------------

{quote}Is this happening because the affected (does not rejoin the cluster 
after restart) node actually does not host any 'redis-ignite-internal-cache-0' 
partitions?
{quote}
I am not sure yet about the root cause. However, this is not the reason. I can 
see 'redis-ignite-internal-cache-0' in the persistent storage directory.

 

> Restarted node fails with NullPointerException
> ----------------------------------------------
>
>                 Key: IGNITE-23551
>                 URL: https://issues.apache.org/jira/browse/IGNITE-23551
>             Project: Ignite
>          Issue Type: Bug
>          Components: cache
>    Affects Versions: 2.15, 2.16
>         Environment: OS: Ubuntu/debian
> Java: Openjdk version "17.0.13" 2024-10-15
>            Reporter: Ashu Pachauri
>            Priority: Major
>         Attachments: ignite-config.xml, ignite.log
>
>
> We are using Ignite as a persistant caching system primarily to write KVs 
> using the redis interface; we define redis caches statically in the xml 
> config. 
> We have been plagued by an issue where restarting a node in an existing 
> stable cluster does not work and the node fails every time trying to join the 
> cluster giving a NullPointerException. This happens with any and every node 
> in the cluster and persists no matter how many times the node is started up.  
> After a full cluster restart the issue goes away. 
>  
> Following is the stacktrace we see in the logs of the failed node:
> {code:java}
> [10:51:34,769][SEVERE][tcp-disco-msg-worker-[fa915882 
> 10.132.0.114:47500]-#2-#57][TcpDiscoverySpi] TcpDiscoverSpi's message worker 
> thread failed abnormally. S
> topping the node in order to prevent cluster wide instability.
> java.lang.NullPointerException: Cannot invoke 
> "org.apache.ignite.internal.managers.discovery.GridDiscoveryManager$CachePredicate.addClientNode(java.util.UUID,
>  boolean)" because "p" is null
>         at 
> org.apache.ignite.internal.managers.discovery.GridDiscoveryManager.addClientNode(GridDiscoveryManager.java:428)
>         at 
> org.apache.ignite.internal.processors.cache.ClusterCachesInfo.addReceivedClientNodesToDiscovery(ClusterCachesInfo.java:1600)
>         at 
> org.apache.ignite.internal.processors.cache.ClusterCachesInfo.onGridDataReceived(ClusterCachesInfo.java:1519)
>         at 
> org.apache.ignite.internal.processors.cache.GridCacheProcessor.onGridDataReceived(GridCacheProcessor.java:3137)
>         at 
> org.apache.ignite.internal.managers.discovery.GridDiscoveryManager$4.onExchange(GridDiscoveryManager.java:1019)
>         at 
> org.apache.ignite.spi.discovery.tcp.TcpDiscoverySpi.onExchange(TcpDiscoverySpi.java:2197)
>         at 
> org.apache.ignite.spi.discovery.tcp.ServerImpl$RingMessageWorker.processNodeAddFinishedMessage(ServerImpl.java:5359)
>         at 
> org.apache.ignite.spi.discovery.tcp.ServerImpl$RingMessageWorker.processMessage(ServerImpl.java:3242)
>         at 
> org.apache.ignite.spi.discovery.tcp.ServerImpl$RingMessageWorker.processMessage(ServerImpl.java:2918)
>         at 
> org.apache.ignite.spi.discovery.tcp.ServerImpl$MessageWorker.body(ServerImpl.java:8048)
>         at 
> org.apache.ignite.spi.discovery.tcp.ServerImpl$RingMessageWorker.body(ServerImpl.java:3089)
>         at 
> org.apache.ignite.internal.util.worker.GridWorker.run(GridWorker.java:125)
>         at 
> org.apache.ignite.spi.discovery.tcp.ServerImpl$MessageWorkerThread.body(ServerImpl.java:7979)
>         at org.apache.ignite.spi.IgniteSpiThread.run(IgniteSpiThread.java:58) 
> {code}
> Attaching the config and logs for a test cluster for reference.
> [^ignite-config.xml]
> [^ignite.log]



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to