Mark,thanks for your response. We hit the same problem last night ( before making any of your suggested changes ). Thankfully this time we did a whole load of analysis that may be of use.
We had one data node that got a java heap size error and left the pool of our elasticsearch cluster. We looked at this one node and found a few things. 1) It had a heavy load prior to it becoming unusable. 2) It started to garbage collect often and frequently. 3) The problematic node had high cpu and garbage collected every 30 seconds or so for 7 hours before we finally saw a heap size error in our logs, at the point the node was unresponsive and it left the cluster. 4) The node contained shards of our largest indices. The interesting and worrying thing was that during this 7 hour period the elasticsearch cluster itself was unusable ( not accepting any reads) and only recovered once the node had left. We think that the cluster should have recovered sooner perhaps and not got into such a state. Anyone else seeing similar issues? Dip -- You received this message because you are subscribed to the Google Groups "elasticsearch" group. To unsubscribe from this group and stop receiving emails from it, send an email to [email protected]. To view this discussion on the web visit https://groups.google.com/d/msgid/elasticsearch/39b418ac-bc87-4d6d-ad92-ca85da79fdbf%40googlegroups.com. For more options, visit https://groups.google.com/d/optout.
