We've had a very similar issue, but haven't been able to figure out what the problem is. How do you "fix" the problem? Will a node restart fix the problem immediately or do you need to restart the whole machine?
On Tuesday, July 29, 2014 1:59:52 PM UTC-7, [email protected] wrote: > > Hey guys, > > We've been running a 5 node cluster for our index (5 shards, 1 replica, > evenly distributed on 5 nodes), and are running into a problem with one of > the nodes in the cluster. It is not unique to any specific node, and can > happen sporadically on any of the nodes. > > One of the machines starts spiking up close to 100% CPU Load, and close to > 8 OS Load (which is amusing, considering there are only 4 CPU cores on the > machine), while all the other machines operate normally way below those > figures. Naturally, this behavior is accompanied by extremely high write > times, and read times, as well. > > Here's what Marvel looks like: > > > <https://lh5.googleusercontent.com/-bxUFPhqAnVk/U9gKg4c19nI/AAAAAAAAABE/S_w68vZ63Uo/s1600/Marvel+-+Node+Statistics.png> > > Here's all the information we could gather: > > - Full thread dump from while this occurred: > https://gist.github.com/danielschonfeld/ff6c3744197f2c748632 > - GET _nodes/stats: > https://gist.github.com/schonfeld/693c8dbf0dd57e4cff7c > - GET _nodes/hot_threads: > https://gist.github.com/schonfeld/766d771d211e452a7100 > - GET _cluster/stats: > https://gist.github.com/schonfeld/d5395f97e3a87745cc1f > > > Thoughts? insights? Any clues would be greatly appreciated. > -- You received this message because you are subscribed to the Google Groups "elasticsearch" group. To unsubscribe from this group and stop receiving emails from it, send an email to [email protected]. To view this discussion on the web visit https://groups.google.com/d/msgid/elasticsearch/6660d09b-98f9-41c4-87d0-9ee56890c7b9%40googlegroups.com. For more options, visit https://groups.google.com/d/optout.
