On Fri, 11 Dec 2009, Adam Nielsen wrote: >> No errors that I can find, omsa doesn't obviously show anything as degraded. > > Do you have a utility like the old megamgr that can get you controller > stats? That will report actual disk errors that are seen by the RAID > controller, which may not make it all the way through to the OS.
No, I was relying on omsa. I'll take a look at what else I can use. > Do you mean that when there is a 'stall', you can still access the disk > locally? In your original e-mail you suggested this wasn't the case. > It's probably worth confirming which it is, because that would indicate > it is either an NFS or a hardware issue. The effect is felt locally, but so far it's only been triggered by a particular source of NFS traffic. So no, local i/o is buggered when it's in this state/ >> I need to find a reliable way of triggering this behaviour. >> >> A looping once-per-second sync, made the machine reliably available, >> although I suspect this was just treating the symptoms. > > This doesn't seem to suggest a hardware problem. It's possible that > your disk caches are just set too large, and once it reaches a critical > point they get flushed to disk, but because they're so big the disks > can't keep up. Imagine if you had a 1GB writeback cache which suddenly > the kernel decided needed to be flushed ASAP - that 1GB of disk write > activity could cause all programs doing disk IO to stall for a few > seconds until it was complete. Sure. > If a once-per-second sync avoids the issue, I would suggest shrinking > your disk writeback cache. Specifically what are you suggesting I adjust? I'd already had a poke through /proc and not acheived anything useful. Thanks for the pointers, jh _______________________________________________ Linux-PowerEdge mailing list [email protected] https://lists.us.dell.com/mailman/listinfo/linux-poweredge Please read the FAQ at http://lists.us.dell.com/faq
