On Fri, 11 Dec 2009, Adam Nielsen wrote:

>> No errors that I can find, omsa doesn't obviously show anything as degraded.
>
> Do you have a utility like the old megamgr that can get you controller
> stats?  That will report actual disk errors that are seen by the RAID
> controller, which may not make it all the way through to the OS.

No, I was relying on omsa.  I'll take a look at what else I can use.

> Do you mean that when there is a 'stall', you can still access the disk
> locally?  In your original e-mail you suggested this wasn't the case.
> It's probably worth confirming which it is, because that would indicate
> it is either an NFS or a hardware issue.

The effect is felt locally, but so far it's only been triggered by a
particular source of NFS traffic.  So no, local i/o is buggered when it's in
this state/

>> I need to find a reliable way of triggering this behaviour.
>>
>> A looping once-per-second sync, made the machine reliably available,
>> although I suspect this was just treating the symptoms.
>
> This doesn't seem to suggest a hardware problem.  It's possible that
> your disk caches are just set too large, and once it reaches a critical
> point they get flushed to disk, but because they're so big the disks
> can't keep up.  Imagine if you had a 1GB writeback cache which suddenly
> the kernel decided needed to be flushed ASAP - that 1GB of disk write
> activity could cause all programs doing disk IO to stall for a few
> seconds until it was complete.

Sure.

> If a once-per-second sync avoids the issue, I would suggest shrinking
> your disk writeback cache.

Specifically what are you suggesting I adjust?  I'd already had a poke through
/proc and not acheived anything useful.

Thanks for the pointers,

jh

_______________________________________________
Linux-PowerEdge mailing list
[email protected]
https://lists.us.dell.com/mailman/listinfo/linux-poweredge
Please read the FAQ at http://lists.us.dell.com/faq

Reply via email to