Chris,

You are the clear-seeing genius they say you are... ;) In anycase, these
have been some of the weirdest hardware moments I've ever had: it seems that
an existing 32MB DIMM was right on the edge of giving up at the moment that
I installed a new 128MB DIMM in the firewall and converted it to reiserfs.

It was on the edge in such a manner that it survived 12 consecutive passes
of complete memtest86 test-sets (and 40 3-parallel process kernel compiles
without a sig11).... this morning however, the BIOS caught a memory error at
bootup (the machine had crashed during the night); a subsequent run of
memtest86 caught the error within the first 30 seconds.  I've established
with some DIMM-swapping that the problem is with the DIMM and not the slot
it was in.

So, the DIMM was in the process of dying as it were... it was being
sporadically flakey at a time when many variables were being added to the
equation, which made the fault all the more difficult to confirm.

Thanks everyone,
Charl

On Mon, Aug 06, 2001 at 10:37:25AM -0400, Chris Mason wrote:
On Thursday, August 02, 2001 10:58:33 AM +0200 "Charl P. Botha"
> <[EMAIL PROTECTED]> wrote:
> > hoping the call-traces (which are symbolic in my attached oopses) would
> > tell somebody something.
> 
> You've got 2 oopses hitting 2 different hash tables, which will generally
> point to ram or the CPU.  Neither spot was involved with pulling stuff off
> disk, so I don't expect it to be related to disk corruption, or the
> drives/controllers. 
> 
> I saw later on that you turned off the unmask_irq option, but if you are
> still having problems I would start swaping CPUs/ram.

-- 
charl p. botha      | computer graphics and cad/cam 
http://cpbotha.net/ | http://www.cg.its.tudelft.nl/

Reply via email to