On Monday 04 February 2008 10:17:37 am Brock Palen wrote: > The > cluster IS to big, but there isn't a person at the university who is > willing to pay for anything other than more cluster nodes. Enough > with politics.
That's the first time I hear a cluster is too big, people usually complain about the contrary. :) But the second part sounds very very familiar, though... Anyway. > I just had another node get evicted while running code causing the > code to lock up. This time it was the MDS that evicted it. Pinging > work though: > > [EMAIL PROTECTED] ~]# lctl ping [EMAIL PROTECTED] > [EMAIL PROTECTED] > [EMAIL PROTECTED] Ok. > I have attached the output of lctl dk from the client and some > syslog messages from the MDS. (recover.c:188:ptlrpc_request_handle_notconn()) import nobackup-MDT0000-mdc-000001012bd27c00 of [EMAIL PROTECTED]@tcp abruptly disconnected: reconnecting (import.c:133:ptlrpc_set_import_discon()) nobackup-MDT0000-mdc-000001012bd27c00: Connection to service nobackup-MDT0000 via nid [EMAIL PROTECTED] was lost; I will let Lustre people comment on this, but this sure looks like a network problem to me. Is there any information you can get out of the switches (logs, dropped packets, retries, stats, anything)? > Nope both servers have 2GB ram, and load is almost 0. No swapping. Do you see dropped packets or errors in your ifconfig output, on the servers and/or clients? Cheers, -- Kilian _______________________________________________ Lustre-discuss mailing list [email protected] http://lists.lustre.org/mailman/listinfo/lustre-discuss
