> Which version of lustre do you use? > Server and clients same version and same os? which one?
lustre-1.6.4.1 The servers (oss and mds/mgs) use the RHEL4 rpm from lustre.org: 2.6.9-55.0.9.EL_lustre.1.6.4.1smp The clients run patchless RHEL4 2.6.9-67.0.1.ELsmp One set of clients are on a 10.x network while the servers and other half of clients are on a 141. network, because we are using the tcp network type, we have not setup any lnet routes. I don't think should cause a problem but I include the information for clarity. We do route 10.x on campus. > > Harald > > On Monday 04 February 2008 04:11 pm, Brock Palen wrote: >> on our cluster that has been running lustre for about 1 month. I have >> 1 MDT/MGS and 1 OSS with 2 OST's. >> >> Our cluster uses all Gige and has about 608 nodes 1854 cores. >> >> We have allot of jobs that die, and/or go into high IO wait, strace >> shows processes stuck in fstat(). >> >> The big problem is (i think) I would like some feedback on it that of >> these 608 nodes 209 of them have in dmesg the string >> >> "This client was evicted by" >> >> Is this normal for clients to be dropped like this? Is there some >> tuning that needs to be done to the server to carry this many nodes >> out of the box? We are using default lustre install with Gige. >> >> >> Brock Palen >> Center for Advanced Computing >> [EMAIL PROTECTED] >> (734)936-1985 >> >> >> _______________________________________________ >> Lustre-discuss mailing list >> [email protected] >> http://lists.lustre.org/mailman/listinfo/lustre-discuss > > -- > Harald van Pee > > Helmholtz-Institut fuer Strahlen- und Kernphysik der Universitaet Bonn > _______________________________________________ > Lustre-discuss mailing list > [email protected] > http://lists.lustre.org/mailman/listinfo/lustre-discuss > > _______________________________________________ Lustre-discuss mailing list [email protected] http://lists.lustre.org/mailman/listinfo/lustre-discuss
