on our cluster that has been running lustre for about 1 month. I have 1 MDT/MGS and 1 OSS with 2 OST's.
Our cluster uses all Gige and has about 608 nodes 1854 cores. We have allot of jobs that die, and/or go into high IO wait, strace shows processes stuck in fstat(). The big problem is (i think) I would like some feedback on it that of these 608 nodes 209 of them have in dmesg the string "This client was evicted by" Is this normal for clients to be dropped like this? Is there some tuning that needs to be done to the server to carry this many nodes out of the box? We are using default lustre install with Gige. Brock Palen Center for Advanced Computing [EMAIL PROTECTED] (734)936-1985 _______________________________________________ Lustre-discuss mailing list [email protected] http://lists.lustre.org/mailman/listinfo/lustre-discuss
