[nfs-discuss] NFS hanging with RPC timeout.
Sounds like it is possibly different, but in the end once I knew what to look for, we could test for it quite easily: # nmap -p 512-1023 172.20.12.5 623/tcp filtered unknown 664/tcp filtered unknown (The IP in this case in NFS client, but I ran nmap against both to look for other stolen ports). We then did the Inetd hack to occupy ports 623/664, and no hung NFS since then. (knock on wood). Kyle McDonald wrote: > > Ok. I got a matching snoop (from the server) and tcpdump (from the linux > client) and the replies seen leaving the server are definitely not seen > by the client: > > From the server: > releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C ACCESS3 > FH=6C39 (read,lookup,modify,extend,delete) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 > S=2049 Ack=1184956508 Seq=1590845549 Len=0 Win=53688 > Options= > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R ACCESS3 > OK (read,lookup) > releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM TCP D=2049 > S=800 Ack=1590845673 Seq=1184956508 Len=0 Win=512 > Options= > releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C > READDIRPLUS3 FH=6C39 Cookie=0 for 512/4096 > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C > READDIRPLUS3 FH=6C39 Cookie=0 for 512/4096 (retransmit) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 > S=2049 Ack=1184956668 Seq=1590847685 Len=0 Win=53688 > Options= > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C > READDIRPLUS3 FH=6C39 Cookie=0 for 512/4096 (retransmit) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 > S=2049 Ack=1184956828 Seq=1590847685 Len=0 Win=53688 > Options= > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R > READDIRPLUS3 OK 12 entries (No more) > releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM TCP D=2049 > S=800 Fin Ack=1590845673 Seq=1184956828 Len=0 Win=512 > Options= > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 > S=2049 Ack=1184956829 Seq=1590849697 Len=0 Win=53688 > Options= > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 > S=2049 Fin Ack=1184956829 Seq=1590849697 Len=0 Win=53688 > Options= > releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM TCP D=2049 > S=800 Rst Seq=1184956829 Len=0 Win=0 > > From the client: > 15:09:24.165259 IP releng1.862107800 > Galileo.RelEng.Egenera.COM.nfs: > 140 access [|nfs] > 15:09:24.165611 IP Galileo.RelEng.Egenera.COM.nfs > releng1.800: . ack > 277 win 53688 > 15:09:24.165855 IP Galileo.RelEng.Egenera.COM.nfs > releng1.862107800: > reply ok 124 access [|nfs] > 15:09:24.165867 IP releng1.800 > Galileo.RelEng.Egenera.COM.nfs: . ack > 293 win 512 > 15:09:24.166069 IP releng1.878885016 > Galileo.RelEng.Egenera.COM.nfs: > 160 readdirplus [|nfs] > 15:09:24.379488 IP releng1.878885016 > Galileo.RelEng.Egenera.COM.nfs: > 160 readdirplus [|nfs] > 15:09:24.379772 IP Galileo.RelEng.Egenera.COM.nfs > releng1.800: . ack > 437 win 53688 > 15:10:24.156981 IP releng1.878885016 > Galileo.RelEng.Egenera.COM.nfs: > 160 readdirplus [|nfs] > 15:10:24.157285 IP Galileo.RelEng.Egenera.COM.nfs > releng1.800: . ack > 597 win 53688 > 15:15:39.184608 IP releng1.800 > Galileo.RelEng.Egenera.COM.nfs: F > 597:597(0) ack 293 win 512 > 15:15:39.185420 IP Galileo.RelEng.Egenera.COM.nfs > releng1.800: . ack > 598 win 53688 > 15:15:39.185548 IP Galileo.RelEng.Egenera.COM.nfs > releng1.800: F > 4317:4317(0) ack 598 win 53688 > 15:15:39.185572 IP releng1.800 > Galileo.RelEng.Egenera.COM.nfs: R > 1184956829:1184956829
[nfs-discuss] NFS hanging with RPC timeout.
Ok. I got a matching snoop (from the server) and tcpdump (from the linux client) and the replies seen leaving the server are definitely not seen by the client: From the server: releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C ACCESS3 FH=6C39 (read,lookup,modify,extend,delete) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 S=2049 Ack=1184956508 Seq=1590845549 Len=0 Win=53688 Options= Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R ACCESS3 OK (read,lookup) releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM TCP D=2049 S=800 Ack=1590845673 Seq=1184956508 Len=0 Win=512 Options= releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C READDIRPLUS3 FH=6C39 Cookie=0 for 512/4096 Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C READDIRPLUS3 FH=6C39 Cookie=0 for 512/4096 (retransmit) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 S=2049 Ack=1184956668 Seq=1590847685 Len=0 Win=53688 Options= Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C READDIRPLUS3 FH=6C39 Cookie=0 for 512/4096 (retransmit) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 S=2049 Ack=1184956828 Seq=1590847685 Len=0 Win=53688 Options= Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM TCP D=2049 S=800 Fin Ack=1590845673 Seq=1184956828 Len=0 Win=512 Options= Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 S=2049 Ack=1184956829 Seq=1590849697 Len=0 Win=53688 Options= Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 S=2049 Fin Ack=1184956829 Seq=1590849697 Len=0 Win=53688 Options= releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM TCP D=2049 S=800 Rst Seq=1184956829 Len=0 Win=0 From the client: 15:09:24.165259 IP releng1.862107800 > Galileo.RelEng.Egenera.COM.nfs: 140 access [|nfs] 15:09:24.165611 IP Galileo.RelEng.Egenera.COM.nfs > releng1.800: . ack 277 win 53688 15:09:24.165855 IP Galileo.RelEng.Egenera.COM.nfs > releng1.862107800: reply ok 124 access [|nfs] 15:09:24.165867 IP releng1.800 > Galileo.RelEng.Egenera.COM.nfs: . ack 293 win 512 15:09:24.166069 IP releng1.878885016 > Galileo.RelEng.Egenera.COM.nfs: 160 readdirplus [|nfs] 15:09:24.379488 IP releng1.878885016 > Galileo.RelEng.Egenera.COM.nfs: 160 readdirplus [|nfs] 15:09:24.379772 IP Galileo.RelEng.Egenera.COM.nfs > releng1.800: . ack 437 win 53688 15:10:24.156981 IP releng1.878885016 > Galileo.RelEng.Egenera.COM.nfs: 160 readdirplus [|nfs] 15:10:24.157285 IP Galileo.RelEng.Egenera.COM.nfs > releng1.800: . ack 597 win 53688 15:15:39.184608 IP releng1.800 > Galileo.RelEng.Egenera.COM.nfs: F 597:597(0) ack 293 win 512 15:15:39.185420 IP Galileo.RelEng.Egenera.COM.nfs > releng1.800: . ack 598 win 53688 15:15:39.185548 IP Galileo.RelEng.Egenera.COM.nfs > releng1.800: F 4317:4317(0) ack 598 win 53688 15:15:39.185572 IP releng1.800 > Galileo.RelEng.Egenera.COM.nfs: R 1184956829:1184956829(0) win 0 I'll have to try to setup a port mirror on the Switch I'm using. I don't know if that switch can mirror the LACP group though - I'll have to disable it. Any other suggestions? -Kyle Kyle McDonald wrote: > I'm having a similiar problem from a linux client to a sNV_b103 server. > > For me though the mount works fine, it's the NFS accesses that hang. > > Here's a snoop that shows what the server is seeing: > > releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C > GETATTR3 FH=6C39 > Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 > S=2049 Ack=659557
[nfs-discuss] NFS hanging with RPC timeout.
I'm having a similiar problem from a linux client to a sNV_b103 server. For me though the mount works fine, it's the NFS accesses that hang. Here's a snoop that shows what the server is seeing: releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C GETATTR3 FH=6C39 Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 S=2049 Ack=659557530 Seq=471989561 Len=0 Win=53688 Options= Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R GETATTR3 OK releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM TCP D=2049 S=800 Ack=471989677 Seq=659557530 Len=0 Win=512 Options= releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C ACCESS3 FH=6C39 (read,lookup,modify,extend,delete) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R ACCESS3 OK (read,lookup) releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C READDIRPLUS3 FH=6C39 Cookie=0 for 512/4096 Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C READDIRPLUS3 FH=6C39 Cookie=0 for 512/4096 (retransmit) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 S=2049 Ack=659557830 Seq=471991813 Len=0 Win=53688 Options= Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C READDIRPLUS3 FH=6C39 Cookie=0 for 512/4096 (retransmit) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 S=2049 Ack=659557990 Seq=471991813 Len=0 Win=53688 Options= Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C FSSTAT3 FH=6C39 Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R FSSTAT3 OK releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM TCP D=2049 S=800 Ack=471989801 Seq=659558126 Len=0 Win=512 Options= Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C READDIRPLUS3 FH=6C39 Cookie=0 for 512/4096 (retransmit) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM NFS R READDIRPLUS3 OK 12 entries (No more) releng1.RelEng.Egenera.COM -> Galileo.RelEng.Egenera.COM NFS C READDIRPLUS3 FH=6C39 Cookie=0 for 512/4096 (retransmit) Galileo.RelEng.Egenera.COM -> releng1.RelEng.Egenera.COM TCP D=800 S=2049 Ack=659558286 Seq=471996009 Len=0 Win=53688 Options= This (to my uneducated eye) shows the server repling multiple times, and the client retransmitting the READDIR3 multiple times. I'm not familiar enough with Linux (yet) to run the equivalent of snoop (what is it? ethereal?) or I'd include traces from the client also. The client (and server) are both IBM x346 eServers. Like the link below, both have BMC(like an LOM) modules to manage the machine. Also like the link below these modules share the ethernet port with one of the broadcom (not intel) ethernet interfaces built into the motherboard. However in this case: 1) Neither the server nor the client are using the shared broadcom interface on the motherboard. 2) The client is using the other broadcom interface on the MB. 3) The Server is using a LACP aggr group (setup with dladm [with mtu=9000]) built up from 4 intel e1000g interfaces on a PCI card. So if packets are being lost on the return trip from the server to the client, I don't think it's for the same reason, though it may be similiar. Note this on a ZFS filesystem, but from the traces above I'm not inclied to think that has anything to do with the problem. The server does have other ethernet interfaces on other subnets. However the testing above was careful to do the NFS mount with only the IP of the one interface, that one interface is also the one used for the default route, and snoop was running on the others and showed zero incoming or outgoing traffic (to or from the clients IP) durring this same period. Anyone got any ideas? -Kyle Jorgen Lundman wrote: > > I stumbled across this entry: > > http://blogs.sun.com/shepler/entry/port_623_or_the_mount > > and even though we do not see this issue with port 623, but rather > 664. But sure enough, it was sending SYN/ACK, then timeout until RST. > > I waited for the port to the released, told inetd to listen on port > 664 and voila, mount works fine again. > > We use Supermicros with Intel? 82573V and 82573L. > > I would send Shepler my thanks but comments are disabled. > > > > Useless logs: > > > # mount -o proto=tcp,vers=3 172.20.12.226:/export/src /mn
[nfs-discuss] NFS hanging with RPC timeout.
I stumbled across this entry: http://blogs.sun.com/shepler/entry/port_623_or_the_mount and even though we do not see this issue with port 623, but rather 664. But sure enough, it was sending SYN/ACK, then timeout until RST. I waited for the port to the released, told inetd to listen on port 664 and voila, mount works fine again. We use Supermicros with Intel? 82573V and 82573L. I would send Shepler my thanks but comments are disabled. Useless logs: # mount -o proto=tcp,vers=3 172.20.12.226:/export/src /mnt 172.20.12.6 -> 172.20.12.226 PORTMAP C GETPORT prog=15 (MOUNT) vers=3 proto=UDP 172.20.12.226 -> 172.20.12.6 PORTMAP R GETPORT port=39049 172.20.12.6 -> 172.20.12.226 MOUNT3 C Null 172.20.12.226 -> 172.20.12.6 MOUNT3 R Null 172.20.12.6 -> 172.20.12.226 MOUNT3 C Mount /export/src 172.20.12.226 -> 172.20.12.6 MOUNT3 R Mount OK FH=076E Auth=unix 172.20.12.6 -> 172.20.12.226 PORTMAP C GETPORT prog=13 (NFS) vers=3 proto=TCP 172.20.12.226 -> 172.20.12.6 PORTMAP R GETPORT port=2049 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=38337 Syn Seq=592414549 Len=0 Win=49640 Options= 172.20.12.226 -> 172.20.12.6 TCP D=38337 S=2049 Syn Ack=592414550 Seq=2210245643 Len=0 Win=49640 Options= 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=38337 Ack=2210245644 Seq=592414550 Len=0 Win=49640 172.20.12.6 -> 172.20.12.226 NFS C NULL3 172.20.12.226 -> 172.20.12.6 TCP D=38337 S=2049 Ack=592414670 Seq=2210245644 Len=0 Win=49520 172.20.12.226 -> 172.20.12.6 NFS R NULL3 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=38337 Ack=2210245672 Seq=592414670 Len=0 Win=49640 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=38337 Fin Ack=2210245672 Seq=592414670 Len=0 Win=49640 172.20.12.226 -> 172.20.12.6 TCP D=38337 S=2049 Ack=592414671 Seq=2210245672 Len=0 Win=49640 172.20.12.226 -> 172.20.12.6 TCP D=38337 S=2049 Fin Ack=592414671 Seq=2210245672 Len=0 Win=49640 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=38337 Ack=2210245673 Seq=592414671 Len=0 Win=49640 172.20.12.6 -> 172.20.12.226 PORTMAP C GETPORT prog=13 (NFS) vers=3 proto=TCP 172.20.12.226 -> 172.20.12.6 PORTMAP R GETPORT port=2049 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=38338 Syn Seq=3614232918 Len=0 Win=49640 Options= 172.20.12.226 -> 172.20.12.6 TCP D=38338 S=2049 Syn Ack=3614232919 Seq=2210460804 Len=0 Win=49640 Options= 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=38338 Ack=2210460805 Seq=3614232919 Len=0 Win=49640 172.20.12.6 -> 172.20.12.226 NFS C NULL3 172.20.12.226 -> 172.20.12.6 TCP D=38338 S=2049 Ack=3614233039 Seq=2210460805 Len=0 Win=49520 172.20.12.226 -> 172.20.12.6 NFS R NULL3 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=38338 Ack=2210460833 Seq=3614233039 Len=0 Win=49640 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=38338 Fin Ack=2210460833 Seq=3614233039 Len=0 Win=49640 172.20.12.226 -> 172.20.12.6 TCP D=38338 S=2049 Ack=3614233040 Seq=2210460833 Len=0 Win=49640 172.20.12.226 -> 172.20.12.6 TCP D=38338 S=2049 Fin Ack=3614233040 Seq=2210460833 Len=0 Win=49640 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=38338 Ack=2210460834 Seq=3614233040 Len=0 Win=49640 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=664 Rst Ack=0 Seq=3456416233 Len=0 Win=49640 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=664 Syn Seq=3507413975 Len=0 Win=49640 Options= 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=664 Syn Seq=3507413975 Len=0 Win=49640 Options= # netstat 172.20.12.6.664 172.20.12.226.2049 0 0 49640 0 SYN_SENT After inetd hack: # netstat *.664*.*0 0 49152 0 BOUND # mount -o proto=tcp,vers=3 172.20.12.226:/export/src /mnt 172.20.12.6 -> 172.20.12.226 TCP D=2049 S=661 Syn Seq=1448210229 Len=0 Win=49640 Options= # df 172.20.12.226:/export/src 24T11G24T 1%/mnt Jorgen Lundman wrote: > > Ok, it still happens even when not using aliases, it just took longer to > turn up. > > Attempting to mount (snoop running on NFS client) -- Jorgen Lundman | Unix Administrator | +81 (0)3 -5456-2687 ext 1017 (work) Shibuya-ku, Tokyo| +81 (0)90-5578-8500 (cell) Japan| +81 (0)3 -3375-1767 (home)
[nfs-discuss] NFS hanging with RPC timeout.
0.0.0.0.0.111 rpcbindsuperuser 103tcp 0.0.0.0.0.111 rpcbindsuperuser 102tcp 0.0.0.0.0.111 rpcbindsuperuser 104udp 0.0.0.0.0.111 rpcbindsuperuser 103udp 0.0.0.0.0.111 rpcbindsuperuser 102udp 0.0.0.0.0.111 rpcbindsuperuser 1000241udp 0.0.0.0.128.10 status superuser 1000241tcp 0.0.0.0.128.3 status superuser 1000241ticlts\021\000\000\000status superuser 1000241ticotsord \024\000\000\000status superuser 1000241ticots\027\000\000\000status superuser 1001331udp 0.0.0.0.128.10 - superuser 1001331tcp 0.0.0.0.128.3 - superuser 1001331ticlts\021\000\000\000- superuser 1001331ticotsord \024\000\000\000- superuser 1001331ticots\027\000\000\000- superuser 1000211udp 0.0.0.0.15.205 nlockmgr 1 10737418241tcp 0.0.0.0.128.4 - 1 1000212udp 0.0.0.0.15.205 nlockmgr 1 1000213udp 0.0.0.0.15.205 nlockmgr 1 1000214udp 0.0.0.0.15.205 nlockmgr 1 1000211tcp 0.0.0.0.15.205 nlockmgr 1 1000212tcp 0.0.0.0.15.205 nlockmgr 1 1000213tcp 0.0.0.0.15.205 nlockmgr 1 1000214tcp 0.0.0.0.15.205 nlockmgr 1 1001551ticotsord l\000\000\000 smserverd superuser 1000111ticltso\000\000\000 rquotadsuperuser 1000111udp 0.0.0.0.128.18 rquotadsuperuser 1002311ticltsx4500-05.unix.nfsauth - superuser 1002311ticotsord x4500-05.unix.nfsauth - superuser 1002311ticotsx4500-05.unix.nfsauth - superuser 151udp 0.0.0.0.128.19 mountd superuser 151ticlts\203\000\000\000mountd superuser 151tcp 0.0.0.0.128.13 mountd superuser 151ticotsord \210\000\000\000mountd superuser 151ticots\213\000\000\000mountd superuser 152udp 0.0.0.0.128.19 mountd superuser 152ticlts\203\000\000\000mountd superuser 152tcp 0.0.0.0.128.13 mountd superuser 152ticotsord \210\000\000\000mountd superuser 152ticots\213\000\000\000mountd superuser 153udp 0.0.0.0.128.19 mountd superuser 153ticlts\203\000\000\000mountd superuser 153tcp 0.0.0.0.128.13 mountd superuser 153ticotsord \210\000\000\000mountd superuser 153ticots\213\000\000\000mountd superuser 132udp 0.0.0.0.8.1 nfs1 133udp 0.0.0.0.8.1 nfs1 1002272udp 0.0.0.0.8.1 nfs_acl1 1002273udp 0.0.0.0.8.1 nfs_acl1 132tcp 0.0.0.0.8.1 nfs1 133tcp 0.0.0.0.8.1 nfs1 134tcp 0.0.0.0.8.1 nfs1 1002272tcp 0.0.0.0.8.1 nfs_acl1 1002273tcp 0.0.0.0.8.1 nfs_acl1 Why would the NFS client not be able to talk to the server? 3 minutes later, rpcinfo got unstuck and NFS was back again. Without me doing anything but snoop. Lund Robert van Veelen wrote: > I will try this out on my test hosts. Are you using NFSv3 exclusively? There > are no v4 clients in your env? If the clients are exclusively ro, have you > tried mointing with ro flag? > > -rob > > > -Original Message- > From: Jorgen Lundman [mailto:lundman at gmo.jp] > Sent: Monday, February 16, 2009 12:11 AM Eastern Standard Time > To: Robert van Veelen > Subject: Re: [nfs-discuss] NFS hanging with RPC timeout. > > > I have not had time to prove this, but by asking the other admins which > NFS server mounts hung, nobody could remember x4500-01 ever hanging. The > reason I asked them was because x4500-01 is the only one where the > "alias" is lower IP than the real IP. > > 01-alias: .220 > 01-real: .221 > 02-real: .222 > 02-alias: .223 > 03-real: .224 > 03-alias: .225 > 04-real: .226 > 04-alias: .227 > > Now, IP "value" should not matter, I know, but it just "felt" like it > was related. :) &g
[nfs-discuss] NFS hanging with RPC timeout.
You make it sounds like it might still hang, even using the real IP. ;) The www, ssl, cgi and navi clusters experience the NFS problem the most, about 2-3 times a day, so I have remounted those servers to use the real IP. This should tell if it makes any difference to not use the alias. If any (other) servers experience the NFS problem, I will run the suggested commands. Lund Dai Ngo wrote: > It's good that you now have a work-around without rebooting the client > or server. > IP alias might, or might not, be a problem. However the real problem is > why the > hang occurs after it has been working for awhile with the server > configured with > IP alias. > > I think the mount with the real IP worked because the client used a > different > (source) port for new connection, 620. If you try to mount using the IP > alias > I think the client will use port 664, which already hang (the original > problem), > and this is why the mount failed. The reason the client uses port 664 to do > the mount because this connection was already established to the server > using > the IP alias. > > You can run these commands on the server to get a little more info on > port 664: > > # ps -ef |grep nfsd --> get the nfsd PID > # pfiles nfsd_PID ---> to see all sockets nfsd are using > # pstack nfsd_PID --> to see what the nfsd threads are doing > # netstat -P tcp -f inet --> to see what state the TCP sockets are in > > -Dai > > Jorgen Lundman wrote: >> >> Ok, a server was already hung when I got to work today. >> >> >> ** >> >> x4500-04: NFS Server, Sol 10 5/08 >> Server IP (real) 172.20.12.226 netmask ff00 >> NFS IP (alias) 172.20.12.227 netmask ff00 >> >> x4500-04:~# netstat -in ; netstat -rn >> Name Mtu Net/Dest AddressIpkts Ierrs Opkts Oerrs >> Collis Queue >> lo0 8232 127.0.0.0 127.0.0.1 1411 0 1411 0 0 0 >> e1000g0 1500 172.20.12.0 172.20.12.226 2762497849 0 1789082372 >> 0 0 0 >> e1000g1 1500 172.20.19.0 172.20.19.226 96059758 0 52485074 0 >> 0 0 >> >> >> Routing Table: IPv4 >> Destination Gateway Flags Ref Use >> Interface >> - - -- >> - >> default 172.20.12.1 UG1 20456 >> 172.20.12.0 172.20.12.226U 1 45968 e1000g0 >> 172.20.12.0 172.20.12.227U 1 0 >> e1000g0:1 >> 172.20.19.0 172.20.19.226U 1 1662 e1000g1 >> 224.0.0.0172.20.12.226U 1 0 e1000g0 >> 127.0.0.1127.0.0.1UH5316 lo0 >> >> >> ** >> >> NFS client: Sol 10 5/08 >> Client IP172.20.12.6 netmask ff00 >> >> # netstat -in ; netstat -rn >> Name Mtu Net/Dest AddressIpkts Ierrs Opkts Oerrs >> Collis Queue >> lo0 8232 127.0.0.0 127.0.0.1 2175 0 2175 0 0 0 >> e1000g0 1500 172.20.12.0 172.20.12.643315618 0 41987515 0 >> 0 0 >> e1000g1 1500 172.20.11.0 172.20.11.619673254 0 13928826 0 >> 0 0 >> >> >> Routing Table: IPv4 >> Destination Gateway Flags Ref Use >> Interface >> - - -- >> - >> default 172.20.11.4 UG1 52386 >> 10.0.0.0 172.20.12.1 UG1 0 >> 172.16.0.0 172.20.12.1 UG1193 >> 172.20.11.0 172.20.11.6 U 1 2406 e1000g1 >> 172.20.12.0 172.20.12.6 U 1 3163 e1000g0 >> 192.168.0.0 172.20.12.1 UG1120 >> 224.0.0.0172.20.12.6 U 1 0 e1000g0 >> 127.0.0.1127.0.0.1UH4 2046 lo0 >> >> >> >> * >> >> >> >> Snoop running on NFS Client 172.20.12.6 attempting to (re)mount volume >> with TCP: >> >> # snoop -r host 172.20.12.227 or host 172.20.12.226 & >> # mount /export/www >> 172.20.12.6 -> 172.20.12.227 PORTMAP C GETPORT prog=15 (MOUNT) >> vers=3 proto=UDP >> 172.20.12.226 -> 172.20.12.6 PORTMAP R GETPORT port=39049 >> 172.20.12.6 -> 172.20.12.227 MOUNT3 C Null >> 172.20.12.226 -> 172.20.12.6 MOUNT3 R Null >> 172.20.12.6 -> 172.20.12.227 MOUNT3 C Mount /export/www >> 172.20.12.226 -> 172.20.12.6 MOUNT3 R Mount OK FH=D402 Auth=unix >> 172.20.12.6 -> 172.20.12.227 PORTMAP C GETPORT prog=13 (NFS) >> vers=3 proto=TCP >> 172.20.12.226 -> 172.20.12.6 PORTMAP R GETPORT port=2049 >> 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=63800 Syn Seq=788700586 >> Len=0 Win=49640 Options= >> 172.20.12.227 -> 172.20.12.6 TCP D=63800 S=2049 Syn Ack=788700587 >> Seq=3596066619 Len=0 Win=49640 Options=>
[nfs-discuss] NFS hanging with RPC timeout.
Ok, a server was already hung when I got to work today. ** x4500-04: NFS Server, Sol 10 5/08 Server IP (real) 172.20.12.226 netmask ff00 NFS IP (alias) 172.20.12.227 netmask ff00 x4500-04:~# netstat -in ; netstat -rn Name Mtu Net/Dest AddressIpkts Ierrs Opkts Oerrs Collis Queue lo0 8232 127.0.0.0 127.0.0.1 1411 0 1411 0 0 0 e1000g0 1500 172.20.12.0 172.20.12.226 2762497849 0 1789082372 0 0 0 e1000g1 1500 172.20.19.0 172.20.19.226 96059758 0 52485074 0 0 0 Routing Table: IPv4 Destination Gateway Flags Ref Use Interface - - -- - default 172.20.12.1 UG1 20456 172.20.12.0 172.20.12.226U 1 45968 e1000g0 172.20.12.0 172.20.12.227U 1 0 e1000g0:1 172.20.19.0 172.20.19.226U 1 1662 e1000g1 224.0.0.0172.20.12.226U 1 0 e1000g0 127.0.0.1127.0.0.1UH5316 lo0 ** NFS client: Sol 10 5/08 Client IP172.20.12.6 netmask ff00 # netstat -in ; netstat -rn Name Mtu Net/Dest AddressIpkts Ierrs Opkts Oerrs Collis Queue lo0 8232 127.0.0.0 127.0.0.1 2175 0 2175 0 0 0 e1000g0 1500 172.20.12.0 172.20.12.643315618 0 41987515 0 0 0 e1000g1 1500 172.20.11.0 172.20.11.619673254 0 13928826 0 0 0 Routing Table: IPv4 Destination Gateway Flags Ref Use Interface - - -- - default 172.20.11.4 UG1 52386 10.0.0.0 172.20.12.1 UG1 0 172.16.0.0 172.20.12.1 UG1193 172.20.11.0 172.20.11.6 U 1 2406 e1000g1 172.20.12.0 172.20.12.6 U 1 3163 e1000g0 192.168.0.0 172.20.12.1 UG1120 224.0.0.0172.20.12.6 U 1 0 e1000g0 127.0.0.1127.0.0.1UH4 2046 lo0 * Snoop running on NFS Client 172.20.12.6 attempting to (re)mount volume with TCP: # snoop -r host 172.20.12.227 or host 172.20.12.226 & # mount /export/www 172.20.12.6 -> 172.20.12.227 PORTMAP C GETPORT prog=15 (MOUNT) vers=3 proto=UDP 172.20.12.226 -> 172.20.12.6 PORTMAP R GETPORT port=39049 172.20.12.6 -> 172.20.12.227 MOUNT3 C Null 172.20.12.226 -> 172.20.12.6 MOUNT3 R Null 172.20.12.6 -> 172.20.12.227 MOUNT3 C Mount /export/www 172.20.12.226 -> 172.20.12.6 MOUNT3 R Mount OK FH=D402 Auth=unix 172.20.12.6 -> 172.20.12.227 PORTMAP C GETPORT prog=13 (NFS) vers=3 proto=TCP 172.20.12.226 -> 172.20.12.6 PORTMAP R GETPORT port=2049 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=63800 Syn Seq=788700586 Len=0 Win=49640 Options= 172.20.12.227 -> 172.20.12.6 TCP D=63800 S=2049 Syn Ack=788700587 Seq=3596066619 Len=0 Win=49640 Options= 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=63800 Ack=3596066620 Seq=788700587 Len=0 Win=49640 172.20.12.6 -> 172.20.12.227 NFS C NULL3 172.20.12.227 -> 172.20.12.6 TCP D=63800 S=2049 Ack=788700707 Seq=3596066620 Len=0 Win=49520 172.20.12.227 -> 172.20.12.6 NFS R NULL3 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=63800 Ack=3596066648 Seq=788700707 Len=0 Win=49640 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=63800 Fin Ack=3596066648 Seq=788700707 Len=0 Win=49640 172.20.12.227 -> 172.20.12.6 TCP D=63800 S=2049 Ack=788700708 Seq=3596066648 Len=0 Win=49640 172.20.12.227 -> 172.20.12.6 TCP D=63800 S=2049 Fin Ack=788700708 Seq=3596066648 Len=0 Win=49640 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=63800 Ack=3596066649 Seq=788700708 Len=0 Win=49640 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=664 Syn Seq=2946510831 Len=0 Win=49640 Options= 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=664 Syn Seq=2946510831 Len=0 Win=49640 Options= 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=664 Syn Seq=2946510831 Len=0 Win=49640 Options= Interesting, looks like x4500-04 is replying with the wrong IP. Packet capture on x4500-04: # snoop -r host 172.20.12.6 Using device /dev/e1000g0 (promiscuous mode) 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=664 Rst Ack=0 Seq=2924968134 Len=0 Win=49640 172.20.12.227 -> 172.20.12.6 TCP D=664 S=2049 Rst Win=49640 172.20.12.6 -> 172.20.12.227 PORTMAP C GETPORT prog=15 (MOUNT) vers=3 proto=UDP 172.20.12.226 -> 172.20.12.6 PORTMAP R GETPORT port=39049 172.20.12.6 -> 172.20.12.227 MOUNT3 C Null 172.20.12.226 -> 172.20.12.6 MOUNT3 R Null 172.20.12.6 -> 172.20.12.227 MOUNT3 C Mount /export/www 172.20.12.226 -> 172.20.12.6 MOUNT3 R Mount OK FH=D402 Auth=unix 172.20.12.6
[nfs-discuss] NFS hanging with RPC timeout.
It's good that you now have a work-around without rebooting the client or server. IP alias might, or might not, be a problem. However the real problem is why the hang occurs after it has been working for awhile with the server configured with IP alias. I think the mount with the real IP worked because the client used a different (source) port for new connection, 620. If you try to mount using the IP alias I think the client will use port 664, which already hang (the original problem), and this is why the mount failed. The reason the client uses port 664 to do the mount because this connection was already established to the server using the IP alias. You can run these commands on the server to get a little more info on port 664: # ps -ef |grep nfsd --> get the nfsd PID # pfiles nfsd_PID ---> to see all sockets nfsd are using # pstack nfsd_PID --> to see what the nfsd threads are doing # netstat -P tcp -f inet --> to see what state the TCP sockets are in -Dai Jorgen Lundman wrote: > > Ok, a server was already hung when I got to work today. > > > ** > > x4500-04: NFS Server, Sol 10 5/08 > Server IP (real) 172.20.12.226 netmask ff00 > NFS IP (alias) 172.20.12.227 netmask ff00 > > x4500-04:~# netstat -in ; netstat -rn > Name Mtu Net/Dest AddressIpkts Ierrs Opkts Oerrs > Collis Queue > lo0 8232 127.0.0.0 127.0.0.1 1411 0 1411 0 0 0 > e1000g0 1500 172.20.12.0 172.20.12.226 2762497849 0 1789082372 > 0 0 0 > e1000g1 1500 172.20.19.0 172.20.19.226 96059758 0 52485074 0 > 0 0 > > > Routing Table: IPv4 > Destination Gateway Flags Ref Use > Interface > - - -- > - > default 172.20.12.1 UG1 20456 > 172.20.12.0 172.20.12.226U 1 45968 e1000g0 > 172.20.12.0 172.20.12.227U 1 0 > e1000g0:1 > 172.20.19.0 172.20.19.226U 1 1662 e1000g1 > 224.0.0.0172.20.12.226U 1 0 e1000g0 > 127.0.0.1127.0.0.1UH5316 lo0 > > > ** > > NFS client: Sol 10 5/08 > Client IP172.20.12.6 netmask ff00 > > # netstat -in ; netstat -rn > Name Mtu Net/Dest AddressIpkts Ierrs Opkts Oerrs > Collis Queue > lo0 8232 127.0.0.0 127.0.0.1 2175 0 2175 0 0 0 > e1000g0 1500 172.20.12.0 172.20.12.643315618 0 41987515 0 > 0 0 > e1000g1 1500 172.20.11.0 172.20.11.619673254 0 13928826 0 > 0 0 > > > Routing Table: IPv4 > Destination Gateway Flags Ref Use > Interface > - - -- > - > default 172.20.11.4 UG1 52386 > 10.0.0.0 172.20.12.1 UG1 0 > 172.16.0.0 172.20.12.1 UG1193 > 172.20.11.0 172.20.11.6 U 1 2406 e1000g1 > 172.20.12.0 172.20.12.6 U 1 3163 e1000g0 > 192.168.0.0 172.20.12.1 UG1120 > 224.0.0.0172.20.12.6 U 1 0 e1000g0 > 127.0.0.1127.0.0.1UH4 2046 lo0 > > > > * > > > > Snoop running on NFS Client 172.20.12.6 attempting to (re)mount volume > with TCP: > > # snoop -r host 172.20.12.227 or host 172.20.12.226 & > # mount /export/www > 172.20.12.6 -> 172.20.12.227 PORTMAP C GETPORT prog=15 (MOUNT) > vers=3 proto=UDP > 172.20.12.226 -> 172.20.12.6 PORTMAP R GETPORT port=39049 > 172.20.12.6 -> 172.20.12.227 MOUNT3 C Null > 172.20.12.226 -> 172.20.12.6 MOUNT3 R Null > 172.20.12.6 -> 172.20.12.227 MOUNT3 C Mount /export/www > 172.20.12.226 -> 172.20.12.6 MOUNT3 R Mount OK FH=D402 Auth=unix > 172.20.12.6 -> 172.20.12.227 PORTMAP C GETPORT prog=13 (NFS) > vers=3 proto=TCP > 172.20.12.226 -> 172.20.12.6 PORTMAP R GETPORT port=2049 > 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=63800 Syn Seq=788700586 > Len=0 Win=49640 Options= > 172.20.12.227 -> 172.20.12.6 TCP D=63800 S=2049 Syn Ack=788700587 > Seq=3596066619 Len=0 Win=49640 Options= 0,nop,nop,sackOK> > 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=63800 Ack=3596066620 > Seq=788700587 Len=0 Win=49640 > 172.20.12.6 -> 172.20.12.227 NFS C NULL3 > 172.20.12.227 -> 172.20.12.6 TCP D=63800 S=2049 Ack=788700707 > Seq=3596066620 Len=0 Win=49520 > 172.20.12.227 -> 172.20.12.6 NFS R NULL3 > 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=63800 Ack=3596066648 > Seq=788700707 Len=0 Win=49640 > 172.20.12.6 -> 172.20.12.227 TCP D=2049 S=63800 Fin Ack=3596066648 > Seq=788700707 Len=0 Win=49640 > 172.20.12.227 -> 172.20.12.6 TCP D=63800 S=
[nfs-discuss] NFS hanging with RPC timeout.
Jorgen Lundman wrote: > > > Which works without issue. So it is not an NFS problem, it seems to be > related to alias IPs. > > Do you know a way around this? Or perhaps you can suggest a place > where I can go to ask. As a quick solution we will just forgo the > Alias IP and mount directly on the "real" IP. Why can I change > protocol (TCP->UDP and vv) to get around it, why can I reboot the NFS > client as well. Did we create the aliases wrong? > > I apologise for the noise in NFS discussion list. > > Lund Jorgen, Noise, no, this is not noise. I just think most people are overloaded right now with getting ready for Connectathon 2009. (www.connectathon.org) You have an interesting problem, please follow-up to the list with a summary of what you find. That will eventually help the next person who hits this issue and understands how to use google as well. Thanks, Tom PS: One of the things I used to get struck by as a developer who got customer escalations was their definition of what was NFS and my colleagues definition of NFS. After all, we worked deep in the protocol, the system, etc. Our definition is very narrow, very focused. And a customer tends to think in broader systems - NFS is an application to them. What a developer has to realize is that their subsystem can not exist in a vacuum - nfs means nothing if the DNS server does not do reverse name mapping in a timely fashion. And what the customer has to realize is that the developer may be an expert in the subsystem, but clueless as to how it is used in real world systems. I remember once releasing a new feature that we planned on deprecating (at a previous company). In 6 months, when we went to deprecate it, we found we couldn't as customers were using it to solve problems other than the ones I envisioned. Going to the problem at hand, I wonder how many developers commonly use an IP alias in their day to day testing? Within the NFS group, I'm confident in saying none. But someone in say a clustering group may do so daily. (No clue if they are the ones who can solve your problem.)
[nfs-discuss] NFS hanging with RPC timeout.
Thank you for your reply,
The x4500s uses a real IP, and an IP alias. The NFS mounts are connected
to the alias, so it would be "easier" to fail-over to a different x4500
should there be a need. We have not yet got that far as we are exploring
ways to replicate the data from the active x4500 to a passive x4500 first.
But I will test the idea that the alias might be involved, see if it
mounts the real IP etc. I will report back.
Dai Ngo wrote:
> The problem seems to be on the TCP connection between the client and the
> nfsd on
> the server. The portmap and mount requests used UDP and they went OK.
>
> There are a number TCP RST packets sent from both the client and server,
> this indicated
> there might be problem with packets lost causing both sides to be out of
> sync.
>
> Looks like the server has 2 NICs on the same subnet, 172.20.12.221 and
> 172.20.12.220.
> Have you tried disable 172.20.12.220 and just use 172.20.12.221 to see
> if it helps.
> What the output of the 'netstat -in' and 'netstat -rn' on the server and
> the client look like?
>
> By the way, where were the packets captured from? on the server or the
> client. It's more
> useful if you can capture the packets on both sides and attach the raw
> capture files so
> they can be compared and examined in more details.
>
> -Dai
>
> Jorgen Lundman wrote:
>> (Resent due to wrong sender, sorry)
>>
>>
>> Hello list!
>>
>> *** NFS Servers:
>>
>> x4500-01 to x4500-05
>> : Solaris 10 5/08, ZFS and "UFS on ZVOL" exported.
>> : NFSD_SERVER=1024, LOCKD_SERVER=128 average use about 900 / 20 threads.
>> : "bufhwm_pct,maxusers,ndquot,ncsize,ufs_ninode,clnt_max_conns,
>> : rpcmod:cotsmaxdupreqs,rpcmod:maxdupreqs" tweaked in /etc/system.
>>
>> *** NFS Clients:
>>
>> Supermicro 1U * 40
>> : Solaris 10 5/08
>> : No tweaks, Mounted as
>> : x4500-03:/export/mail - /export/mail nfs - yes vers=3,hard,intr,quota
>> : x4500-02:/export/preview - /export/preview nfs - yes vers=3,hard,intr
>>
>>
>> *** Background
>>
>> Using vers=3 to have uid mapping, without the need for UID lookups. UFS
>> on ZVOL are mounted with "quota". ZFS exported filesystems are mounted
>> without. The system is live and generally works very well.
>>
>> However, NFS will periodically hang. Usually to just one of the x4500
>> servers at a time, the solution currently is just to reboot the client.
>> I have attempted to fully umount all filesystems, and terminate the NFS
>> and RPC processes, in an attempt to remount. This will not fix it. I can
>> not really restart the NFSD/RPC processes on the x4500s.
>>
>> Usually looks like:
>>
>> # df -h
>> [snip]
>> x4500-03:/export/preview
>> 23T 3.9M23T 1%/export/preview
>> NFS server x4500-01 not responding still trying
>> ^C
>>
>> Note that during this time, x4500-01 is still functioning correctly to
>> the other 39 servers, and x4500-02,03,04,05 are still mounted correctly
>> on this NFS client.
>>
>> # umount /export/www
>> # mount /export/www
>> NFS server x4500-01-vip not responding still trying
>>
>> Truss of the mount says:
>> 23102: 0. getpid()= 23102
>> [23101]
>> 23102: 0. door_call(5, 0x080475A0)= 0
>> 23102: 0.0001 close(5)= 0
>> NFS server x4500-01-vip not responding still trying
>> ^C23102:69.0780 mount("x4500-01-vip:/export/www", "/export/www",
>> MS_DATA|MS_OPTIONSTR, "nfs3", 0x0806D400, 76, 0x0804777C, 1024) Err#4
>> EINTR
>>
>> Snoop says (x4500-01 is 172.20.12.220, NFS Client is 172.20.12.16)
>>
>> 172.20.12.16 -> 172.20.12.220 PORTMAP C GETPORT prog=15 (MOUNT)
>> vers=3 proto=UDP
>> 172.20.12.221 -> 172.20.12.16 PORTMAP R GETPORT port=39967
>> 172.20.12.16 -> 172.20.12.220 MOUNT3 C Null
>> 172.20.12.221 -> 172.20.12.16 MOUNT3 R Null
>> 172.20.12.16 -> 172.20.12.220 MOUNT3 C Mount /export/www
>> 172.20.12.221 -> 172.20.12.16 MOUNT3 R Mount OK FH=D502 Auth=unix
>> 172.20.12.16 -> 172.20.12.220 PORTMAP C GETPORT prog=13 (NFS) vers=3
>> proto=TCP
>> 172.20.12.221 -> 172.20.12.16 PORTMAP R GETPORT port=2049
>> 172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Syn Seq=2255048579
>> Len=0 Win=49640 Options=
>> 172.20.12.220 -> 172.20.12.16 TCP D=54091 S=2049 Syn Ack=2255048580
>> Seq=611591914 Len=0 Win=49640 Options=> 0,nop,nop,sackOK>
>> 172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Ack=611591915
>> Seq=2255048580 Len=0 Win=49640
>> 172.20.12.16 -> 172.20.12.220 NFS C NULL3
>> 172.20.12.220 -> 172.20.12.16 TCP D=54091 S=2049 Ack=2255048700
>> Seq=611591915 Len=0 Win=49520
>> 172.20.12.220 -> 172.20.12.16 NFS R NULL3
>> 172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Ack=611591943
>> Seq=2255048700 Len=0 Win=49640
>> 172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Fin Ack=611591943
>> Seq=2255048700 Len=0 Win=49640
>> 172.20.12.220 -> 172.20.12.16 TCP D=54091 S=2049 Ack=2255048701
>> Seq=611591943 Len=0 Win=49640
>> 172.20.12.220 -> 172.20.12.16
[nfs-discuss] NFS hanging with RPC timeout.
The problem seems to be on the TCP connection between the client and the
nfsd on
the server. The portmap and mount requests used UDP and they went OK.
There are a number TCP RST packets sent from both the client and server,
this indicated
there might be problem with packets lost causing both sides to be out of
sync.
Looks like the server has 2 NICs on the same subnet, 172.20.12.221 and
172.20.12.220.
Have you tried disable 172.20.12.220 and just use 172.20.12.221 to see
if it helps.
What the output of the 'netstat -in' and 'netstat -rn' on the server and
the client look like?
By the way, where were the packets captured from? on the server or the
client. It's more
useful if you can capture the packets on both sides and attach the raw
capture files so
they can be compared and examined in more details.
-Dai
Jorgen Lundman wrote:
> (Resent due to wrong sender, sorry)
>
>
> Hello list!
>
> *** NFS Servers:
>
> x4500-01 to x4500-05
> : Solaris 10 5/08, ZFS and "UFS on ZVOL" exported.
> : NFSD_SERVER=1024, LOCKD_SERVER=128 average use about 900 / 20 threads.
> : "bufhwm_pct,maxusers,ndquot,ncsize,ufs_ninode,clnt_max_conns,
> : rpcmod:cotsmaxdupreqs,rpcmod:maxdupreqs" tweaked in /etc/system.
>
> *** NFS Clients:
>
> Supermicro 1U * 40
> : Solaris 10 5/08
> : No tweaks, Mounted as
> : x4500-03:/export/mail - /export/mail nfs - yes vers=3,hard,intr,quota
> : x4500-02:/export/preview - /export/preview nfs - yes vers=3,hard,intr
>
>
> *** Background
>
> Using vers=3 to have uid mapping, without the need for UID lookups. UFS
> on ZVOL are mounted with "quota". ZFS exported filesystems are mounted
> without. The system is live and generally works very well.
>
> However, NFS will periodically hang. Usually to just one of the x4500
> servers at a time, the solution currently is just to reboot the client.
> I have attempted to fully umount all filesystems, and terminate the NFS
> and RPC processes, in an attempt to remount. This will not fix it. I can
> not really restart the NFSD/RPC processes on the x4500s.
>
> Usually looks like:
>
> # df -h
> [snip]
> x4500-03:/export/preview
> 23T 3.9M23T 1%/export/preview
> NFS server x4500-01 not responding still trying
> ^C
>
> Note that during this time, x4500-01 is still functioning correctly to
> the other 39 servers, and x4500-02,03,04,05 are still mounted correctly
> on this NFS client.
>
> # umount /export/www
> # mount /export/www
> NFS server x4500-01-vip not responding still trying
>
> Truss of the mount says:
> 23102: 0. getpid()= 23102
> [23101]
> 23102: 0. door_call(5, 0x080475A0)= 0
> 23102: 0.0001 close(5)= 0
> NFS server x4500-01-vip not responding still trying
> ^C23102:69.0780 mount("x4500-01-vip:/export/www", "/export/www",
> MS_DATA|MS_OPTIONSTR, "nfs3", 0x0806D400, 76, 0x0804777C, 1024) Err#4 EINTR
>
> Snoop says (x4500-01 is 172.20.12.220, NFS Client is 172.20.12.16)
>
> 172.20.12.16 -> 172.20.12.220 PORTMAP C GETPORT prog=15 (MOUNT)
> vers=3 proto=UDP
> 172.20.12.221 -> 172.20.12.16 PORTMAP R GETPORT port=39967
> 172.20.12.16 -> 172.20.12.220 MOUNT3 C Null
> 172.20.12.221 -> 172.20.12.16 MOUNT3 R Null
> 172.20.12.16 -> 172.20.12.220 MOUNT3 C Mount /export/www
> 172.20.12.221 -> 172.20.12.16 MOUNT3 R Mount OK FH=D502 Auth=unix
> 172.20.12.16 -> 172.20.12.220 PORTMAP C GETPORT prog=13 (NFS) vers=3
> proto=TCP
> 172.20.12.221 -> 172.20.12.16 PORTMAP R GETPORT port=2049
> 172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Syn Seq=2255048579
> Len=0 Win=49640 Options=
> 172.20.12.220 -> 172.20.12.16 TCP D=54091 S=2049 Syn Ack=2255048580
> Seq=611591914 Len=0 Win=49640 Options=
> 172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Ack=611591915
> Seq=2255048580 Len=0 Win=49640
> 172.20.12.16 -> 172.20.12.220 NFS C NULL3
> 172.20.12.220 -> 172.20.12.16 TCP D=54091 S=2049 Ack=2255048700
> Seq=611591915 Len=0 Win=49520
> 172.20.12.220 -> 172.20.12.16 NFS R NULL3
> 172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Ack=611591943
> Seq=2255048700 Len=0 Win=49640
> 172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Fin Ack=611591943
> Seq=2255048700 Len=0 Win=49640
> 172.20.12.220 -> 172.20.12.16 TCP D=54091 S=2049 Ack=2255048701
> Seq=611591943 Len=0 Win=49640
> 172.20.12.220 -> 172.20.12.16 TCP D=54091 S=2049 Fin Ack=2255048701
> Seq=611591943 Len=0 Win=49640
> 172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Ack=611591944
> Seq=2255048701 Len=0 Win=49640
> 172.20.12.16 -> 172.20.12.220 TCP D=2049 S=664 Syn Seq=1161480442 Len=0
> Win=49640 Options=
> 172.20.12.220 -> 172.20.12.16 TCP D=664 S=2049 Ack=1118215538
> Seq=4284552307 Len=0 Win=49640
> 172.20.12.16 -> 172.20.12.220 TCP D=2049 S=664 Syn Seq=1161480442 Len=0
> Win=49640 Options=
> 172.20.12.220 -> 172.20.12.16 TCP D=664 S=2049 Ack=1118215538
> Seq=4284552307 Len=0 Win=49640
> [delay]
> 172.20.12.16 -> 172.20.12.220 TCP D=2049 S=664 Syn Se
[nfs-discuss] NFS hanging with RPC timeout.
(Resent due to wrong sender, sorry)
Hello list!
*** NFS Servers:
x4500-01 to x4500-05
: Solaris 10 5/08, ZFS and "UFS on ZVOL" exported.
: NFSD_SERVER=1024, LOCKD_SERVER=128 average use about 900 / 20 threads.
: "bufhwm_pct,maxusers,ndquot,ncsize,ufs_ninode,clnt_max_conns,
: rpcmod:cotsmaxdupreqs,rpcmod:maxdupreqs" tweaked in /etc/system.
*** NFS Clients:
Supermicro 1U * 40
: Solaris 10 5/08
: No tweaks, Mounted as
: x4500-03:/export/mail - /export/mail nfs - yes vers=3,hard,intr,quota
: x4500-02:/export/preview - /export/preview nfs - yes vers=3,hard,intr
*** Background
Using vers=3 to have uid mapping, without the need for UID lookups. UFS
on ZVOL are mounted with "quota". ZFS exported filesystems are mounted
without. The system is live and generally works very well.
However, NFS will periodically hang. Usually to just one of the x4500
servers at a time, the solution currently is just to reboot the client.
I have attempted to fully umount all filesystems, and terminate the NFS
and RPC processes, in an attempt to remount. This will not fix it. I can
not really restart the NFSD/RPC processes on the x4500s.
Usually looks like:
# df -h
[snip]
x4500-03:/export/preview
23T 3.9M23T 1%/export/preview
NFS server x4500-01 not responding still trying
^C
Note that during this time, x4500-01 is still functioning correctly to
the other 39 servers, and x4500-02,03,04,05 are still mounted correctly
on this NFS client.
# umount /export/www
# mount /export/www
NFS server x4500-01-vip not responding still trying
Truss of the mount says:
23102: 0. getpid()= 23102
[23101]
23102: 0. door_call(5, 0x080475A0)= 0
23102: 0.0001 close(5)= 0
NFS server x4500-01-vip not responding still trying
^C23102:69.0780 mount("x4500-01-vip:/export/www", "/export/www",
MS_DATA|MS_OPTIONSTR, "nfs3", 0x0806D400, 76, 0x0804777C, 1024) Err#4 EINTR
Snoop says (x4500-01 is 172.20.12.220, NFS Client is 172.20.12.16)
172.20.12.16 -> 172.20.12.220 PORTMAP C GETPORT prog=15 (MOUNT)
vers=3 proto=UDP
172.20.12.221 -> 172.20.12.16 PORTMAP R GETPORT port=39967
172.20.12.16 -> 172.20.12.220 MOUNT3 C Null
172.20.12.221 -> 172.20.12.16 MOUNT3 R Null
172.20.12.16 -> 172.20.12.220 MOUNT3 C Mount /export/www
172.20.12.221 -> 172.20.12.16 MOUNT3 R Mount OK FH=D502 Auth=unix
172.20.12.16 -> 172.20.12.220 PORTMAP C GETPORT prog=13 (NFS) vers=3
proto=TCP
172.20.12.221 -> 172.20.12.16 PORTMAP R GETPORT port=2049
172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Syn Seq=2255048579
Len=0 Win=49640 Options=
172.20.12.220 -> 172.20.12.16 TCP D=54091 S=2049 Syn Ack=2255048580
Seq=611591914 Len=0 Win=49640 Options=
172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Ack=611591915
Seq=2255048580 Len=0 Win=49640
172.20.12.16 -> 172.20.12.220 NFS C NULL3
172.20.12.220 -> 172.20.12.16 TCP D=54091 S=2049 Ack=2255048700
Seq=611591915 Len=0 Win=49520
172.20.12.220 -> 172.20.12.16 NFS R NULL3
172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Ack=611591943
Seq=2255048700 Len=0 Win=49640
172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Fin Ack=611591943
Seq=2255048700 Len=0 Win=49640
172.20.12.220 -> 172.20.12.16 TCP D=54091 S=2049 Ack=2255048701
Seq=611591943 Len=0 Win=49640
172.20.12.220 -> 172.20.12.16 TCP D=54091 S=2049 Fin Ack=2255048701
Seq=611591943 Len=0 Win=49640
172.20.12.16 -> 172.20.12.220 TCP D=2049 S=54091 Ack=611591944
Seq=2255048701 Len=0 Win=49640
172.20.12.16 -> 172.20.12.220 TCP D=2049 S=664 Syn Seq=1161480442 Len=0
Win=49640 Options=
172.20.12.220 -> 172.20.12.16 TCP D=664 S=2049 Ack=1118215538
Seq=4284552307 Len=0 Win=49640
172.20.12.16 -> 172.20.12.220 TCP D=2049 S=664 Syn Seq=1161480442 Len=0
Win=49640 Options=
172.20.12.220 -> 172.20.12.16 TCP D=664 S=2049 Ack=1118215538
Seq=4284552307 Len=0 Win=49640
[delay]
172.20.12.16 -> 172.20.12.220 TCP D=2049 S=664 Syn Seq=1161480442 Len=0
Win=49640 Options=
172.20.12.220 -> 172.20.12.16 TCP D=664 S=2049 Ack=1118215538
Seq=4284552307 Len=0 Win=49640
172.20.12.16 -> 172.20.12.220 TCP D=2049 S=664 Syn Seq=1161480442 Len=0
Win=49640 Options=
172.20.12.220 -> 172.20.12.16 TCP D=664 S=2049 Ack=1118215538
Seq=4284552307 Len=0 Win=49640
[repeat, delay]
*** truss of mountd on x4500-01 while attempting mount:
# truss -Dfip 28717
28717: 6.8156 pollsys(0x080CAE38, 9, 0x, 0x) = 1
28717: 0.0002 lwp_kill(788, SIG#0)Err#3 ESRCH
28717: 0.0001 lwp_create(0x08047B90, LWP_DETACHED|LWP_SUSPENDED,
0x08047DB0) = 791
28717/1: 0.0002 lwp_continue(791) = 0
28717/791: 6.8159 lwp_create()(returning as new lwp ...) = 0
28717/1: 0.0001 fxstat(2, 7, 0x08047CB0)= 0
28717/791: 0.0003 setustack(0xFECD1A60)
28717/1: 0. getmsg(7, 0x08047D8C, 0x080CC018, 0x08047DAC) = 0
28717/791: 0.0001 schedctl()
= 0xFEFB2010
2
