Hi Szabolcs, On 10/25/2016 04:07 PM, Szabolcs F. wrote: > Hi Alwin, > > the Cisco 4948 switches don't have jumbo frames enabled. Global Ethernet > MTU is 1500 bytes. Port security is not enabled. > > When the issue happens the hosts are able to ping each other without any > packet loss. > > On Tue, Oct 25, 2016 at 3:02 PM, Alwin Antreich <[email protected]> > wrote: > >> Hi Szabolcs, >> >> On 10/25/2016 12:24 PM, Szabolcs F. wrote: >>> Hi Alwin, >>> >>> bond0 is on two Cisco 4948 switches and bond1 is on two Cisco N3K-3064PQ >>> switches. They worked fine for about two months in this setup. But last >>> week (after I started to have these issues) I powered down one Cisco 4948 >>> and one N3K-3064PQ switch (in both cases the designated backup switches >>> were powered down). This is to make sure all servers use the same switch >> as >>> their active link. After that I stopped the Proxmox cluster (all nodes) >> and >>> started them again, but the issue occurred again. >> >> Ok, so one thing less to check. How are your remaining switch configured, >> especially, where the pve cluster is on? Do >> they use jumbo frames? Or some network/port security? >> >>> >>> I've just added the 'bond_primary ethX' option to the interfaces file. >> I'll >>> reboot everything once again and see if it helps. >> >> That's only going to be used, when you have all links connected and want >> to prefer a link to be the primary, eg. 10GbE >> as primary and 1GbE as backup. >> >>> >>> syslog: http://pastebin.com/MsuCcNx8 >>> dmesg: http://pastebin.com/xUPMKDJR >>> pveproxy (I can only see access.log for pveproxy, so this is the service >>> status): http://pastebin.com/gPPb4F3x >> >> I couldn't find anything unusual, but that doesn't mean there isn't. >> >>> >>> What other logs should I be reading? >>> >>> Thanks >>> >>> On Tue, Oct 25, 2016 at 11:23 AM, Alwin Antreich < >> [email protected]> >>> wrote: >>> >>>> Hi Szabolcs, >>>> >>>> On 10/25/2016 10:01 AM, Szabolcs F. wrote: >>>>> Hi Alwin, >>>>> >>>>> thanks for your hints. >>>>> >>>>>> On which interface is proxmox running on? Are these interfaces clogged >>>>> because, there is some heavy network IO going on? >>>>> I've got my two Intel Gbps network interfaces bonded together (bond0) >> as >>>>> active-backup and vmbr0 is bridged on this bond, then Proxmox is >> running >>>> on >>>>> this interface. I.e. http://pastebin.com/WZKQ02Qu >>>>> All nodes are configured like this. There is no heavy IO on these >>>>> interfaces, because the storage network uses the separate 10Gbps fiber >>>>> Intel NICs (bond1). >>>> >>>> Is your bond working properly? Is the bond on the same switch or two >>>> different? >>>> >>>> Usually I add the "bond_primary ethX" option to set the interface that >>>> should be primarily used in active-backup >>>> configuration - side note. :-) >>>> >>>> What are the logs on the server showing? You know, syslog, dmesg, >>>> pveproxy, etc. ;-) >>>> >>>>> >>>>>> Another guess, are all servers synchronizing with a NTP server and >> have >>>>> the correct time? >>>>> Yes, NTP is working properly, the firewall lets all NTP request go >>>> through. >>>>> >>>>> >>>>> On Mon, Oct 24, 2016 at 5:19 PM, Alwin Antreich < >>>> [email protected]> >>>>> wrote: >>>>> >>>>>> Hello Szabolcs, >>>>>> >>>>>> On 10/24/2016 03:16 PM, Szabolcs F. wrote: >>>>>>> Hello, >>>>>>> >>>>>>> I've got a Proxmox VE 4.3 cluster of 12 nodes. All of them are Dell >>>> C6220 >>>>>>> sleds. Each has 2x Intel Xeon E5-2670 CPU and 64GB RAM. I've got two >>>>>>> separate networks: 1Gbps LAN (Cisco 4948 switch) and 10Gbps storage >>>>>> (Cisco >>>>>>> N3K-3064PQ fiber switch). The Dell nodes use the integrated Intel >> Gbit >>>>>>> adapters for LAN and Intel PCI-E 10Gbps cards for the fiber network >>>>>> (ixgbe >>>>>>> driver). The storage servers are separate, they run FreeNAS and >> export >>>>>> the >>>>>>> shares with NFS. My virtual machines (I've made about 40 of them so >>>> far) >>>>>>> are KVM/QCOW2 and they are stored on the FreeNAS storage. So far so >>>> good. >>>>>>> I've been using this environment as a test and was almost ready to >> push >>>>>>> into production. >>>>>> On which interface is proxmox running on? Are these interfaces clogged >>>>>> because, there is some heavy network IO going on? >>>>>>> >>>>>>> But I have a problem with the cluster. From time to time the pveproxy >>>>>>> service dies on the nodes or the web UI lists all nodes (except the >> one >>>>>> I'm >>>>>>> actually logged into) as unreachable (red cross). Sometimes all nodes >>>> are >>>>>>> listed as working (green status) but if I try to connect to a virtual >>>>>>> machine I get a 'connection refused' error. When the cluster acts up >> I >>>>>>> can't do any VM migration and any other VM management (i.e. console, >>>>>>> start/stop/reset, new VM, etc). When it happens the only way to >> recover >>>>>> is >>>>>>> powering down all 12 nodes and starting them one after another. Then >>>>>>> everything works properly for a random amount of time: sometimes for >>>>>> weeks, >>>>>>> sometimes for only a few days. >>>>>> Another guess, are all servers synchronizing with a NTP server and >> have >>>>>> the correct time? >>>>>>> >>>>>>> I followed the network troubleshooting guide with omping, multicast, >>>> etc >>>>>>> and confirmed I've got multicase enabled and the troubleshooting >> didn't >>>>>>> return any error. The /etc/hosts file is configured on all nodes with >>>> the >>>>>>> proper hostname/IP list of all nodes. >>>>>>> When trying to do 'service pve-cluster restart' I get these errors: >>>>>>> http://pastebin.com/NXnEf4rd (running pmxcsf manually mounts the >>>>>> /etc/pve >>>>>>> properly, but doesn't fix the cluster/proxy issue) >>>>>>> pvecm status : http://pastebin.com/jsDFkqu3 (I powered down one >> node, >>>>>>> that's why it's missing) >>>>>>> pvecm nodes : http://pastebin.com/1WR8Yij8 >>>>>>> Corosync has a lot of these in the /var/logs/daemon.log : >>>>>>> http://pastebin.com/ajhE8Rb9 >>>>>>> >>>>>>> Someone please help! >>>>>>> >>>>>>> Thanks, >>>>>>> Szabolcs >>>>>>> _______________________________________________ >>>>>>> pve-user mailing list >>>>>>> [email protected] >>>>>>> http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user >>>>>>> >>>>>> >>>>>> -- >>>>>> Cheers, >>>>>> Alwin >>>>>> _______________________________________________ >>>>>> pve-user mailing list >>>>>> [email protected] >>>>>> http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user >>>>>> >>>>> _______________________________________________ >>>>> pve-user mailing list >>>>> [email protected] >>>>> http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user >>>>> >>>> >>>> -- >>>> Cheers, >>>> Alwin >>>> _______________________________________________ >>>> pve-user mailing list >>>> [email protected] >>>> http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user >>>> >>> _______________________________________________ >>> pve-user mailing list >>> [email protected] >>> http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user >>> >> >> When that happens, is the network working correctly between hosts? >> >> -- >> Cheers, >> Alwin >> _______________________________________________ >> pve-user mailing list >> [email protected] >> http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user >> > _______________________________________________ > pve-user mailing list > [email protected] > http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user >
http://pve.proxmox.com/pipermail/pve-user/2013-March/005358.html https://forum.proxmox.com/threads/constantly-losing-quorum.10755/ I still suspect that there might be a network issue. Maybe it helps to put traffic into separate VLANs. -- Cheers, Alwin _______________________________________________ pve-user mailing list [email protected] http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user
