Hi Alwin, thanks for the links. Do you mean VLAN tagging on trunk ports or completely separated, untagged, dedicated ports?
ps: I forgot to ask about the jumbo frames. Should I enable them? Thanks, Szabolcs On Tue, Oct 25, 2016 at 6:09 PM, Alwin Antreich <[email protected]> wrote: > Hi Szabolcs, > > On 10/25/2016 04:07 PM, Szabolcs F. wrote: > > Hi Alwin, > > > > the Cisco 4948 switches don't have jumbo frames enabled. Global Ethernet > > MTU is 1500 bytes. Port security is not enabled. > > > > When the issue happens the hosts are able to ping each other without any > > packet loss. > > > > On Tue, Oct 25, 2016 at 3:02 PM, Alwin Antreich < > [email protected]> > > wrote: > > > >> Hi Szabolcs, > >> > >> On 10/25/2016 12:24 PM, Szabolcs F. wrote: > >>> Hi Alwin, > >>> > >>> bond0 is on two Cisco 4948 switches and bond1 is on two Cisco > N3K-3064PQ > >>> switches. They worked fine for about two months in this setup. But last > >>> week (after I started to have these issues) I powered down one Cisco > 4948 > >>> and one N3K-3064PQ switch (in both cases the designated backup switches > >>> were powered down). This is to make sure all servers use the same > switch > >> as > >>> their active link. After that I stopped the Proxmox cluster (all nodes) > >> and > >>> started them again, but the issue occurred again. > >> > >> Ok, so one thing less to check. How are your remaining switch > configured, > >> especially, where the pve cluster is on? Do > >> they use jumbo frames? Or some network/port security? > >> > >>> > >>> I've just added the 'bond_primary ethX' option to the interfaces file. > >> I'll > >>> reboot everything once again and see if it helps. > >> > >> That's only going to be used, when you have all links connected and want > >> to prefer a link to be the primary, eg. 10GbE > >> as primary and 1GbE as backup. > >> > >>> > >>> syslog: http://pastebin.com/MsuCcNx8 > >>> dmesg: http://pastebin.com/xUPMKDJR > >>> pveproxy (I can only see access.log for pveproxy, so this is the > service > >>> status): http://pastebin.com/gPPb4F3x > >> > >> I couldn't find anything unusual, but that doesn't mean there isn't. > >> > >>> > >>> What other logs should I be reading? > >>> > >>> Thanks > >>> > >>> On Tue, Oct 25, 2016 at 11:23 AM, Alwin Antreich < > >> [email protected]> > >>> wrote: > >>> > >>>> Hi Szabolcs, > >>>> > >>>> On 10/25/2016 10:01 AM, Szabolcs F. wrote: > >>>>> Hi Alwin, > >>>>> > >>>>> thanks for your hints. > >>>>> > >>>>>> On which interface is proxmox running on? Are these interfaces > clogged > >>>>> because, there is some heavy network IO going on? > >>>>> I've got my two Intel Gbps network interfaces bonded together (bond0) > >> as > >>>>> active-backup and vmbr0 is bridged on this bond, then Proxmox is > >> running > >>>> on > >>>>> this interface. I.e. http://pastebin.com/WZKQ02Qu > >>>>> All nodes are configured like this. There is no heavy IO on these > >>>>> interfaces, because the storage network uses the separate 10Gbps > fiber > >>>>> Intel NICs (bond1). > >>>> > >>>> Is your bond working properly? Is the bond on the same switch or two > >>>> different? > >>>> > >>>> Usually I add the "bond_primary ethX" option to set the interface that > >>>> should be primarily used in active-backup > >>>> configuration - side note. :-) > >>>> > >>>> What are the logs on the server showing? You know, syslog, dmesg, > >>>> pveproxy, etc. ;-) > >>>> > >>>>> > >>>>>> Another guess, are all servers synchronizing with a NTP server and > >> have > >>>>> the correct time? > >>>>> Yes, NTP is working properly, the firewall lets all NTP request go > >>>> through. > >>>>> > >>>>> > >>>>> On Mon, Oct 24, 2016 at 5:19 PM, Alwin Antreich < > >>>> [email protected]> > >>>>> wrote: > >>>>> > >>>>>> Hello Szabolcs, > >>>>>> > >>>>>> On 10/24/2016 03:16 PM, Szabolcs F. wrote: > >>>>>>> Hello, > >>>>>>> > >>>>>>> I've got a Proxmox VE 4.3 cluster of 12 nodes. All of them are Dell > >>>> C6220 > >>>>>>> sleds. Each has 2x Intel Xeon E5-2670 CPU and 64GB RAM. I've got > two > >>>>>>> separate networks: 1Gbps LAN (Cisco 4948 switch) and 10Gbps storage > >>>>>> (Cisco > >>>>>>> N3K-3064PQ fiber switch). The Dell nodes use the integrated Intel > >> Gbit > >>>>>>> adapters for LAN and Intel PCI-E 10Gbps cards for the fiber network > >>>>>> (ixgbe > >>>>>>> driver). The storage servers are separate, they run FreeNAS and > >> export > >>>>>> the > >>>>>>> shares with NFS. My virtual machines (I've made about 40 of them so > >>>> far) > >>>>>>> are KVM/QCOW2 and they are stored on the FreeNAS storage. So far so > >>>> good. > >>>>>>> I've been using this environment as a test and was almost ready to > >> push > >>>>>>> into production. > >>>>>> On which interface is proxmox running on? Are these interfaces > clogged > >>>>>> because, there is some heavy network IO going on? > >>>>>>> > >>>>>>> But I have a problem with the cluster. From time to time the > pveproxy > >>>>>>> service dies on the nodes or the web UI lists all nodes (except the > >> one > >>>>>> I'm > >>>>>>> actually logged into) as unreachable (red cross). Sometimes all > nodes > >>>> are > >>>>>>> listed as working (green status) but if I try to connect to a > virtual > >>>>>>> machine I get a 'connection refused' error. When the cluster acts > up > >> I > >>>>>>> can't do any VM migration and any other VM management (i.e. > console, > >>>>>>> start/stop/reset, new VM, etc). When it happens the only way to > >> recover > >>>>>> is > >>>>>>> powering down all 12 nodes and starting them one after another. > Then > >>>>>>> everything works properly for a random amount of time: sometimes > for > >>>>>> weeks, > >>>>>>> sometimes for only a few days. > >>>>>> Another guess, are all servers synchronizing with a NTP server and > >> have > >>>>>> the correct time? > >>>>>>> > >>>>>>> I followed the network troubleshooting guide with omping, > multicast, > >>>> etc > >>>>>>> and confirmed I've got multicase enabled and the troubleshooting > >> didn't > >>>>>>> return any error. The /etc/hosts file is configured on all nodes > with > >>>> the > >>>>>>> proper hostname/IP list of all nodes. > >>>>>>> When trying to do 'service pve-cluster restart' I get these errors: > >>>>>>> http://pastebin.com/NXnEf4rd (running pmxcsf manually mounts the > >>>>>> /etc/pve > >>>>>>> properly, but doesn't fix the cluster/proxy issue) > >>>>>>> pvecm status : http://pastebin.com/jsDFkqu3 (I powered down one > >> node, > >>>>>>> that's why it's missing) > >>>>>>> pvecm nodes : http://pastebin.com/1WR8Yij8 > >>>>>>> Corosync has a lot of these in the /var/logs/daemon.log : > >>>>>>> http://pastebin.com/ajhE8Rb9 > >>>>>>> > >>>>>>> Someone please help! > >>>>>>> > >>>>>>> Thanks, > >>>>>>> Szabolcs > >>>>>>> _______________________________________________ > >>>>>>> pve-user mailing list > >>>>>>> [email protected] > >>>>>>> http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user > >>>>>>> > >>>>>> > >>>>>> -- > >>>>>> Cheers, > >>>>>> Alwin > >>>>>> _______________________________________________ > >>>>>> pve-user mailing list > >>>>>> [email protected] > >>>>>> http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user > >>>>>> > >>>>> _______________________________________________ > >>>>> pve-user mailing list > >>>>> [email protected] > >>>>> http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user > >>>>> > >>>> > >>>> -- > >>>> Cheers, > >>>> Alwin > >>>> _______________________________________________ > >>>> pve-user mailing list > >>>> [email protected] > >>>> http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user > >>>> > >>> _______________________________________________ > >>> pve-user mailing list > >>> [email protected] > >>> http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user > >>> > >> > >> When that happens, is the network working correctly between hosts? > >> > >> -- > >> Cheers, > >> Alwin > >> _______________________________________________ > >> pve-user mailing list > >> [email protected] > >> http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user > >> > > _______________________________________________ > > pve-user mailing list > > [email protected] > > http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user > > > > http://pve.proxmox.com/pipermail/pve-user/2013-March/005358.html > https://forum.proxmox.com/threads/constantly-losing-quorum.10755/ > > I still suspect that there might be a network issue. Maybe it helps to put > traffic into separate VLANs. > > -- > Cheers, > Alwin > _______________________________________________ > pve-user mailing list > [email protected] > http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user > _______________________________________________ pve-user mailing list [email protected] http://pve.proxmox.com/cgi-bin/mailman/listinfo/pve-user
