Hi,

On Fri, Jun 05, 2009 at 12:06:11PM +0100, [email protected] wrote:
> > 
> > Hi,
> > 
> > On Thu, Jun 04, 2009 at 03:47:03PM +0100, [email protected]
> wrote:
> > > Hello,
> > > this is a little bit off topic but maybe anybody can help me. The
> > physical setup of our Debian etch heartbeat cluster looks like this:
> > >
> > > Cluster of two nodes connected with bonding to the network, with a
> > crossover connection to each other (interconnect) and with dual port
> > Qlogic HBAs to our SAN. The SAN consists of two independent (physical)
> > SANs with different VSANs (just to be more human readable) connected
> to
> > some EMC clariions CX3-20f.
> > >
> > > We use multipathd, LVM2 and ocfs2. This is the multipath.conf:
> > >
> > > defaults {
> > >         rr_min_io                       1000
> > >         polling_interval                2
> > >         no_path_retry                   5
> > >         user_friendly_names             yes
> > > }
> > >
> > > blacklist {
> > >         devnode "^(ram|raw|loop|fd|md|dm-|sr|scd|st)[0-9]*"
> > >         devnode "^hd[a-z][[0-9]*]"
> > >         devnode "^cciss!c[0-9]d[0-9]*[p[0-9]*]"
> > >         device {
> > >                 vendor "DGC"
> > >                 product "LUNZ" #EMC Clariion management LUN
> > >         }
> > >         device {
> > >                 vendor "ATA"  #We do not need mutlipathing for local
> > drives
> > >                 product "*"
> > >         }
> > >         device {
> > >                 vendor "AMI" # No multipathing for SUN Virtual
> devices
> > >                 product "*"
> > >         }
> > >         device {
> > >                 vendor "HITACHI" # No multipathing for local scsi
> disks
> > >                 product "H101414SCSUN146G"
> > >         }
> > > }
> > >
> > > devices {
> > >         ## Device attributes for EMC CLARiiON
> > >         device {
> > >                 vendor                  "DGC"
> > >                 product                 "*"
> > >                 path_grouping_policy    group_by_prio
> > >                 getuid_callout          "/sbin/scsi_id -g -u -s
> > /block/%n"
> > >                 prio_callout            "/sbin/mpath_prio_emc
> /dev/%n"
> > >                 hardware_handler        "1 emc"
> > >                 features                "1 queue_if_no_path"
> > >                 no_path_retry           fail
> > >                 path_checker            emc_clariion
> > >                 path_selector           "round-robin 0"
> > >                 failback                immediate
> > >         }
> > > }
> > >
> > > multipaths {
> > >         multipath {
> > >                 wwid
> > 3600601603ac511001c7c92fec775dd11
> > >                 alias                   stosan01_lun070
> > >         }
> > > }
> > >
> > >
> > > And we changed the "filter" of lvm2 to not scan /dev/sd* devices.
> After
> > booting up the system multipath -ll looks like this:
> > >
> > > stosan01_lun070 (3600601603ac511001c7c92fec775dd11) dm-7 DGC,RAID 5
> > > [size=133G][features=0][hwhandler=1 emc]
> > > \_ round-robin 0 [prio=2][active]
> > >  \_ 0:0:1:1 sdd 8:48  [active][ready]
> > >  \_ 3:0:1:1 sdh 8:112 [active][ready]
> > > \_ round-robin 0 [prio=0][enabled]
> > >  \_ 0:0:0:1 sdb 8:16  [active][ready]
> > >  \_ 3:0:0:1 sdf 8:80  [active][ready]
> > >
> > > If we now kill one of the connections by disable the zoning on one
> of
> > our SANs this changes to:
> > >
> > > stosan01_lun070 (3600601603ac511001c7c92fec775dd11) dm-7 DGC,RAID 5
> > > [size=133G][features=0][hwhandler=1 emc]
> > > \_ round-robin 0 [prio=1][active]
> > >  \_ 0:0:1:1 sdd 8:48  [active][ready]
> > > \_ round-robin 0 [prio=0][enabled]
> > >  \_ 0:0:0:1 sdb 8:16  [active][ready]
> > >
> > > And /var/log/messages
> > >
> > > Jun  4 14:53:58 defr1elcbtd02 kernel: o2net: no longer connected to
> node
> > defr1elcbtd01 (num 0) at 192.168.0.101:7777
> > > Jun  4 14:54:01 defr1elcbtd02 kernel:
> > (4741,0):dlm_wait_for_node_death:374 5EE89BC01EFC405E9197C198DEEAE678:
> > waiting 5000ms for notification of death of node 0
> > > Jun  4 14:54:16 defr1elcbtd02 last message repeated 3 times
> > > Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: multipath:
> Failing
> > path 8:80.
> > > Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: multipath:
> Failing
> > path 8:112.
> > > Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: multipath emc:
> long
> > trespass command will be send
> > > Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: multipath emc:
> > honor reservation bit will not be set (default)
> > > Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: ioctl: error
> adding
> > target to table
> > 
> > ?
> > 
> > > Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: multipath emc:
> long
> > trespass command will be send
> > > Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: multipath emc:
> > honor reservation bit will not be set (default)
> > > Jun  4 14:54:21 defr1elcbtd02 kernel:
> > (4741,0):dlm_wait_for_node_death:374 5EE89BC01EFC405E9197C198DEEAE678:
> > waiting 5000ms for notification of death of node 0
> > > Jun  4 14:54:26 defr1elcbtd02 kernel: scsi 1:0:1:1: SCSI error:
> return
> > code = 0x00020000
> > 
> > This looks like a multipath issue. AFAIK, the SCSI layer
> > shouldn't perceive a change in path availability. Perhaps better
> > to post to another list where such problems are discussed.
> > 
> > Thanks,
> > 
> > Dejan
> 
> I will post this to the ocfs2-users mailinglist. Just to note on Debian
> lenny we never had this error. Looks like a bug in Debian etch to me.

Hmm, I doubt that it's the problem of ocfs2, to me it looks more
like multipathing, i.e. a on some level below.

Thanks,

Dejan

> 
> 
> > 
> > > Jun  4 14:54:26 defr1elcbtd02 kernel: end_request: I/O error, dev
> sdh,
> > sector 1672
> > > Jun  4 14:54:26 defr1elcbtd02 kernel: device-mapper: multipath emc:
> > emc_pg_init: sending switch-over command
> > > Jun  4 14:54:26 defr1elcbtd02 kernel:
> > (4741,0):dlm_wait_for_node_death:374 5EE89BC01EFC405E9197C198DEEAE678:
> > waiting 5000ms for notification of death of node 0
> > > Jun  4 14:54:57 defr1elcbtd02 last message repeated 6 times
> > >
> > > After 60 seconds the kernel panics because of ocfs2 is fencing this
> > node:
> > >
> > > o2quo_make_decision:143 ERROR: fencing this node because it is
> connected
> > to a half-quorum of 1 out of 2 nodes which doesn't include the lowest
> > active node 0
> > >
> > > The o2cb settings are:
> > > O2CB_HEARTBEAT_THRESHOLD=31
> > > O2CB_IDLE_TIMEOUT_MS=30000
> > > O2CB_KEEPALIVE_DELAY_MS=2000
> > > O2CB_RECONNECT_DELAY_MS=2000
> > >
> > > If we just set the FC Port on one of the switches to shutdown
> everything
> > works fine. Qla2xxx detects port down, multipathd detects the missing
> path
> > and ocfs2 is not recognizing anything - just working fine like it
> > should...
> > >
> > > Whats going on? Is this a bug?
> > >
> > > Regards Florian
> > >
> > >
> > > Florian Engelmann | Unix System and Storage Administrator | IT
> Services
> > Germany | BT (Germany) GmbH & Co. oHG | Tel:+49 (0)89 2600 8173 |
> Mob:+49
> > (0)176 1073 1081 | Fax:+49 (0)89 2600 4955 | Email:
> > [email protected] | Sitz und Registergericht der oHG: M?nchen,
> HRA
> > 77639 | Pers?nlich haftende Gesellschafterin der oHG: BT Deutschland
> GmbH
> > | Sitz und Registergericht: M?nchen, HRB 132307 | Gesch?ftsf?hrung:
> > Karsten Lereuth, Joachim Piroth, Dr. Sebastian Brandis, Martin Glock,
> Neil
> > Smith, Judith Sasse | www.bt.com/globalservices
> > >
> > > Diese E-Mail enth?lt Informationen von BT (Germany) GmbH & Co. oHG.
> > Diese sind m?glicherweise vertraulich und ausschlie?lich f?r den
> > Adressaten bestimmt. Jeglicher Zugriff auf diese E-Mail durch andere
> > Personen als den Adressaten ist untersagt. Sollten Sie nicht der f?r
> diese
> > E-Mail bestimmte Adressat sein, ist Ihnen jede Ver?ffentlichung,
> > Vervielf?ltigung oder Weitergabe untersagt. Sollten Sie diese
> > elektronische Nachricht irrt?mlicherweise erhalten haben, so
> informieren
> > Sie uns bitte unverz?glich telefonisch oder per E-Mail (Kontaktdaten
> > obenstehend). Das E-Mail-System wird von BT (Germany) GmbH & Co. oHG
> im
> > Rahmen der gesetzlichen Bestimmungen laufend ?berwacht, insbesondere
> um
> > einen st?rungsfreien Betrieb zu gew?hrleisten. Auch die vom System
> > ?bertragenen Nachrichten werden im Rahmen der gesetzlichen
> Bestimmungen
> > durch BT (Germany) GmbH & Co. oHG ?berwacht und k?nnen aufgezeichnet
> oder
> > ausgewertet werden.
> > > This electronic message contains information from BT (Germany) GmbH
> &
> > Co. oHG which may be privileged or confidential. The information is
> > intended to be for the use of the individual(s) or entity named above.
> If
> > you are not the intended recipient be aware that any disclosure,
> copying,
> > distribution or use of the contents of this information is prohibited.
> If
> > you have received this electronic message in error, please notify us
> by
> > telephone or e-mail (to the numbers or address above) immediately.
> > Activity and use of the BT (Germany) GmbH & Co. oHG e-mail system is
> > monitored to secure its effective operation and for other lawful
> business
> > purposes. Communications using this system will also be monitored and
> may
> > be recorded to secure effective operation and for other lawful
> business
> > purposes.
> > >
> > >
> > > _______________________________________________
> > > Linux-HA mailing list
> > > [email protected]
> > > http://lists.linux-ha.org/mailman/listinfo/linux-ha
> > > See also: http://linux-ha.org/ReportingProblems
> > _______________________________________________
> > Linux-HA mailing list
> > [email protected]
> > http://lists.linux-ha.org/mailman/listinfo/linux-ha
> > See also: http://linux-ha.org/ReportingProblems
> _______________________________________________
> Linux-HA mailing list
> [email protected]
> http://lists.linux-ha.org/mailman/listinfo/linux-ha
> See also: http://linux-ha.org/ReportingProblems
_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems

Reply via email to