Hi,

On Tue, Nov 13, 2007 at 10:19:38AM -0800, Peregrino wrote:
> 
> --- [EMAIL PROTECTED] wrote:
> 
> > Send Linux-HA mailing list submissions to
> >     [email protected]
> > 
> > To subscribe or unsubscribe via the World Wide Web,
> > visit
> >     http://lists.linux-ha.org/mailman/listinfo/linux-ha
> > or, via email, send a message with subject or body
> > 'help' to
> >     [EMAIL PROTECTED]
> > 
> > You can reach the person managing the list at
> >     [EMAIL PROTECTED]
> > 
> > When replying, please edit your Subject line so it
> > is more specific
> > than "Re: Contents of Linux-HA digest..."
> > 
> > 
> > Today's Topics:
> > 
> >    1. Frustrating Problems With Active/Passive
> > Config (Peregrino)
> >    2. Re: pingd removed transient attr from node
> > attributes  after
> >       short network outage, and did not recreated it
> > (Dejan Muhamedagic)
> >    3. Re: Frustrating Problems With Active/Passive
> > Config
> >       (Dejan Muhamedagic)
> >    4. Re: RE: Missing gratious ARP (Alan Robertson)
> >    5. Failure to start resource makes it impossible
> > to fail     back
> >       (Anders Brownworth)
> >    6. Can't reload grouped resources (An?bal
> > Monsalve Salazar)
> >    7. clone behavior when set on-fail to "restart"
> > (Junko IKEDA)
> > 
> > 
> >
> ----------------------------------------------------------------------
> > 
> > Message: 1
> > Date: Mon, 12 Nov 2007 10:49:41 -0800 (PST)
> > From: Peregrino <[EMAIL PROTECTED]>
> > Subject: [Linux-HA] Frustrating Problems With
> > Active/Passive Config
> > To: [email protected]
> > Message-ID:
> > <[EMAIL PROTECTED]>
> > Content-Type: text/plain; charset=iso-8859-1
> > 
> > Hello all,
> > 
> > I have an Active/Passive configuration that I have
> > been trying to get working that is very frustrating.
> > This may have already been answered somewhere, and
> > if
> > so, just point me to the place.
> > 
> > Overview:
> > I have a system (PUBLIC_WEB) with a public viewable
> > IP
> > address (128.249.65.X) that runs a web service. I
> > have
> > a VMware Host on another system with a public
> > viewable
> > address (VMWARE_HOST) (128.249.65.X) that contains a
> > virtual system (VIRT_PUBLIC_WEB) with a private ip
> > (172.30.10.X). 
> > I want the virtual system (VIRT_PUBLIC_WEB) with the
> > private IP (172.30.10.X) to provide simple failover
> > for the public system (PUBLIC_WEB).
> > I have this working with another node, but in that
> > case both nodes are on the private ip (172.30.10.X),
> > so I know it is possible for VMware virtual systems
> > to
> > act as failover devices.
> > 
> > In all cases the host files are correct, the ha.cf
> > files match as do the haresources files. All systems
> > can ping one another. I am using heartbeat v2, but
> > configured as a v1 (no crm or anything fancy).
> > 
> > Questions:
> > First of all, is it possible for a heartbeat node on
> > one net to failover and assume a different nets
> > address. That seems to be the hangup here. 
> > If anyone else has done a setup like this I would
> > appreciate some assistance.
> > 
> > Second, might it be a routing issue? Each system can
> > ping across its net to the other, but does heartbeat
> > have its own way of doing things that may not
> > include
> > using the standard linux net tools?
> > 
> > When the master node (PUBLIC_WEB) is taken down, the
> > virtual system (VIRT_PUBLIC_WEB) detects it and
> > craps
> > out running the resource.d/IPaddr <ip address>
> > start.
> > 
> > Any help would be appreciated. Once I have been
> > assured that this configuration is possible, I will
> > post any config files that are needed to help me fix
> > this problem. 
> > 
> > Many Thank In Advance.
> > 
> > __________________________________________________
> > Do You Yahoo!?
> > Tired of spam?  Yahoo! Mail has the best spam
> > protection around 
> > http://mail.yahoo.com 
> > 
> > 
> Here is the error log when I shutdown the Active node
> on the public IP. Remember, the Active node is on a
> public IP, and the Passive node is on a private IP.
> 
> =====================
> Nov  9 13:55:07 opalVM heartbeat: [2019]: WARN: node
> opal.brc.bcm.edu: is dead
> Nov  9 13:55:07 opalVM heartbeat: [2019]: info:
> Comm_now_up(): updating status to active
> Nov  9 13:55:07 opalVM heartbeat: [2019]: info: Local
> status now set to: 'active'
> Nov  9 13:55:07 opalVM heartbeat: [2019]: WARN: No
> STONITH device configured.
> Nov  9 13:55:07 opalVM heartbeat: [2019]: WARN: Shared
> disks are not protected.
> Nov  9 13:55:07 opalVM heartbeat: [2019]: info:
> Resources being acquired from opal.brc.bcm.edu.
> Nov  9 13:55:07 opalVM heartbeat: [2408]: info: No
> local resources
> [/usr/local/HA-2.0.8/lib/heartbeat/ResourceManager
> listkeys opalvm.cluster.brc.bcm.edu] to acquire.
> Nov  9 13:55:07 opalVM heartbeat: [2408]: info:
> Writing type [resource] message to FIFO
> Nov  9 13:55:07 opalVM heartbeat: [2408]: info: FIFO
> message [type resource] written rc=79
> Nov  9 13:55:07 opalVM harc[2407]: info: Running
> /usr/local/HA-2.0.8/etc/ha.d/rc.d/status status
> Nov  9 13:55:07 opalVM heartbeat: [2019]: info:
> AnnounceTakeover(local 0, foreign 1, reason
> 'T_RESOURCES' (0))
> Nov  9 13:55:07 opalVM heartbeat: [2019]: info:
> AnnounceTakeover(local 1, foreign 1, reason
> 'T_RESOURCES(us)' (0))
> Nov  9 13:55:07 opalVM heartbeat: [2019]: info:
> Initial resource acquisition complete
> (T_RESOURCES(us))
> Nov  9 13:55:07 opalVM heartbeat: [2019]: info: STATE
> 1 => 3
> Nov  9 13:55:07 opalVM heartbeat: [2019]: info:
> Exiting req_our_resources process 2408 returned rc 0.
> Nov  9 13:55:07 opalVM heartbeat: [2019]: info:
> AnnounceTakeover(local 1, foreign 1, reason
> 'req_our_resources' (1))
> Nov  9 13:55:07 opalVM mach_down[2427]: info: Taking
> over resource group 128.249.65.204
> Nov  9 13:55:07 opalVM ResourceManager[2447]: info:
> Acquiring resource group: opal.brc.bcm.edu
> 128.249.65.204 apache::/etc/httpd/conf/httpd.conf
> Nov  9 13:55:07 opalVM IPaddr[2471]: INFO:  Resource
> is stopped
> Nov  9 13:55:07 opalVM ResourceManager[2447]: info:
> Running /usr/local/HA-2.0.8/etc/ha.d/resource.d/IPaddr
> 128.249.65.204 start
> Nov  9 13:55:07 opalVM IPaddr[2525]: ERROR:
> /usr/local/HA-2.0.8/lib/heartbeat/findif
> 128.249.65.204 failed [rc=1].
> Nov  9 13:55:07 opalVM IPaddr[2516]: ERROR:  Generic
> error
> Nov  9 13:55:08 opalVM ResourceManager[2447]: ERROR:
> Return code 1 from
> /usr/local/HA-2.0.8/etc/ha.d/resource.d/IPaddr
> Nov  9 13:55:08 opalVM ResourceManager[2447]: CRIT:
> Giving up resources due to failure of 128.249.65.204
> Nov  9 13:55:08 opalVM ResourceManager[2447]: info:
> Releasing resource group: opal.brc.bcm.edu
> 128.249.65.204 apache::/etc/httpd/conf/httpd.conf
> Nov  9 13:55:08 opalVM ResourceManager[2447]: info:
> Running /usr/local/HA-2.0.8/etc/ha.d/resource.d/apache
> /etc/httpd/conf/httpd.conf stop
> Nov  9 13:55:08 opalVM apache[2584]: INFO: apache is
> not running.
> Nov  9 13:55:08 opalVM apache[2581]: INFO:  Success
> Nov  9 13:55:08 opalVM ResourceManager[2447]: info:
> Running /usr/local/HA-2.0.8/etc/ha.d/resource.d/IPaddr
> 128.249.65.204 stop
> Nov  9 13:55:09 opalVM IPaddr[2664]: INFO:  Success
> Nov  9 13:55:09 opalVM mach_down[2427]: info:
> /usr/local/HA-2.0.8/lib/heartbeat/mach_down:
> nice_failback: foreign resources acquired
> Nov  9 13:55:09 opalVM mach_down[2427]: info:
> mach_down takeover complete for node opal.brc.bcm.edu.
> Nov  9 13:55:09 opalVM heartbeat: [2019]: info:
> AnnounceTakeover(local 1, foreign 1, reason
> 'T_RESOURCES(us)' (1))
> Nov  9 13:55:09 opalVM heartbeat: [2019]: info:
> mach_down takeover complete.
> Nov  9 13:55:09 opalVM heartbeat: [2019]: info:
> AnnounceTakeover(local 1, foreign 1, reason
> 'mach_down' (1))
> Nov  9 13:55:09 opalVM heartbeat: [2019]: info:
> Exiting status process 2407 returned rc 0.
> Nov  9 13:55:09 opalVM setroubleshoot:      SELinux is
> preventing /sbin/ifconfig (ifconfig_t) "read write" to
> socket:[6686] (initrc_t).      For complete SELinux
> messages. run sealert -l
> 2def91cb-1661-4e67-a845-ff44f62aa3eb
> Nov  9 13:55:17 opalVM heartbeat: [2019]: info: Local
> Resource acquisition completed. (none)
> Nov  9 13:55:17 opalVM heartbeat: [2019]: info: local
> resource transition completed.
> Nov  9 13:55:17 opalVM heartbeat: [2019]: info:
> AnnounceTakeover(local 1, foreign 1, reason
> 'T_RESOURCES(us)' (1))
> Nov  9 13:55:39 opalVM hb_standby[2717]: Going standby
> [foreign].
> Nov  9 13:55:39 opalVM heartbeat: [2019]: info:
> opalvm.cluster.brc.bcm.edu wants to go standby
> [foreign]
> Nov  9 13:55:39 opalVM heartbeat: [2019]: info:
> i_hold_resources: 3
> Nov  9 13:55:39 opalVM heartbeat: [2019]: info: New
> standby state: 1
> Nov  9 13:55:40 opalVM heartbeat: [2019]: WARN:
> Standby timer has 9610 ms left
> ===============
> 
> Here is my ha.cf file from the passive node on the
> private IP (it matches the active nodes ha.cf)
> 
> ==========================
> logfacility     daemon
> keepalive 1
> deadtime 10
> warntime 5
> initdead 120 # depend on your hardware
> udpport 694
> ucast eth0 192.168.1.22
> auto_failback off
> node opal.brc.bcm.edu
> node opalVM.cluster.brc.bcm.edu
> crm no
> /usr/local/HA-2.0.8/lib/heartbeat/ipfail
> debug 2
> 
> =====================
> And here is my haresources file from the passive node
> on the private IP.
> 
> =====================
> opal.brc.bcm.edu 128.249.65.204
> apache::/etc/httpd/conf/httpd.conf
> =====================
> 
> This Active/Passive config works on a test system
> where both nodes are on the same private network
> (172.30.10.X), but it is not working here with each
> node on a separate network (128.249.65.X and
> 172.30.10.X).
> 
> Note the following errors when heartbeat tries to
> reassign the IP address:
> 
> Nov  9 13:55:07 opalVM ResourceManager[2447]: info:
> Running /usr/local/HA-2.0.8/etc/ha.d/resource.d/IPaddr
> 128.249.65.204 start
> Nov  9 13:55:07 opalVM IPaddr[2525]: ERROR:
> /usr/local/HA-2.0.8/lib/heartbeat/findif
> 128.249.65.204 failed [rc=1].

You can also try to specify the network interface (see
http://www.linux-ha.org/HeartbeatResourceAgent/IPaddr). Can't
say why, but findif can't find it.

> Nov  9 13:55:07 opalVM IPaddr[2516]: ERROR:  Generic
> error
> Nov  9 13:55:08 opalVM ResourceManager[2447]: ERROR:
> Return code 1 from
> /usr/local/HA-2.0.8/etc/ha.d/resource.d/IPaddr
> Nov  9 13:55:08 opalVM ResourceManager[2447]: CRIT:
> Giving up resources due to failure of 128.249.65.204
> 
> ---------------------
> I have 3 physical NIC's (eth0, eth1, eth2).
> eth0 is the heartbeat detection.
> eth1 is on a VLAN for the public IP
> eth2 is on a VLAN for the private IP.
> 
> Both nodes can ping all other nodes in both networks
> (Public=128.249.65.X, Private=172.30.10.X, and
> Heartbeat=192.10.1.X)
> ----------------------
> If there is anything else anyone needs to help, please
> let me know. 
> 

Thanks,

Dejan

> 
>       
> ____________________________________________________________________________________
> Get easy, one-click access to your favorites. 
> Make Yahoo! your homepage.
> http://www.yahoo.com/r/hs 
> _______________________________________________
> Linux-HA mailing list
> [email protected]
> http://lists.linux-ha.org/mailman/listinfo/linux-ha
> See also: http://linux-ha.org/ReportingProblems
_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems

Reply via email to