On 9/30/26 8:37 AM, Jaygue Lee wrote:
> Hi Dumitru, Mairtin,
> 

Hi Jaygue.

> Thanks for the review and for applying patch 1.  I sent a v2 of patch 2
> with the fix moved into if_status_mgr_release_iface().
> 
> I would like to come back to the question at the end of the cover letter,
> because the series only covers the tail of the incident.  The dead backend

Sorry, I missed that part of the v1 cover letter, thanks for pointing it
out!

> kept receiving traffic for about 95 minutes.  The first 60 were the host
> itself being gone: ovn-controller stalled, then the whole host died.
> Port_Binding.chassis and up do not change in that state, so neither patch
> applies.  The series removes the remaining 35 minutes, after the host came
> back without its VMs.
> 
> In my deployment the first part is the one that matters: any hypervisor
> that hosts a backend and dies leaves that backend "online" until the host
> reboots and probes again.  BFD on the gateway chassis reported the host
> down within seconds and moved the cr-lrps, but nothing connects that
> knowledge to Service_Monitor.
> 

But isn't this true for anything that was "owned" by the chassis that is
currently down?  E.g., port bindings bound on that chassis will appear
as "reachable" to all other chassis.  The same with all dynamically
learned things, e.g. Learned_routes.

> So the question still stands: is it acceptable for OVN to set a
> Service_Monitor offline based on chassis liveness, rather than on a probe
> by the owning chassis?  If so, which shape would you prefer:
> 
>   A. the gateway ovn-controllers write "offline" for rows whose chassis
>      they see BFD-down (no schema change, but a non-owner writes the
>      status), or
>   B. each chassis publishes its BFD-down peers and northd decides with a
>      majority (single writer and a real quorum, but a schema change)?
> 
> I am happy to write an RFC for either, behind an option if you prefer.
> 

Isn't the right way to address this to make sure, on the CMS side, that
a "dead" chassis is removed?  That is, add an out of band (outside of
OVN) chassis liveness mechanism and if a chassis (or its ovn-controller
is stalled) then the CMS would just remove that host from the OVN
cluster and remove the SB.Chassis record.

With your current fixes that would address the problem in general,
wouldn't it?

Regards,
Dumitru

_______________________________________________
dev mailing list
[email protected]
https://mail.openvswitch.org/mailman/listinfo/ovs-dev

Reply via email to