On 9/30/26 8:37 AM, Jaygue Lee wrote: > Hi Dumitru, Mairtin, > Hi Jaygue.
> Thanks for the review and for applying patch 1. I sent a v2 of patch 2 > with the fix moved into if_status_mgr_release_iface(). > > I would like to come back to the question at the end of the cover letter, > because the series only covers the tail of the incident. The dead backend Sorry, I missed that part of the v1 cover letter, thanks for pointing it out! > kept receiving traffic for about 95 minutes. The first 60 were the host > itself being gone: ovn-controller stalled, then the whole host died. > Port_Binding.chassis and up do not change in that state, so neither patch > applies. The series removes the remaining 35 minutes, after the host came > back without its VMs. > > In my deployment the first part is the one that matters: any hypervisor > that hosts a backend and dies leaves that backend "online" until the host > reboots and probes again. BFD on the gateway chassis reported the host > down within seconds and moved the cr-lrps, but nothing connects that > knowledge to Service_Monitor. > But isn't this true for anything that was "owned" by the chassis that is currently down? E.g., port bindings bound on that chassis will appear as "reachable" to all other chassis. The same with all dynamically learned things, e.g. Learned_routes. > So the question still stands: is it acceptable for OVN to set a > Service_Monitor offline based on chassis liveness, rather than on a probe > by the owning chassis? If so, which shape would you prefer: > > A. the gateway ovn-controllers write "offline" for rows whose chassis > they see BFD-down (no schema change, but a non-owner writes the > status), or > B. each chassis publishes its BFD-down peers and northd decides with a > majority (single writer and a real quorum, but a schema change)? > > I am happy to write an RFC for either, behind an option if you prefer. > Isn't the right way to address this to make sure, on the CMS side, that a "dead" chassis is removed? That is, add an out of band (outside of OVN) chassis liveness mechanism and if a chassis (or its ovn-controller is stalled) then the CMS would just remove that host from the OVN cluster and remove the SB.Chassis record. With your current fixes that would address the problem in general, wouldn't it? Regards, Dumitru _______________________________________________ dev mailing list [email protected] https://mail.openvswitch.org/mailman/listinfo/ovs-dev
