Hi Dumitru, Mairtin,
Thanks for the review and for applying patch 1. I sent a v2 of patch 2
with the fix moved into if_status_mgr_release_iface().
I would like to come back to the question at the end of the cover letter,
because the series only covers the tail of the incident. The dead backend
kept receiving traffic for about 95 minutes. The first 60 were the host
itself being gone: ovn-controller stalled, then the whole host died.
Port_Binding.chassis and up do not change in that state, so neither patch
applies. The series removes the remaining 35 minutes, after the host came
back without its VMs.
In my deployment the first part is the one that matters: any hypervisor
that hosts a backend and dies leaves that backend "online" until the host
reboots and probes again. BFD on the gateway chassis reported the host
down within seconds and moved the cr-lrps, but nothing connects that
knowledge to Service_Monitor.
So the question still stands: is it acceptable for OVN to set a
Service_Monitor offline based on chassis liveness, rather than on a probe
by the owning chassis? If so, which shape would you prefer:
A. the gateway ovn-controllers write "offline" for rows whose chassis
they see BFD-down (no schema change, but a non-owner writes the
status), or
B. each chassis publishes its BFD-down peers and northd decides with a
majority (single writer and a real quorum, but a schema change)?
I am happy to write an RFC for either, behind an option if you prefer.
Regards,
Jaygue
_______________________________________________
dev mailing list
[email protected]
https://mail.openvswitch.org/mailman/listinfo/ovs-dev