aanogueira opened a new pull request, #13897:
URL: https://github.com/apache/apisix/pull/13897
## Problem
A freshly created health-check target defaults to `internal_health =
healthy` with zero probes (`add_target`'s hardcoded `is_healthy=true`), so a
pod that restarts while its backend is already unhealthy briefly routes real
traffic to it until enough active probes correct the target's state.
Separately, checker creation itself is entirely lazy -- seeded only by
`fetch_checker()` on the live request path -- so an idle-but-critical upstream
with no prior traffic would never get a checker built at all ahead of a
readiness check.
There is currently no way to tell "healthy" (a real check passed) apart from
"healthy" (the zero-probe default) from outside the checker, so nothing can
gate readiness on "has this actually been checked yet."
## Changes
- `apisix/healthcheck_manager.lua`: adds two accessors for a readiness
plugin to use --
- `ensure_checker(resource_path)`: proactively seeds a checker for a
resource even with zero prior traffic, reusing the existing
`timer_create_checker` construction path. Also resolves domain-name upstream
nodes via `parse_domain_in_up` up front -- that resolution otherwise only
happens on the live request path, so a checker built ahead of traffic would
start probing under an unresolved identity and get silently rebuilt (wiping its
probe count) the moment real traffic first resolved the domain.
- `is_resource_probed(resource_path)`: true only once every target of the
resource's checker has had enough real active-check attempts for its state to
have actually converged -- a single attempt is not always enough: with e.g.
`unhealthy.http_failures = 2` configured, `internal_health` only converges
after two consecutive attempts. Computes the required attempt threshold from
the checker's own config (`max(unhealthy.http_failures, .tcp_failures,
.timeouts, healthy.successes)`) and delegates to a companion module function,
`resty.healthcheck.all_targets_probed(name, shm_name, min_attempts)`, proposed
as a separate change against lua-resty-healthcheck-api7 (the vendored library
this repo depends on): a per-target probe-attempt counter in shm, incremented
each time an active check is actually dispatched for a target (success,
failure, or timeout all count -- attempted, not "healthy"), queryable from any
worker.
- `t/node/healthcheck-fresh-node-default-healthy.t`: new tests covering lazy
checker creation (`fetch_checker` returns `false` until the next timer tick),
`ensure_checker` building a checker with zero prior traffic,
`all_targets_probed` flipping only after a real probe, and the multi-attempt
threshold behavior specifically (stays `false` after 1 attempt when
`min_attempts=2`, flips `true` only after the 2nd).
Validated end to end in a local kind cluster: pod-restart-while-unhealthy
shows zero leaked requests before or after the readiness transition.
## Dependency note
The probe-counter shm mechanism lives in the vendored
`lua-resty-healthcheck-api7` library, not this repo. The library-side change is
being proposed as a standalone PR against api7/lua-resty-healthcheck.
## Testing
- `t/node/healthcheck-fresh-node-default-healthy.t` (this repo)
- Local kind end-to-end trial: pod restart while backend already unhealthy,
comparing leaked traffic and the readiness transition, stock vs. patched
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]