Andrey Yarovoy created HDDS-16360:
-------------------------------------
Summary: Add a health check endpoint to S3 Gateway for load
balancer probes
Key: HDDS-16360
URL: https://issues.apache.org/jira/browse/HDDS-16360
Project: Apache Ozone
Issue Type: Improvement
Components: s3gateway
Reporter: Andrey Yarovoy
h3. Problem
S3 Gateway is horizontally scaled and fronted by a load balancer (the
distribution ships {{s3-haproxy.cfg}} for exactly this). Today S3G exposes no
purpose-built health endpoint, so operators fall back to probing an incidental
URL:
* The S3 listener on {{9878}} — but {{GET /}} is {{ListBuckets}} and any other
path is interpreted as a bucket/object, so there's no path that reliably means
"am I healthy" without colliding with a bucket name, and in secure mode it
requires AWS SigV4.
* An admin servlet on {{19878}} ({{{}/conf{}}}, {{{}/jmx{}}}, {{{}/prom{}}}) —
but these report only that the Jetty process is up; a 200 there does *not*
confirm the gateway can actually serve S3 (e.g. its {{{}OzoneClient{}}}/OM
connection may be broken). The bundled {{s3-haproxy.cfg}} in fact does no HTTP
health check at all — it balances raw TCP.
The result: a load balancer can keep routing traffic to an S3G instance that is
listening but cannot reach OM, or that is still starting up.
h3. Goal
Provide a stable, documented, unauthenticated HTTP endpoint that a load
balancer can poll to decide whether an S3G instance should receive traffic,
reflecting both process liveness and the gateway's ability to serve S3 requests.
h3. Proposed approach
Add a lightweight health servlet on {{S3GatewayWebAdminServer}} (the plain
Jetty admin server, port {{{}19878{}}}) at a fixed path (e.g. {{{}/health{}}}),
returning {{200}} when healthy and a non-2xx (e.g. {{{}503{}}}) when not, with
a tiny plaintext/JSON body for humans.
Serving it on the admin server rather than the S3 listener avoids the S3
path-router bucket-name collision entirely and keeps it off the
SigV4-authenticated path. Load balancers can still gate the S3 traffic port on
this endpoint: HAProxy in {{mode tcp}} performs Layer 7 HTTP health checks
independently of the traffic mode and can target a different port than the one
it balances ({{{}check port 19878{}}}), so it balances raw S3 TCP on {{9878}}
while HTTP-probing {{/health}} on {{{}19878{}}}.
The endpoint reports *Readiness* — the gateway can serve S3: its cached
{{OzoneClient}} is initialized and a cheap OM reachability check succeeds. This
is what the LB actually needs to gate traffic on.
The endpoint must be reachable by an LB without Kerberos/SPNEGO. Requires
explicitly excluding the health path from the S3G admin-server auth filters.
New {{@Config}} keys (e.g. enable/disable, probe timeout), defaulting to
enabled.
h3. Scope
* New health servlet + registration on {{{}S3GatewayWebAdminServer{}}}.
* Auth-filter exclusion for the health path in secure mode.
* {{@Config}} keys with getters (self-documenting; no {{ozone-default.xml}}
edits).
* Update {{s3-haproxy.cfg}} (and {{{}common/s3-haproxy.cfg{}}}) to add an
{{option httpchk}} against {{/health}} on the admin port, as a worked example.
h3. Out of scope
* Per-bucket / OM-side health, cluster-wide readiness aggregation.
* A health path on the S3 listener ({{{}9878{}}}) itself — unnecessary given
HAProxy can check the admin port.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]