Meng-Shuan Tsai created HDDS-16172:
--------------------------------------
Summary: scaling up OM replicas leaves the new OM pod stuck in
Init CrashLoopBackOff
Key: HDDS-16172
URL: https://issues.apache.org/jira/browse/HDDS-16172
Project: Apache Ozone
Issue Type: Sub-task
Reporter: Meng-Shuan Tsai
h3. Problem
Scaling up OM replicas (e.g. {{om.replicas: 2 -> 3}} with
{{{}om.persistence.enabled=true{}}}) leaves the new OM pod stuck in
{{Init:0/1}} CrashLoopBackOff forever, and {{helm upgrade --wait}} times out.
The {{om-bootstrap}} init container script
(charts/ozone/templates/om/om-bootstrap-configmap.yaml) runs:
{code:java}
ozone admin om --set "ozone.om.nodes.$OZONE_CLUSTER_ID=$overwriteCmd"
--bootstrap
{code}
but {{ozone admin om}} has no {{{}-{}}}set or {{-bootstrap}} option (neither in
2.0.0, the chart's appVersion when the script was introduced in HDDS-11618, nor
in the current 2.2.0). The command fails with exit code 2, the script retries 3
times and gives up, and the init container crash-loops:
{code:java}
Unknown options: '--set',
'ozone.om.nodes.cluster1=ozone-om-0,ozone-om-1,ozone-om-2', '--bootstrap'
Usage: ozone admin om [-hV] [--verbose] [COMMAND]
Bootstrap failed after 3 attempts with exit code 2
{code}
Note that simply fixing the command to {{ozone om --bootstrap}} is not enough:
that command keeps running as the OM daemon on success, so it can never
complete inside an init container, and the config check it performs
({{{}checkConfigBeforeBootstrap{}}}) requires the existing OMs to be restarted
with the new node list first - a rolling restart which is itself blocked by the
new pod being stuck in Init. Details and fix proposals will follow in the
comments/PR.
h3. Steps to reproduce (kind)
# Install with {{{}om.replicas=3{}}}, persistence enabled for om/scm/datanode
# {{helm upgrade}} to {{om.replicas=2}} (scale-down works)
# {{helm upgrade}} back to {{om.replicas=3}}
# {{ozone-om-2}} is stuck in {{Init:0/1}} CrashLoopBackOff with the log above
This path has never worked since it was introduced in HDDS-11618; chart CI
({{{}ct lint{}}} / {{{}ct install{}}}) never exercises a scale-up against a
live cluster, so it went unnoticed.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]