First thing coming to mind is: ceph mgr fail

If you already did that something else is going on.

Zitat von Marek Szuba via ceph-users <[email protected]>:

Hello,

After the latest run of system updates and reboots on our prototype Ceph
cluster (Ubuntu Jammy + cephadm-managed Ceph 20.2.0 on Podman), we have
ended up with something seriously weird. Specifically, at a glance the
cluster appears to be in service:

* Ceph reports the health of the cluster as OK
* according to both 'ceph -s' and 'ceph orch ps', all the daemons which
should be up and running ARE up and running. In particular, all the
expected OSDs appear to be operational
* 'ceph -s' shows no recovery I/O

However:

* despite all of the above, 'ceph -s' reports 1448 remapped PGs and
about 24% objects classified as misplaced
* 'ceph pg ls remapped' does report a long list of PGs with misplaced
objects
* on said list, most if not all PGs seem (I haven't checked all 1448
lines) to have 3 OSDs each listed under ACTING but only either 1 or 2
listed under UP - but not in a consistent way, in that an OSD missing
from the UP list for one PG does appear on the UP lists of other PGs

Any idea what might be going on here, and how to make it stop? We were
about to upgrade that cluster to 20.2.1 following the system updates,
however we are very much hesitant to do so while things look so odd in
case the upgrade makes things worse.

--
MS

_______________________________________________
ceph-users mailing list -- [email protected]
To unsubscribe send an email to [email protected]


_______________________________________________
ceph-users mailing list -- [email protected]
To unsubscribe send an email to [email protected]

Reply via email to