Can you please talk more about how you’re running the cluster and what data
is already on it? There’s nothing about growing a stretch cluster that
naively would lead to this behavior, but it does have a lot more
constraints than a regular Ceph cluster does about which OSDs it can use to
go active with (because it needs a current copy in both sites, and it’s
trying to keep 4 copies up to date rather than 3).
-Greg

On Mon, Apr 6, 2026 at 7:26 AM Anthony D'Atri via ceph-users <
[email protected]> wrote:

> Are you using upmap-remapped to mitigate the thundering herd? And, if
> you're using the PG autoscaler, to mitigate the thundering herd of backfill
> that it adds?
>
> > On Apr 6, 2026, at 7:33 AM, Michal Strnad via ceph-users <
> [email protected]> wrote:
> >
> > Hi.
> >
> > We operate a two‑site stretch cluster (version 20.2.0) with cca 2000
> OSDs and another 1000 OSDs to be added. We are concerned with issues
> related to a large number of peering connections and remappings when a node
> with 38 OSDs is added at once. Adding 38 OSDs triggers remapping of
> roughly 2000 PGs and can leave tens of minutes of PGs inactive, which we
> see reflected in the ceph status output before the cluster finally
> stabilizes.
> >
> > In a classic (non‑stretch) Ceph cluster the convergence is noticeably
> quicker (inactive PGs appear only for a very short interval). In our
> stretch deployment, however, the window when PGs are inactive is
> considerably longer, consistent with our observations. The network latency
> between the two sites is consistently below 10 ms with no packet loss.
> >
> > We are currently trying to determine whether this impacts client‑side
> read/write simulations via S3, but this has not yet been verified. Our next
> step is to test with RBD, which should be more resilient and demonstrable.
> Naturally, we would like to control the process in a way that minimizes the
> impact on clients. Could you please confirm that the periods of PG
> inactivity are short enough to prevent client visible glitches, or even
> better, could you recommend an optimal protocol for adding new OSDs to a
> cluster of this size to ensure a smooth convergence? Do you have any tested
> or recommended timeout adjustments for the stretch configuration?
> >
> > Thank you for any feedback
> > --
> > Michal
> > _______________________________________________
> > ceph-users mailing list -- [email protected]
> > To unsubscribe send an email to [email protected]
> _______________________________________________
> ceph-users mailing list -- [email protected]
> To unsubscribe send an email to [email protected]
>
_______________________________________________
ceph-users mailing list -- [email protected]
To unsubscribe send an email to [email protected]

Reply via email to