Prithvi, that looks mostly AI generated.  You should disclose that.  FWIW
my post was a 50/50 many-turn collaboration with Opus.

A possible option I didn't say is *not* to add this information to Solr
responses, leaving tracing as the place to get it.  I may well just do that.

On Sun, Sep 20, 2026 at 2:58 PM Prithvi S <[email protected]>
wrote:

> Hi David,
>
> Thanks for writing this up :)  so my understanding: after a distributed
> query you already need shards.info to know which replica served each
> shard,
> and today you still have to go correlate Luke / core admin / metrics after
> the fact, by which time the searcher may have rolled. In-band on that
> request is the right idea, especially searcher open time.
>
> +1 on putting the same facts on the shard sub-request spans no matter where
> the response body lands. That's the production path that doesn't depend on
> someone having thought to add a debug parameter, and sampling is already
> how tracing controls volume. I'd treat tracing as part of the same change,
> not a follow-up.
>
> On the response body I lean slightly the other way: #1, then #2, and I'd
> drop #3
> These aren't facts about how the query was parsed or scored. They're facts
> about the replica that served this request. That's the job shards.info
> already has, shardAddress is there specifically so you can tell which
> replica ran. Putting openedAt / segmentCount / numDocs / maxDocs next to
> shardAddress is the natural join. Under debug, operators still have to
> correlate two trees to answer the question you started with.
> A nested `searcher` block, included whenever shards.info is on, feels like
> enough of an opt-in to me. shards.info is already off by default, the
> payload is a handful of numbers, and the rest of that section is already
> diagnostic (time, error, shardAddress). A second parameter would make the
> common case, "add shards.info to this slow query", miss the new fields
> unless you already knew to ask twice. debug=all picking this up "for free"
> isn't actually the production win it sounds like, because debug=all also
> turns on explains.
>
> If we do go #2, two existing DebugComponent behaviors are worth planning
> around so it isn't quite the isolated add it first appears:
>   - modifyRequest currently strips debug off every shard request that isn't
> PURPOSE_GET_FIELDS, then re-adds timing and track. query / results never
> ride on GET_TOP_IDS. Searcher info needs to come from the replica that
> executed the query, including rows=0 (no GET_FIELDS). A new
> debug=indexsearcher would need to be forwarded the way track is, not the
> way query/results are.
>   - finishStage merges per-shard debug (parsed query, explains, …).
> Searcher facts have to stay per-replica, keyed by who produced them, the
> track shape, not the query/results shape. And debug=all is not quite free
> either: setDebug() / isDebug() / isDebugAll() are a closed set of four
> flags today, so a fifth value has to be wired through
> SolrPluginUtils.getDebugInterests as well.
>
> There's also a small precedent for #1 that keeps the producer local to the
> replica: QueryComponent already copies segmentTerminatedEarly /
> maxHitsTerminatedEarly / approximateTotalHits off the shard's
> responseHeader into shards.info. A SearcherDiagnostics helper on
> SolrIndexSearcher could fill that same role for both the response and the
> span attributes, so the grouping copies (including shards.info.firstPhase)
> stay mechanical.
>
> Small correction: debug=track is documented, briefly, in
> solrcloud-distributed-requests.adoc. I agree it shouldn't become a second
> shards.info.
>
> I'd skip "stale" for a first cut; open timestamp already answers "how old
> is this view." isLeader is cluster-state rather than SolrIndexSearcher, but
> it's cheap from CloudDescriptor and useful sitting next to the rest.
>
> I could live with #2 if that's the consensus, the debug namespace is a
> genuine plus, but I'd rather not add a third per-shard tree for replica
> identity we already have a home for.
>
> Cheers,
> Prithvi S
>
> On Sun, Sep 20, 2026 at 4:18 AM David Smiley <[email protected]> wrote:
>
> > Hi all,
> >
> > I'm looking to augment Solr responses with more detail about the searcher
> > to help diagnose a performance issue.  We could expose some cheap facts
> per
> > shard:
> >
> > - searcher open timestamp - when the currently-registered searcher was
> > opened (what I'm most interested in right now)
> > - segment count
> > - maxDocs
> > - numDocs
> > - isLeader?
> > - stale? (changes are pending visibility by a new searcher, e.g. a
> commit.
> > Maybe expose staleness as a time duration instead.  Disclaimer: I tacked
> > this on last minute; purely hypothetical addition I probably won't
> actually
> > do)
> >
> > All but the last are easily retrieved from the SolrIndexSearcher.
> Together
> > they make it easy to spot "shard2 is running a searcher from 5 minutes
> ago"
> > (therefore may be warm / performant... or return ~stale results) or
> "these
> > two shards disagree on doc counts" - currently you have to go correlate
> > metrics or core admin calls out-of-band, per replica, after the fact. And
> > in a distributed query the coordinator picks one replica per shard, so
> > afterward, you need shards.info to know which replica actually served
> you.
> >
> > Producing the data is the easy part. The conundrum is where to put it.
> Solr
> > already has multiple overlapping locations for per-shard diagnostics, so
> > potentially adding a third requires careful consideration. I see three
> > plausible homes.  A dimension that distinguishes them is how you request
> > the information.
> >
> > Option 1: shards.info
> >
> > Add the fields to the existing shards.info section.
> >
> > This is the established spot for per-shard information - it already
> carries
> > numFound, maxScore, per-shard timing and errors, and it covers
> single-shard
> > collections too, not just multi-shard ones. The main cost is that
> > shards.info is currently assembled in several independent code paths, so
> > any new field has to be added in each of them (the grouping path even
> uses
> > a different top-level key).
> >
> > If we go this way there's a follow-on question: does the new information
> > simply always appear whenever shards.info is requested, or does it need
> > its
> > own parameter to opt in? These fields are more debug-flavored than the
> rest
> > of what shards.info carries, so reasonable people may differ.
> >
> >
> > Option 2: a new debug value, e.g. debug=indexsearcher
> >
> > Alongside the existing debug=query / timing / results / track.
> >
> > This has a single implementation point, composes with the other debug
> > values, and debug=all picks it up for free. Notably it also answers the
> > opt-in question for free - debug is already a multi-valued parameter, so
> > asking for this and not asking for it is a solved problem needing no new
> > parameter. It keeps diagnostic payload out of shards.info too.
> >
> > The small drawback is that it adds yet another *per-shard* structure to
> the
> > response, on top of shards.info and debug=track - which is awkward given
> > how much those two already overlap.  Perhaps debug=track may not "count"
> > since it's undocumented and perhaps should stay that way.
> >
> >
> > Option 3: a nested block in each shard's response header (new)
> >
> > Each shard could report its own searcher info in its responseHeader, and
> > the coordinator surfaces the per-shard copies.
> >
> > The appeal is that a non-distributed request would see it with no extra
> > machinery at all, since the response header is always returned. The
> > drawback is it would need a new parameter to ask for it explicitly, since
> > we wouldn't want this on by default for every request.  And of course
> it's
> > existence may beg the question of why not instead add onto shards.info
> or
> > debug.
> >
> >
> > Regardless of which option wins: distributed tracing
> >
> > Whichever home we pick for the response body, I'd like the same
> information
> > to also be set as attributes on the distributed tracing spans. Each shard
> > sub-request already gets its own span, so these facts belong naturally on
> > the span for the shard that produced them.
> >
> > Arguably that's where the data is most valuable. Tracing already
> correlates
> > the whole request tree for you, including which replica served each
> shard,
> > and it doesn't require anyone to have thought to add a debug parameter
> to a
> > production query - by which time the searcher may well have rolled over
> and
> > the evidence is gone.  Tracing already incorporates the notion of
> > "sampling" to avoid excess information.
> >
> >
> > Finally, I don't have a strong preference, but I do lean a little in this
> > order #2, #1, #3.  I like debug because (a) I like the structured
> namespace
> > for choosing different debug categories -- just add another (b) the
> > implementation is more straight-forward from what I'm seeing instead of
> > shards.info which has some gotchas and is more spread out instead of
> > isolated to a component.
> >
> > ~ David
> >
>

Reply via email to