Hi Julian,

Thanks for the detailed report. I ran each item on a build of main
(11.0.0-SNAPSHOT, 78e219d), on a two-shard SolrCloud node or a user-managed
node, with a control request or a control run for each.

1. Reproduced. `fl=aliasId:id` with `distrib.singlePass=true` returns a 500
with the NPE in `mergeIds`; the combined query handler fails the same way
(`"fields": ["key:id"]` with `distrib.singlePass`). Filed as SOLR-18496,
with a fix and tests in https://github.com/apache/solr/pull/5019.

2. Not reproduced.

   - `SchemaManager`: the stream from ZooKeeper is a
`ByteArrayInputStream`, so closing it changes nothing. 150 Schema API
writes left the number of open file descriptors unchanged, and with
`managed-schema.xml` removed from ZooKeeper the call fails before any
stream is opened.

   - `SystemIdResolver`: the JDK parser closes the stream the resolver
returns (1 opened, 1 closed).

3. The NPE happens only on a core without an `<updateLog>`. The Ref Guide
says the transaction log is required for Real-time Get, so I did not file
it; other components fail there with "updateLog must be enabled", and RTG
could do the same.

4. (a) Reproduced: eight concurrent RELOADs of a core that failed to load
create the core eight times, one succeeds, seven fail on the index lock,
and the core stays in `initFailures` while serving. Filed as SOLR-18497
with a fix and a test in https://github.com/apache/solr/pull/5020.

   (b)-(f): I ran concurrent stress for reload against unload (100 rounds),
lazy load against unload (50), rename against create (60), swap against
unload (60) and two creates into one instance directory (80). None showed
an inconsistency. That does not rule them out; rename and swap take no
reservation in the code, as you describe.

The code behind 1, 3 and 4(a) is the same in the 9.10.1 and 10.0.0 sources;
I only ran it on main.

If you would like to take over either fix, say so on the issue.

Serhiy



On Sat, 3 Oct 2026 at 19:30, Eric Pugh via dev <[email protected]> wrote:

> Thanks for contributing..  Also, I live in Cville and am a Wahoo, so nice
> to see a Virginia.edu email address ;-).
>
> Solr is a volunteer project, so just because a bug may exist, doesn't mean
> fixing it ends up being top of the list!
>
> So, the best way to move forward is to pick one of these bugs and see if
> you can make a reproducing script that people can run.   If you can provide
> a script or unit test that shows say the renaming a unique key via FL is a
> bug, then that is the best first step.    One thing that I like about a
> script that reproduces the problem is that it confirms that it’s not some
> theoretical bug.   If you can use “bin/solr start -e techproducts” and then
> show the steps that show the bug, that is great.    Open up a JIRA and then
> a pull request with your reproducing bug fix, and then supply the fix.
>
> That makes it much easier for our very busy set of committers to quickly
> evaluate the potential fix and then give feedback or merge it.
>
> Eric
>
>
>
>
> > On Oct 3, 2026, at 11:49 AM, Julian Chandra Sutadi <
> [email protected]> wrote:
> >
> > Hello, I am a student from Bandung Institute of Technology, currently
> doing
> > a research collaboration with a PhD student from University of Virginia
> > about distributed systems. During our research, we found these potential
> > bugs in Solr’s current main branch (b5c71bc) that are similar to bugs
> that
> > were previously fixed in Solr. I would really appreciate it if you are
> able
> > to help us confirm whether these are actual undiscovered bugs or not.
> >
> >
> > 1. Renaming the unique key via fl can cause an NPE during merging in
> > QueryComponent and CombinedQueryComponent (similar to SOLR-15273)
> > SOLR-15273 fixed an NPE during the response construction of distributed
> > grouped queries when the unique key is renamed via an fl alias. It added
> a
> > rename lookup to StoredFieldsShardResponseProcessor, which is used only
> for
> > grouped queries:
> >
> >    if (rb.rsp.getReturnFields().getFieldRenames().get(uniqueIdFieldName)
> > != null) {
> >      // if id was renamed we need to use the new name
> >      uniqueIdFieldName =
> > rb.rsp.getReturnFields().getFieldRenames().get(uniqueIdFieldName);
> >    }
> >
> > Non-grouped queries can still fail the same way when all of the following
> > apply:
> >  a. distrib.singlePass=true is enabled.
> >  b. The schema’s unique key field is renamed through an fl field alias
> > (e.g. fl=aliasId:id) and not also requested as id.
> >  c. The query returns at least one document.
> >
> > If the schema unique key is id, shard responses contain the field as
> > aliasId. However, on a non-group query, the program will take the
> > QueryComponent.handleRegularResponses path that is not covered by
> > SOLR-15273 fix, which will invoke QueryComponent.mergeIds and then
> > QueryComponent.returnFields. The returnFields method already uses the
> > renamed key (SOLR-6744), but mergeIds runs first and still reads id,
> which
> > will cause NPE since the id is renamed. CombinedQueryComponent overrides
> > mergeIds and has the same lookup.
> >
> > The suggested fix is to apply the same rename lookup in
> > QueryComponent.mergeIds and CombinedQueryComponent.mergeIds, but only for
> > single-pass shard requests.
> >
> >
> > 2. InputStreams opened by SystemIdResolver and SchemaManager can leak
> > (similar to SOLR-16628)
> > SOLR-16628 reported that resource leak (leaking InputStream) may occur
> > during XML config file parsing. The issue was fixed by adding explicit
> > cleanup of the InputStream. However, there are still two code paths in
> the
> > codebase where Solr opens a stream and does not close it itself.
> >
> > i) SchemaManager.getFreshManagedSchema
> > SchemaManager.getFreshManagedSchema opens a schema stream and, in one
> > branch, passes it to IndexSchemaFactory.getParsedSchema, the same way
> > IndexSchemaFactory.loadConfig does. SOLR-16628 wrapped the stream in
> > IndexSchemaFactory.loadConfig with try-with-resources, but
> > getFreshManagedSchema still opens it without try-with-resources or a
> > finally block:
> >
> >    // IndexSchemaFactory.loadConfig (fixed by SOLR-16628)
> >    try (InputStream is =
> >        (schemaInputStream == null ? loader.openResource(name) :
> > schemaInputStream)) {
> >      ConfigNode node = getParsedSchema(is, loader, name);
> >      ...
> >    }
> >
> >    // SchemaManager.getFreshManagedSchema
> >    InputStream in = resourceLoader.openResource(schemaResourceName);
> >    if (in instanceof ZkSolrResourceLoader.ZkByteArrayInputStream) {
> >      ... () -> IndexSchemaFactory.getParsedSchema(in, zkLoader, ...) ...
> > // closed only by the parser
> >    } else {
> >      return (ManagedIndexSchema) core.getLatestSchema();              //
> > `in` never closed
> >    }
> >
> >  - if branch: the stream comes from ZooKeeper, and only the parser closes
> > it. This is the pattern where SOLR-16628 would add an explicit cleanup
> >  - else branch: the schema wasn't in ZooKeeper, so
> > ZkSolrResourceLoader.openResource fell back to
> > classLoader.getResourceAsStream. The method returns the in-memory schema
> > via core.getLatestSchema() and drops `in` without ever closing it.
> >
> > The suggested fix would be to open the stream with try-with-resources, as
> > in IndexSchemaFactory.loadConfig
> >
> > ii) SystemIdResolver
> > The XML parser calls back into SystemIdResolver to open resources for
> > external entities, XIncludes, and xsl:import/xsl:include. The resolver
> > opens the resource and hands the stream to the parser without keeping a
> > reference to it:
> >
> >    // Inside SystemIdResolver.resolveEntity()
> >    final InputSource is = new InputSource(loader.openResource(path));
> > // L144  stream opened
> >    is.setSystemId(uri.toASCIIString());
> >    is.setPublicId(publicId);
> >    return is;                         // no reference kept
> >
> > Hence, after the call returns, Solr has nothing it can use to close the
> > stream.
> >
> >
> > 3. Possible UpdateLog NPE in RealtimeGetComponent.java
> >
> >    // RealTimeGetComponent.java
> >    105  public class RealTimeGetComponent extends SearchComponent {
> >    117    public void process(ResponseBuilder rb) throws IOException {
> >           ...
> >    224      SolrDocumentList docList = new SolrDocumentList();
> >    225      UpdateLog ulog = core.getUpdateHandler().getUpdateLog();
> >  <-- null if no <updateLog>
> >             ...
> >    245        for (String idStr : reqIds.allIds) {
> >    246          fieldType.readableToIndexed(idStr, idBytes);
> >    247          // if _route_ is passed, id is a child doc.  TODO remove
> > in SOLR-15064
> >    248          if (!opennedRealtimeSearcher &&
> > !params.get(ShardParams._ROUTE_, idStr).equals(idStr)) {
> >    249            searcherInfo.clear();
> >    250            resultContext = null;
> >    251            ulog.openRealtimeSearcher(); // force open a new
> > realtime searcher   <-- unguarded
> >    252            opennedRealtimeSearcher = true;
> >    253          } else if (ulog != null) {
> >           <-- guarded
> >    254            Object o = ulog.lookup(idBytes.get());
> >
> > RealTimeGetComponent.process gets ulog from getUpdateLog() (L225), which
> is
> > null when no <updateLog> is configured. One branch checks it (L253), but
> > the branch before it, taken when the request's _route_ differs from the
> id
> > (L251), calls ulog.openRealtimeSearcher() unguarded. On a core without an
> > update log, this may throw an NPE.
> >
> > Suggested fix: add a simple check (ulog != null) check at L251.
> >
> >
> > 4. Race conditions on core lifecycle operations (follow up to SOLR-14969)
> > SOLR-14969 fixed the race condition that occurred during concurrent
> > CoreContainer.create operations, causing the second create’s
> > CoreContainer.createFromDescriptor invocation to fail. This was because
> the
> > core existence check was run before the waitAddPendingCoreOps reservation
> > and was not repeated after it, allowing concurrent creates to both pass
> the
> > check. It was fixed by adding the inFlightCreations reservation
> mechanism,
> > but only for create operation.
> >
> > Other core lifecycle paths have similar races:
> >  a. Reload of a failed core. reload checks if the core is not loaded
> > (solrCores.getCoreFromAnyList) and has an initialization failure
> > (coreInitFailures.get(name)) before waitAddPendingCoreOps. It then calls
> > createFromDescriptor without rechecking. Hence, a concurrent reload() may
> > recreate a core which has been recovered, very similar to SOLR-14969.
> >  b. Reload of a loaded core. reload calls solrCores.addCoreDescriptor
> > before waitAddPendingCoreOps. If a concurrent unload completes in
> between,
> > the descriptor may be added back for a core that no longer exists
> >  c. getCore reads the descriptor (solrCores.getCoreDescriptor) before
> > waitAddPendingCoreOps, and doesn’t reread it afterwards before invoking
> > createFromDescriptor. This way, a concurrent unload of a not-yet-loaded
> > core can be undone as the core is created and the descriptor
> re-registered.
> >  d. Rename (say a → b) core operation does not have any reservation (i.e.
> > no waitAddPendingCoreOps). It can race with a concurrent create of b. If
> > the concurrent create passes the existence check before rename succeeds,
> > both rename and create may race and invoke registerCore under the name b.
> >  e. The swap operation is done without any waitAddPendingCoreOps on
> either
> > name. It may race with other operations as illustrated above.
> >  f. coresLocator.create (CorePropertiesLocator’s method) does a
> non-atomic
> > Files.exists() then write of core.properties. Two concurrent creates with
> > different names but the same instance dir can both pass the check and
> race
> > on the same core.properties, causing the one who loses the lock
> contention
> > to delete the file.
> >
> >
> > If any of these are confirmed, I would be happy to open JIRA issues and
> > submit a PR for the fixes. I'm also glad to provide more details or run
> > additional checks if that would help.
> >
> > Thank you for your time.
> >
> > Best regards,
> > Julian
>
> Disclaimer
>
> The information contained in this communication from the sender is
> confidential. It is intended solely for use by the recipient and others
> authorized to receive it. If you are not the recipient, you are hereby
> notified that any disclosure, copying, distribution or taking action in
> relation of the contents of this information is strictly prohibited and may
> be unlawful.
>
> This email has been scanned for viruses and malware, and may have been
> automatically archived by Mimecast, a leader in email security and cyber
> resilience. Mimecast integrates email defenses with brand protection,
> security awareness training, web security, compliance and other essential
> capabilities. Mimecast helps protect large and small organizations from
> malicious activity, human error and technology failure; and to lead the
> movement toward building a more resilient world. To find out more, visit
> our website.
>

Reply via email to