[
https://issues.apache.org/jira/browse/SOLR-18307?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121783#comment-18121783
]
David Smiley commented on SOLR-18307:
-------------------------------------
Excellent Mikhail! I used this today on a fairly realistic data set at work,
and found it indeed outperforms the built-in join handily. I have something I
shared in the last meetup that I have not yet publicly shared. Based on the
numbers I'm seeing across the 3, this is what I found (again, my query, my data
– YMMV):
Measured on one shard replica (~1.03M docs, 19 segments), `distrib=false`,
`cache=false`, query: _redacted_
|| ||{!compositeNested} + \{!child} + \{!parent}x3||{!join method=topLevelDV}
×3||{!auxIndexJoin} ×3 ||
||Warm query time|*~6 ms*|~190 ms|~12-15 ms|
||One-time cost|~430 ms to build the merged view (250-500 ms range, 4
builds)|~500 ms first query (top-level ordinal map)|1.5-2 s first query|
||What the one-time cost covers|All queries until the next commit (independent
of the query)|All joins on that field until the next commit|Only the segments
the inner query reaches; a new query touching other segments paid 1.5 s again|
||Can be moved to commit-time warming|Yes, one warming query|Yes|Yes, but the
warming query must match every doc on the inner side for full coverage (not yet
measured)|
||Index or schema changes|Use nested docs + index sorting. Needs a
concatenated nest path field.|None|Sidecar index on disk beside the main index|
||Status|Prototype|Stock Solr|Experimental in Solr 10|
Notes:
- All three return the same result. The join variants link entries to their
parent by `_parent_document_id`. The ` \{!compositeNested}` variant uses block
joins over a merged view where each doc's entries are nested under it.
- In this run, nothing was warmed at commit, so every one-time cost was paid
by the first user query after a commit.
- ` \{!auxIndexJoin}` ran with Solr's search executor threads at their default
(-1). Its full-coverage warming cost is still to be measured.
🤖 This comment was *partially* generated by AI (Claude Code).
> segment-parallel query time join with sidecar index
> ----------------------------------------------------
>
> Key: SOLR-18307
> URL: https://issues.apache.org/jira/browse/SOLR-18307
> Project: Solr
> Issue Type: Improvement
> Components: query
> Reporter: Mikhail Khludnev
> Assignee: Mikhail Khludnev
> Priority: Major
> Labels: pull-request-available
> Fix For: 10.1
>
> Attachments: Screenshot from 2026-08-29 18-14-22.png, hotspot.md
>
> Time Spent: 50m
> Remaining Estimate: 0h
>
> Hereby I propose a new query-time join implementation
> * it's segment parallel
> * it's lazy - utilizes Two-Phase searching and beneficial from highly
> selective filter in "to"-side, as well as leap over whole "to" segments.
> * uses Lucene DocValues in sidecar - just an addition, there's no hard
> surgery on existing codebase. It's just a QParserPlugin.
> * joins crosscore, and works in cloud mode
> * may use Memory Directory for sidecar
> * writes sidecar lazily, but you can write it ahead
> * sweeps segments which are not needed anymore
> * the [simple benchmark|https://github.com/mkhludnev/aijoin-benchmark]
> demonstrates a few times gain
> WIP https://github.com/apache/solr/pull/4749
> benchmark: https://github.com/mkhludnev/aijoin-benchmark
> Caveat: during the demo 7/15/26 we evidence nearly the same size of sidecar
> index. it should be further investigated.
> Note: the name comes from Auxiliary Index Join.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]