[ 
https://issues.apache.org/jira/browse/SOLR-18307?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18124564#comment-18124564
 ] 

Mikhail Khludnev edited comment on SOLR-18307 at 10/7/26 3:23 PM:
------------------------------------------------------------------

Just recording an idea for the next {{\{!auxIndexJoin}}} iteration:
* still stick to the sidecar index. I have a few ideas on how to fold it into 
the main index, but they look too invasive to be worth it.
* take the join leaf, i.e. the {{from -> to}} column, {{to_docs[from_doc]}}.
* break this column into a few stripes along disjoint {{to}} ranges.
* encode each stripe of {{to_docs[from_doc]}} into a blob, picking a dense or a 
sparse format per stripe, and store it in a BinaryDocValues (BinDV) field.
* add those stripe blobs as documents into the aux-index, along with the 
from-seg and to-seg ids, the stripe ordinal, and the bounding box — the {{to}} 
and {{from}} ranges.
* such an index is enough to get an approximation, which we then refine by 
draining the stripes.
* merging and sweeping become much easier this way.


was (Author: mkhludnev):
Just recording an idea for the next {{{!auxIndexJoin}}} iteration:
* still stick to the sidecar index. I have a few ideas on how to fold it into 
the main index, but they look too invasive to be worth it.
* take the join leaf, i.e. the {{from -> to}} column, {{to_docs[from_doc]}}.
* break this column into a few stripes along disjoint {{to}} ranges.
* encode each stripe of {{to_docs[from_doc]}} into a blob, picking a dense or a 
sparse format per stripe, and store it in a BinaryDocValues (BinDV) field.
* add those stripe blobs as documents into the aux-index, along with the 
from-seg and to-seg ids, the stripe ordinal, and the bounding box — the {{to}} 
and {{from}} ranges.
* such an index is enough to get an approximation, which we then refine by 
draining the stripes.
* merging and sweeping become much easier this way.

> segment-parallel query time join with sidecar index 
> ----------------------------------------------------
>
>                 Key: SOLR-18307
>                 URL: https://issues.apache.org/jira/browse/SOLR-18307
>             Project: Solr
>          Issue Type: Improvement
>          Components: query
>            Reporter: Mikhail Khludnev
>            Assignee: Mikhail Khludnev
>            Priority: Major
>              Labels: pull-request-available
>             Fix For: 10.1
>
>         Attachments: Screenshot from 2026-08-29 18-14-22.png, hotspot.md
>
>          Time Spent: 50m
>  Remaining Estimate: 0h
>
> Hereby I propose a new query-time join implementation
> * it's segment parallel 
> * it's lazy - utilizes Two-Phase searching and beneficial from highly 
> selective filter in "to"-side, as well as leap over whole "to" segments. 
> * uses Lucene DocValues in sidecar - just an addition, there's no hard 
> surgery on existing codebase. It's just a QParserPlugin.  
> * joins crosscore, and works in cloud mode
> * may use Memory Directory for sidecar
> * writes sidecar lazily, but you can write it ahead
> * sweeps segments which are not needed anymore 
> * the [simple benchmark|https://github.com/mkhludnev/aijoin-benchmark] 
> demonstrates a few times gain 
> WIP https://github.com/apache/solr/pull/4749
> benchmark: https://github.com/mkhludnev/aijoin-benchmark
> Caveat: during the demo 7/15/26 we evidence nearly the same size of sidecar 
> index. it should be further investigated. 
> Note: the name comes from Auxiliary Index Join. 



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to