Hi all,

Apologies if there's a better place. Still fairly new to Calcite.

We're starting to use Calcite for query planning, just with the HepPlanner
for now, with custom RelNodes and rules.  Two of our patterns need several
roots to share one planner, so equal subtrees resolve to the same canonical
RelNode. What's the intended way to do this?

The patterns:
1. Batch requests. One top level request may carry independent queries that
can share large common subtrees.
2. Deferred resolution. Sometimes part of a query depends on an
asynchronous lookup in order to resolve what the RelNodes should be
underneath, so we plan the rest of the tree around a placeholder leaf and
expand once the asynchronous lookup returns. Ideally we'd plan that with
the same planner so it reuses nodes already in the graph.

We have a workaround for #1 - we can wrap each independent query together
in a synthetic multi-input root and pass that to HepPlanner, plan all at
once with .setRoot+findBestExpr, and read the result back from walking the
inputs of the result node.

This doesn't help for pattern #2 . In HepPlanner, the graph is
single-rooted.  findBestExp() (and GC elsewhere outside large-plan mode)
sweeps everything not reachable from the current root, which discards
earlier roots.

Ideas we've considered
1. Drive the planner by hand and delay GC until the end: call setRoot and
executeProgram per root, then extract each plan at the end. This breaks
when one query contains another, because a later pass can discard an
earlier query's saved root. Working around that brings back the synthetic
root, and it still doesn't help deferred resolution.
2. Re-plan with each expansion: plan the extracted plan plus the expansion
under a new synthetic root in the same planner. This works, but because
extraction rewrites the graph in place, only leaves are shared with the
plan already running and everything above them is re-planned and
re-evaluated.

Is a synthetic multi-input the best way to deal with this kind of
batch-query problem, or is/should there be a canonical alternative? How
would you recommend approaching the deferred-query case, where we want new
queries to dedupe against the plan we've extracted and already started
executing? Would something here be welcome as a contribution?

Thanks!

Michael

Reply via email to