Hi all, Apologies if there's a better place. Still fairly new to Calcite.
We're starting to use Calcite for query planning, just with the HepPlanner for now, with custom RelNodes and rules. Two of our patterns need several roots to share one planner, so equal subtrees resolve to the same canonical RelNode. What's the intended way to do this? The patterns: 1. Batch requests. One top level request may carry independent queries that can share large common subtrees. 2. Deferred resolution. Sometimes part of a query depends on an asynchronous lookup in order to resolve what the RelNodes should be underneath, so we plan the rest of the tree around a placeholder leaf and expand once the asynchronous lookup returns. Ideally we'd plan that with the same planner so it reuses nodes already in the graph. We have a workaround for #1 - we can wrap each independent query together in a synthetic multi-input root and pass that to HepPlanner, plan all at once with .setRoot+findBestExpr, and read the result back from walking the inputs of the result node. This doesn't help for pattern #2 . In HepPlanner, the graph is single-rooted. findBestExp() (and GC elsewhere outside large-plan mode) sweeps everything not reachable from the current root, which discards earlier roots. Ideas we've considered 1. Drive the planner by hand and delay GC until the end: call setRoot and executeProgram per root, then extract each plan at the end. This breaks when one query contains another, because a later pass can discard an earlier query's saved root. Working around that brings back the synthetic root, and it still doesn't help deferred resolution. 2. Re-plan with each expansion: plan the extracted plan plus the expansion under a new synthetic root in the same planner. This works, but because extraction rewrites the graph in place, only leaves are shared with the plan already running and everything above them is re-planned and re-evaluated. Is a synthetic multi-input the best way to deal with this kind of batch-query problem, or is/should there be a canonical alternative? How would you recommend approaching the deferred-query case, where we want new queries to dedupe against the plan we've extracted and already started executing? Would something here be welcome as a contribution? Thanks! Michael
