Hi Ian and Suryaa, Thanks for reading through it. Suryaa covered it well, I'll just add a few details.
On 1, the gateway itself is all Python and only goes through the HTTP API. The one place it depends on AsterixDB code is function_metadata() (change 21483), which isn't merged yet. So against the nightly, the two function lookup tools won't work but everything else will. On 2, most of this is already set up in the GitHub repo, and it only takes a few minutes per run, so I don't think it's overkill. Every PR runs lint, strict type checks and the unit tests on Python 3.10 to 3.12. There's also a live job that boots a real single-node cluster and runs integration tests against it. Right now that job uses the 0.9.9 release, and moving it to the nightly is a small change. What's still missing is the Apache side: license headers in the source files, RAT, and a release process. On 3, one more detail. Besides queries, the gateway can cancel a query it started, which goes through DELETE /admin/requests/running. That's the only non-read call it makes. In the shared HTTP mode you can use a bearer token or OAuth, and it binds to localhost by default. On 4, Suryaa's steps are what I'd give too. To try the function tools or the dashboard panel, you'd need to build AsterixDB with 21483 applied instead of using the nightly build. I agree on keeping it modular. The MCP side changes a lot faster than AsterixDB does, so it makes sense for the gateway to release on its own schedule and just note which AsterixDB versions each release works with. Apache Doris does the same thing with a separate doris-mcp-server repo, so there's precedent within the ASF. And no rush at all on the review, I understand other things come first. GSoC is over but I'm planning to keep working on this and contributing to the repo either way. Suryaa, happy to fix nits/issues. Can you share them or just raise an issue in GitHub? Thanks, Vivek On Fri, 25 Sept 2026 at 21:49, Suryaa Charan Shivakumar via dev < [email protected]> wrote: > Hello Ian, > > Thanks for the questions and review. Short answers below; Vivek, please > correct anything I get wrong. > > 1) Language: Python (3.10+), built on the official MCP Python SDK. It talks > to AsterixDB only over the existing HTTP API (/query/service and /admin/*), > so it doesn't touch any Java code in AsterixDB. > > 2) CI: Yes, but it can stay light. Unit tests, lint and type checks run in > a couple of minutes on GitHub Actions. What could be of more value is a > small live-integration job against the AsterixDB nightly. In my testing > that's where the issues showed up: mocked tests pass, but the gateway > behaves differently against a real cluster. It also ties into your point > about modularity. The MCP side can release on its own schedule, and CI > against the nightly tells us when an AsterixDB change breaks it. GitHub > Actions is also fine. Needs Apache RAT license checks, signed source > releases, and the Apache PyPI naming convention I believe. > > 3) What "sidecar" means here: a small separate process that runs next to > the cluster. An LLM client (Claude Desktop, an IDE, etc.) connects to the > sidecar over MCP, and the sidecar forwards bounded, read-only queries to > the CC. AsterixDB itself isn't changed; the database still enforces > read-only, because the sidecar forces readonly=true on every request. It > can run locally (stdio) or as a shared HTTP service with token auth. > > 4) Walkthrough: The short version against Vivek's GitHub repo: > > # AsterixDB (you can download the nightly release from our website > https://asterixdb.apache.org/download.html) > unzip asterix-server-0.9.10-SNAPSHOT-binary-assembly.zip > cd apache-asterixdb-0.9.10-SNAPSHOT/opt/local && > bin/start-sample-cluster.sh > > # MCP server > git clone https://github.com/Vivek1106-04/asterixdb-mcp-server && cd > asterixdb-mcp-server > python -m venv .venv && source .venv/bin/activate && pip install -e > ".[dev]" > ASTERIXDB_MCP_CC_BASE_URL=http://localhost:19002 asterixdb-mcp-server > > Then point any MCP client at the asterixdb-mcp-server command (the README > has the Claude Desktop config). > > I ran it end-to-end against our snapshot locally. The core query and schema > tools work well. I found a few nits/issues to fix before we propose moving > it into an asterixdb-mcp repo, and I'll share those with Vivek separately > or fix it. If we choose to maintain it under apache license, we might have > to add license headers. > > @Vivek, Thank you for your effort and sorry that things are taking time! > Great work and output over the summer. > > Thanks, > Suryaa > > On Fri, Sep 25, 2026 at 8:18 AM Ian Maxon <[email protected]> wrote: > > > Hi Vivek, > > Thanks for sharing this, and condensing everything into an APE. It > > looks really neat overall, and well designed to be relatively > > self-contained. I think one of the concerns that I had, at least, is > > that a lot of things in the area of LLMs move very quickly. Since > > AsterixDB is a relatively mature and stable codebase at this point, > > keeping it modular to the extent possible lets this move at its own > > pace, rather than being held up by practices and standards intended to > > constrain change for longstanding things within AsterixDB. > > > > I have some very high level questions about the MCP side of the repo. > > You will have to forgive me for my ignorance; I only really know about > > what has been going on in this area from a few brief words here and > > there either from you or Suryaa off-list. Maybe even one of you have > > answered some of these in the past and I have simply forgotten. > > 1 ) What language is it written in? > > 2 ) Does it need CI for each new change, or would that kind of be > overkill? > > 3 ) It's mentioned it is a sidecar of sorts. How does this usually work? > > 4 ) Is there a possibility to get a walkthrough on how to compile and > > run everything, end-to-end? I know the goal I have talked about with > > Suryaa and you is that it should all go in the asterixdb-mcp repo, but > > until we close that loop, just whatever exists in a GitHub repo or > > fork would be fine. > > > > Thanks for all your hard work on this so far. All the demos I have > > seen of it look super neat. The fact things are taking a long time > > isn't a reflection on the quality of your effort, it's just that some > > things around here tend to get the lion's share of attention, and > > other things that are very interesting but not perceived as urgent can > > get less attention than they deserve. > > - Ian > > > > > > On Fri, Aug 28, 2026 at 10:27 AM Vivek Gangavarapu > > <[email protected]> wrote: > > > > > > Hi all, > > > > > > Initiating discussion on APE 36, an MCP gateway for AsterixDB. > > > > > > Feature: MCP Gateway > > > > > > The gateway lets agent clients talk to a cluster over MCP instead of > > being > > > handed /query/service and a prompt. It exposes the catalog, dataset > > > schemas, optimizer and Hyracks plans, ADVISE index recommendations and > > > cluster diagnostics as MCP tools, with results bounded so a SELECT * > > > doesn't fill the model's context. It runs as a standalone sidecar, not > > > inside the CC: MCP to the client, plain HTTP to the CC, no cluster > state > > > and no caching. It also never decides for itself whether a statement > > > mutates — it sets readonly=true and lets the engine reject, rather than > > > keeping a deny-list that would drift from the grammar. > > > > > > Inside AsterixDB the change is small: function_metadata() (21483) so > the > > > gateway reads the live function registry, and an optional /mcp proxy > plus > > > dashboard panel (21492), off by default. > > > > > > APE: > > > > > > https://cwiki.apache.org/confluence/spaces/ASTERIXDB/pages/449286300/APE+36+AsterixDB+MCP > > > > > > Thanks, > > > Vivek > > >
