Hello Ian,

Thanks for the questions and review. Short answers below; Vivek, please
correct anything I get wrong.

1) Language: Python (3.10+), built on the official MCP Python SDK. It talks
to AsterixDB only over the existing HTTP API (/query/service and /admin/*),
so it doesn't touch any Java code in AsterixDB.

2) CI: Yes, but it can stay light. Unit tests, lint and type checks run in
a couple of minutes on GitHub Actions. What could be of more value is a
small live-integration job against the AsterixDB nightly. In my testing
that's where the issues showed up: mocked tests pass, but the gateway
behaves differently against a real cluster. It also ties into your point
about modularity. The MCP side can release on its own schedule, and CI
against the nightly tells us when an AsterixDB change breaks it. GitHub
Actions is also fine. Needs Apache RAT license checks, signed source
releases, and the Apache PyPI naming convention I believe.

3) What "sidecar" means here: a small separate process that runs next to
the cluster. An LLM client (Claude Desktop, an IDE, etc.) connects to the
sidecar over MCP, and the sidecar forwards bounded, read-only queries to
the CC. AsterixDB itself isn't changed; the database still enforces
read-only, because the sidecar forces readonly=true on every request. It
can run locally (stdio) or as a shared HTTP service with token auth.

4) Walkthrough: The short version against Vivek's GitHub repo:

# AsterixDB (you can download the nightly release from our website
https://asterixdb.apache.org/download.html)
unzip asterix-server-0.9.10-SNAPSHOT-binary-assembly.zip
cd apache-asterixdb-0.9.10-SNAPSHOT/opt/local && bin/start-sample-cluster.sh

# MCP server
git clone https://github.com/Vivek1106-04/asterixdb-mcp-server && cd
asterixdb-mcp-server
python -m venv .venv && source .venv/bin/activate && pip install -e ".[dev]"
ASTERIXDB_MCP_CC_BASE_URL=http://localhost:19002 asterixdb-mcp-server

Then point any MCP client at the asterixdb-mcp-server command (the README
has the Claude Desktop config).

I ran it end-to-end against our snapshot locally. The core query and schema
tools work well. I found a few nits/issues to fix before we propose moving
it into an asterixdb-mcp repo, and I'll share those with Vivek separately
or fix it. If we choose to maintain it under apache license, we might have
to add license headers.

@Vivek, Thank you for your effort and sorry that things are taking time!
Great work and output over the summer.

Thanks,
Suryaa

On Fri, Sep 25, 2026 at 8:18 AM Ian Maxon <[email protected]> wrote:

> Hi Vivek,
> Thanks for sharing this, and condensing everything into an APE. It
> looks really neat overall, and well designed to be relatively
> self-contained. I think one of the concerns that I had, at least, is
> that a lot of things in the area of LLMs move very quickly. Since
> AsterixDB is a relatively mature and stable codebase at this point,
> keeping it modular to the extent possible lets this move at its own
> pace, rather than being held up by practices and standards intended to
> constrain change for longstanding things within AsterixDB.
>
> I have some very high level questions about the MCP side of the repo.
> You will have to forgive me for my ignorance; I only really know about
> what has been going on in this area from a few brief words here and
> there either from you or Suryaa off-list. Maybe even one of you have
> answered some of these in the past and I have simply forgotten.
> 1 ) What language is it written in?
> 2 ) Does it need CI for each new change, or would that kind of be overkill?
> 3 ) It's mentioned it is a sidecar of sorts. How does this usually work?
> 4 ) Is there a possibility to get a walkthrough on how to compile and
> run everything, end-to-end? I know the goal I have talked about with
> Suryaa and you is that it should all go in the asterixdb-mcp repo, but
> until we close that loop, just whatever exists in a GitHub repo or
> fork would be fine.
>
> Thanks for all your hard work on this so far. All the demos I have
> seen of it look super neat. The fact things are taking a long time
> isn't a reflection on the quality of your effort, it's just that some
> things around here tend to get the lion's share of attention, and
> other things that are very interesting but not perceived as urgent can
> get less attention than they deserve.
> - Ian
>
>
> On Fri, Aug 28, 2026 at 10:27 AM Vivek Gangavarapu
> <[email protected]> wrote:
> >
> > Hi all,
> >
> > Initiating discussion on APE 36, an MCP gateway for AsterixDB.
> >
> > Feature: MCP Gateway
> >
> > The gateway lets agent clients talk to a cluster over MCP instead of
> being
> > handed /query/service and a prompt. It exposes the catalog, dataset
> > schemas, optimizer and Hyracks plans, ADVISE index recommendations and
> > cluster diagnostics as MCP tools, with results bounded so a SELECT *
> > doesn't fill the model's context. It runs as a standalone sidecar, not
> > inside the CC: MCP to the client, plain HTTP to the CC, no cluster state
> > and no caching. It also never decides for itself whether a statement
> > mutates — it sets readonly=true and lets the engine reject, rather than
> > keeping a deny-list that would drift from the grammar.
> >
> > Inside AsterixDB the change is small: function_metadata() (21483) so the
> > gateway reads the live function registry, and an optional /mcp proxy plus
> > dashboard panel (21492), off by default.
> >
> > APE:
> >
> https://cwiki.apache.org/confluence/spaces/ASTERIXDB/pages/449286300/APE+36+AsterixDB+MCP
> >
> > Thanks,
> >  Vivek
>

Reply via email to