Shreya Chakraborty created IMPALA-15220:
-------------------------------------------
Summary: [FE] Maintain an ANTLR4 SQL grammar to enable
multi-language parser generation (Python, C++, Go, Rust)
Key: IMPALA-15220
URL: https://issues.apache.org/jira/browse/IMPALA-15220
Project: IMPALA
Issue Type: Improvement
Components: Frontend
Reporter: Shreya Chakraborty
*Problem Statement*
Currently, Apache Impala's SQL parser is implemented in Java using JFlex / CUP
inside the `fe` (frontend) module (`org.apache.impala.analysis.Parser`).
While this works well for internal query planning inside the JVM coordinator,
it creates significant friction for external tools, client drivers, and
non-Java ecosystems:
1. *JVM Dependency:* Applications written in Python, Rust, Go, or C++ must
either spin up a heavyweight JVM process or rely on complex IPC wrappers just
to parse/validate Impala SQL statements.
2. *Ecosystem Fragmentation:* Third-party tools (CI/CD linters, query
formatters, IDE extensions like VS Code, and semantic routing proxies) cannot
easily re-use Impala's dialect parser natively.
*Proposed Solution*
Introduce an ANTLR4 grammar file (`ImpalaLexer.g4` and `ImpalaParser.g4`) for
the Impala SQL dialect along with standard Visitor/Listener base classes. While
ANTLR4 generates a Concrete Syntax Tree (CST) / Parse Tree out-of-the-box, its
multi-target code generation allows developers in any supported backend (C++,
Python, Go, Rust, Java) to implement generated Visitors or Listeners to walk
the CST and construct their own domain-specific ASTs or IR representations.
ANTLR4 generates standalone, high-performance parsers across multiple target
languages, including C++, Java, Python, Go, Rust, and JavaScript.
*Key Deliverables:*
1. *Grammar Specification (`.g4` files):* A formal ANTLR4 grammar matching
Impala’s SQL dialect (reserved keywords, DDL, DML, SELECT expressions, and
built-in functions).
2. *Build Integration:* Integrate ANTLR4 code generation into the build
pipeline to keep the grammar in sync with dialect updates.
3. *Reference Parsers/AST:* Provide reference target implementations or
artifacts for popular client languages (e.g., Python/Javascript).
*Key Use Cases & Benefits*
1. *Multi-Language Client Backends:* Native client-side parsing and query
validation in Python (e.g., `impala-python`), C++, Go, or Rust without
requiring a Java runtime.
2. *Tooling & IDE Support:* Enables syntax highlighting, auto-completion, and
static linting in editors (VS Code, JetBrains) and CI/CD SQL validation
pipelines.
3. *Lightweight AST Transformations:* Allows tools to extract tables, columns,
aliases, and predicates from queries before sending them to the cluster.
*Reference Implementation*
A prototype implementation demonstrating this pattern is being explored in Hue:
1. *Parse Tree Visitors:* Used specifically for extracting syntax structures to
feed structured feedback loops into LLMs and run schema validation against
metadata.
2. *Autocomplete Grammar Separation:* Note that while the core Impala parser
requires a strict grammar, editor interactive tools may utilize a
tailored/error-tolerant variant (`ImpalaAutocomplete.g4`) optimized for partial
token streams.
[PR in Hue|https://github.infra.cloudera.com/CDH/hue/pull/1096] -
grammar/impala/Common.g4 (for the impala grammar files)
Would love to get feedback from the maintainers on whether an ANTLR4 grammar
could serve as an official repository component or a sub-module.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)