Shreya Chakraborty created IMPALA-15220:
-------------------------------------------

             Summary: [FE] Maintain an ANTLR4 SQL grammar to enable 
multi-language parser generation (Python, C++, Go, Rust)
                 Key: IMPALA-15220
                 URL: https://issues.apache.org/jira/browse/IMPALA-15220
             Project: IMPALA
          Issue Type: Improvement
          Components: Frontend
            Reporter: Shreya Chakraborty


*Problem Statement*
Currently, Apache Impala's SQL parser is implemented in Java using JFlex / CUP 
inside the `fe` (frontend) module (`org.apache.impala.analysis.Parser`). 

While this works well for internal query planning inside the JVM coordinator, 
it creates significant friction for external tools, client drivers, and 
non-Java ecosystems:
1. *JVM Dependency:* Applications written in Python, Rust, Go, or C++ must 
either spin up a heavyweight JVM process or rely on complex IPC wrappers just 
to parse/validate Impala SQL statements.
2. *Ecosystem Fragmentation:* Third-party tools (CI/CD linters, query 
formatters, IDE extensions like VS Code, and semantic routing proxies) cannot 
easily re-use Impala's dialect parser natively.

*Proposed Solution*
Introduce an ANTLR4 grammar file (`ImpalaLexer.g4` and `ImpalaParser.g4`) for 
the Impala SQL dialect along with standard Visitor/Listener base classes. While 
ANTLR4 generates a Concrete Syntax Tree (CST) / Parse Tree out-of-the-box, its 
multi-target code generation allows developers in any supported backend (C++, 
Python, Go, Rust, Java) to implement generated Visitors or Listeners to walk 
the CST and construct their own domain-specific ASTs or IR representations.

ANTLR4 generates standalone, high-performance parsers across multiple target 
languages, including C++, Java, Python, Go, Rust, and JavaScript.

*Key Deliverables:*
1. *Grammar Specification (`.g4` files):* A formal ANTLR4 grammar matching 
Impala’s SQL dialect (reserved keywords, DDL, DML, SELECT expressions, and 
built-in functions).
2. *Build Integration:* Integrate ANTLR4 code generation into the build 
pipeline to keep the grammar in sync with dialect updates.
3. *Reference Parsers/AST:* Provide reference target implementations or 
artifacts for popular client languages (e.g., Python/Javascript).

*Key Use Cases & Benefits*
1. *Multi-Language Client Backends:* Native client-side parsing and query 
validation in Python (e.g., `impala-python`), C++, Go, or Rust without 
requiring a Java runtime.
2. *Tooling & IDE Support:* Enables syntax highlighting, auto-completion, and 
static linting in editors (VS Code, JetBrains) and CI/CD SQL validation 
pipelines.
3. *Lightweight AST Transformations:* Allows tools to extract tables, columns, 
aliases, and predicates from queries before sending them to the cluster.

*Reference Implementation* 
A prototype implementation demonstrating this pattern is being explored in Hue: 
1. *Parse Tree Visitors:* Used specifically for extracting syntax structures to 
feed structured feedback loops into LLMs and run schema validation against 
metadata. 
2. *Autocomplete Grammar Separation:* Note that while the core Impala parser 
requires a strict grammar, editor interactive tools may utilize a 
tailored/error-tolerant variant (`ImpalaAutocomplete.g4`) optimized for partial 
token streams. 

[PR in Hue|https://github.infra.cloudera.com/CDH/hue/pull/1096] - 
grammar/impala/Common.g4 (for the impala grammar files)

Would love to get feedback from the maintainers on whether an ANTLR4 grammar 
could serve as an official repository component or a sub-module.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to