Hi,
Kay Röpke wrote:
Hi!
On Apr 7, 2009, at 4:27 AM, Stewart Smith wrote:
On Mon, Apr 06, 2009 at 12:03:15PM -0700, Eric Day wrote:
Not sure if I agree here. I think a pre-parser plugin/module would
be better than sticking this in the client library.
What about post-parser? i.e. trap the parse error and then try
remapping. This way wouldn't affect the 99.99% of real queries.
I wouldn't do it in a pre-parser stage, preprocessors are evil. In
reality you almost always want to stay inside the language and do
rewriting on trees rather than strings.
Just look at how hard it can be to understand C/C++ syntax errors when
they result from complex macro expansions (the problem only is that cpp
is not operating on C syntax trees but on strings).
How about a recognition pipeline?
the good thing is that it would provide a way for people to add on their own
parsers
while reusing the core facilities of existing parser(s), as opposed to:
a- writing their own parser from scratch
b- merging their parser with that of drizzle which in an extreme case would mean
going through lengthy code reviews, testing and still screwing everything up.
c- using a hack like sql pass through to a db engine
as an example, imagine that someone wants to implement some nice memory only
OLAP db engine
that prefers to deal with OLAP cubes rather than tables. Assuming that they
actually
succeeded in creating the engine, they would like to do:
CREATE CUBE .... sucks data from InnoDB tables into a cube and sets up some
magic to maintain the cube in sync
ROLLUP ... rollups some data in the cube, OLAP engine specific syntax
CREATE TABLE ... engine=InnoDB AS ROLLUP ... rollups some data and saves it
on disk in an InnoDB table.
a simplified example is specifying compression size for InnoDB Plugin which had
to be done by sharing the KEY_BLOCK_SIZE
syntax, but not the semantics, with MyISAM.
Thanks,
Michael Izioumtchenko
my opinions don't necessarily coincide with those of my employer
I can see two different approaches:
The first idea would be to chain the recognizers (= lexer + parser),
where the main SQL recognizer pair is the first one to have a shot at an
incoming query.
Should recognition fail, the entire thing get's passed to the next one
(multiple parsers could share lexers, thus avoiding expensive
re-tokenization, but they don't have to).
In case none of the recognizers can make sense of the input, the first
one would supply the error message, because it's the "main one".
In this approach both efficiency and quality of error messages are
likely to be problems, though. However, it's easier to implement than
the next one.
The second idea would be to have multiple recognizers that recognize
_disjoint_ sets of query syntax. They would form a tree, with the
easiest layout having a tree depth of 1.
Here the main problem is to efficiently pick a parser (not
lexer+parser!). This can be done very efficiently by retrieving the
possible start symbols from all the parsers and then delegating to the
one that knows about the symbol (e.g. a token "SHOW" would cause the
"SHOW" parser to run, the real SQL parser would have "SELECT", "INSERT",
... as start symbols).
The implementation would simply be a big switch based on the first token
(or even deeper into the token stream, if that's what you want). It can
be implemented by hand as a recursive descent parser with one token
lookahead (LL(1)) pretty easily.
The downside is that you need to have exactly one lexer to yield the
token used in picking the right parser. This lexer doesn't not need to
be the one used by the invoked recognizer, but then it would have to be
generated or manually kept in sync as the superset of all individual
lexers. Doing this at runtime, esp looking at supporting dynamically
loaded recognizers, is not impossible, but might not be the most
efficient nor most trivial code ever written ;)
What's the point of all this?
Well, given this architecture, you could support SHOW commands by having
the SHOW-parser generate a tree that looks exactly like the
corresponding information_schema query. This is completely transparent
for the client and eases maintenance should the I_S schema ever change:
Doing the preprocessing on the client side would mean that an old client
would constantly generate invalid I_S queries. Having it server side you
can keep both the SHOW commands and I_S schema compatible.
Another upside to the second approach are the improved error messages
you can generate. There's no confusion which recognizer is responsible
for what type of query.
cheers,
-k
P.S.: The concept of recognition delegation could easily extend to other
plugins, too, as long as the token that triggers this delegation is part
of the "main" lexer. Token stream implementation permitting you could
also stack lexers that then feed off of the same character stream.
Ambiguities are the biggest problem, especially when allowing plugins to
"inject" token rules into the main lexer. However, given it's upsides
for modularity and extensibility I think it's worth the risk of
backtracking a lot.
Have a look at:
- http://www.program-transformation.org/Transform/IslandGrammars
-
http://www.antlr.org/wiki/display/ANTLR3/Island+Grammars+Under+Parser+Control
_______________________________________________
Mailing list: https://launchpad.net/~drizzle-discuss
Post to : [email protected]
Unsubscribe : https://launchpad.net/~drizzle-discuss
More help : https://help.launchpad.net/ListHelp