alessandro-mizzaro-sonarsource opened a new issue, #2603:
URL: https://github.com/apache/datafusion-sqlparser-rs/issues/2603
## Description
Several `Display` implementations do not restore delimiters consistently
when serializing decoded AST values or tokens.
As a result, SQL that initially parses as one statement can be displayed as
SQL that reparses as multiple statements.
Reproduced with `sqlparser` 0.63.0.
## Finding 1
`escape_quoted_string` applies the same backslash and doubled-quote
assumptions to SQL representations that use different escaping rules.
For example, it considers a quote following a backslash to be already
escaped without accounting for the parity of the preceding backslash run or the
rules of the specific literal.
### Reproducer
```rust
use sqlparser::{dialect::GenericDialect, parser::Parser};
fn main() {
let dialect = GenericDialect {};
let input = r#"SELECT 'safe\''; SELECT 2; --' AS harmless"#;
let first = Parser::parse_sql(&dialect, input).unwrap();
assert_eq!(first.len(), 1);
let serialized = first[0].to_string();
println!("serialized: {serialized}");
let second = Parser::parse_sql(&dialect, &serialized).unwrap();
assert_eq!(second.len(), 2);
}
```
Output:
```text
serialized: SELECT 'safe\'; SELECT 2; --' AS harmless
```
The input parses as one statement. Its displayed representation reparses as
two statements.
Confirmed affected representations include:
- Ordinary single-quoted strings with `GenericDialect` and `BigQueryDialect`
- Ordinary double-quoted strings with `BigQueryDialect`
- National strings with `BigQueryDialect`
- Single- and double-quoted byte strings with `BigQueryDialect`
- Single- and double-quoted raw strings with `BigQueryDialect`
- Hexadecimal strings with `GenericDialect`
- Double-quoted identifiers with `GenericDialect`
- Backtick-quoted identifiers with `MySqlDialect`
## Finding 2
Some AST formatters place decoded values directly between SQL delimiters
without restoring the required escaping.
Examples include:
- Triple-single- and triple-double-quoted strings
- Bracket-quoted identifiers
- PostgreSQL `COMMENT` values
- PostgreSQL `NOTIFY` payloads
### Reproducer
```rust
use sqlparser::{dialect::PostgreSqlDialect, parser::Parser};
fn main() {
let dialect = PostgreSqlDialect {};
let input =
r#"COMMENT ON TABLE harmless IS 'safe''; SELECT 2; --'"#;
let first = Parser::parse_sql(&dialect, input).unwrap();
assert_eq!(first.len(), 1);
let serialized = first[0].to_string();
println!("serialized: {serialized}");
let second = Parser::parse_sql(&dialect, &serialized).unwrap();
assert_eq!(second.len(), 2);
}
```
Output:
```text
serialized: COMMENT ON TABLE harmless IS 'safe'; SELECT 2; --'
```
The original `COMMENT` value contains the quote and the remaining text.
During serialization, the decoded value is inserted directly between single
quotes, causing the displayed representation to contain a second statement.
The same behavior was reproduced with triple-quoted BigQuery strings,
bracket-quoted MSSQL identifiers, and PostgreSQL `NOTIFY` payloads.
## Finding 3
`Token::Display` places the decoded contents of string tokens directly
between delimiters without restoring their escaping.
This affects token-based round trips even for representations whose AST
formatter may escape the value correctly.
### Reproducer
```rust
use sqlparser::{
dialect::GenericDialect,
parser::Parser,
tokenizer::Tokenizer,
};
fn main() {
let dialect = GenericDialect {};
let input = r#"SELECT 'safe''; SELECT 2; --' AS harmless"#;
let first = Parser::parse_sql(&dialect, input).unwrap();
assert_eq!(first.len(), 1);
let tokens = Tokenizer::new(&dialect, input).tokenize().unwrap();
let serialized = tokens
.iter()
.map(ToString::to_string)
.collect::<String>();
println!("serialized: {serialized}");
let second = Parser::parse_sql(&dialect, &serialized).unwrap();
assert_eq!(second.len(), 2);
}
```
Output:
```text
serialized: SELECT 'safe'; SELECT 2; --' AS harmless
```
The tokenizer decodes the doubled quote in the original literal.
`Token::Display` then emits that decoded quote without escaping it again.
Confirmed affected token representations include ordinary, national,
escaped, Unicode, hexadecimal, byte, raw, and triple-quoted string tokens.
## Expected behavior
`Display` may normalize the spelling or formatting of SQL, but its output
should remain valid SQL representing an equivalent AST or token value.
In particular, serialization should not change the number or type of
statements.
## Possible direction
- Escape delimiters according to the specific literal, identifier, statement
field, or token representation.
- Do not infer that decoded values are already escaped.
- Account for odd and even backslash runs where backslash escaping applies.
- Escape complete triple delimiters.
- Serialize `]` as `]]` inside bracket-quoted identifiers.
- Route quoted statement fields through appropriate typed formatters.
- Apply equivalent escaping rules to `Token::Display`, or provide a lossless
token reconstruction API.
- Add round-trip tests comparing statement count, statement type, and
normalized AST.
## Related work
- https://github.com/apache/datafusion-sqlparser-rs/issues/2409 covers the
bracket-quoted identifier case.
- https://github.com/apache/datafusion-sqlparser-rs/pull/2542 fixes the
basic hexadecimal literal case, but the shared escaping behavior remains
reproducible with even-length backslash runs.
Credit: Sonar
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]