zhangshenghang opened a new pull request, #12526:
URL: https://github.com/apache/seatunnel/pull/12526

   ## Purpose of this pull request
   
   ClickHouse sink SQL templates use sink field names as named parameters 
(`:fieldName`). `FieldNamedPreparedStatement.parseNamedStatement` scans 
parameter names character by character with `Character.isJavaIdentifierPart`, 
so when a column name contains characters that are not valid Java identifier 
parts — for example `#` (as in JDE-style columns like `GLREG#`), spaces, or `-` 
— the parameter is truncated at that character:
   
   - `:GLREG#` is parsed as parameter `GLREG` followed by a stray `#`, 
producing a broken `?#` placeholder
   - the sink then fails with `GLREG# doesn't exist in the parameters of SQL 
statement: ...`
   
   The jdbc connector already handles these column names since #8222 switched 
its placeholder scan to a regex-based match; the ClickHouse connector still 
uses the identifier scan and keeps this bug.
   
   ## Solution
   
   Pass the sink field names into `parseNamedStatement` and try whole-name 
matching against that list before falling back to the legacy identifier scan:
   
   - longest match wins, and a candidate is rejected when the character right 
after it would extend the identifier, so `GLREG` and `GLREG#` coexist as 
distinct parameters
   - only field names the legacy scan cannot represent (containing `#`, space, 
`-`, ...) enter the allow-list path; plain identifiers and dotted ClickHouse 
identifiers keep the old parsing
   - candidate names containing `:` are skipped, so a field name cannot swallow 
the next placeholder
   - the existing two-arg `parseNamedStatement` and the empty-name fail-fast 
check keep their exact behavior
   
   ## Verification
   
   - Unit tests (`mvn -pl seatunnel-connectors-v2/connector-clickhouse test`): 
new `SqlUtilsTest` cases cover `#`/space/`-` column names, the `GLREG` vs 
`GLREG#` prefix collision, legacy-behavior fallback on the old two-arg 
signature, and a guard against swallowing the next placeholder.
   - E2E: new 
`ClickhouseIT#testClickhouseSinkWithSpecialCharactersInColumnNames` copies a 
table with `GLREG`, `GLREG#`, `MY COL`, `COL-1` columns from a ClickHouse 
source to a ClickHouse sink and asserts the written rows field by field. The 
schema comes from the source table, so the special column names never appear as 
keys in the job config.
     - with this patch: job FINISHED, all rows written and verified
     - control run without the patch: job FAILED with 
`IllegalArgumentException` from `FieldNamedPreparedStatement.prepareStatement` 
(the same signature described above), confirming the E2E covers the regression
   
   Note: defining such column names directly inside a HOCON job config (e.g. in 
a FakeSource `schema.fields` block) currently fails earlier in 
`ConfigShadeUtils.decryptConfig`, because HOCON path parsing rejects `#` in 
keys. That is a separate pre-existing issue in the config layer; 
database-catalog-driven schemas used here are unaffected.
   
   ## Check list
   
   - [x] Code changed are covered with tests, or it does not need tests for 
some reasons.
   - [ ] If any new added or updated dependencies need to change license doc.
   - [ ] If necessary documents are added to the PR.
   - [ ] If this is a connector change, has the document of connector been 
updated.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to