NestDream opened a new pull request, #9717:
URL: https://github.com/apache/paimon/pull/9717

   ### Purpose
   
   fix #9716
   
   When the catalog is case-insensitive, 
`ComputedColumnUtils.sortComputedColumnArgs` (introduced in #5972, shipped in 
1.3.0) upper-cases the whole `--computed_column` argument before parsing it. 
Three things break:
   
   - Literals change case. `dt=date_format(create_time,yyyy-MM-dd)` is built 
with `YYYY-MM-DD` (week year, day of year): a record with `create_time = 
2023-03-23 10:15:00` lands in `dt=2023-03-82` with no error. `cast(hello, 
STRING)` becomes `HELLO`.
   - The field reference is upper-cased and the parsers look it up in the 
source record by exact name, so it only matches upper-case source columns. With 
lower-case columns the computed value is null; for a partition key the job 
fails on every record with `Cannot write null to non-null column(dt)`.
   - A computed column referencing another one fails with `Referenced field 
'_year' is not in given fields`: the type is registered under the upper-cased 
name while `ReferencedField` looks it up lower-cased.
   
   This is the path where the Paimon table already exists and the schema cannot 
be read from the source (Kafka/Pulsar with an empty topic at startup), the only 
caller that passes `caseSensitive` to `buildComputedColumns`. Hive catalogs are 
case-insensitive by default, JDBC catalogs always. The Javadoc of 
`buildComputedColumns` already says field names are not changed at building 
phase; #5972 broke that.
   
   Fix:
   - `sortComputedColumnArgs` lower-cases only the keys used for the dependency 
sort (the form `ReferencedField` uses) and keeps names and arguments as typed. 
`buildComputedColumns` registers a computed column's type under the lower-cased 
name.
   - `ComputedColumn.evalFromRecord(Map)` matches the referenced field by exact 
name, then ignoring case when the catalog is case-insensitive. The record's 
column case is only known at runtime, and without this step removing the 
upper-casing would break a lower-case reference against upper-case source 
columns, which works today. The four parsers call it; MySQL, Postgres and 
MongoDB always build with `caseSensitive=true`, so nothing changes for them.
   
   ### Tests
   
   - `ComputedColumnUtilsTest`: literals kept as typed, cross-reference between 
computed columns written in different case, `evalFromRecord` for both catalog 
modes. The first two fail on master.
   - `KafkaCanalSyncTableActionITCase#testComputedColumnWithCaseInsensitive` 
adds `_DATE_STR=date_format(_DATE,yyyy-MM-dd)`; the 
`triggerSchemaRetrievalException=true` case times out on master and passes here.
   - The existing computed column ITCases for Kafka, MySQL, Postgres and 
MongoDB pass locally.
   - Reproduced on a Flink 1.20.1 standalone cluster with Kafka: table created 
first in a `case-sensitive=false` catalog, `kafka_sync_table` started on an 
empty topic, one canal-json record. Master writes `dt=2023-03-82` (upper-case 
source columns) or restarts on every record with `Cannot write null to non-null 
column(dt)` (lower-case source columns). With the fix both write 
`dt=2023-03-23`.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to