tzphh opened a new issue, #10102: URL: https://github.com/apache/paimon/issues/10102
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon/issues) and found nothing similar. ### Paimon version 2.0.0 ### Compute Engine Flink 1.19.3 (downstream build), Java 17 ### Minimal reproduce step 1. Create an append-only table with a NOT NULL column and Data Evolution enabled: ```sql CREATE TABLE test_db.t ( id BYTES NOT NULL, feature STRING ) WITH ( 'bucket' = '-1', 'file.format' = 'parquet', 'row-tracking.enabled' = 'true', 'data-evolution.enabled' = 'true' ); ``` 2. Write complete rows, then use the Data Evolution partial-column writer to update only `feature` for the same row IDs. The update files contain only `feature`; `id` remains in the base files. This is not an ordinary INSERT that omits `id`. 3. Run database compaction in batch/divided mode: ```sql SET 'execution.runtime-mode' = 'batch'; CALL sys.compact_database( including_databases => 'test_db', mode => 'divided', including_tables => 'test_db[.]t' ); ``` Compaction takes the ordinary Append path and fails because the partial-column files do not contain the NOT NULL column. ### What doesn't meet your expectations? **Expected behavior** Database compaction should use the Data Evolution compaction path, merging partial-column files by row ID before writing complete rows. **Actual behavior** Compaction uses `AppendCompactTask` / `BaseAppendFileStoreWrite.compactRewrite` instead. The partial-column Parquet file does not contain a NOT NULL column that is stored in the base file. The ordinary Append compaction path does not merge these files by row ID, and rewriting fails with the following error (column name replaced with `id`): ```text java.lang.IllegalArgumentException: Field 'id' expected not null but found null value. at ParquetRowDataWriter$RowWriter.write(...) at BaseAppendFileStoreWrite.compactRewrite(...) at AppendCompactTask.doCompact(...) ``` **Local verification** With `paimon-bundle:2.0.0`, the ordinary Append compaction task fails earlier during reading: ```text Required column is missing in data file. Col: [id] ``` The dedicated `DataEvolutionNormalCompactTask` succeeds and preserves the original non-null IDs and updated values. The production and local failures occur at different stages, but both involve the ordinary Append compaction path. ### Anything else? **Possible cause** In divided mode, `CompactDatabaseAction` routes `BUCKET_UNAWARE` tables to `AppendTableCompact` without checking `dataEvolutionEnabled()`. Data Evolution tables also use this bucket mode, so they enter the ordinary Append compaction path. That path reads files separately instead of merging partial-column files by row ID. Single-table `CompactAction` already checks `dataEvolutionEnabled()` and selects the dedicated compactor. Relevant code: - [Database compaction routing](https://github.com/apache/paimon/blob/b07f7ffbb21ffdc2853bec6757545e8d2b9ea092/paimon-flink/paimon-flink-common/src/main/java/org/apache/paimon/flink/action/CompactDatabaseAction.java#L189-L204) - [Single-table compaction routing](https://github.com/apache/paimon/blob/b07f7ffbb21ffdc2853bec6757545e8d2b9ea092/paimon-flink/paimon-flink-common/src/main/java/org/apache/paimon/flink/action/CompactAction.java#L154-L165) **Suggested fix** Use `DataEvolutionTableCompact` for Data Evolution tables in batch/divided mode. Combined mode should also be handled explicitly to prevent Data Evolution tables from entering the ordinary Append compaction path. **Regression test** Write complete rows containing a NOT NULL column, then update only another column through Data Evolution. Verify that compaction succeeds and preserves the row count, original non-null IDs, and updated values. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
