voonhous opened a new issue, #19770:
URL: https://github.com/apache/hudi/issues/19770

   ## Bug Description
   
   **What happened:**
   
   On a table with `hoodie.populate.meta.fields=false`, a row-writer `INSERT 
OVERWRITE` in dynamic partition mode goes through while a clustering plan is 
pending on the very file groups it replaces. No 
`HoodieClusteringUpdateException`, the plan stays pending, and the partition's 
data is replaced underneath it. The same statement on a table with meta fields 
populated is rejected, as #18829 intended.
   
   **To reproduce** (Spark SQL, COW):
   
   ```sql
   set hoodie.spark.sql.insert.into.operation=bulk_insert;
   set hoodie.datasource.overwrite.mode=dynamic;
   create table t (id int, name string, price double, dt string) using hudi
     tblproperties (primaryKey = 'id', 'hoodie.populate.meta.fields' = 'false') 
partitioned by (dt) location '/tmp/t';
   insert into t values (1, 'a1', 10, '2021-07-18');
   insert into t values (2, 'a2', 20, '2021-07-18');
   insert into t values (3, 'a3', 30, '2021-07-19');
   call run_clustering(table => 't', op => 'schedule', selected_partitions => 
'dt=2021-07-18');   -- 1 pending plan
   insert overwrite table t partition (dt) values (1, 'a1_new', 11, 
'2021-07-18');                  -- succeeds; expected: rejected
   select id, name, dt from t;   -- (1, a1_new, 2021-07-18): row 2 is gone, 
plan still pending
   ```
   
   Drop the `'hoodie.populate.meta.fields' = 'false'` property and the last 
`insert overwrite` throws `Not allowed to update the clustering file group ...`.
   
   **Cause:**
   
   
`BaseDatasetBulkInsertCommitActionExecutor.rejectIfOverlappingPendingClustering`
 (added by #18829) resolves the partitions being replaced through 
`DatasetBulkInsertOverwriteCommitActionExecutor.resolveTargetPartitions`, whose 
dynamic arm reads `_hoodie_partition_path` off the prepared dataset:
   
   ```java
   // DatasetBulkInsertOverwriteCommitActionExecutor.java:94-100
   // Dynamic partition path: read the populated _hoodie_partition_path meta 
field. The base
   // class invokes this hook after 
HoodieDatasetBulkInsertHelper.prepareForBulkInsert, so the
   // field is guaranteed to be present and populated by the configured key 
generator.
   return 
preparedRecords.select(HoodieRecord.PARTITION_PATH_METADATA_FIELD).distinct()...
   ```
   
   That guarantee does not hold when meta fields are off: 
`HoodieDatasetBulkInsertHelper.prepareForBulkInsert` stubs every meta column as 
a null literal (`HoodieDatasetBulkInsertHelper.scala:145`, `metaFieldsStubs = 
metaFields.map(f => Alias(Literal.create(null, StringType), f.name)())`). 
`resolveTargetPartitions` therefore returns `[null]`, 
`getLatestFileSlices(null)` is asked of the file-system view (with the embedded 
timeline server the request is `slices/partition/latest/?partition` with an 
empty value), nothing comes back, `fileGroupsToBeReplaced` is empty and the 
check returns without consulting the update strategy. The write then proceeds, 
and `getPartitionToReplacedFileIds` records the real partition from the write 
statuses, so the replacement itself is committed correctly; only the guard is 
skipped.
   
   The static arm is unaffected (it reads `STATIC_OVERWRITE_PARTITION_PATHS` 
from config), and the RDD path is unaffected 
(`SparkInsertOverwriteCommitActionExecutor` derives partitions from the 
records' keys).
   
   A second, adjacent inconsistency in the same method: `execute()` picks the 
partitioner from 
`writeConfig.getBoolean(HoodieTableConfig.POPULATE_META_FIELDS)` 
(`BaseDatasetBulkInsertCommitActionExecutor.java:117`) while 
`prepareForBulkInsert` decides whether to stub the columns from 
`config.populateMetaFields()` (`HoodieDatasetBulkInsertHelper.scala:77`), which 
since #19205 is derived from `hoodie.meta.fields.mode` on v10 tables. #19378 
listed this site under "guards bypassable by setting only the mode" and was 
closed; the two reads can still disagree.
   
   **Expected behavior:**
   
   The pending-clustering check should reject (or, with 
`SparkAllowUpdateStrategy`, defer) the overwrite regardless of whether meta 
fields are populated. When `_hoodie_partition_path` is not populated, 
`resolveTargetPartitions` needs another source for the partitions: the key 
generator output that `prepareForBulkInsert` already computes, or the distinct 
partition values of the incoming rows via the configured partition path field.
   
   **Related:**
   
   - #18830 (open) describes the general symptom this is a remaining case of; 
#18829 fixed it for tables with meta fields populated.
   - #19378 (closed) touched the same raw-boolean read but not the 
partition-path resolution.
   
   ## Environment
   
   Master `18ae8c349058` (the executor and helper code are unchanged at head), 
Spark 3.5 / Scala 2.12 profile, local filesystem, metadata table enabled, 
embedded timeline server on. Found while adding row-writer pending-clustering 
coverage on #19163; the new tests there cover only the populated-meta-fields 
case, which passes.
   
   ## Logs and Stack Trace
   
   Timeline-server requests issued by the overwrite in the failing run (the 
first is the guard's lookup, the second the post-write replacement bookkeeping):
   
   ```
   GET 
/v1/hoodie/view/slices/partition/latest/?partition&lastinstantts=20260827103228088&...
   GET 
/v1/hoodie/view/slices/partition/latest/?partition=dt%3D2021-07-18&lastinstantts=20260827103228088&...
   ```
   
   Probe output:
   
   ```
   PENDING=1
   NONOVERLAP RESULT=NO_EXCEPTION
   OVERLAP RESULT=NO_EXCEPTION
   ROWS=(1,a1_new,2021-07-18) (4,b1,2021-07-19)
   ```
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to