Akash3121 opened a new pull request, #9978:
URL: https://github.com/apache/paimon/pull/9978

   ### Purpose
   Allow callers of `TableWrite` to delete a row from a primary-key table 
without supplying values for every non-key `NOT NULL` column.
    
    Previously, `TableWriteImpl` applied the table's complete nullability 
validation to every row kind. A direct delete such as the following was 
therefore rejected when `v` was declared `NOT NULL`, even though the 
deduplicate merge engine only needs the key to identify the row being deleted:
    
    ```java
    // Schema: pt INT NOT NULL, k INT NOT NULL, v BIGINT NOT NULL
    // Primary key: (pt, k)
    write.write(GenericRow.ofKind(RowKind.DELETE, 1, 1, null));
   ```
   This differs from SQL delete paths that first read the existing row and 
consequently submit a fully populated delete record. Direct  TableWrite  users 
should not need that additional read when the merge semantics do not consume 
the omitted fields.
   
   Closes #4702.
   
   ### Changes
   
   Row-kind-aware nullability validation
   
    `TableWriteImpl`  now determines the effective row kind before selecting 
the nullability checks to apply.
   
   For normal rows, the existing behavior remains unchanged: every field 
declared  NOT NULL  must still be present.
   
   For eligible  `DELETE`  rows,  `TableWriteImpl`  can use a restricted set of 
required fields instead of validating all non-null columns. The original 
constructor is preserved and continues to perform full validation, so existing 
callers do not change behavior unless they explicitly provide the 
delete-specific field set.
   
   The effective row kind is derived from the default-wrapped row, which 
preserves support for  rowkind.field . Nullability is still checked against the 
original input row so configured defaults do not unintentionally bypass input 
validation.
   
   #### Conservative scope
   
   The relaxed behavior is enabled only for primary-key tables that satisfy 
both conditions:
    - The merge engine is  DEDUPLICATE .
    - Cross-partition updates are disabled.
   
   For an eligible delete, the following fields remain required when declared  
NOT NULL :
    - All primary-key fields.
    - All configured sequence fields.
   
   Sequence fields remain mandatory because they participate in ordering. A 
delete with an older sequence value must not remove a newer row.
   
   Full nullability validation is retained for:
   
    - Inserts.
    -  UPDATE_BEFORE  and  UPDATE_AFTER  rows.
    - Partial-update tables.
    - Aggregate and other non-deduplicate merge engines.
    - Cross-partition update tables, where partition information is required 
for routing.
   
   #### Internal serialization support
   
   Skipping the explicit nullability check alone is insufficient. The 
serializer generated from a logical  NOT NULL  primitive type directly invokes 
getters such as  getInt  and  getLong ; attempting to serialize a Java  null  
through those getters can fail.
   
   For eligible deduplicate tables, the internal key-value value type therefore 
marks fields that may be omitted from a delete as nullable. Primary keys and 
sequence fields retain their original nullability.
   
   This is an internal storage/serialization type only. The table's public 
logical schema and its  NOT NULL  constraints are unchanged, and non-delete 
rows continue to be validated against the complete logical schema.
   
   The same internal nullable value type is also supplied to lookup merge 
processing so buffered or spilled delete records are serialized consistently.
   
   #### Behavior examples
   
   Given this table:
   
    pt  INT     NOT NULL
    k   INT     NOT NULL
    v   BIGINT  NOT NULL
    PRIMARY KEY (pt, k)
   
   ```
   A deduplicate table now accepts: 
    write.write(GenericRow.ofKind(RowKind.DELETE, 1, 1, null));
   
   It still rejects a delete with a missing key:
    write.write(GenericRow.ofKind(RowKind.DELETE, 1, null, null));
   
   It also continues to reject null payloads for non-delete row kinds:
    write.write(GenericRow.ofKind(RowKind.UPDATE_AFTER, 1, 1, null));
   
   For a table with a sequence field, the sequence value must still be supplied:
   
    // Rejected because seq is missing.
    write.write(GenericRow.ofKind(RowKind.DELETE, 1, 1, null, null));
    
    // Accepted: key and sequence are supplied; the payload is omitted.
    write.write(GenericRow.ofKind(RowKind.DELETE, 1, 1, 11, null));
   ```
   
   #### Changelog considerations
   
   This change does not reconstruct the previous non-key values for a direct 
delete. A changelog producer that exposes the input delete record may therefore 
contain null values for omitted non-key fields.
   
   That is expected for this API path: the final table semantics are preserved, 
but callers that require a complete before-image must still provide the 
previous values or use a changelog mode capable of reconstructing them.
   
   ### Tests
   Added coverage for:
   
    - A deduplicate primary-key table accepting a delete with only its primary 
key fields populated.
    - Missing primary-key fields still failing nullability validation.
    - Inserts,  UPDATE_BEFORE , and  UPDATE_AFTER  retaining complete 
nullability checks.
    - Partial-update tables retaining complete delete validation.
    - Sequence fields remaining required for deletes.
    - An older-sequence delete not removing a newer row.
    - A newer-sequence delete removing the row.
    - Deletes generated through  rowkind.field  receiving the same behavior.
    - Generated update row kinds retaining complete nullability checks.
    - Lookup merge and its spill-capable buffering path supporting partial 
delete payloads.
    - Cross-partition primary-key tables retaining partition-field validation.
      
   ### Notes for reviewers
   
   The restriction to non-cross-partition  DEDUPLICATE  tables is intentional:
   
    - Deduplicate merge semantics select the latest key-value record and do not 
need non-key values to compute a deletion.
    - Partial-update and aggregate merge functions may consume non-key retract 
values.
    - Cross-partition updates require partition information for routing and 
index maintenance.
    - Sequence fields affect record ordering and therefore cannot be omitted 
safely.
   
   The internal nullable value type is necessary to serialize the omitted 
fields safely; changing only  TableWriteImpl.checkNullability  would allow the 
row through validation but could fail later in primitive field serialization.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to