Akash3121 opened a new pull request, #9977:
URL: https://github.com/apache/paimon/pull/9977

   This pr is on https://github.com/apache/paimon/issues/9276
   
   ### Purpose
   `PartitionIndex` stores the mapping from a key hash to its assigned bucket 
in an `Int2ShortHashMap`:
    
    ```java 
    hash2Bucket.put(hash, (short) bucket);
   ```
   
   Bucket IDs are non-negative, so this representation supports IDs from  0  
through  32767 , or at most  32768  buckets.
   
   However,  dynamic-bucket.max-buckets  was an unrestricted integer. When it 
was configured above  32768 ,  PartitionIndex  could allocate bucket ID  32768  
or higher and then narrow it to a signed  short .
   
   For example, bucket ID  32768  becomes  -32768  after narrowing:
   
    First assignment:  32768
    Stored short value: -32768
    Later assignment:  -32768
   
   As a result, the first record for a key and subsequent records for the same 
key could be sent to different buckets, violating the requirement that a key 
always retains its bucket assignment.
   
   The same limit also applies to PyPaimon because Python-written 
dynamic-bucket metadata must remain readable and writable by Java.
   ### Before and after
   
   #### Before
   
    dynamic-bucket.max-buckets = 40000
    
    assign(hash) -> 32768
    store mapping -> (short) 32768 == -32768
    assign(same hash again) -> -32768
   
   A single key could therefore be written to two different buckets.
   
   The unbounded path also stopped at bucket ID  32766 , even though ID  32767  
is representable by the memo.
   
   #### After
   
    dynamic-bucket.max-buckets = 40000
    
    writer initialization / assignment -> clear IllegalArgumentException or 
ValueError
   
    dynamic-bucket.max-buckets = 32768
    
    highest valid bucket ID -> 32767
    repeated assignment       -> 32767
   
   All bucket IDs are checked before entering the short-backed map, so an 
out-of-range ID can no longer wrap into a negative bucket.
   ### Changes
   
   Java
   
    - Define  32768  as the maximum supported dynamic bucket count.
    - Accept  dynamic-bucket.max-buckets  only when it is:
     -  -1 , meaning no configured cap within the implementation limit; or
     - between  1  and  32768 , inclusive.
    - Validate the configured maximum when regular and overwrite bucket 
assigners are initialized.
    - Validate bucket IDs before every conversion to  short .
    - Reject restored HASH-index entries containing bucket IDs outside  
0..32767  instead of narrowing them to negative values.
    - Allow bucket ID  32767  in both explicit  32768  mode and  -1  mode.
   
   PyPaimon
   
    - Apply the same  -1  or  1..32768  configuration boundary.
    - Allow bucket ID  32767  in regular and overwrite assignment.
    - Reject restored HASH-index entries whose bucket IDs are outside the 
Java-compatible range.
    - Keep Java and Python dynamic-bucket behavior interoperable.
   
   Documentation
   
    - Document the supported  dynamic-bucket.max-buckets  range in the 
dynamic-bucket guide.
    - Regenerate the core configuration reference.
   
   Notes To Reviewer: The main design choice is to retain the compact 
short-backed hash-to-bucket memo and reject unsupported configurations rather 
than widening the value type for every dynamic-bucket table.
   
   Validation is intentionally performed on writer/assigner initialization 
rather than general table schema loading. This prevents new corruption while 
preserving read access to existing tables that have an out-of-range option but 
have not necessarily created an out-of-range bucket.
   
   The maximum is a bucket count, not a maximum bucket ID:
   
    maximum bucket count = 32768
    valid bucket IDs     = 0..32767


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to