KiteSoar opened a new pull request, #20099:
URL: https://github.com/apache/hudi/pull/20099

   ### Describe the issue this Pull Request addresses
   
   Closes #19639.
   
   The clean procedures read archived clean instants from 
`metaClient.getArchivedTimeline`, but the archived timeline does not have the 
cleaner plan or clean metadata payloads loaded by default. As a result, 
`show_clean_plans` returned null plan fields, while `show_cleans` and 
`show_cleans_metadata` could fail with an `IOException` when they tried to 
deserialize archived clean metadata.
   
   ### Summary and Changelog
   
   Users can now query archived clean plans, clean summaries, and per-partition 
clean metadata through the three clean procedures.
   
   - Load the required plan or metadata payloads for archived clean instants 
and expose them through a temporary timeline used by the procedures.
   - Decode archived records for both timeline layout V1 and V2. V1 cleaner 
records are converted from their Avro representation, while V2 payloads are 
copied from their byte buffers.
   - Apply the procedure limit before loading archived payloads for 
`show_clean_plans` and `show_cleans`. Preserve `show_cleans_metadata` row-limit 
behavior because one clean instant can produce multiple partition rows.
   - Add coverage for archived clean plans and metadata on table versions 6 and 
9.
   
   Validation:
   
   - `TestArchivedTimelineV1`: 24 tests passed.
   - `TestShowCleansProcedures`: 11 tests passed.
   
   ### Impact
   
   This fixes the existing `show_clean_plans`, `show_cleans`, and 
`show_cleans_metadata` behavior when `show_archived` is enabled. It does not 
change public APIs, storage formats, or configuration. Archived payload loading 
is bounded by `limit` where the procedure returns one row per clean instant.
   
   ### Risk Level
   
   Low. The change is limited to Spark SQL clean procedures and uses the 
existing archived timeline loader. Tests cover both legacy and current timeline 
layouts.
   
   ### Documentation Update
   
   None.
   
   ### Contributor's checklist
   
   - [x] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [x] Enough context is provided in the sections above
   - [x] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to