hudi-agent commented on code in PR #20099:
URL: https://github.com/apache/hudi/pull/20099#discussion_r4127237962


##########
hudi-spark-datasource/hudi-spark/src/main/scala/org/apache/spark/sql/hudi/command/procedures/ShowCleansProcedure.scala:
##########
@@ -256,6 +266,48 @@ object ShowCleansProcedure {
   val NAME = "show_cleans"
 
   def builder: Supplier[ProcedureBuilder] = () => new 
ShowCleansProcedure(false)
+
+  private[procedures] def getArchivedCleanTimeline(metaClient: 
HoodieTableMetaClient,

Review Comment:
   🤖 nit: the misaligned continuation indentation for `loadPlans`/`limit` 
params stands out from the rest of the file's style — worth running through 
scalafmt/checkstyle before merging.
   
   <sub><i>⚠️ AI-generated; verify before applying. React 👍/👎 to flag 
quality.</i></sub>



##########
hudi-spark-datasource/hudi-spark/src/main/scala/org/apache/spark/sql/hudi/command/procedures/ShowCleansProcedure.scala:
##########
@@ -165,10 +172,13 @@ class ShowCleansProcedure(includePartitionMetadata: 
Boolean) extends BaseProcedu
       getCleans(metaClient.getActiveTimeline, limit)
     }
     val finalResults = if (showArchived) {
+      val archivedCleanLimit = if (includePartitionMetadata) Int.MaxValue else 
limit

Review Comment:
   🤖 One way to get this bound without guessing a batch size: fill the 
temporary timeline's `HoodieInstantReader.getContentStream` lazily, loading 
payloads on the first miss for a descending window of instants such as the next 
`limit`. Then the existing `takeWhile(rowCount < limit)` in 
`getCleansWithPartitionMetadata` decides how many archived cleans get read. 
Empty (0-row) cleans just pull in the next window, so it stays exact on both 
points raised here and in the `limit` suggestion below. On V2, 
`ArchivedTimelineLoaderV2.getFilteredFiles` prunes LSM files by the time range, 
so each window only reads the files that overlap it. On V1 each window rescans 
archive files, so the window should stay reasonably large there.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to