hudi-agent commented on code in PR #20099:
URL: https://github.com/apache/hudi/pull/20099#discussion_r4127237962
##########
hudi-spark-datasource/hudi-spark/src/main/scala/org/apache/spark/sql/hudi/command/procedures/ShowCleansProcedure.scala:
##########
@@ -256,6 +266,48 @@ object ShowCleansProcedure {
val NAME = "show_cleans"
def builder: Supplier[ProcedureBuilder] = () => new
ShowCleansProcedure(false)
+
+ private[procedures] def getArchivedCleanTimeline(metaClient:
HoodieTableMetaClient,
Review Comment:
🤖 nit: the misaligned continuation indentation for `loadPlans`/`limit`
params stands out from the rest of the file's style — worth running through
scalafmt/checkstyle before merging.
<sub><i>⚠️ AI-generated; verify before applying. React 👍/👎 to flag
quality.</i></sub>
##########
hudi-spark-datasource/hudi-spark/src/main/scala/org/apache/spark/sql/hudi/command/procedures/ShowCleansProcedure.scala:
##########
@@ -165,10 +172,13 @@ class ShowCleansProcedure(includePartitionMetadata:
Boolean) extends BaseProcedu
getCleans(metaClient.getActiveTimeline, limit)
}
val finalResults = if (showArchived) {
+ val archivedCleanLimit = if (includePartitionMetadata) Int.MaxValue else
limit
Review Comment:
🤖 One way to get this bound without guessing a batch size: fill the
temporary timeline's `HoodieInstantReader.getContentStream` lazily, loading
payloads on the first miss for a descending window of instants such as the next
`limit`. Then the existing `takeWhile(rowCount < limit)` in
`getCleansWithPartitionMetadata` decides how many archived cleans get read.
Empty (0-row) cleans just pull in the next window, so it stays exact on both
points raised here and in the `limit` suggestion below. On V2,
`ArchivedTimelineLoaderV2.getFilteredFiles` prunes LSM files by the time range,
so each window only reads the files that overlap it. On V1 each window rescans
archive files, so the window should stay reasonably large there.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]