KiteSoar commented on code in PR #20099:
URL: https://github.com/apache/hudi/pull/20099#discussion_r4172552981
##########
hudi-spark-datasource/hudi-spark/src/main/scala/org/apache/spark/sql/hudi/command/procedures/ShowCleansProcedure.scala:
##########
@@ -256,6 +266,48 @@ object ShowCleansProcedure {
val NAME = "show_cleans"
def builder: Supplier[ProcedureBuilder] = () => new
ShowCleansProcedure(false)
+
+ private[procedures] def getArchivedCleanTimeline(metaClient:
HoodieTableMetaClient,
+ loadPlans: Boolean,
+ limit: Int = Int.MaxValue):
HoodieTimeline = {
+ val cleanInstants =
metaClient.getArchivedTimeline.getCleanerTimeline.filterCompletedInstants
+ .getReverseOrderedInstants.iterator().asScala.take(limit).toList
+ val contents = new ConcurrentHashMap[String, Array[Byte]]()
+ val factory = metaClient.getTableFormat.getTimelineFactory
+ if (cleanInstants.nonEmpty) {
+ val legacy = metaClient.getTimelineLayoutVersion.getVersion <
TimelineLayoutVersion.VERSION_2
+ val actionField = if (legacy) "actionType" else "action"
+ val contentField = if (legacy) {
+ if (loadPlans) "hoodieCleanerPlan" else "hoodieCleanMetadata"
+ } else {
+ if (loadPlans) "plan" else "metadata"
+ }
Review Comment:
Moved the format-specific decoding into ArchivedCleanTimelineUtils in
hudi-common and reused the V2 action/plan/metadata field constants in
b231fe4c5612. The helper is now available for CleansCommand to reuse. CLI
integration is not included in this PR, so the CLI-specific V1 issue mentioned
here remains outside this fix.
##########
hudi-spark-datasource/hudi-spark/src/main/scala/org/apache/spark/sql/hudi/command/procedures/ShowCleansProcedure.scala:
##########
@@ -165,10 +172,13 @@ class ShowCleansProcedure(includePartitionMetadata:
Boolean) extends BaseProcedu
getCleans(metaClient.getActiveTimeline, limit)
}
val finalResults = if (showArchived) {
+ val archivedCleanLimit = if (includePartitionMetadata) Int.MaxValue else
limit
Review Comment:
Done in b231fe4c5612. Int.MaxValue now applies only to instant descriptors;
archived metadata payloads are loaded lazily in bounded batches while partition
rows are needed. Empty cleans continue to the next batch. Tests cover empty and
multi-row cleans and verify which payloads are loaded.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]