viirya commented on a change in pull request #25856: [SPARK-29182][Core] Cache 
preferred locations of checkpointed RDD
URL: https://github.com/apache/spark/pull/25856#discussion_r327440215
 
 

 ##########
 File path: core/src/main/scala/org/apache/spark/rdd/ReliableCheckpointRDD.scala
 ##########
 @@ -82,14 +83,28 @@ private[spark] class ReliableCheckpointRDD[T: ClassTag](
     Array.tabulate(inputFiles.length)(i => new CheckpointRDDPartition(i))
   }
 
+  // Cache of preferred locations of checkpointed files.
+  private[spark] val cachedPreferredLocations: mutable.HashMap[Int, 
Seq[String]] =
+    mutable.HashMap.empty
 
 Review comment:
   I think you meant when at step 6, all replicas are placed different hosts in 
our cached, right?
   
   Since A1 is shutdown, Spark won't launch task on it. Say Spark launches task 
at A2 or A3, the task still can access block from B1 or B2. Could you explain 
why it will fail?

----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
 
For queries about this service, please contact Infrastructure at:
[email protected]


With regards,
Apache Git Services

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to