hudi-agent commented on code in PR #20105:
URL: https://github.com/apache/hudi/pull/20105#discussion_r4127211823


##########
hudi-spark-datasource/hudi-spark-common/src/main/scala/org/apache/hudi/HoodieFileIndex.scala:
##########
@@ -104,6 +104,41 @@ case class HoodieFileIndex(spark: SparkSession,
     startCompletionTime = options.get(DataSourceReadOptions.START_COMMIT.key),
     endCompletionTime = options.get(DataSourceReadOptions.END_COMMIT.key)) 
with FileIndex {
 
+  /**
+   * Equality deliberately excludes [[fileStatusCache]].
+   *
+   * It is a per-call handle: FileStatusCache.getOrCreate(session) returns a 
fresh
+   * SharedInMemoryCache view on every invocation, and the class has no 
equals, so it compares by
+   * identity. Being a case-class parameter it would otherwise land in the 
generated equals
+   * (@transient affects serialization, not equality), and two indexes over 
the same table would
+   * never compare equal. Spark's CacheManager matches by plan, so that left 
recacheByPlan unable
+   * to find an existing entry after a refresh: the old materialised dataset 
was stranded rather
+   * than replaced, and spark.catalog.isCached reported false.
+   *
+   * The cache an index happens to hold is a performance detail, not part of 
the table's identity.
+   */
+  override def equals(other: Any): Boolean = other match {

Review Comment:
   🤖 Could this make DataFrame-path reads return stale data? Scenario: 
`spark.read.format("hudi").load(p).cache()`, then an append via the DataFrame 
writer with meta sync off, then a new `spark.read...load(p)`. That third read 
now matches the old cached entry (`ParquetFileFormat.equals` is just 
`isInstanceOf`, so the file format won't break the tie), where today it misses 
and reads fresh. It might be worth adding the DataFrame-writer invalidation 
(e.g. `recacheByPath`) in this PR, or at least a test for this sequence. @yihua
   
   <sub><i>⚠️ AI-generated; verify before applying. React 👍/👎 to flag 
quality.</i></sub>



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to