hudi-agent commented on code in PR #20105:
URL: https://github.com/apache/hudi/pull/20105#discussion_r4127211823
##########
hudi-spark-datasource/hudi-spark-common/src/main/scala/org/apache/hudi/HoodieFileIndex.scala:
##########
@@ -104,6 +104,41 @@ case class HoodieFileIndex(spark: SparkSession,
startCompletionTime = options.get(DataSourceReadOptions.START_COMMIT.key),
endCompletionTime = options.get(DataSourceReadOptions.END_COMMIT.key))
with FileIndex {
+ /**
+ * Equality deliberately excludes [[fileStatusCache]].
+ *
+ * It is a per-call handle: FileStatusCache.getOrCreate(session) returns a
fresh
+ * SharedInMemoryCache view on every invocation, and the class has no
equals, so it compares by
+ * identity. Being a case-class parameter it would otherwise land in the
generated equals
+ * (@transient affects serialization, not equality), and two indexes over
the same table would
+ * never compare equal. Spark's CacheManager matches by plan, so that left
recacheByPlan unable
+ * to find an existing entry after a refresh: the old materialised dataset
was stranded rather
+ * than replaced, and spark.catalog.isCached reported false.
+ *
+ * The cache an index happens to hold is a performance detail, not part of
the table's identity.
+ */
+ override def equals(other: Any): Boolean = other match {
Review Comment:
🤖 Could this make DataFrame-path reads return stale data? Scenario:
`spark.read.format("hudi").load(p).cache()`, then an append via the DataFrame
writer with meta sync off, then a new `spark.read...load(p)`. That third read
now matches the old cached entry (`ParquetFileFormat.equals` is just
`isInstanceOf`, so the file format won't break the tie), where today it misses
and reads fresh. It might be worth adding the DataFrame-writer invalidation
(e.g. `recacheByPath`) in this PR, or at least a test for this sequence. @yihua
<sub><i>⚠️ AI-generated; verify before applying. React 👍/👎 to flag
quality.</i></sub>
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]