comphead commented on code in PR #5634: URL: https://github.com/apache/datafusion-comet/pull/5634#discussion_r4231926329
########## docs/source/user-guide/latest/in-memory-cache.md: ########## @@ -21,18 +21,24 @@ Comet can store Spark's in-memory cache (`CACHE TABLE`, `df.cache()`, `df.persist()`) in an Arrow format that Comet operators read directly. Without it, a cached table is stored in Spark's own -format and every scan of it has to convert each batch before Comet can continue, which shows up in -the plan as a `CometSparkColumnarToColumnar` above the cache scan. +format, which Comet operators cannot read. Under Comet's default settings the operators above the +cache scan then run on Spark. With `spark.comet.convert.inMemoryCache.enabled`, a +`CometSparkColumnarToColumnar` above the scan converts each batch for Comet operators instead. -This feature is **experimental and disabled by default**. Turn it on at startup, alongside the rest -of Comet's configuration: +This feature is **experimental and enabled by default** from Spark 3.5. To turn it off, set the Review Comment: Experimental and enabled by default IMO are mutually exclusive -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
