[GitHub] [spark] MaxGekk opened a new pull request #31308: [SPARK-34215][SQL] Keep tables cached after truncation

GitBox Sun, 24 Jan 2021 11:04:06 -0800


MaxGekk opened a new pull request #31308:
URL: https://github.com/apache/spark/pull/31308



   ### What changes were proposed in this pull request?
   Invoke `CatalogImpl.refreshTable()` instead of combination of 
`SessionCatalog.refreshTable()` + `uncacheQuery()`. This allows to clear cached 
table data while keeping the table cached.
   
   ### Why are the changes needed?
   1. To improve user experience with Spark SQL
   2. To be consistent to other commands, see 
https://github.com/apache/spark/pull/31206
   
   ### Does this PR introduce _any_ user-facing change?
   Yes.
   
   Before:
   ```scala
   scala> sql("CREATE TABLE tbl (c0 int)")
   res1: org.apache.spark.sql.DataFrame = []
   scala> sql("INSERT INTO tbl SELECT 0")
   res2: org.apache.spark.sql.DataFrame = []
   scala> sql("CACHE TABLE tbl")
   res3: org.apache.spark.sql.DataFrame = []
   scala> sql("SELECT * FROM tbl").show(false)
   +---+
   |c0 |
   +---+
   |0  |
   +---+
   scala> spark.catalog.isCached("tbl")
   res5: Boolean = true
   scala> sql("TRUNCATE TABLE tbl")
   res6: org.apache.spark.sql.DataFrame = []
   scala> spark.catalog.isCached("tbl")
   res7: Boolean = false
   ```
   
   After:
   ```scala
   scala> sql("TRUNCATE TABLE tbl")
   res6: org.apache.spark.sql.DataFrame = []
   scala> spark.catalog.isCached("tbl")
   res7: Boolean = true
   ```
   
   ### How was this patch tested?
   Added new test to `CachedTableSuite`:
   ```
   $ build/sbt -Phive -Phive-thriftserver "test:testOnly *CachedTableSuite"
   $ build/sbt -Phive -Phive-thriftserver "test:testOnly *CatalogedDDLSuite"
   ```


----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

For queries about this service, please contact Infrastructure at:
[email protected]



---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

[GitHub] [spark] MaxGekk opened a new pull request #31308: [SPARK-34215][SQL] Keep tables cached after truncation

Reply via email to