HeartSaVioR commented on a change in pull request #31638:
URL: https://github.com/apache/spark/pull/31638#discussion_r594073932
##########
File path:
sql/core/src/main/scala/org/apache/spark/sql/execution/streaming/FileStreamSink.scala
##########
@@ -40,17 +41,31 @@ object FileStreamSink extends Logging {
* be read.
*/
def hasMetadata(path: Seq[String], hadoopConf: Configuration, sqlConf:
SQLConf): Boolean = {
- path match {
- case Seq(singlePath) =>
- val hdfsPath = new Path(singlePath)
- val fs = hdfsPath.getFileSystem(hadoopConf)
- if (fs.isDirectory(hdfsPath)) {
- val metadataPath = getMetadataLogPath(fs, hdfsPath, sqlConf)
- fs.exists(metadataPath)
- } else {
- false
- }
- case _ => false
+ if (sqlConf.getConf(SQLConf.FILE_SINK_FORMAT_CHECK_ENABLED)) {
Review comment:
@sunchao
Thanks for the answer! That is same as my expectation.
@xuanyuanking
According to the answer, I don't think there's a risk on dealing glob path
as wrong input and just returning false. It has been working like so, plus
possibility to throw additional exception. I don't think we need to add flag to
keep the behavior - we can just fix it.
----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]