Thanks, Thakrar, I have tried to check the existence of path before read it, but HDFSCli python package seems not support wildcard. "FileSystem.globStatus" is a java api while I am using python via livy.... Do you know any python api implementing the same function?
Regard, Junfeng Chen On Mon, May 21, 2018 at 9:01 PM, Thakrar, Jayesh < [email protected]> wrote: > Probably you can do some preprocessing/checking of the paths before you > attempt to read it via Spark. > > Whether it is local or hdfs filesystem, you can try to check for existence > and other details by using the "FileSystem.globStatus" method from the > Hadoop API. > > > > *From: *JF Chen <[email protected]> > *Date: *Sunday, May 20, 2018 at 10:30 PM > *To: *user <[email protected]> > *Subject: *How to skip nonexistent file when read files with spark? > > > > Hi Everyone > > I meet a tricky problem recently. I am trying to read some file paths > generated by other method. The file paths are represented by wild card in > list, like [ '/data/*/12', '/data/*/13'] > > But in practice, if the wildcard cannot match any existed path, it will > throw an exception:"pyspark.sql.utils.AnalysisException: 'Path does not > exist: ...'", and the program stops after that. > > Actually I want spark can just ignore and skip these nonexistent file > path, and continues to run. I have tried python HDFSCli api to check the > existence of path , but hdfs cli cannot support wildcard. > > > > Any good idea to solve my problem? Thanks~ > > > > Regard, > Junfeng Chen >
