[ https://issues.apache.org/jira/browse/FLINK-15533?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17012633#comment-17012633 ]
Kostas Kloudas edited comment on FLINK-15533 at 1/10/20 9:58 AM: ----------------------------------------------------------------- Hi [~lirui], I tried it on the {{master}} (without my patch) with Yarn and HDFS and job submission from the command line and I cannot reproduce it. My job is: {code:java} StreamExecutionEnvironment streamEnv = StreamExecutionEnvironment.getExecutionEnvironment(); DataStream dataStream = streamEnv.fromCollection(Arrays.asList(1, 2, 3)); dataStream.writeAsText("hdfs://" + args[0] + ":9000/tmp/output"); streamEnv.execute(); {code} and the command in the CLI : {{./bin/flink run ./MY_JOB.jar HOSTNAME}} Also I tried it with providing parallelism using the {{-p}} option and it still works. Could you provide some more details so that I can reproduce it? was (Author: kkl0u): Hi [~lirui], I tried it on the {{master}} (without my patch) with Yarn and HDFS and job submission from the command line and I cannot reproduce it. My job is: {code:java} StreamExecutionEnvironment streamEnv = StreamExecutionEnvironment.getExecutionEnvironment(); DataStream dataStream = streamEnv.fromCollection(Arrays.asList(1, 2, 3)); dataStream.writeAsText("hdfs://" + args[0] + ":9000/tmp/output"); streamEnv.execute(); {code} and the command in the CLI : {{./bin/flink run ./examples/streaming/MY_JOB.jar HOSTNAME}} Also I tried it with providing parallelism using the {{-p}} option and it still works. Could you provide some more details so that I can reproduce it? > Writing DataStream as text file fails due to output path already exists > ----------------------------------------------------------------------- > > Key: FLINK-15533 > URL: https://issues.apache.org/jira/browse/FLINK-15533 > Project: Flink > Issue Type: Bug > Components: Client / Job Submission > Affects Versions: 1.10.0 > Reporter: Rui Li > Assignee: Kostas Kloudas > Priority: Blocker > Fix For: 1.10.0 > > > The following program reproduces the issue. > {code} > Configuration configuration = GlobalConfiguration.loadConfiguration(); > configuration.set(DeploymentOptions.TARGET, RemoteExecutor.NAME); > StreamExecutionEnvironment streamEnv = new > StreamExecutionEnvironment(configuration); > DataStream dataStream = streamEnv.fromCollection(Arrays.asList(1,2,3)); > dataStream.writeAsText("hdfs://localhost:8020/tmp/output"); > streamEnv.execute(); > {code} > The job will fail with the follow error, even though the output path doesn't > exist before job submission: > {noformat} > org.apache.hadoop.ipc.RemoteException(org.apache.hadoop.fs.FileAlreadyExistsException): > /tmp/output already exists as a directory > {noformat} -- This message was sent by Atlassian Jira (v8.3.4#803005)