[
https://issues.apache.org/jira/browse/PARQUET-1822?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17847455#comment-17847455
]
Ryan Rupp commented on PARQUET-1822:
------------------------------------
Is 1.14.0 to the point that parquet can be used without hadoop-client
dependency? I was playing around with it and observed:
# The compiler complains about the new method overloads for `builder` and
`withConf` - even though I'm using the new Parquet interface overloads, the
compiler will complain about the hadoop classes not being available. I can
"trick" this though by adding hadoop as a provided scope dependency. This is on
Java 11 FWIW.
# Once past that, ParquetWriter if you're not using encryption the null code
path goes down into path that hits Hadoop classes (I temporarily worked around
by just removing the encryption settings as I don't use them in this case)
# After that I hit I believe I hit PARQUET-2353 but didn't dig into it too far
There's some comments on the PR like
[this|https://github.com/apache/parquet-mr/pull/1111#issuecomment-1995928017]
that sound like people are doing this to an extent, maybe dependency on
hadoop-common (instead of hadoop-client) or does anyone have an example of
minimal hadoop dependencies being pulled in? The tests like TestReadWrite
already have Hadoop on the classpath for instance.
Thanks
> Parquet without Hadoop dependencies
> -----------------------------------
>
> Key: PARQUET-1822
> URL: https://issues.apache.org/jira/browse/PARQUET-1822
> Project: Parquet
> Issue Type: Improvement
> Components: parquet-avro
> Affects Versions: 1.11.0
> Environment: Amazon Fargate (linux), Windows development box.
> We are writing Parquet to be read by the Snowflake and Athena databases.
> Reporter: mark juchems
> Assignee: Atour Mousavi Gourabi
> Priority: Minor
> Labels: documentation, newbie
> Fix For: 1.14.0
>
>
> I have been trying for weeks to create a parquet file from avro and write to
> S3 in Java. This has been incredibly frustrating and odd as Spark can do it
> easily (I'm told).
> I have assembled the correct jars through luck and diligence, but now I find
> out that I have to have hadoop installed on my machine. I am currently
> developing in Windows and it seems a dll and exe can fix that up but am
> wondering about Linus as the code will eventually run in Fargate on AWS.
> *Why do I need external dependencies and not pure java?*
> The thing really is how utterly complex all this is. I would like to create
> an avro file and convert it to Parquet and write it to S3, but I am trapped
> in "ParquetWriter" hell!
> *Why can't I get a normal OutputStream and write it wherever I want?*
> I have scoured the web for examples and there are a few but we really need
> some documentation on this stuff. I understand that there may be reasons for
> all this but I can't find them on the web anywhere. Any help? Can't we get
> the "SimpleParquet" jar that does this:
>
> ParquetWriter writer =
> AvroParquetWriter.<GenericData.Record>builder(outputStream)
> .withSchema(avroSchema)
> .withConf(conf)
> .withCompressionCodec(CompressionCodecName.SNAPPY)
> .withWriteMode(Mode.OVERWRITE)//probably not good for prod. (overwrites
> files).
> .build();
>
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]