Hi Gabor,

Thank you very much for the quick reply. We are interested in  separating 
parquet-mr compression API from Hadoop and will try to make a proposal for the 
new API. Will let you know when we are ready and your comments will be 
appreciated.

Thanks,
Xin Dong

-----Original Message-----
From: Gabor Szadovszky <[email protected]> 
Sent: Thursday, June 3, 2021 3:49 PM
To: Parquet Dev <[email protected]>
Subject: Re: Decouple parquet-mr compression API from hadoop compression API

Hi Xin Dong,

There are a couple of open jiras related to this. Like PARQUET-1812 
<https://issues.apache.org/jira/browse/PARQUET-1812> about using the airlift 
implementation of the codecs or your own jiras about the provider-aware codecs. 
I strongly agree on having compression codecs that are independent from Hadoop. 
It would also be required to have our own compression codecs shipped if we want 
to achieve the core features support (
PARQUET-1950 <https://issues.apache.org/jira/browse/PARQUET-1950>).
Feel free to create a jira about this. Also, if you want to invest on it I'm 
happy to review the related PRs.

Cheers,
Gabor

On Thu, Jun 3, 2021 at 4:18 AM Dong, Xin <[email protected]> wrote:

> Hi, All,
> Currently parquet-mr compression logic is using Hadoop compression API 
> which makes parquet-mr compression highly coupled with Hadoop. Does 
> community have any plan to decouple those two APIs? To make the things 
> easier, maybe we can just using api similar to Hadoop compression APIs 
> but belongs to parquet-mr namespace. And simply change current codec 
> to implements the new parquet-mr API. Any thoughts?
> Thanks,
> Xin Dong
>
>

Reply via email to