Hi Junjie,

I'd be glad to have a look at the encryption part. Will add my comments
early next week.

Cheers, Gidon.

On Fri, Jul 5, 2019 at 12:16 PM 俊杰陈 <[email protected]> wrote:

> Sorry, the latest file is
>
> https://github.com/chenjunjiedada/parquet-format/blob/PARQUET-1617/BloomFilter.md
> .
>
> On Fri, Jul 5, 2019 at 5:14 PM 俊杰陈 <[email protected]> wrote:
>
> > Sure, please see this PR
> > <https://github.com/apache/parquet-format/pull/140> or update file here
> > <
> https://github.com/chenjunjiedada/parquet-format/blob/master/BloomFilter.md
> >
> > .
> >
> > Thanks for reviewing spec.
> >
> > On Thu, Jul 4, 2019 at 11:57 PM Zoltan Ivanfi <[email protected]>
> > wrote:
> >
> >> Hi Junjie,
> >>
> >> I read through the specification and while I support the feature in
> >> general, I find that the documentation may not be detailed enough to
> allow
> >> developers of  different language bindings to implement it.
> Specifically,
> >> the Technical Approach section of the docs is very short and refers the
> >> reader to two publications for details. I think the specification would
> >> greatly benefit from including an explanation or a summary of the
> approach
> >> in this section.
> >>
> >> The "Build a Bloom filter" section contains a formula for calculating
> the
> >> optimal filter size for a desired false positive rate, but does not
> >> specify
> >> what false positive rates implementations should target by default and
> >> through what ways should they make it configurable by users. I
> understand
> >> that this may be an intentional omission, since targeting any false
> >> positive rate will result in a specification-compliant result, still I
> >> think it would be best to provide some recommendation for the different
> >> language bindings.
> >>
> >> Since this feature is getting added after encryption, it should be
> briefly
> >> but explicitly mentioned how it interacts with that (basically that it
> has
> >> to be encrypted, otherwise it would leak sensitive information, but by
> >> placing it inside the column chunk metadata, this is automatically taken
> >> care of).
> >>
> >> Finally, as a nitpick, I would prefer in-line links to related materials
> >> instead of numeric references that one must manually look up at the
> bottom
> >> of the page.
> >>
> >> Could you please add these improvements to the specification?
> >>
> >> Thanks,
> >>
> >> Zoltan
> >>
> >> On Wed, Jul 3, 2019 at 4:03 PM 俊杰陈 <[email protected]> wrote:
> >>
> >> > You are welcome, it 's my honor.
> >> >
> >> > I think the PR <https://github.com/apache/parquet-format/pull/139>
> just
> >> > remove murmur3, that should express what I want.
> >> >
> >> >
> >> >
> >> >
> >> > On Wed, Jul 3, 2019 at 9:53 PM Zoltan Ivanfi <[email protected]
> >
> >> > wrote:
> >> >
> >> > > Hi Junjie,
> >> > >
> >> > > Thanks for the update and also for your endruance in going through
> >> this
> >> > > tedious process in order to add bloom filtering to Parquet.
> >> > >
> >> > > I understand that your proposal is to go forward with xxHash instead
> >> of
> >> > the
> >> > > eralier murmur3, which you suggest to deprecate. Since the murmur3
> >> hash
> >> > was
> >> > > never released, I think it could be completely removed from the spec
> >> > > instead of just getting deprecated. What is your opinion on this?
> >> > >
> >> > > Thanks,
> >> > >
> >> > > Zoltan
> >> > >
> >> > > On Wed, Jul 3, 2019 at 3:31 PM 俊杰陈 <[email protected]> wrote:
> >> > >
> >> > > > I see, thanks for guiding on this.
> >> > > >
> >> > > > Per discussion in this thread and some investigation about changes
> >> on
> >> > > > current java and c++ implementation, and I think that is not hard
> to
> >> > > > handle. So I propose to use xxHash (the XXH64 version) as the
> >> default
> >> > > > hash strategy and deprecate previous murmur3 hash.
> >> > > >
> >> > > > I will update vote thread as well to make it clearer to all.
> >> > > >
> >> > > >
> >> > > > On Wed, Jul 3, 2019 at 6:08 PM Zoltan Ivanfi
> >> <[email protected]>
> >> > > > wrote:
> >> > > > >
> >> > > > > Hi Junjie,
> >> > > > >
> >> > > > > I think the vote is ambigous in its current form (can people
> vote
> >> on
> >> > > one
> >> > > > > option only or can they vote on both?) and has a low chance of
> >> > getting
> >> > > > > votes in general because it's not a yes/no question but a
> >> > > > > choose-an-approach question instead. I think most contributors
> >> would
> >> > > > accept
> >> > > > > the hash chosen based on a community discussion but would be
> >> > reluctant
> >> > > to
> >> > > > > make that choice themselves in the form a vote because it
> >> requires a
> >> > > much
> >> > > > > deeper dive into the technical intricacies involved. The
> >> committers
> >> > are
> >> > > > > experienced in the parquet code base but may not be as
> >> experienced in
> >> > > > bloom
> >> > > > > filters as you are.
> >> > > > >
> >> > > > > In my opinion, to get bloom filtering into parquet-mr, you
> should
> >> > > > convince
> >> > > > > the committers that the proposal is viable by addressing their
> >> > concerns
> >> > > > > (which I believe you have done), and not by delegating the task
> of
> >> > > making
> >> > > > > choices to them. I would suggest that you propose which one (or
> >> both)
> >> > > of
> >> > > > > the hashes should be included, summarize your motivations in
> this
> >> > > thread
> >> > > > > and if you don't get any objections for a day or two, call a
> >> YES/NO
> >> > > vote
> >> > > > > for that specific proposal in a separate thread.
> >> > > > >
> >> > > > > Thanks,
> >> > > > >
> >> > > > > Zoltan
> >> > > > >
> >> > > > > On Tue, Jul 2, 2019 at 3:52 AM 俊杰陈 <[email protected]> wrote:
> >> > > > >
> >> > > > > > Any thoughts from other committers and developers?
> >> > > > > >
> >> > > > > > I 'd like to start a vote firstly, you could either provide
> your
> >> > > input
> >> > > > here
> >> > > > > > or on vote thread.
> >> > > > > >
> >> > > > > >
> >> > > > > >
> >> > > > > > On Mon, Jul 1, 2019 at 8:20 PM Zoltan Ivanfi
> >> > <[email protected]
> >> > > >
> >> > > > > > wrote:
> >> > > > > >
> >> > > > > > > Hi,
> >> > > > > > >
> >> > > > > > > I would like to clarify one point of my previous e-mail:
> >> While I
> >> > > > reasoned
> >> > > > > > > that for compressions and encodings we should avoid picking
> >> > > > algorithms
> >> > > > > > > superseded by better ones, I also reasoned that for bloom
> >> filters
> >> > > we
> >> > > > do
> >> > > > > > not
> >> > > > > > > necessarily have to be as strict, because a reader with
> >> missing
> >> > > > > > > implementation will still be able to read data from files
> that
> >> > > > contain
> >> > > > > > > unsupported bloom filter data structures.
> >> > > > > > >
> >> > > > > > > Personally I'm fine with moving forward with the current
> hash
> >> > > > proposal,
> >> > > > > > > even if the chosen algorithm is not considered to be the
> best
> >> of
> >> > > its
> >> > > > > > class.
> >> > > > > > >
> >> > > > > > > Br,
> >> > > > > > >
> >> > > > > > > Zoltan
> >> > > > > > >
> >> > > > > > > On Sun, Jun 30, 2019 at 11:02 PM Jim Apple <
> >> [email protected]>
> >> > > > wrote:
> >> > > > > > >
> >> > > > > > > > On 2019/06/28 16:43:23, Ryan Blue
> <[email protected]
> >> >
> >> > > > wrote:
> >> > > > > > > > > I agree with Zoltan. Since we want to ensure
> >> compatibility,
> >> > it
> >> > > > would
> >> > > > > > be
> >> > > > > > > > > better to choose the best option now instead of making
> >> > everyone
> >> > > > > > support
> >> > > > > > > > two
> >> > > > > > > > > options forever.
> >> > > > > > > >
> >> > > > > > > > I'd guess there probably isn't a single best option. I
> >> suspect
> >> > > > there's
> >> > > > > > a
> >> > > > > > > > tradeoff between ease of implementation and speed, for
> >> > instance,
> >> > > > since
> >> > > > > > I
> >> > > > > > > > expect it's easy to find an MD5 library in most
> programming
> >> > > > languages
> >> > > > > > and
> >> > > > > > > > operating systems, yet MD5 is very slow compared to
> >> > > > non-cryptographic
> >> > > > > > > hash
> >> > > > > > > > functions designed for speed like xxhash.
> >> > > > > > > >
> >> > > > > > > > There's also a significant amount of variability across
> >> > processor
> >> > > > > > > families
> >> > > > > > > > (64-bit multiply-shift in ARM vs x86-64) or even different
> >> > > > versions of
> >> > > > > > > the
> >> > > > > > > > same processor family (CLHash in Haswell vs. Sandy Lake).
> >> There
> >> > > are
> >> > > > > > also
> >> > > > > > > > quality tradeoffs that depend on the average bye length of
> >> the
> >> > > > input
> >> > > > > > (FNV
> >> > > > > > > > vs vhash) or how much L1 cache the user wants to use for
> the
> >> > hash
> >> > > > > > > function
> >> > > > > > > > (tabulation hashing vs. multiply-shift).
> >> > > > > > > >
> >> > > > > > > > To deal with this level of ambiguity, I'd suggest that v1
> >> > should
> >> > > > > > include
> >> > > > > > > a
> >> > > > > > > > hash function that works well for certain common
> >> environments.
> >> > As
> >> > > > far
> >> > > > > > as
> >> > > > > > > I
> >> > > > > > > > know, murmur and xxhash would both fit that bill.
> >> > > > > > > >
> >> > > > > > >
> >> > > > > >
> >> > > > > >
> >> > > > > > --
> >> > > > > > Thanks & Best Regards
> >> > > > > >
> >> > > >
> >> > > >
> >> > > >
> >> > > > --
> >> > > > Thanks & Best Regards
> >> > > >
> >> > >
> >> >
> >> >
> >> > --
> >> > Thanks & Best Regards
> >> >
> >>
> >
> >
> > --
> > Thanks & Best Regards
> >
>
>
> --
> Thanks & Best Regards
>

Reply via email to