[
https://issues.apache.org/jira/browse/HUDI-8390?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
sivabalan narayanan updated HUDI-8390:
--------------------------------------
Fix Version/s: 1.0.0
> Fix cols stats pruning for filtering out a file slice after fetching stats
> and predicate pruning
> ------------------------------------------------------------------------------------------------
>
> Key: HUDI-8390
> URL: https://issues.apache.org/jira/browse/HUDI-8390
> Project: Apache Hudi
> Issue Type: Improvement
> Components: metadata
> Reporter: sivabalan narayanan
> Assignee: sivabalan narayanan
> Priority: Major
> Fix For: 1.0.0
>
>
> h3. Pruning Design:
> * step1 : Fetch latest file slices for pruned partitions (from MDT)
> * step2.a : Fetch stats from Col stats index which outputs in the format
> \{{File1, col1 ➝ stat1}, \{File2, col1 ➝ stat2},...} i.e. one entry per
> file,column combo. Here we are reading using
> HoodieTableMetadata.{*}getRecordsByKeyPrefixes(){*}. just that we are passing
> in just the {*}columns{*}.
> ** step2.b: Apply filter function to prune entries from step 2.a based on
> the list from step 1. col stats value will contain the file name and we
> filter based on that. Output from this step will be latest files looked up
> from col stats partition in MDT.
> ** step2.b : Construct a matrix of the format File1 ➝ \{col1_valuecount,
> col1_minvalue, col1_maxvalue, col2_valuecount, .... } i.e. one entry per file.
> ** step2.c: Get the list of files indexed by col stats.
> ** step2.d: Apply the query predicate and get the list of pruned file names
> over step 2.b.
> ** step3: If there are any files missing to be indexed from col stats (step1
> output - step2.c output), add them back to 2.d to get list of final pruned
> files list. Or in other words, pruned files + missingToIndexFiles are the
> final set of candidate files we return from this step.
> *** lets name the output from step3 as *candidate files.*
> ** step4: For every file slice from step3 => if every file in this file
> slice is missing from the candidate files, we can ignore the file slice(in
> other words, every file in this file slice did not match the predicate from
> col stats, we are safe to ignore the entire file slice). Even if one file is
> present in candidate files, we need to include the file slice in its entirety.
>
> This is regarding step 4. As per current logic, we go through every file in
> the candidate file and check if any of them matches any file in the current
> file slice. If it matches, we include the file slice and move onto next file
> slice for processing. Shouldn't we reverse the lookup. For every file slice ➝
> for every file ➝ check if its part of candidate list, if any match, include
> the file slice. if not ignore the file slice.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)