[
https://issues.apache.org/jira/browse/IMPALA-8205?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16769511#comment-16769511
]
Tim Armstrong commented on IMPALA-8205:
---------------------------------------
The field in thrift is required, so I'm not sure what else we're meant to do
here -
https://github.com/apache/hive/blob/71dfd1d11f239caf8f16bc29db0f959e566f7659/standalone-metastore/metastore-common/src/main/thrift/hive_metastore.thrift#L499
The "1" looks like it's a bug that's been in there for a long time.
> Illegal statistics for numFalse and numTrue
> -------------------------------------------
>
> Key: IMPALA-8205
> URL: https://issues.apache.org/jira/browse/IMPALA-8205
> Project: IMPALA
> Issue Type: Bug
> Reporter: wuchang
> Priority: Major
> Labels: impala, numFalse, numTrue, statistics
>
> When impala compute statistics, it set *numFalse = -1* and *numTrue = 1* when
> the statistic is missing;
> *-1* for *numFalse* will corrupt some query engine like Presto and there
> already exists some PR report and hotfix it :
> [presto-11859|https://github.com/prestodb/presto/pull/11859]
> *1* for *numTrue* is also unreasonable because we are not sure whether it
> indicates the real numTrue statistics or a missing statistics;
> Also, previously , the *nullCount* also use -1 to indicate its absence which
> also caused problem for Presto. Presto has to add a hotfix for
> it([presto-11549|https://github.com/prestodb/presto/pull/11549]) . But it is
> a fortunate that impala has fixed this bug;
> It is necessary to set to null when these statistics are absent instead of -1
> and 1.
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]