[
https://issues.apache.org/jira/browse/FLINK-1297?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14736534#comment-14736534
]
ASF GitHub Bot commented on FLINK-1297:
---------------------------------------
Github user tammymendt commented on the pull request:
https://github.com/apache/flink/pull/605#issuecomment-138852658
I can have a look at it this afternoon. I suspect it might be because the
function getAllAccumulators has been deprecated.
On Sep 9, 2015 11:11 AM, "Fabian Hueske" <[email protected]> wrote:
> After rebasing the PR to the current master, the
> OperatorStatsAccumulatorTest is failing. In the original PR (based on a
> master from begin of August) the test is passing. @mxm
> <https://github.com/mxm>, you are more familiar with Flink's
> accumulators. Can you have a look? Thanks!
>
> —
> Reply to this email directly or view it on GitHub
> <https://github.com/apache/flink/pull/605#issuecomment-138847826>.
>
> Add support for tracking statistics of intermediate results
> -----------------------------------------------------------
>
> Key: FLINK-1297
> URL: https://issues.apache.org/jira/browse/FLINK-1297
> Project: Flink
> Issue Type: Improvement
> Components: Distributed Runtime
> Reporter: Alexander Alexandrov
> Assignee: Alexander Alexandrov
> Fix For: 0.10
>
> Original Estimate: 1,008h
> Remaining Estimate: 1,008h
>
> One of the major problems related to the optimizer at the moment is the lack
> of proper statistics.
> With the introduction of staged execution, it is possible to instrument the
> runtime code with a statistics facility that collects the required
> information for optimizing the next execution stage.
> I would therefore like to contribute code that can be used to gather basic
> statistics for the (intermediate) result of dataflows (e.g. min, max, count,
> count distinct) and make them available to the job manager.
> Before I start, I would like to hear some feedback form the other users.
> In particular, to handle skew (e.g. on grouping) it might be good to have
> some sort of detailed sketch about the key distribution of an intermediate
> result. I am not sure whether a simple histogram is the most effective way to
> go. Maybe somebody would propose another lightweight sketch that provides
> better accuracy.
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)