[
https://issues.apache.org/jira/browse/BEAM-3914?focusedWorklogId=85724&page=com.atlassian.jira.plugin.system.issuetabpanels:worklog-tabpanel#worklog-85724
]
ASF GitHub Bot logged work on BEAM-3914:
----------------------------------------
Author: ASF GitHub Bot
Created on: 29/Mar/18 17:27
Start Date: 29/Mar/18 17:27
Worklog Time Spent: 10m
Work Description: tgroh opened a new pull request #4977: [BEAM-3914]
Deduplicate Unzipped Flattens after Pipeline Fusion
URL: https://github.com/apache/beam/pull/4977
Flattens are implicitly unzipped during fusion, by as aggressively as
possible
fusing them into any stage that produces one of their inputs. This may lead
to multiple stages producing the same output PCollection (in fact, if a
flatten
has inputs that are not all in the same stage, this is guaranteed to happen).
Because utilities expect exactly one PTransform to produce any given
PCollection,
this behavior breaks those utilities due to mismatched preconditions. This
inserts
a synthetic flatten node to flatten together all of the partial PCollections
produced
by each flatten input, executed in the runner, and introduces a synthetic
PCollection
for each of the locations the flatten exists within as the only producer of
that output.
------------------------
Follow this checklist to help us incorporate your contribution quickly and
easily:
- [ ] Make sure there is a [JIRA
issue](https://issues.apache.org/jira/projects/BEAM/issues/) filed for the
change (usually before you start working on it). Trivial changes like typos do
not require a JIRA issue. Your pull request should address just this issue,
without pulling in other changes.
- [ ] Format the pull request title like `[BEAM-XXX] Fixes bug in
ApproximateQuantiles`, where you replace `BEAM-XXX` with the appropriate JIRA
issue.
- [ ] Write a pull request description that is detailed enough to
understand:
- [ ] What the pull request does
- [ ] Why it does it
- [ ] How it does it
- [ ] Why this approach
- [ ] Each commit in the pull request should have a meaningful subject line
and body.
- [ ] Run `mvn clean verify` to make sure basic checks pass. A more
thorough check will be performed on your pull request automatically.
- [ ] If this contribution is large, please file an Apache [Individual
Contributor License Agreement](https://www.apache.org/licenses/icla.pdf).
----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on GitHub and use the
URL above to go to the specific comment.
For queries about this service, please contact Infrastructure at:
[email protected]
Issue Time Tracking
-------------------
Worklog Id: (was: 85724)
Time Spent: 10m
Remaining Estimate: 0h
> 'Unzip' flattens before performing fusion
> -----------------------------------------
>
> Key: BEAM-3914
> URL: https://issues.apache.org/jira/browse/BEAM-3914
> Project: Beam
> Issue Type: Improvement
> Components: runner-core
> Reporter: Thomas Groh
> Assignee: Thomas Groh
> Priority: Major
> Labels: portability
> Time Spent: 10m
> Remaining Estimate: 0h
>
> This consists of duplicating nodes downstream of a flatten that exist within
> an environment, and reintroducing the flatten immediately upstream of a
> runner-executed transform (the flatten should be executed within the runner)
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)