[
https://issues.apache.org/jira/browse/BEAM-6670?focusedWorklogId=212778&page=com.atlassian.jira.plugin.system.issuetabpanels:worklog-tabpanel#worklog-212778
]
ASF GitHub Bot logged work on BEAM-6670:
----------------------------------------
Author: ASF GitHub Bot
Created on: 14/Mar/19 00:16
Start Date: 14/Mar/19 00:16
Worklog Time Spent: 10m
Work Description: jkff commented on issue #7956: [BEAM-6670] Add option
to disable reshuffling of JdbcIO
URL: https://github.com/apache/beam/pull/7956#issuecomment-472655600
Thank you, however I am currently not active on Beam so I don't have time to
do a full review.
R: @chamikaramj perhaps?
Quick note though - generally we avoid adding any kind of tuning parameters
unless there's a very, very good reason (because people almost always end up
setting them to a bad value, or to a value that was good at the time it was set
but is no longer good and nobody updates it) - the standard we've been using is
"some pipeline behaves catastrophically bad in a way that can only be prevented
by adding the tuning parameter, and it can not be chosen automatically". Even
when this standard is met, it is better to let the user specify hints in
declarative rather than operational terms - e.g. how TextIO.read() has a
"withHintMatchesManyFiles()" parameter rather than "withHintGoViaReadAll()"
(which is what that parameter currently actually does).
It would help if you could e.g. elaborate just how much worse is the
performance you observe without specifying this parameter in your pipeline.
----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
For queries about this service, please contact Infrastructure at:
[email protected]
Issue Time Tracking
-------------------
Worklog Id: (was: 212778)
Time Spent: 0.5h (was: 20m)
> Option to disable reparallelization in JdbcIO.Read
> --------------------------------------------------
>
> Key: BEAM-6670
> URL: https://issues.apache.org/jira/browse/BEAM-6670
> Project: Beam
> Issue Type: Wish
> Components: io-java-jdbc
> Reporter: Mike Pedersen
> Priority: Minor
> Time Spent: 0.5h
> Remaining Estimate: 0h
>
> I'm doing approx. 20 JDBC queries against a database and then joining them
> together in a group by. Every single one of these queries does a reshuffle,
> which is sort of useless due to them being fed to a CoGroupByKey immediately
> afterwards.
> Reshuffle by default seems sensible by the principle of least surprise, but
> it would be nice to have a way to disable it when it's not necessary. For
> example a "withReshuffle(boolean)" method.
> This should be an easy addition and I am willing to add this if it sounds
> reasonable enough.
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)