Github user srowen commented on a diff in the pull request:

    https://github.com/apache/spark/pull/4696#discussion_r26078952
  
    --- Diff: docs/programming-guide.md ---
    @@ -728,6 +728,69 @@ def doStuff(self, rdd):
     
     </div>
     
    +### Understanding closures <a name="ClosuresLink"></a>
    +One of the harder things about Spark is understanding the scope and life 
cycle of variables and methods when executing code across a cluster. RDD 
operations that modify variables outside of their scope can be a frequent 
source of confusion. In the example below we'll look at code that uses 
`foreach()` to increment a counter, but similar issues can occur for other 
operations as well.
    +
    +#### Example
    +
    +Consider the naive RDD element sum below, which behaves completely 
differently depending on whether execution is happening within the same JVM. A 
common example of this is when running spark in `local` mode (`--master = 
local[n]`) versus deploying a Spark application to a cluster (e.g. via 
spark-submit to YARN): 
    +
    +<div class="codetabs">
    +
    +<div data-lang="scala"  markdown="1">
    +{% highlight scala %}
    +var counter = 0
    +var rdd = sc.parallelize(data)
    +
    +// Wrong: Don't do this!!
    +rdd.foreach(x => counter += x)
    +
    +println("Counter value: " + counter)
    +{% endhighlight %}
    +</div>
    +
    +<div data-lang="java"  markdown="1">
    +{% highlight java %}
    +int counter = 0;
    +JavaRDD<Integer> rdd = sc.parallelize(data); 
    +
    +// Wrong: Don't do this!!
    +rdd.foreach(x -> counter += x;)
    --- End diff --
    
    This one won't compile and I think the Python equivalent won't either, but 
don't fix it up yet; if that's all there is I can easily do it on merge.


---
If your project is set up for it, you can reply to this email and have your
reply appear on GitHub as well. If your project does not have this feature
enabled and wishes so, or if the feature is enabled but not working, please
contact infrastructure at [email protected] or file a JIRA ticket
with INFRA.
---

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to