GitHub user erenavsarogullari opened a pull request:
[SPARK-17894] [CORE] Ensure uniqueness of TaskSetManager name.
## What changes were proposed in this pull request?
`TaskSetManager` should have unique name to avoid adding duplicate ones to
parent `Pool` via `SchedulableBuilder`. This problem has been surfaced with
following discussion: [[PR: Avoid adding of duplicate
There is 1x1 relationship between `stageAttemptId` and `TaskSetManager` so
`taskSet.Id` covering both `stageId` and `stageAttemptId` looks to be used for
uniqueness of `TaskSetManager` name instead of just `stageId`.
**Current TaskSetManager Name** :
`var name = "TaskSet_" + taskSet.stageId.toString`
**Proposed TaskSetManager Name** :
`val name = "TaskSet_" + taskSet.Id ` `// taskSet.Id = (stageId + "." +
**Sample** : TaskSet_0.0
## How was this patch tested?
Added new Unit Test.
cc @kayousterhout @markhamstra
You can merge this pull request into a Git repository by running:
$ git pull https://github.com/erenavsarogullari/spark SPARK-17894
Alternatively you can review and apply these changes as the patch at:
To close this pull request, make a commit to your master/trunk branch
with (at least) the following in the commit message:
This closes #15463
Author: erenavsarogullari <erenavsarogull...@gmail.com>
Ensure uniqueness of TaskSetManager name.
If your project is set up for it, you can reply to this email and have your
reply appear on GitHub as well. If your project does not have this feature
enabled and wishes so, or if the feature is enabled but not working, please
contact infrastructure at infrastruct...@apache.org or file a JIRA ticket
To unsubscribe, e-mail: reviews-unsubscr...@spark.apache.org
For additional commands, e-mail: reviews-h...@spark.apache.org