[ https://issues.apache.org/jira/browse/FLINK-2131?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15556280#comment-15556280 ]
ASF GitHub Bot commented on FLINK-2131: --------------------------------------- Github user skonto commented on the issue: https://github.com/apache/flink/pull/757 @sachingoel0101 @tillrohrmann any plans for this PR? > Add Initialization schemes for K-means clustering > ------------------------------------------------- > > Key: FLINK-2131 > URL: https://issues.apache.org/jira/browse/FLINK-2131 > Project: Flink > Issue Type: Task > Components: Machine Learning Library > Reporter: Sachin Goel > Assignee: Sachin Goel > > The Lloyd's [KMeans] algorithm takes initial centroids as its input. However, > in case the user doesn't provide the initial centers, they may ask for a > particular initialization scheme to be followed. The most commonly used are > these: > 1. Random initialization: Self-explanatory > 2. kmeans++ initialization: http://ilpubs.stanford.edu:8090/778/1/2006-13.pdf > 3. kmeans|| : http://theory.stanford.edu/~sergei/papers/vldb12-kmpar.pdf > For very large data sets, or for large values of k, the kmeans|| method is > preferred as it provides the same approximation guarantees as kmeans++ and > requires lesser number of passes over the input data. -- This message was sent by Atlassian JIRA (v6.3.4#6332)