[jira] [Commented] (FLINK-2131) Add Initialization schemes for K-means clustering

ASF GitHub Bot (JIRA) Fri, 07 Oct 2016 19:00:04 -0700

    [ 
https://issues.apache.org/jira/browse/FLINK-2131?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15556922#comment-15556922
 ]


ASF GitHub Bot commented on FLINK-2131:
---------------------------------------

Github user sachingoel0101 commented on the issue:

    https://github.com/apache/flink/pull/757
  
    I'll update based on your comments in a few days. ^^
    
    On Oct 8, 2016 06:21, "Stavros Kontopoulos" <[email protected]>
    wrote:
    
    > @sachingoel0101 <https://github.com/sachingoel0101> @tillrohrmann
    > <https://github.com/tillrohrmann> any plans for this PR?
    >
    > —
    > You are receiving this because you were mentioned.
    > Reply to this email directly, view it on GitHub
    > <https://github.com/apache/flink/pull/757#issuecomment-252364165>, or mute
    > the thread
    > 
<https://github.com/notifications/unsubscribe-auth/AIdpFR7VFxxqFaf_74WUxwKsOrGhOi1Eks5qxrfugaJpZM4E0H5N>
    > .
    >



> Add Initialization schemes for K-means clustering
> -------------------------------------------------
>
>                 Key: FLINK-2131
>                 URL: https://issues.apache.org/jira/browse/FLINK-2131
>             Project: Flink
>          Issue Type: Task
>          Components: Machine Learning Library
>            Reporter: Sachin Goel
>            Assignee: Sachin Goel
>
> The Lloyd's [KMeans] algorithm takes initial centroids as its input. However, 
> in case the user doesn't provide the initial centers, they may ask for a 
> particular initialization scheme to be followed. The most commonly used are 
> these:
> 1. Random initialization: Self-explanatory
> 2. kmeans++ initialization: http://ilpubs.stanford.edu:8090/778/1/2006-13.pdf
> 3. kmeans|| : http://theory.stanford.edu/~sergei/papers/vldb12-kmpar.pdf
> For very large data sets, or for large values of k, the kmeans|| method is 
> preferred as it provides the same approximation guarantees as kmeans++ and 
> requires lesser number of passes over the input data.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

[jira] [Commented] (FLINK-2131) Add Initialization schemes for K-means clustering

Reply via email to