Hi,

I am trying to run LDA module on my private hadoop cluster consisiting of 11
datanodes and each of them having approx. 70 GB of storage remaining and has
4 GB of ram and these dual core machines. 

My input dataset consist of around 620k documents with a total size of 2.5GB
on which I want to train LDA. There are 2 issues that I am facing, 

1) It is taking enormous amount of time to learn, i.e. my 1st iteration's
map job itself is not completing in a day
2) My tasktrackers starts giving Spill Fail error after the disk is
completely full and I have no clue as to what kind of temporary storage
mapper is storing.

Can anyone help me out on what the issue could be and how to rectify it. 

Thanks.

--
View this message in context: 
http://lucene.472066.n3.nabble.com/Storage-and-Running-Time-issue-running-LDA-on-Cluster-tp3986579.html
Sent from the Mahout User List mailing list archive at Nabble.com.

Reply via email to