Thanks for the suggestions, Ted. Your idea of treating combined categories as one is something I hadn't thought of. FWIW, looking at the documents we have available for training, multiple categories occur about 13% of the time.
David On Fri, Jun 08, 2012 at 12:46:28AM +0200, Ted Dunning wrote: > There are a variety of methods to use here. I would recommend you try a > variety of them to decide on your best approach. > > 1) first, count all the combinations of labels. If there are not that > many, you may just want to consider each combination a separate category. > Another option is to separate the categories into independent sets of > disjoint categories. > > 2) second, try to determine which categories are most confusable or > similar. One way to do this is to simply build a NB or CNB classifier and > look at the confusion matrix to see which errors get made. Those groups of > categories that get confused are candidates for a super category. > > 3) now start building your classifiers. You might try various tree > structures of categories, including a flat tree, one build up of confusable > classes, and one built up based on your intuitions. At each level, build > either a binary classifier per category or a 1 of n classifier if no two > categories at that level ever are tagged. You may want to build a > secondary binary classifier for each category whose inputs are the outputs > from each of the first level categorizers. > > 4) tune and adjust. tune and adjust. > > On Fri, Jun 8, 2012 at 12:31 AM, David Engel <[email protected]> wrote: > > > Hi, > > > > I've been dabbling with Mahout off and on for a few months preparing > > for a classification project. It's now time to stop experimenting and > > do something for real. I've picked up a lot of things from following > > this list, but would like some advice regarding a few things before > > proceeding. I'll start with a very brief description of the project > > and then follow up with some questions. > > > > We need to classify potentially millions of documents into about 100 > > or so categories. Most documents will probably only belong to 1 > > category, but some will belong to several. It's also possible for > > some documents to not belong to any of the chosen categories. > > > > As noted, we need to handle the case where a document belongs to > > multiple categories. My understanding is the classification > > algorithms are primarily geared to classifying an item into one > > category and we would need run multiple classifiers in parallel to > > match multiple categories. Is that correct? I found something in the > > subversion logs referencing "multilabel" support that sounded > > interesting, but it was removed a few weeks ago. Is that of any > > relevance? > > > > Also as noted, we need to handle the case where a document belongs to > > no categories. Do any of the classification algorithms support the > > concept of an implicit "other" or "none" category or do we need to add > > an explicit one? If the latter, how many training samples do we need > > to use compared to the number of samples for the target categories? > > > > Finally, I recall seeing on this list that some of the classification > > algorithms break down if more than 20 to 30 categories are used and > > that multiple classifiers should be used hierarchically when more > > categories are needed. Is that still correct? If so, is there any > > preferred way to organize the cascaded classifiers? I'm currently > > analyzing the documents we will use for training to see which > > categories often, seldom or never occur together. > > > > David > > -- > > David Engel > > [email protected] > > -- David Engel [email protected]
