Thanks Ted, too. :)
>Hi sebastian, > >That makes sense. Thank you very much. > >--- Young > > > > > >>I did some inspection on the grouplens dataset and it turned out that >>Ted was absolutely right. >> >>I picked some random users and checked how many items >>getAllOtherItemIDs(...) returns for them. Actually the whole dataset is >>the result in most of the cases. >> >>So IMHO it's correct that the current implementation of >>getAllOtherItemIDs(...) is not suited for this specific dataset. >> >>Young, should do some tests with your own data. If I remember correctly, >>you said it's purchases from an onlineshop, these should result in a >>much more sparse user-item-matrix and therefore much faster computations. >> >>--sebastian >> >> >>Anatomy of the data: >> >>number of preferences: 1000209 >>number of items: 3706 >>number of users: 6040 >>number of users with more than 500 prefs: 396 >>number of items with more than 100 prefs: 2006 >> >> >>Random userID, number of candidate items: >> >>1950, 3567 >>2010, 2973 >>4193, 3444 >>734, 3658 >>4655, 3364 >>1569, 3611 >>3717, 3407 >>4313, 3608 >>195, 2884 >>3827, 3516 >>3803, 3671 >>3476, 3001 >>1912, 2759 >>1354, 3022 >>3961, 3475 >>2963, 3661 >>3381, 3661 >>5137, 3583 >>3870, 3675 >>2269, 3671 >>1843, 3586 >>5905, 3553 >>2067, 3506 >>456, 3548 >>477, 3495 >> >> >> >>Am 21.07.2010 16:17, schrieb Ted Dunning: >>> Is it possible that there are some items that all users see/rate/interact >>> with? >>> >>> That can cause problems like this because all users are then somewhat >>> similar and you wind up inspecting the entire rating matrix. >>> >>> Any such items should be added to a kill list. >>> >>> 2010/7/21 Sebastian Schelter <[email protected]> >>> >>> >>>> Well, >>>> >>>> there must be something wrong here, I've seen systems in production with >>>> similar sized data that responded clearly below 1 second all the time. I >>>> really can't imagine >>>> that FastIDSet would be the cause for this. >>>> >>>> Are you sure the JVM has all the machine for itself? No email >>>> application in the background checking mails, no memory swapping of the OS? >>>> >>>> If you want you can make your data and test code available and I can >>>> check it on my notebook. >>>> >>>> --sebastian >>>> >>>> >>>> >>>> >>>> Am 21.07.2010 11:23, schrieb Young: >>>> >>>>> Hi again, >>>>> I use java profiler to find out the latency comes from the >>>>> >>>> FastIDSet.addAll() and FastIDSet.add().. >>>> >>>>> Blows are source code. >>>>> >>>>> protected FastIDSet getAllOtherItems(long theUserID) throws >>>>> >>>> TasteException { >>>> >>>>> ...... >>>>> for (int j = 0; j < size2; j++) { >>>>> >>>>> >>>> possibleItemsIDs.addAll(dataModel.getItemIDsFromUser(prefs2.getUserID(j))); >>>> >>>>> } >>>>> ...... >>>>> } >>>>> public FastIDSet getItemIDsFromUser(long userID) throws TasteException >>>>> >>>> { >>>> >>>>> PreferenceArray prefs = getPreferencesFromUser(userID); >>>>> int size = prefs.length(); >>>>> FastIDSet result = new FastIDSet(size); >>>>> for (int i = 0; i < size; i++) { >>>>> result.add(prefs.getItemID(i)); >>>>> } >>>>> return result; >>>>> } >>>>> >>>>> >>>>> >>>>> >>>>> >>>>> >>>>> >>>>> >>>>> >>>>> >>>>>> Yes, I use the genericdatamodel which is in-memory. >>>>>> >>>>>> >>>>>> >>>>>> >>>>>> >>>>>> >>>>>> >>>>>>> Is only the similarity matrix in-memory? The crucial thing here is the >>>>>>> data model not the similarity matrix, are you using an in-memory data >>>>>>> >>>> model? >>>> >>>>>>> Am 21.07.2010 08:33, schrieb Young: >>>>>>> >>>>>>> >>>>>>>> Yes, I am pretty sure. I have stored the similarity matrix in-memory >>>>>>>> >>>> and I print out the time spent in getAllOtherItems() and this is the only >>>> one time-consuming method in the recommendation. My laptop CPU is Intel >>>> P8600 2.4G, and the memory used for JVM is 1GB. >>>> >>>>>>>> >>>>>>>> >>>>>>>> >>>>>>>> >>>>>>>> >>>>>>>> >>>>>>>> >>>>>>>>> Hi Young, >>>>>>>>> >>>>>>>>> I would disagree that a response time of 6 seconds is OK for online >>>>>>>>> recommendations, the time should be something like < 100ms. >>>>>>>>> I'm really surprised that you would see such response times with an >>>>>>>>> in-memory data model, I have experience with in-memory models of >>>>>>>>> >>>> roughly >>>> >>>>>>>>> the same size >>>>>>>>> and usually the computations are blazingly fast. >>>>>>>>> >>>>>>>>> Are you absolutely sure that the time is spent in this method and not >>>>>>>>> later in the similarity computation? >>>>>>>>> >>>>>>>>> --sebastian >>>>>>>>> >>>>>>>>> Am 21.07.2010 07:54, schrieb Young: >>>>>>>>> >>>>>>>>> >>>>>>>>> >>>>>>>>>> So based on the 1M dataset, the time spent in >>>>>>>>>> >>>> getAllOtherItems(userID) is among the 2 and 10 seconds. >>>> >>>>>>>>>> for example, >>>>>>>>>> If one user rates 200 items and for each item, the time spent in >>>>>>>>>> >>>> calculating the neighbors is expected to 30ms. >>>> >>>>>>>>>> So that makes 6 seconds. It is generally okay. But if the dataset is >>>>>>>>>> >>>> expanded to 100M dataset, I think 30ms may grow up to 30 * 100 ms and that >>>> will be a long time. >>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>>> It still seems strange to observe such a bottleneck, I'm not sure >>>>>>>>>>> what's going on. >>>>>>>>>>> You are using an in-memory model like GenericDataModel? >>>>>>>>>>> We could look at ways to optimize that method, though it looks >>>>>>>>>>> >>>> reasonably tight. >>>> >>>>>>>>>>> Where within that method do you see time spent? >>>>>>>>>>> >>>>>>>>>>> 2010/7/20 Young <[email protected]>: >>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>>>> Hi again, >>>>>>>>>>>> When I do the itembased recommendation, I find there are some >>>>>>>>>>>> >>>> latency in getAllOtherItems(long userID). Because it is calculating the >>>> items' neighbors and merge these neighbors together. So I am thinking if I >>>> precompute each item's neighbors and store in the database, then when I >>>> getAllOtherItems(), I could merge these neighbors directly. Is this useful >>>> for reducing the latency? >>>> >>>>>>>>>>>> Or is there other way to make the online-recommendation much >>>>>>>>>>>> >>>> faster? >>>> >>>>>>>>>>>> Thank you. >>>>>>>>>>>> >>>>>>>>>>>> >>>>>>>>>>>> >>>>>>>>>>>> >>>>>>>>>>>> >>>>>>>>>>>> >>>>>>>>>>>> >>>>>>>>>>>> >>>>>>>>>>>>> Yes you probably want a new, separate table. You have an extra >>>>>>>>>>>>> >>>> step of >>>> >>>>>>>>>>>>> computing some notion of similarity anyway, and you probably want >>>>>>>>>>>>> >>>> to >>>> >>>>>>>>>>>>> separate this table from your main data table anyhow for reasons >>>>>>>>>>>>> >>>> of >>>> >>>>>>>>>>>>> performance and business logic separation. >>>>>>>>>>>>> >>>>>>>>>>>>> 2010/7/19 Young <[email protected]>: >>>>>>>>>>>>> >>>>>>>>>>>>> >>>>>>>>>>>>> >>>>>>>>>>>>> >>>>>>>>>>>>>> So my prpblem is that I want to build datamodel based on what >>>>>>>>>>>>>> >>>> user has bought or added to their favorite or rated. >>>> >>>>>>>>>>>>>> You mean I need a table describe all these user behavior. For >>>>>>>>>>>>>> >>>> example, if user buys one item, I guess the user preference is 4 and add >>>> into this table? >>>> >>>>>>>>>>>>>> >>>>>>>>>>>>>> >>>>>>>>>>>>>> >>>>>>>>>>>>>> >>>>>>>>>>>>>> >>>>>>>>>>>>>> >>>>>>>>>>>>>> >>>>>>>>>>>>>>> No, you need one table (or view if you like) containing all >>>>>>>>>>>>>>> >>>> data. If >>>> >>>>>>>>>>>>>>> you can't do this, you could write your own copy of a >>>>>>>>>>>>>>> >>>> JDBCDataModel >>>> >>>>>>>>>>>>>>> that can query multiple tables, or, that changes its SQL >>>>>>>>>>>>>>> >>>> queries to >>>> >>>>>>>>>>>>>>> use UNION statements. I imagine it will slow down a lot. >>>>>>>>>>>>>>> >>>>>>>>>>>>>>> If you mean, can you use a table with preferences with a model >>>>>>>>>>>>>>> >>>> that >>>> >>>>>>>>>>>>>>> ignores preferences, sure you can. The extra column is ignored. >>>>>>>>>>>>>>> >>>>>>>>>>>>>>> 2010/7/19 Young <[email protected]>: >>>>>>>>>>>>>>> >>>>>>>>>>>>>>> >>>>>>>>>>>>>>> >>>>>>>>>>>>>>> >>>>>>>>>>>>>>>> Hi, >>>>>>>>>>>>>>>> I have three tables, one is with preference and another two >>>>>>>>>>>>>>>> >>>> are without preference. Does mahout have some algorithm to integret these >>>> tables into one datamodel? >>>> >>>>>>>>>>>>>>>> Thank you >>>>>>>>>>>>>>>> >>>>>>>>>>>>>>>> >>>>>>>>>>>>>>>> >>>>>>>>>>>>>>>> >>>>>>>>>>>>>> >>>>>>>>>>>>>> >>>>>>>>>>>>>> >>>>>>>>>>>> >>>>>>>>>>>> >>>>>>>>>>>> >>>>>>>>> >>>>>>>>> >>>>>>> >>>> >>>> >>> >>
