Hi sebastian,
 
That makes sense. Thank you very much.
 
--- Young




>I did some inspection on the grouplens dataset and it turned out that
>Ted was absolutely right.
>
>I picked some random users and checked how many items
>getAllOtherItemIDs(...) returns for them. Actually the whole dataset is
>the result in most of the cases.
>
>So IMHO it's correct that the current implementation of
>getAllOtherItemIDs(...) is not suited for this specific dataset.
>
>Young, should do some tests with your own data. If I remember correctly,
>you said it's purchases from an onlineshop, these should result in a
>much more sparse user-item-matrix and therefore much faster computations.
>
>--sebastian
>
>
>Anatomy of the data:
>
>number of preferences:    1000209
>number of items:        3706
>number of users:        6040
>number of users with more than 500 prefs:    396
>number of items with more than 100 prefs:  2006
>
>
>Random userID, number of candidate items:
>
>1950,    3567
>2010,    2973
>4193,    3444
>734,    3658
>4655,    3364
>1569,    3611
>3717,    3407
>4313,    3608
>195,    2884
>3827,    3516
>3803,    3671
>3476,    3001
>1912,    2759
>1354,    3022
>3961,    3475
>2963,    3661
>3381,    3661
>5137,    3583
>3870,    3675
>2269,    3671
>1843,    3586
>5905,    3553
>2067,    3506
>456, 3548
>477, 3495
>
>
>
>Am 21.07.2010 16:17, schrieb Ted Dunning:
>> Is it possible that there are some items that all users see/rate/interact
>> with?
>>
>> That can cause problems like this because all users are then somewhat
>> similar and you wind up inspecting the entire rating matrix.
>>
>> Any such items should be added to a kill list.
>>
>> 2010/7/21 Sebastian Schelter <[email protected]>
>>
>>   
>>> Well,
>>>
>>> there must be something wrong here, I've seen systems in production with
>>> similar sized data that responded clearly below 1 second all the time. I
>>> really can't imagine
>>> that FastIDSet would be the cause for this.
>>>
>>> Are you sure the JVM has all the machine for itself? No email
>>> application in the background checking mails, no memory swapping of the OS?
>>>
>>> If you want you can make your data and test code available and I can
>>> check it on my notebook.
>>>
>>> --sebastian
>>>
>>>
>>>
>>>
>>> Am 21.07.2010 11:23, schrieb Young:
>>>     
>>>> Hi again,
>>>> I use java profiler to find out the latency comes from the
>>>>       
>>> FastIDSet.addAll() and FastIDSet.add()..
>>>     
>>>> Blows are source code.
>>>>
>>>>  protected FastIDSet getAllOtherItems(long theUserID) throws
>>>>       
>>> TasteException {
>>>     
>>>>   ......
>>>>       for (int j = 0; j < size2; j++) {
>>>>
>>>>       
>>> possibleItemsIDs.addAll(dataModel.getItemIDsFromUser(prefs2.getUserID(j)));
>>>     
>>>>       }
>>>>   ......
>>>> }
>>>>   public FastIDSet getItemIDsFromUser(long userID) throws TasteException
>>>>       
>>> {
>>>     
>>>>     PreferenceArray prefs = getPreferencesFromUser(userID);
>>>>     int size = prefs.length();
>>>>     FastIDSet result = new FastIDSet(size);
>>>>     for (int i = 0; i < size; i++) {
>>>>       result.add(prefs.getItemID(i));
>>>>     }
>>>>     return result;
>>>>   }
>>>>
>>>>
>>>>
>>>>
>>>>
>>>>
>>>>
>>>>
>>>>
>>>>       
>>>>> Yes, I use the genericdatamodel which is in-memory.
>>>>>
>>>>>
>>>>>
>>>>>
>>>>>
>>>>>
>>>>>         
>>>>>> Is only the similarity matrix in-memory? The crucial thing here is the
>>>>>> data model not the similarity matrix, are you using an in-memory data
>>>>>>           
>>> model?
>>>     
>>>>>> Am 21.07.2010 08:33, schrieb Young:
>>>>>>
>>>>>>           
>>>>>>> Yes, I am pretty sure. I have stored the similarity matrix in-memory
>>>>>>>             
>>> and I print out the time spent in getAllOtherItems() and this is the only
>>> one time-consuming method in the recommendation. My laptop CPU is Intel
>>> P8600 2.4G, and the memory used for JVM is 1GB.
>>>     
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>             
>>>>>>>> Hi Young,
>>>>>>>>
>>>>>>>> I would disagree that a response time of 6 seconds is OK for online
>>>>>>>> recommendations, the time should be something like < 100ms.
>>>>>>>> I'm really surprised that you would see such response times with an
>>>>>>>> in-memory data model, I have experience with in-memory models of
>>>>>>>>               
>>> roughly
>>>     
>>>>>>>> the same size
>>>>>>>> and usually the computations are blazingly fast.
>>>>>>>>
>>>>>>>> Are you absolutely sure that the time is spent in this method and not
>>>>>>>> later in the similarity computation?
>>>>>>>>
>>>>>>>> --sebastian
>>>>>>>>
>>>>>>>> Am 21.07.2010 07:54, schrieb Young:
>>>>>>>>
>>>>>>>>
>>>>>>>>               
>>>>>>>>> So based on the 1M dataset, the time spent in
>>>>>>>>>                 
>>> getAllOtherItems(userID) is among the 2 and 10 seconds.
>>>     
>>>>>>>>> for example,
>>>>>>>>> If one user rates 200 items and for each item, the time spent in
>>>>>>>>>                 
>>> calculating the neighbors is expected to 30ms.
>>>     
>>>>>>>>> So that makes 6 seconds. It is generally okay. But if the dataset is
>>>>>>>>>                 
>>> expanded to 100M dataset, I think 30ms may grow up to 30 * 100 ms and that
>>> will be a long time.
>>>     
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>                 
>>>>>>>>>> It still seems strange to observe such a bottleneck, I'm not sure
>>>>>>>>>> what's going on.
>>>>>>>>>> You are using an in-memory model like GenericDataModel?
>>>>>>>>>> We could look at ways to optimize that method, though it looks
>>>>>>>>>>                   
>>> reasonably tight.
>>>     
>>>>>>>>>> Where within that method do you see time spent?
>>>>>>>>>>
>>>>>>>>>> 2010/7/20 Young <[email protected]>:
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>                   
>>>>>>>>>>> Hi again,
>>>>>>>>>>> When I do the itembased recommendation, I find there are some
>>>>>>>>>>>                     
>>> latency in getAllOtherItems(long userID). Because it is calculating the
>>> items' neighbors and merge these neighbors together. So I am thinking if I
>>> precompute each item's neighbors and store in the database, then when I
>>> getAllOtherItems(), I could merge these neighbors directly. Is this useful
>>> for reducing the latency?
>>>     
>>>>>>>>>>> Or is there other way to make the online-recommendation much
>>>>>>>>>>>                     
>>> faster?
>>>     
>>>>>>>>>>> Thank you.
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>                     
>>>>>>>>>>>> Yes you probably want a new, separate table. You have an extra
>>>>>>>>>>>>                       
>>> step of
>>>     
>>>>>>>>>>>> computing some notion of similarity anyway, and you probably want
>>>>>>>>>>>>                       
>>> to
>>>     
>>>>>>>>>>>> separate this table from your main data table anyhow for reasons
>>>>>>>>>>>>                       
>>> of
>>>     
>>>>>>>>>>>> performance and business logic separation.
>>>>>>>>>>>>
>>>>>>>>>>>> 2010/7/19 Young <[email protected]>:
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>                       
>>>>>>>>>>>>> So my prpblem is that I want to build datamodel based on what
>>>>>>>>>>>>>                         
>>> user has bought or added to their favorite or rated.
>>>     
>>>>>>>>>>>>> You mean I need a table describe all these user behavior. For
>>>>>>>>>>>>>                         
>>> example, if user buys one item, I guess the user preference is 4 and add
>>> into this table?
>>>     
>>>>>>>>>>>>>
>>>>>>>>>>>>>
>>>>>>>>>>>>>
>>>>>>>>>>>>>
>>>>>>>>>>>>>
>>>>>>>>>>>>>
>>>>>>>>>>>>>                         
>>>>>>>>>>>>>> No, you need one table (or view if you like) containing all
>>>>>>>>>>>>>>                           
>>> data. If
>>>     
>>>>>>>>>>>>>> you can't do this, you could write your own copy of a
>>>>>>>>>>>>>>                           
>>> JDBCDataModel
>>>     
>>>>>>>>>>>>>> that can query multiple tables, or, that changes its SQL
>>>>>>>>>>>>>>                           
>>> queries to
>>>     
>>>>>>>>>>>>>> use UNION statements. I imagine it will slow down a lot.
>>>>>>>>>>>>>>
>>>>>>>>>>>>>> If you mean, can you use a table with preferences with a model
>>>>>>>>>>>>>>                           
>>> that
>>>     
>>>>>>>>>>>>>> ignores preferences, sure you can. The extra column is ignored.
>>>>>>>>>>>>>>
>>>>>>>>>>>>>> 2010/7/19 Young <[email protected]>:
>>>>>>>>>>>>>>
>>>>>>>>>>>>>>
>>>>>>>>>>>>>>
>>>>>>>>>>>>>>                           
>>>>>>>>>>>>>>> Hi,
>>>>>>>>>>>>>>> I have three tables, one is with preference and another two
>>>>>>>>>>>>>>>                             
>>> are without preference. Does mahout have some algorithm to integret these
>>> tables into one datamodel?
>>>     
>>>>>>>>>>>>>>> Thank you
>>>>>>>>>>>>>>>
>>>>>>>>>>>>>>>
>>>>>>>>>>>>>>>
>>>>>>>>>>>>>>>                             
>>>>>>>>>>>>>
>>>>>>>>>>>>>
>>>>>>>>>>>>>                         
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>                     
>>>>>>>>
>>>>>>>>               
>>>>>>           
>>>
>>>     
>>   
>

Reply via email to