I did some inspection on the grouplens dataset and it turned out that
Ted was absolutely right.

I picked some random users and checked how many items
getAllOtherItemIDs(...) returns for them. Actually the whole dataset is
the result in most of the cases.

So IMHO it's correct that the current implementation of
getAllOtherItemIDs(...) is not suited for this specific dataset.

Young, should do some tests with your own data. If I remember correctly,
you said it's purchases from an onlineshop, these should result in a
much more sparse user-item-matrix and therefore much faster computations.

--sebastian


Anatomy of the data:

number of preferences:    1000209
number of items:        3706
number of users:        6040
number of users with more than 500 prefs:    396
number of items with more than 100 prefs:  2006


Random userID, number of candidate items:

1950,    3567
2010,    2973
4193,    3444
734,    3658
4655,    3364
1569,    3611
3717,    3407
4313,    3608
195,    2884
3827,    3516
3803,    3671
3476,    3001
1912,    2759
1354,    3022
3961,    3475
2963,    3661
3381,    3661
5137,    3583
3870,    3675
2269,    3671
1843,    3586
5905,    3553
2067,    3506
456, 3548
477, 3495



Am 21.07.2010 16:17, schrieb Ted Dunning:
> Is it possible that there are some items that all users see/rate/interact
> with?
>
> That can cause problems like this because all users are then somewhat
> similar and you wind up inspecting the entire rating matrix.
>
> Any such items should be added to a kill list.
>
> 2010/7/21 Sebastian Schelter <[email protected]>
>
>   
>> Well,
>>
>> there must be something wrong here, I've seen systems in production with
>> similar sized data that responded clearly below 1 second all the time. I
>> really can't imagine
>> that FastIDSet would be the cause for this.
>>
>> Are you sure the JVM has all the machine for itself? No email
>> application in the background checking mails, no memory swapping of the OS?
>>
>> If you want you can make your data and test code available and I can
>> check it on my notebook.
>>
>> --sebastian
>>
>>
>>
>>
>> Am 21.07.2010 11:23, schrieb Young:
>>     
>>> Hi again,
>>> I use java profiler to find out the latency comes from the
>>>       
>> FastIDSet.addAll() and FastIDSet.add()..
>>     
>>> Blows are source code.
>>>
>>>  protected FastIDSet getAllOtherItems(long theUserID) throws
>>>       
>> TasteException {
>>     
>>>   ......
>>>       for (int j = 0; j < size2; j++) {
>>>
>>>       
>> possibleItemsIDs.addAll(dataModel.getItemIDsFromUser(prefs2.getUserID(j)));
>>     
>>>       }
>>>   ......
>>> }
>>>   public FastIDSet getItemIDsFromUser(long userID) throws TasteException
>>>       
>> {
>>     
>>>     PreferenceArray prefs = getPreferencesFromUser(userID);
>>>     int size = prefs.length();
>>>     FastIDSet result = new FastIDSet(size);
>>>     for (int i = 0; i < size; i++) {
>>>       result.add(prefs.getItemID(i));
>>>     }
>>>     return result;
>>>   }
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>>       
>>>> Yes, I use the genericdatamodel which is in-memory.
>>>>
>>>>
>>>>
>>>>
>>>>
>>>>
>>>>         
>>>>> Is only the similarity matrix in-memory? The crucial thing here is the
>>>>> data model not the similarity matrix, are you using an in-memory data
>>>>>           
>> model?
>>     
>>>>> Am 21.07.2010 08:33, schrieb Young:
>>>>>
>>>>>           
>>>>>> Yes, I am pretty sure. I have stored the similarity matrix in-memory
>>>>>>             
>> and I print out the time spent in getAllOtherItems() and this is the only
>> one time-consuming method in the recommendation. My laptop CPU is Intel
>> P8600 2.4G, and the memory used for JVM is 1GB.
>>     
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>             
>>>>>>> Hi Young,
>>>>>>>
>>>>>>> I would disagree that a response time of 6 seconds is OK for online
>>>>>>> recommendations, the time should be something like < 100ms.
>>>>>>> I'm really surprised that you would see such response times with an
>>>>>>> in-memory data model, I have experience with in-memory models of
>>>>>>>               
>> roughly
>>     
>>>>>>> the same size
>>>>>>> and usually the computations are blazingly fast.
>>>>>>>
>>>>>>> Are you absolutely sure that the time is spent in this method and not
>>>>>>> later in the similarity computation?
>>>>>>>
>>>>>>> --sebastian
>>>>>>>
>>>>>>> Am 21.07.2010 07:54, schrieb Young:
>>>>>>>
>>>>>>>
>>>>>>>               
>>>>>>>> So based on the 1M dataset, the time spent in
>>>>>>>>                 
>> getAllOtherItems(userID) is among the 2 and 10 seconds.
>>     
>>>>>>>> for example,
>>>>>>>> If one user rates 200 items and for each item, the time spent in
>>>>>>>>                 
>> calculating the neighbors is expected to 30ms.
>>     
>>>>>>>> So that makes 6 seconds. It is generally okay. But if the dataset is
>>>>>>>>                 
>> expanded to 100M dataset, I think 30ms may grow up to 30 * 100 ms and that
>> will be a long time.
>>     
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>                 
>>>>>>>>> It still seems strange to observe such a bottleneck, I'm not sure
>>>>>>>>> what's going on.
>>>>>>>>> You are using an in-memory model like GenericDataModel?
>>>>>>>>> We could look at ways to optimize that method, though it looks
>>>>>>>>>                   
>> reasonably tight.
>>     
>>>>>>>>> Where within that method do you see time spent?
>>>>>>>>>
>>>>>>>>> 2010/7/20 Young <[email protected]>:
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>                   
>>>>>>>>>> Hi again,
>>>>>>>>>> When I do the itembased recommendation, I find there are some
>>>>>>>>>>                     
>> latency in getAllOtherItems(long userID). Because it is calculating the
>> items' neighbors and merge these neighbors together. So I am thinking if I
>> precompute each item's neighbors and store in the database, then when I
>> getAllOtherItems(), I could merge these neighbors directly. Is this useful
>> for reducing the latency?
>>     
>>>>>>>>>> Or is there other way to make the online-recommendation much
>>>>>>>>>>                     
>> faster?
>>     
>>>>>>>>>> Thank you.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>                     
>>>>>>>>>>> Yes you probably want a new, separate table. You have an extra
>>>>>>>>>>>                       
>> step of
>>     
>>>>>>>>>>> computing some notion of similarity anyway, and you probably want
>>>>>>>>>>>                       
>> to
>>     
>>>>>>>>>>> separate this table from your main data table anyhow for reasons
>>>>>>>>>>>                       
>> of
>>     
>>>>>>>>>>> performance and business logic separation.
>>>>>>>>>>>
>>>>>>>>>>> 2010/7/19 Young <[email protected]>:
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>                       
>>>>>>>>>>>> So my prpblem is that I want to build datamodel based on what
>>>>>>>>>>>>                         
>> user has bought or added to their favorite or rated.
>>     
>>>>>>>>>>>> You mean I need a table describe all these user behavior. For
>>>>>>>>>>>>                         
>> example, if user buys one item, I guess the user preference is 4 and add
>> into this table?
>>     
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>                         
>>>>>>>>>>>>> No, you need one table (or view if you like) containing all
>>>>>>>>>>>>>                           
>> data. If
>>     
>>>>>>>>>>>>> you can't do this, you could write your own copy of a
>>>>>>>>>>>>>                           
>> JDBCDataModel
>>     
>>>>>>>>>>>>> that can query multiple tables, or, that changes its SQL
>>>>>>>>>>>>>                           
>> queries to
>>     
>>>>>>>>>>>>> use UNION statements. I imagine it will slow down a lot.
>>>>>>>>>>>>>
>>>>>>>>>>>>> If you mean, can you use a table with preferences with a model
>>>>>>>>>>>>>                           
>> that
>>     
>>>>>>>>>>>>> ignores preferences, sure you can. The extra column is ignored.
>>>>>>>>>>>>>
>>>>>>>>>>>>> 2010/7/19 Young <[email protected]>:
>>>>>>>>>>>>>
>>>>>>>>>>>>>
>>>>>>>>>>>>>
>>>>>>>>>>>>>                           
>>>>>>>>>>>>>> Hi,
>>>>>>>>>>>>>> I have three tables, one is with preference and another two
>>>>>>>>>>>>>>                             
>> are without preference. Does mahout have some algorithm to integret these
>> tables into one datamodel?
>>     
>>>>>>>>>>>>>> Thank you
>>>>>>>>>>>>>>
>>>>>>>>>>>>>>
>>>>>>>>>>>>>>
>>>>>>>>>>>>>>                             
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>                         
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>                     
>>>>>>>
>>>>>>>               
>>>>>           
>>
>>     
>   

Reply via email to