Hi Young,

I would disagree that a response time of 6 seconds is OK for online
recommendations, the time should be something like < 100ms.
I'm really surprised that you would see such response times with an
in-memory data model, I have experience with in-memory models of roughly
the same size
and usually the computations are blazingly fast.

Are you absolutely sure that the time is spent in this method and not
later in the similarity computation?

--sebastian

Am 21.07.2010 07:54, schrieb Young:
> So based on the 1M dataset, the time spent in getAllOtherItems(userID) is 
> among the 2 and 10 seconds. 
> for example,
> If one user rates 200 items and for each item, the time spent in calculating 
> the neighbors is expected to 30ms. 
> So that makes 6 seconds. It is generally okay. But if the dataset is expanded 
> to 100M dataset, I think 30ms may grow up to 30 * 100 ms and that will be a 
> long time. 
>
>
>
>
>   
>> It still seems strange to observe such a bottleneck, I'm not sure
>> what's going on.
>> You are using an in-memory model like GenericDataModel?
>> We could look at ways to optimize that method, though it looks reasonably 
>> tight.
>> Where within that method do you see time spent?
>>
>> 2010/7/20 Young <[email protected]>:
>>     
>>> Hi again,
>>> When I do the itembased recommendation, I find there are some latency in 
>>> getAllOtherItems(long userID). Because it is calculating the items' 
>>> neighbors and merge these neighbors together. So I am thinking if I 
>>> precompute each item's neighbors and store in the database, then when I 
>>> getAllOtherItems(), I could merge these neighbors directly. Is this useful 
>>> for reducing the latency?
>>> Or is there other way to make the online-recommendation much faster?
>>> Thank you.
>>>
>>>
>>>
>>>
>>>       
>>>> Yes you probably want a new, separate table. You have an extra step of
>>>> computing some notion of similarity anyway, and you probably want to
>>>> separate this table from your main data table anyhow for reasons of
>>>> performance and business logic separation.
>>>>
>>>> 2010/7/19 Young <[email protected]>:
>>>>         
>>>>> So my prpblem is that I want to build datamodel based on what user has 
>>>>> bought or added to their favorite or rated.
>>>>> You mean I need a table describe all these user behavior. For example, if 
>>>>> user buys one item, I guess the user preference is 4 and add into this 
>>>>> table?
>>>>>
>>>>>
>>>>>
>>>>>
>>>>>           
>>>>>> No, you need one table (or view if you like) containing all data. If
>>>>>> you can't do this, you could write your own copy of a JDBCDataModel
>>>>>> that can query multiple tables, or, that changes its SQL queries to
>>>>>> use UNION statements. I imagine it will slow down a lot.
>>>>>>
>>>>>> If you mean, can you use a table with preferences with a model that
>>>>>> ignores preferences, sure you can. The extra column is ignored.
>>>>>>
>>>>>> 2010/7/19 Young <[email protected]>:
>>>>>>             
>>>>>>> Hi,
>>>>>>> I have three tables, one is with preference and another two are without 
>>>>>>> preference. Does mahout have some algorithm to integret these tables 
>>>>>>> into one datamodel?
>>>>>>>
>>>>>>> Thank you
>>>>>>>               
>>>>>           
>>>       

Reply via email to