Yes, I am pretty sure. I have stored the similarity matrix in-memory and I print out the time spent in getAllOtherItems() and this is the only one time-consuming method in the recommendation. My laptop CPU is Intel P8600 2.4G, and the memory used for JVM is 1GB.
>Hi Young, > >I would disagree that a response time of 6 seconds is OK for online >recommendations, the time should be something like < 100ms. >I'm really surprised that you would see such response times with an >in-memory data model, I have experience with in-memory models of roughly >the same size >and usually the computations are blazingly fast. > >Are you absolutely sure that the time is spent in this method and not >later in the similarity computation? > >--sebastian > >Am 21.07.2010 07:54, schrieb Young: >> So based on the 1M dataset, the time spent in getAllOtherItems(userID) is >> among the 2 and 10 seconds. >> for example, >> If one user rates 200 items and for each item, the time spent in calculating >> the neighbors is expected to 30ms. >> So that makes 6 seconds. It is generally okay. But if the dataset is >> expanded to 100M dataset, I think 30ms may grow up to 30 * 100 ms and that >> will be a long time. >> >> >> >> >> >>> It still seems strange to observe such a bottleneck, I'm not sure >>> what's going on. >>> You are using an in-memory model like GenericDataModel? >>> We could look at ways to optimize that method, though it looks reasonably >>> tight. >>> Where within that method do you see time spent? >>> >>> 2010/7/20 Young <[email protected]>: >>> >>>> Hi again, >>>> When I do the itembased recommendation, I find there are some latency in >>>> getAllOtherItems(long userID). Because it is calculating the items' >>>> neighbors and merge these neighbors together. So I am thinking if I >>>> precompute each item's neighbors and store in the database, then when I >>>> getAllOtherItems(), I could merge these neighbors directly. Is this useful >>>> for reducing the latency? >>>> Or is there other way to make the online-recommendation much faster? >>>> Thank you. >>>> >>>> >>>> >>>> >>>> >>>>> Yes you probably want a new, separate table. You have an extra step of >>>>> computing some notion of similarity anyway, and you probably want to >>>>> separate this table from your main data table anyhow for reasons of >>>>> performance and business logic separation. >>>>> >>>>> 2010/7/19 Young <[email protected]>: >>>>> >>>>>> So my prpblem is that I want to build datamodel based on what user has >>>>>> bought or added to their favorite or rated. >>>>>> You mean I need a table describe all these user behavior. For example, >>>>>> if user buys one item, I guess the user preference is 4 and add into >>>>>> this table? >>>>>> >>>>>> >>>>>> >>>>>> >>>>>> >>>>>>> No, you need one table (or view if you like) containing all data. If >>>>>>> you can't do this, you could write your own copy of a JDBCDataModel >>>>>>> that can query multiple tables, or, that changes its SQL queries to >>>>>>> use UNION statements. I imagine it will slow down a lot. >>>>>>> >>>>>>> If you mean, can you use a table with preferences with a model that >>>>>>> ignores preferences, sure you can. The extra column is ignored. >>>>>>> >>>>>>> 2010/7/19 Young <[email protected]>: >>>>>>> >>>>>>>> Hi, >>>>>>>> I have three tables, one is with preference and another two are >>>>>>>> without preference. Does mahout have some algorithm to integret these >>>>>>>> tables into one datamodel? >>>>>>>> >>>>>>>> Thank you >>>>>>>> >>>>>> >>>> >
