Is only the similarity matrix in-memory? The crucial thing here is the data model not the similarity matrix, are you using an in-memory data model?
Am 21.07.2010 08:33, schrieb Young: > Yes, I am pretty sure. I have stored the similarity matrix in-memory and I > print out the time spent in getAllOtherItems() and this is the only one > time-consuming method in the recommendation. My laptop CPU is Intel P8600 > 2.4G, and the memory used for JVM is 1GB. > > > > > > >> Hi Young, >> >> I would disagree that a response time of 6 seconds is OK for online >> recommendations, the time should be something like < 100ms. >> I'm really surprised that you would see such response times with an >> in-memory data model, I have experience with in-memory models of roughly >> the same size >> and usually the computations are blazingly fast. >> >> Are you absolutely sure that the time is spent in this method and not >> later in the similarity computation? >> >> --sebastian >> >> Am 21.07.2010 07:54, schrieb Young: >> >>> So based on the 1M dataset, the time spent in getAllOtherItems(userID) is >>> among the 2 and 10 seconds. >>> for example, >>> If one user rates 200 items and for each item, the time spent in >>> calculating the neighbors is expected to 30ms. >>> So that makes 6 seconds. It is generally okay. But if the dataset is >>> expanded to 100M dataset, I think 30ms may grow up to 30 * 100 ms and that >>> will be a long time. >>> >>> >>> >>> >>> >>> >>>> It still seems strange to observe such a bottleneck, I'm not sure >>>> what's going on. >>>> You are using an in-memory model like GenericDataModel? >>>> We could look at ways to optimize that method, though it looks reasonably >>>> tight. >>>> Where within that method do you see time spent? >>>> >>>> 2010/7/20 Young <[email protected]>: >>>> >>>> >>>>> Hi again, >>>>> When I do the itembased recommendation, I find there are some latency in >>>>> getAllOtherItems(long userID). Because it is calculating the items' >>>>> neighbors and merge these neighbors together. So I am thinking if I >>>>> precompute each item's neighbors and store in the database, then when I >>>>> getAllOtherItems(), I could merge these neighbors directly. Is this >>>>> useful for reducing the latency? >>>>> Or is there other way to make the online-recommendation much faster? >>>>> Thank you. >>>>> >>>>> >>>>> >>>>> >>>>> >>>>> >>>>>> Yes you probably want a new, separate table. You have an extra step of >>>>>> computing some notion of similarity anyway, and you probably want to >>>>>> separate this table from your main data table anyhow for reasons of >>>>>> performance and business logic separation. >>>>>> >>>>>> 2010/7/19 Young <[email protected]>: >>>>>> >>>>>> >>>>>>> So my prpblem is that I want to build datamodel based on what user has >>>>>>> bought or added to their favorite or rated. >>>>>>> You mean I need a table describe all these user behavior. For example, >>>>>>> if user buys one item, I guess the user preference is 4 and add into >>>>>>> this table? >>>>>>> >>>>>>> >>>>>>> >>>>>>> >>>>>>> >>>>>>> >>>>>>>> No, you need one table (or view if you like) containing all data. If >>>>>>>> you can't do this, you could write your own copy of a JDBCDataModel >>>>>>>> that can query multiple tables, or, that changes its SQL queries to >>>>>>>> use UNION statements. I imagine it will slow down a lot. >>>>>>>> >>>>>>>> If you mean, can you use a table with preferences with a model that >>>>>>>> ignores preferences, sure you can. The extra column is ignored. >>>>>>>> >>>>>>>> 2010/7/19 Young <[email protected]>: >>>>>>>> >>>>>>>> >>>>>>>>> Hi, >>>>>>>>> I have three tables, one is with preference and another two are >>>>>>>>> without preference. Does mahout have some algorithm to integret these >>>>>>>>> tables into one datamodel? >>>>>>>>> >>>>>>>>> Thank you >>>>>>>>> >>>>>>>>> >>>>>>> >>>>>>> >>>>> >>>>> >>
