On Fri, Dec 14, 2012 at 7:24 AM, Konrad Hinsen <[email protected]> wrote: > Raul Miller writes: > > > Even then your dominant cost is reading the data off of disk. If the > > data is converted on the way in from the disk to RAM, and most of your > > CPU time is spent waiting for the disk, your conversion time might not > > even be measurable. > > My situation would be that the data is read to memory, ending up in a > NumPy array, and then I have to copy it to a J array in a subsequent > step. That's why I am interested in exploring the shared data approach.
Yes... I was thinking that this aspect might be redesigned. > > Note that if your data structure occupies half of your physical > > memory, J intermediate results might be expensive -- my rule of thumb > > (just a guess based on raw data size, until I get real measurements) > > is to expect 5x memory overhead from J. > > That's also a problem with NumPy, but there I know how to get around > it. With J, that remains to be explored. Typically, I get my algorithms working with smaller data sets and then introduce some kind of mechanism to break the content up in blocks. Initially, I start with very small blocks (like 3 elements or something) to make sure that I have my edge conditions set up properly, and then I bump the blocks up to something useful (like something in the 10s of megabytes). Thanks, -- Raul ---------------------------------------------------------------------- For information about J forums see http://www.jsoftware.com/forums.htm
