Hi, having some wierd problems with random network errors. Cut over to 2.0.10 Monday morning and things appeared to go well, but then the random glitches started hitting. Would occasionally get one of these on 1.6.0.8 but rare enough folks just worked around it, now it is really hampering getting things done. Googled on the messages and found some help in the mailing list archives for a problem that seems to match ours in that the errors were the same and only seemed to happen in virtual machines. So I manually applied these patches to OpenSRF:
719b5344d5cc5a85bb506af0432ef0c563800fff ab63faa85e349020774e0b3297a84370281d573f Here is the link to the thread that seemed to apply: http://list.georgialibraries.org/pipermail/open-ils-general/2011-October/005639.html Today things are better but we are still getting hit pretty hard. We have massive overkill in our hardware provisioning, system load never hits 1.0, 0.30 is more typical, even when a transaction that is about to fail is in progress. Have two physical hosts with a six core Opteron and 16GB of RAM in each. Each is currently inhabited by a single libvirt/KVM based VM, one the front end and one running Postgresql. Each VM has more than enough ram to prevent swapping and the DB should be essentially running from ram, and anyway the discs are SSDs. There is enough computing grunt here to easily handle the work. The real and virtual OS on all is Debian 6.0 with the 2.6.32 kernel backported to get needed hardware support in the real machines. So has anyone else seem this problem and what else can I try?
signature.asc
Description: This is a digitally signed message part
