> > > Hi Danny, > > > Thanks very much for the references and suggestion. > > > Our correlator is 192 inputs, 125 MHz bandwidth with 1024 frequency > channel. > And I want to do some test about how fast the computer can capture packet > from the NIC. Now I have one NIC with dual 10GbE port made by Mellanox.
Hi Lin. A few things to watch for when trying this out: You have to use jumbo packets to get near to the line rate (~9.6 Gb/sec). You have to tune the kernel to increase the buffer sizes, and other parameters. See below for specifics for our systems and our Mellanox and solar flare NICs. Your best results may need different settings. You also have to set the CPU and IRQ affinities to keep the drivers and your applications from moving around the CPU. If you have a dual CPU machine, you have to make sure to set up the affinities to eliminate any NUMA problems of applications being in the "wrong" memory banks. John # NRAO additions to network settings net.ipv4.tcp_tw_recycle = 1 net.ipv4.tcp_fin_timeout = 10 net.core.rmem_max = 16777216 net.core.wmem_max = 16777216 net.ipv4.tcp_rmem = 4096 87380 16777216 net.ipv4.tcp_wmem = 4096 65536 16777216 #net.ipv4.tcp_sack = 0 net.ipv4.tcp_no_metrics_save = 1 net.core.netdev_max_backlog = 3000 # VEGAS/DIBAS shared memory kernel.sem = 1024 32000 32 32767 # Mellanox recommends the following net.ipv4.tcp_timestamps = 0 net.ipv4.tcp_sack = 0 net.core.netdev_max_backlog = 250000 net.core.rmem_default = 16777216 net.core.wmem_default = 16777216 net.core.optmem_max = 16777216 net.ipv4.tcp_mem = 16777216 16777216 16777216 net.ipv4.tcp_low_latency = 1 > > > Best Wishes, > Lin > > > > > > > > > > At 2015-04-01 23:50:22, "Danny Price" <[email protected]> wrote: >>Hi Shu Lean >> >>The short answer is: it depends, but you should probably buy a good Xeon >>CPU if you are not sure. How many antennas and what bandwidth will you >>have? >> >>The LEDA correlator (that I work with) uses xGPU for the X-engine, >>and PSRDADA for packet capture / shared memory buffers. Here's some >>papers for reference: >>http://psrdada.sourceforge.net/ >>http://www.worldscientific.com/doi/abs/10.1142/S2251171715500038?src=recsys >>http://www.worldscientific.com/doi/abs/10.1142/S2251171714500020 >> >>For reference, our server specifications are: >> >>Ubuntu 12.04.2 LTS (GNU/Linux 3.5.0-23-generic x86_64) >>Supermicro 1027TR-TQF >>2x Intel(R) Xeon(R) CPU E5-2670 0 @ 2.60GHz >>128 GB RAM >>2x NVIDIA K20X GPU >>1x Mellanox ConnectX-3 40GbE NIC >> >>Each server (11 total) is capturing 21.4 Gb/s, for 512 inputs, with 2.6 >>MHz bandwidth per server. >> >>The MWA correlator (http://arxiv.org/abs/1501.05992) and the PAPER >>correlator also use xGPU. >>I believe PAPER are using HASHPIPE instead of PSRDADA for their pipeline >>https://github.com/david-macmahon/hashpipe >> >>The CHIME project have recently posted some pretty darn impressive info >>on their correlator: >>http://arxiv.org/abs/1503.06189 >>http://arxiv.org/abs/1503.06203 >>http://arxiv.org/abs/1503.06202 >> >>CHIME are not using xGPU as they are using AMD GPUs, which don't support >>CUDA. I think LOFAR >>might be doing something with AMD GPUs also. From the CHIME papers, >>they seem to be getting amazing packet capture results by using NTOP >>PF_RING ZC (Zero Copy): >>http://www.ntop.org/products/pf_ring/pf_ring-zc-zero-copy/ >> >>Most NIC vendors now have "offload engines" that try and reduce load by >>bypassing the kernel. For example, Mellanox has libvma, and Solarflare >>has TCP OE. If you're about to build something big, I would highly >>recommend you purchase some cards from different vendors and do some lab >>tests. It may turn out that one card in particular has a killer feature. >>Please let us all know if you do! >> >>- Danny >> >>shu.lean wrote: >>> Hi all, >>> >>> if using GPU to do X Enigne, what the specification of computer should >>> be, like the performance of CPU, DDR memory? So the CPU can transfer >>> the data NIC recieved to GPU in time? >>> >>> Thanks very much. >>> >>> Lin >>> >>> >

