Dear Nick, Thank you very much for your time to answer questions specific to my case. Re (1), I agree with the 3*T_eq + 2*(time on GFGGF), which is very clear. I guess in general it should be 2*T_eq + 2*(time on GFGGF) + T_neq, where the relation between T_neq and T_eq depends on the relation between N_neq and N_eq?
Re (2), I should have said that for my zero-bias calculations, TS_comm indeed is <1-2% for all my systems and the calculation is very fast, and it only increases to 1/3 when I turn on finite-bias calculation - I will definitely check with my cluster administrator, but do you happen to have any idea from top of your head why this is so only for finite-bias? And regarding getSFE, no, I don't see getSFE percentage changes by that significant amount. The node on our cluster is job-dedicated, so no other heavy jobs can occupy the same node. The reason why I cannot fully occupy a node is because I have huge systems (8000-10000 orbitals per unit cell), and I usually use all cores on the nodes for siesta part and a lot of nodes, but they will die for the transiesta part due to out-of-memory killer (I tested my case and I need ~8GB/core to run through even in zero-bias). Then I will use partially filled (usually ~<50%) nodes to continue my calculation to give more memory per core in order to have the transiesta part run. Another reason why I avoid use same number of cores for transiesta and siesta part is that I use many cores (say ~128) in siesta run and I never have that many of energy grids in transiesta so some cores are wasted. The reason why I use same number of cores as number of energy grids is because I don't want a core to do two energy grids (N_neq or N_eq/cores > 1) since I naively guess that the computing time would double? But you are right that by doing that I would save communication time between cores, which I will definitely try. If you have any additional comments, I would appreciate that :) Best regards, Zhenfei Liu On Monday, June 23, 2014 1:38 PM, Nick Papior Andersen <[email protected]> wrote: 2014-06-23 18:44 GMT+00:00 Bruin Liu <[email protected]>: Dear Nick, > > >Thank you very much for this extremely clear and useful answers! It helped me >a lot to understand the code. I have two more questions/comments: > > >(1) From your answers, I understand that for a _single_ energy grid, whether >or not it's an equilibrium one (in the upper half plane) or it's a >non-equilibrium one (on the real axis with a small imag part) does NOT >actually affect the computing time. The reason why non-eq takes more time is >mainly because of GFGammaGF subroutine (Eq. 22 in Brandbyge paper) for each >SCF cycle. I also understand that Eqs. 30/33 in Brandbyge paper is a single >sum (loop over energy grid in the code), so for each SCF cycle, it does not >require much time, but due to convergence issues explained in the paper, >weighting (Eq. 36) and more SCF cycles than zero-bias calculation are needed. >So the real axis integral itself only "indirectly" affects _total_ computing >time but not the computing time for a single SCF cycle. Please let me know if >some of the words do not make sense; Well, Eq. 30 and 33 are individual sums, not a single sum (hence two of the additional arrays are from these), so for each real axis energy point it has two GFGammaGF calls. A product for left and right scattering matrix. I would not say that the real axis integral "indirectly" affects total computation time, it really does so directly, or maybe I do not follow what you are trying to say :). Say you use 30 energy points in the equilibrium contour, and 30 on the real axis, then you have two cases: 1) Voltage == 0 You only have 30 energy points. And you only have two matrices. Here the speed is directly related to the Gf calculation. This takes T_eq. 2) Voltage /= 0 You have the equivalent of 60 equilibrium points, and then 30 real axis points. So here you have a timing that is 3 * T_eq + 2 * time spend on Gf.G.Gf. And you have 6 matrices. It should be clear from the Brandbyge paper why this is so. When I consider performance I only regard a single SCF cycle, how many SCF cycles are dependent on the system. But yes, high bias requires many real-axis points and also has slower convergence. > >(2) I compared some of my .times files for zero-bias and finite-bias >calculations. I found GFT+getSFE, GFGGF, and TS_comm each consists of roughly >1/3 of TS_calc, respectively. For different systems, the percentage varies, >but not by too much. It surprised me that TS_comm costs that much. I checked >the code, and found that between the two call timer("TS_comm") are only the >MPI_AllReduce for the density matrices. Since for DMs, the difference between >the zero-bias and finite-bias calculations are the four additional DM >matrices. Can I say that although the computing of the additional DM does not >cost much, the distribution and collection of data does? FYI, this calculation >has 21 eq. energy grids for each leads, and 21 non-eq. energy grids. I used 21 >cores (3 nodes, 7 cores/node). Do you think increasing memory/core for this >calculation would decrease TS_comm for this particular calculation? Well, this is something that you need to check with your cluster administrator. Generally TS_comm should be a fraction (<1-2%) of the total calculation time, else you probably have something that can be tuned with your hardware and/or software stack. TS_comm relates directly to how the MPI processors communicate, if you have infiniband it should be extremely fast, if you rely on some old ethernet connection it is REALLY slow. ;) In fact getSFE is also dependent on communication, so if TS_comm is bad, then getSFE is probably also. It also depends on your system size (number of orbitals), for very large systems, say around 5000+ orbitals, this might increase, but I would be worried if it took up 1/3 of the time, it should definitely not increase above 5% (from my experience)! I would try and cut down on nodes to only 1 node. It seems like you have bad interconnect between nodes, also if you only occupy 21 out of 24 you may have problems occupying a node with some other heavy jobs (but this is just a guess, I don't know how your cluster works). Let me be more clear on how you should split your nodes: You should not enforce splitting cores up with number of grid points. In particular if you have a small system then having many nodes is not that smart, the siesta part will be very slow. Decide how many processors you use for siesta and then use the same for transiesta. Only change number of cores if you think the transiesta part is extremely slow and you have enough points to populate the extra processors. My before mentioning of "best performance" does not mean that it is much faster if you enforce the exact criteria. But I suggest you to test this with your system. As was also mentioned, you should have: N_neq / cores = integer, where integer does not necessarily mean 1. :) And no, increasing memory per core does not help you. Only if you are using swap space instead of memory, then indeed yes you would benefit heavily from having more memory per core. If you are worried of this, test transiesta with a smaller system to see if your timings drastically improve. > >Again I appreciate your kind help very much! > > >Best regards, > > > > >Zhenfei Liu > > > >On Monday, June 23, 2014 12:52 AM, Nick Papior Andersen <[email protected]> >wrote: > > > >Dear Zhenfei > > > > >2014-06-19 20:05 GMT+02:00 Bruin Liu <[email protected]>: > >Hello, >> >>>From my experience, the finite bias calculation in the open-boundary part is >>>much slower (takes in general more than an order of magnitude than the >>>zero-bias case for my systems). I am very curious to know what is the reason >>>behind this. >> >>I looked at the code (I am using trunk-433), and I thought about two >>possibilities. (1) when IsVolt=T, I have four more matrices than zero >>bias to deal with in the code: DMRCplx, DMneqLCplx, DMneqRCplx, and EDMRCplx, >>provided that I have more than one k point. So I am wondering if this is one >>of the reasons for longer computing time; (2) I notice all the energy grids >>(equilibrium + non-equilibrium) used for contour integration are looped on an >>equal footing in every SCF cycle, but the non-equilibrium energy grids are on >>real axis with a very small imaginary part. Does the evaluation of quantities >>on real axis cause the long computing time? The Green's functions are >>calculated at the beginning of the run, and I don't see longer computing time >>there, so presumably the getSFE (self-energy calculation) is more expensive >>for an energy grid on the real axis than those in the upper half plane? If >>so, I would like to learn why is the case. I also noticed that in the >>m_ts_options.F90 file, there is a comment saying that "the voltage contour >>point is more "heavy" in computation". I searched the mailing list, and there are suggestions of using a slightly larger imaginary part for the non-equilibrium energy grids, but I should not use a too large value for the imaginary part, right? >> >You are correct in many of the things. >1) There are more matrices for non-equilibrium using bias calculations, >however, the additional matrices does not affect the computation time by a >huge amount, the loops traversing the matrices are rather fast. Yet they >contribute to additional computations + the weight calculation of the L/R >contribution, see eq. 37 in the Brandbyge paper. >2) The equilbrium contour is calculated using equation 24 and 25 in the paper. >I.e. the equilibrium contour only involves calculating the Green's function. >The non-equilibrium contour is calculated using eq. 22 which means that a >triple matrix product is necessary. Basically: >Eq) only calculate the Green's function >Non-eq) calculate the Green's function, then calculate G.Gamma.G^\dagger >The getSFE takes the same time in equilibrium and non-equilibrium, for now, >remark that they are only calculated once, and then read in. :) >The real grid is "only" more expensive in the sense that it is required to >have denser integral grids as the triple product is not a smooth function (the >Green's function is smooth far in the complex plane) and an additional triple >product is necessary. >The larger imaginary part smooths outs the DOS along the real axis which can >help improve convergence at the possible cost of precision. > >>Finally I have a technical question. I know 36 energy grid is the default for >>zero bias calculation. So if I have finite bias, and 10 >>bias contour energy grids, that would give me 36*2+10=82 energy grids. Should >>I use an optimum of 82 cores? But if the 10 bias contour grids take more time >>to compute than the 72 equilibrium energy grids, do I actually "waste" the 72 >>cores when they are waiting for the >>non-equilibrium energy grids to finish? >> >No, for optimal performance of transiesta you do this, N = number of cores. >Neq = total equilibrium contour points, Nneq = real grid points. > > >Neq / N = integer >Nneq / N = integer > > >You do not have to have as many cores as energy points. >As you mention last, yes, the 72 cores wait for the 10 cores to finish. > >>Thank you very much for your time and effort! >> >> >I hope this answered some/all of your questions. :) > >> >>Zhenfei Liu >> >> >> > > >-- > >Kind regards Nick > > -- Kind regards Nick
