2014-06-23 18:44 GMT+00:00 Bruin Liu <[email protected]>:

> Dear Nick,
>
> Thank you very much for this extremely clear and useful answers! It helped
> me a lot to understand the code. I have two more questions/comments:
>
> (1) From your answers, I understand that for a _single_ energy grid,
> whether or not it's an equilibrium one (in the upper half plane) or it's a
> non-equilibrium one (on the real axis with a small imag part) does NOT
> actually affect the computing time. The reason why non-eq takes more time
> is mainly because of GFGammaGF subroutine (Eq. 22 in Brandbyge paper) for
> each SCF cycle. I also understand that Eqs. 30/33 in Brandbyge paper is a
> single sum (loop over energy grid in the code), so for each SCF cycle, it
> does not require much time, but due to convergence issues explained in the
> paper, weighting (Eq. 36) and more SCF cycles than zero-bias calculation
> are needed. So the real axis integral itself only "indirectly" affects
> _total_ computing time but not the computing time for a single SCF cycle.
> Please let me know if some of the words do not make sense;
>
Well, Eq. 30 and 33 are individual sums, not a single sum (hence two of the
additional arrays are from these), so for each real axis energy point it
has two GFGammaGF calls. A product for left and right scattering matrix. I
would not say that the real axis integral "indirectly" affects total
computation time, it really does so directly, or maybe I do not follow what
you are trying to say :). Say you use 30 energy points in the equilibrium
contour, and 30 on the real axis, then you have two cases:
1) Voltage == 0
You only have 30 energy points. And you only have two matrices. Here the
speed is directly related to the Gf calculation.
This takes T_eq.
2) Voltage /= 0
You have the equivalent of 60 equilibrium points, and then 30 real axis
points. So here you have a timing that is 3 * T_eq + 2 * time spend on
Gf.G.Gf.
And you have 6 matrices. It should be clear from the Brandbyge paper why
this is so.
When I consider performance I only regard a single SCF cycle, how many SCF
cycles are dependent on the system. But yes, high bias requires many
real-axis points and also has slower convergence.

>
> (2) I compared some of my .times files for zero-bias and finite-bias
> calculations. I found GFT+getSFE, GFGGF, and TS_comm each consists of
> roughly 1/3 of TS_calc, respectively. For different systems, the percentage
> varies, but not by too much. It surprised me that TS_comm costs that much.
> I checked the code, and found that between the two call timer("TS_comm")
> are only the MPI_AllReduce for the density matrices. Since for DMs, the
> difference between the zero-bias and finite-bias calculations are the four
> additional DM matrices. Can I say that although the computing of the
> additional DM does not cost much, the distribution and collection of data
> does? FYI, this calculation has 21 eq. energy grids for each leads, and 21
> non-eq. energy grids. I used 21 cores (3 nodes, 7 cores/node). Do you think
> increasing memory/core for this calculation would decrease TS_comm for this
> particular calculation?
>
Well, this is something that you need to check with your cluster
administrator. Generally TS_comm should be a fraction (<1-2%) of the total
calculation time, else you probably have something that can be tuned with
your hardware and/or software stack. TS_comm relates directly to how the
MPI processors communicate, if you have infiniband it should be extremely
fast, if you rely on some old ethernet connection it is REALLY slow. ;)
In fact getSFE is also dependent on communication, so if TS_comm is bad,
then getSFE is probably also.
It also depends on your system size (number of orbitals), for very large
systems, say around 5000+ orbitals, this might increase, but I would be
worried if it took up 1/3 of the time, it should definitely not increase
above 5% (from my experience)! I would try and cut down on nodes to only 1
node. It seems like you have bad interconnect between nodes, also if you
only occupy 21 out of 24 you may have problems occupying a node with some
other heavy jobs (but this is just a guess, I don't know how your cluster
works).
Let me be more clear on how you should split your nodes:
You should not enforce splitting cores up with number of grid points. In
particular if you have a small system then having many nodes is not that
smart, the siesta part will be very slow.
Decide how many processors you use for siesta and then use the same for
transiesta. Only change number of cores if you think the transiesta part is
extremely slow and you have enough points to populate the extra processors.
My before mentioning of "best performance" does not mean that it is much
faster if you enforce the exact criteria. But I suggest you to test this
with your system.
As was also mentioned, you should have:
N_neq / cores = integer, where integer does not necessarily mean 1. :)
And no, increasing memory per core does not help you.
Only if you are using swap space instead of memory, then indeed yes you
would benefit heavily from having more memory per core.
If you are worried of this, test transiesta with a smaller system to see if
your timings drastically improve.

>
> Again I appreciate your kind help very much!
>
> Best regards,
>
>
> Zhenfei Liu
>
>
>   On Monday, June 23, 2014 12:52 AM, Nick Papior Andersen <
> [email protected]> wrote:
>
>
>  Dear Zhenfei
>
>
> 2014-06-19 20:05 GMT+02:00 Bruin Liu <[email protected]>:
>
> Hello,
>
> From my experience, the finite bias calculation in the open-boundary part
> is much slower (takes in general more than an order of magnitude than the
> zero-bias case for my systems). I am very curious to know what is the
> reason behind this.
>
> I looked at the code (I am using trunk-433), and I thought about two 
> possibilities.
> (1) when IsVolt=T, I have four more matrices than zero
> bias to deal with in the code: DMRCplx, DMneqLCplx, DMneqRCplx, and EDMRCplx,
> provided that I have more than one k point. So I am wondering if this is
> one of the reasons for longer computing time; (2) I notice all the energy
> grids (equilibrium + non-equilibrium) used for contour integration are
> looped on an equal footing in every SCF cycle, but the non-equilibrium
> energy grids are on real axis with a very small imaginary part. Does the
> evaluation of quantities on real axis cause the long computing time? The
> Green's functions are calculated at the beginning of the run, and I don't
> see longer computing time there, so presumably the getSFE (self-energy
> calculation) is more expensive for an energy grid on the real axis than
> those in the upper half plane? If so, I would like to learn why is the
> case. I also noticed that in the m_ts_options.F90 file, there is a
> comment saying that "the voltage contour point is more "heavy" in
> computation". I searched the mailing list, and there are suggestions of
> using a slightly larger imaginary part for the non-equilibrium energy
> grids, but I should not use a too large value for the imaginary part,
> right?
>
> You are correct in many of the things.
> 1) There are more matrices for non-equilibrium using bias calculations,
> however, the additional matrices does not affect the computation time by a
> huge amount, the loops traversing the matrices are rather fast. Yet they
> contribute to additional computations + the weight calculation of the L/R
> contribution, see eq. 37 in the Brandbyge paper.
> 2) The equilbrium contour is calculated using equation 24 and 25 in the
> paper. I.e. the equilibrium contour only involves calculating the Green's
> function. The non-equilibrium contour is calculated using eq. 22 which
> means that a triple matrix product is necessary. Basically:
> Eq) only calculate the Green's function
> Non-eq) calculate the Green's function, then calculate G.Gamma.G^\dagger
> The getSFE takes the same time in equilibrium and non-equilibrium, for
> now, remark that they are only calculated once, and then read in. :)
> The real grid is "only" more expensive in the sense that it is required to
> have denser integral grids as the triple product is not a smooth function
> (the Green's function is smooth far in the complex plane) and an additional
> triple product is necessary.
> The larger imaginary part smooths outs the DOS along the real axis which
> can help improve convergence at the possible cost of precision.
>
>
> Finally I have a technical question. I know 36 energy grid is the default
> for zero bias calculation. So if I have finite bias, and 10
> bias contour energy grids, that would give me 36*2+10=82 energy grids. Should
> I use an optimum of 82 cores? But if the 10 bias contour grids take more
> time to compute than the 72 equilibrium energy grids, do I actually
> "waste" the 72 cores when they are waiting for the
> non-equilibrium energy grids to finish?
>
> No, for optimal performance of transiesta you do this, N = number of
> cores. Neq = total equilibrium contour points, Nneq = real grid points.
>
> Neq / N = integer
> Nneq / N = integer
>
> You do not have to have as many cores as energy points.
> As you mention last, yes, the 72 cores wait for the 10 cores to finish.
>
>
> Thank you very much for your time and effort!
>
> I hope this answered some/all of your questions. :)
>
>
>
> Zhenfei Liu
>
>
>
>
>
> --
> Kind regards Nick
>
>
>


-- 
Kind regards Nick

Responder a