Hello, From my experience, the finite bias calculation in the open-boundary part is much slower (takes in general more than an order of magnitude than the zero-bias case for my systems). I am very curious to know what is the reason behind this.
I looked at the code (I am using trunk-433), and I thought about two possibilities. (1) when IsVolt=T, I have four more matrices than zero bias to deal with in the code: DMRCplx, DMneqLCplx, DMneqRCplx, and EDMRCplx, provided that I have more than one k point. So I am wondering if this is one of the reasons for longer computing time; (2) I notice all the energy grids (equilibrium + non-equilibrium) used for contour integration are looped on an equal footing in every SCF cycle, but the non-equilibrium energy grids are on real axis with a very small imaginary part. Does the evaluation of quantities on real axis cause the long computing time? The Green's functions are calculated at the beginning of the run, and I don't see longer computing time there, so presumably the getSFE (self-energy calculation) is more expensive for an energy grid on the real axis than those in the upper half plane? If so, I would like to learn why is the case. I also noticed that in the m_ts_options.F90 file, there is a comment saying that "the voltage contour point is more "heavy" in computation". I searched the mailing list, and there are suggestions of using a slightly larger imaginary part for the non-equilibrium energy grids, but I should not use a too large value for the imaginary part, right? Finally I have a technical question. I know 36 energy grid is the default for zero bias calculation. So if I have finite bias, and 10 bias contour energy grids, that would give me 36*2+10=82 energy grids. Should I use an optimum of 82 cores? But if the 10 bias contour grids take more time to compute than the 72 equilibrium energy grids, do I actually "waste" the 72 cores when they are waiting for the non-equilibrium energy grids to finish? Thank you very much for your time and effort! Zhenfei Liu
