diff --git a/docs/advanced/input_files/input-main.md b/docs/advanced/input_files/input-main.md index 4e5456899d3..c72351a49f3 100644 --- a/docs/advanced/input_files/input-main.md +++ b/docs/advanced/input_files/input-main.md @@ -574,6 +574,11 @@ - [nocc](#nocc) - [nvirt](#nvirt) - [lr\_nstates](#lr_nstates) + - [lr\_target\_state](#lr_target_state) + - [lr\_degen\_thr](#lr_degen_thr) + - [lr\_degen\_mode](#lr_degen_mode) + - [lr\_grad\_solver](#lr_grad_solver) + - [lr\_target\_spin](#lr_target_spin) - [lr\_unrestricted](#lr_unrestricted) - [abs\_wavelen\_range](#abs_wavelen_range) - [out\_wfc\_lr](#out_wfc_lr) @@ -679,6 +684,8 @@ - ks-lr: Kohn-Sham density functional theory + LR-TDDFT (Under Development Feature) - lr: LR-TDDFT with given KS orbitals (Under Development Feature) - dfpt: density functional perturbation theory (Under Development Feature) + + > Note: Excited-state forces (`cal_force = 1`) and atomic relaxation with `ks-lr` or `lr` currently require `gamma_only = 1` and pseudopotentials without nonlinear core correction (NLCC). NLCC core-density response-force and XC-kernel derivative terms are not implemented; these force requests are rejected after reading the pseudopotentials, before force evaluation. Multi-k spectra and spectra with NLCC pseudopotentials remain available with `cal_force = 0`. Ground-state forces are unaffected by this LR restriction. - **Default**: ksdft ### symmetry @@ -5131,7 +5138,7 @@ ### xc_kernel - **Type**: String -- **Description**: The exchange-correlation kernel used in the calculation. Currently supported: RPA, LDA, PBE, HSE, HF. +- **Description**: The exchange-correlation kernel used in the calculation. Currently supported: RPA, LDA, PWLDA, PBE, and the hybrids HF, PBE0, HSE, B3LYP, CAM_PBEH, LC_PBE, LC_WPBE, LRC_WPBE, LRC_WPBEH. A hybrid kernel needs the ground state to use the same functional: the exact-exchange operator $[\alpha+\beta\,\mathrm{erfc}(\mu r)]/r$ is built from exx_fock_alpha ($\alpha$), exx_erfc_alpha ($\beta$) and exx_erfc_omega ($\mu$), which are keyed off dft_functional, not off this parameter. - **Default**: LDA ### lr_init_xc_kernel @@ -5161,9 +5168,10 @@ ### nocc - **Type**: Integer -- **Description**: The number of occupied orbitals (up to HOMO) used in the LR-TDDFT calculation. - - Note: If the value is illegal ( > nelec/2 or <= 0), it will be autoset to nelec/2. -- **Default**: nband +- **Description**: The number of occupied orbitals (up to HOMO) retained in the majority-spin LR-TDDFT window. A positive value selects a shared core prefix to discard from both spin channels; it does not change the ground-state occupations. + - If omitted, non-positive, or larger than the occupied majority-spin channel, all occupied orbitals are used. + - The full occupied window is determined by the effective electron number (including nelec_delta once) and the ground-state spin populations. For nspin=2, the minority-spin window has abs(N_up-N_down) fewer occupied orbitals. +- **Default**: all occupied orbitals ### nvirt @@ -5177,6 +5185,73 @@ - **Description**: The number of 2-particle states to be solved. - **Default**: 0 +### lr_target_state + +- **Type**: Integer +- **Description**: Initial excited-state index for `calculation = relax`, counted from 0 within the spin channel selected by `lr_target_spin`. + + The gradient of the followed state (or its degenerate multiplet selected by `lr_degen_mode`) is computed, since solving the Z-vector equation dominates the cost of an excited-state gradient. It also selects the state whose excitation energy is added to the ground-state total energy, which is the quantity the energy-based relaxation algorithms (`cg`, `bfgs`, `lbfgs`) line-search on. + + Ignored outside `calculation = relax`: a single-point run solves and reports the gradients of every state. + + > Note: This index seeds the first ionic step. Subsequent steps compute cross-geometry AO overlaps and transform them with the saved and current KS orbitals to compare excitation amplitudes in a common occupied/virtual basis. Orbital sign changes and rotations within those subspaces do not change the overlap criterion. Index changes and low overlaps are reported. In JT mode the selected normalized multiplet mixture and its orbital basis become the reference. The projection is not renormalized, so window leakage remains visible. Tracking currently requires a fixed cell and unchanged orbital windows; degeneracy and a state leaving the solved window still limit state identification. Increase `lr_nstates` or reduce the ionic step when the overlap is low. +- **Default**: 0 + +### lr_degen_thr + +- **Type**: Real +- **Description**: Excited states whose excitation energies lie within this threshold of each other are treated as one degenerate multiplet, and the full gradient matrix $G^{(A\alpha)}_{kl}=\langle X_k|\partial A/\partial R_{A\alpha}|X_l\rangle$ is computed for it in addition to the per-state gradients. Zero (the default) disables this and leaves the per-state gradients as the only output. + + At a $d$-fold degeneracy no single state has a gradient vector: the branch slopes along a displacement $u$ are the eigenvalues of $\sum_{A\alpha}u_{A\alpha}G^{(A\alpha)}$, and the eigenvectors that diagonalise it depend on $u$. The per-state gradients are the diagonal of $G$ in whichever basis the eigensolver happened to return, so only their sum (the trace) is basis-independent, while $G$ itself is the complete first-order information -- it is the linear vibronic coupling Hamiltonian of the multiplet. The extra cost is $d(d-1)/2$ further Z-vector solves per multiplet. + + The threshold proposes candidates; it cannot tell a true degeneracy from an accidental near-degeneracy, where the states have genuinely different excitation energies and the construction does not apply. Each multiplet's actual energy spread and the orthonormality of its eigenvectors are reported in the running log so the distinction can be made there. + + > Note: A sensible value is a few times the eigensolver threshold `lr_thr`, so that states split by real physics are not merged. +- **Default**: 0 +- **Unit**: Ry + +### lr_degen_mode + +- **Type**: String +- **Description**: What `calculation = relax` follows when `lr_target_state` sits inside a degenerate multiplet, as identified by `lr_degen_thr`. It has no effect when the target state is non-degenerate. + + - state: follow the gradient of that one state, as returned by the eigensolver. This is the historical behaviour and is what reproduces earlier results, but inside a multiplet it is not a well-defined quantity: the per-state gradients are the diagonal of the subspace gradient matrix in whichever basis the eigensolver happened to return, so they depend on numerical details of the diagonalisation rather than on physics. + - average: follow the multiplet average $\bar\Omega=\frac{1}{d}\sum_k\Omega_k$, whose gradient is $\operatorname{Tr}G/d$. Unlike the individual states this is a smooth, basis-independent surface, and by symmetry its gradient is totally symmetric, so following it keeps the geometry on the symmetric configuration. Both the reported energy and the reported gradient switch to the average together, which the energy-based optimisers (`cg`, `bfgs`, `lbfgs`) require -- a gradient of one surface line-searched against the energy of another does not converge. This mode deliberately does NOT find the Jahn-Teller distortion, which is orthogonal to the totally symmetric average gradient. + - jt: descend the Jahn-Teller branch. Solves $\min_{\|u\|=1}\lambda_{\min}(\sum_{A\alpha}u_{A\alpha}G^{(A\alpha)})$ -- a joint optimisation over the displacement and the mixing inside the multiplet, since the two are determined together -- and follows the force of the resulting branch. This needs the off-diagonal part of the gradient matrix, so it costs $d(d-1)/2$ further Z-vector solves per step on top of the $d$ diagonal ones. The running log reports the branch's force, its mixing coefficients, and its split into the part common to the multiplet and the part that actually breaks the degeneracy. + + > Note: The usual sequence is `average` first, to reach the symmetric stationary point, then `jt` from there: at a stationary point of the average surface the common part vanishes and the whole force is Jahn-Teller. `jt` is self-limiting -- once a step has split the multiplet there is no group left and the ordinary single-state gradient takes over. + + > Note: `jt` gives the first-order DIRECTION. The distortion amplitude also needs the harmonic term, and the step norm is Cartesian rather than mass-weighted. A linear molecule has no first-order term at all (the effect is second-order Renner-Teller) and the log says so. +- **Default**: state + +### lr_grad_solver + +- **Type**: String +- **Description**: The method to solve the Z-vector (relaxed-density) equation $(A+B)Z=R$ in LR-TDDFT force and relaxation calculations, the linear-equation counterpart of `lr_solver`. Its dimension is $n_k n_{occ} n_{virt}$ summed over spin, where $n_{virt}$ counts every virtual band of the ground state, not only the `nvirt` window of the excitation. + - cg: Solve iteratively with the conjugate-gradient method, applying the orbital Hessian $A+B$ to a vector at each step. The matrix is never built. + - lapack: Construct the full matrix and solve directly with LAPACK (LU). Every MPI process holds the whole matrix and solves the same system. + - scalapack: Construct the matrix distributed over the MPI processes (2D block-cyclic) and solve with ScaLAPACK (LU). + - scalapack_chol: Construct the matrix distributed as for scalapack and solve by a ScaLAPACK Cholesky factorization, about half the flops of the LU and in place. + - elpa: Construct the matrix distributed as for scalapack and solve by an ELPA Cholesky factorization. + + > Note: The direct solvers build the matrix column by column, at the cost of one application of $A+B$ per column, which usually dominates the cost of the solve itself; scalapack, scalapack_chol and elpa need an MPI build, elpa also an ELPA build. + + > Note: scalapack_chol and elpa require $A+B$ to be positive definite, which holds at a stable ground state. If it is not, scalapack_chol stops with an error, while elpa does so only in a single-process run and hangs in a multi-process one; scalapack (LU) has no such requirement. +- **Default**: cg + +### lr_target_spin + +- **Type**: String +- **Description**: Which spin channel `lr_target_state` indexes. + + - singlet / triplet: the two closed-shell channels solved at `nspin = 2`. At `nspin = 1` only `singlet` exists. + - updown: the single spin-conserving channel of an open-shell calculation (`lr_unrestricted`, or a spin-polarised ground state with a non-zero moment). + + An open-shell calculation has only one channel, so any value is accepted there and relaxes that channel; an explicit `triplet` is reported as ignored. A closed-shell calculation rejects `updown`, since singlet and triplet are separate states with separate gradients. + + Ignored outside `calculation = relax`. +- **Default**: singlet + ### lr_unrestricted - **Type**: Boolean diff --git a/docs/parameters.yaml b/docs/parameters.yaml index 813d25e3b0a..5d51435081d 100644 --- a/docs/parameters.yaml +++ b/docs/parameters.yaml @@ -77,6 +77,8 @@ parameters: * ks-lr: Kohn-Sham density functional theory + LR-TDDFT (Under Development Feature) * lr: LR-TDDFT with given KS orbitals (Under Development Feature) * dfpt: density functional perturbation theory (Under Development Feature) + + [NOTE] Excited-state forces (`cal_force = 1`) and atomic relaxation with `ks-lr` or `lr` currently require `gamma_only = 1` and pseudopotentials without nonlinear core correction (NLCC). NLCC core-density response-force and XC-kernel derivative terms are not implemented; these force requests are rejected after reading the pseudopotentials, before force evaluation. Multi-k spectra and spectra with NLCC pseudopotentials remain available with `cal_force = 0`. Ground-state forces are unaffected by this LR restriction. default_value: ksdft unit: "" availability: "" @@ -2940,7 +2942,7 @@ parameters: category: Linear Response TDDFT type: String description: | - The exchange-correlation kernel used in the calculation. Currently supported: RPA, LDA, PBE, HSE, HF. + The exchange-correlation kernel used in the calculation. Currently supported: RPA, LDA, PWLDA, PBE, and the hybrids HF, PBE0, HSE, B3LYP, CAM_PBEH, LC_PBE, LC_WPBE, LRC_WPBE, LRC_WPBEH. A hybrid kernel needs the ground state to use the same functional: the exact-exchange operator $[\alpha+\beta\,\mathrm{erfc}(\mu r)]/r$ is built from exx_fock_alpha ($\alpha$), exx_erfc_alpha ($\beta$) and exx_erfc_omega ($\mu$), which are keyed off dft_functional, not off this parameter. default_value: LDA unit: "" availability: "" @@ -2978,9 +2980,10 @@ parameters: category: Linear Response TDDFT type: Integer description: | - The number of occupied orbitals (up to HOMO) used in the LR-TDDFT calculation. - * Note: If the value is illegal ( > nelec/2 or <= 0), it will be autoset to nelec/2. - default_value: nband + The number of occupied orbitals (up to HOMO) retained in the majority-spin LR-TDDFT window. A positive value selects a shared core prefix to discard from both spin channels; it does not change the ground-state occupations. + * If omitted, non-positive, or larger than the occupied majority-spin channel, all occupied orbitals are used. + * The full occupied window is determined by the effective electron number (including nelec_delta once) and the ground-state spin populations. For nspin=2, the minority-spin window has abs(N_up-N_down) fewer occupied orbitals. + default_value: all occupied orbitals unit: "" availability: "" - name: nvirt @@ -2999,6 +3002,82 @@ parameters: default_value: "0" unit: "" availability: "" + - name: lr_target_state + category: Linear Response TDDFT + type: Integer + description: | + Initial excited-state index for `calculation = relax`, counted from 0 within the spin channel selected by `lr_target_spin`. + + The gradient of the followed state (or its degenerate multiplet selected by `lr_degen_mode`) is computed, since solving the Z-vector equation dominates the cost of an excited-state gradient. It also selects the state whose excitation energy is added to the ground-state total energy, which is the quantity the energy-based relaxation algorithms (`cg`, `bfgs`, `lbfgs`) line-search on. + + Ignored outside `calculation = relax`: a single-point run solves and reports the gradients of every state. + + [NOTE] This index seeds the first ionic step. Subsequent steps compute cross-geometry AO overlaps and transform them with the saved and current KS orbitals to compare excitation amplitudes in a common occupied/virtual basis. Orbital sign changes and rotations within those subspaces do not change the overlap criterion. Index changes and low overlaps are reported. In JT mode the selected normalized multiplet mixture and its orbital basis become the reference. The projection is not renormalized, so window leakage remains visible. Tracking currently requires a fixed cell and unchanged orbital windows; degeneracy and a state leaving the solved window still limit state identification. Increase `lr_nstates` or reduce the ionic step when the overlap is low. + default_value: "0" + unit: "" + availability: "" + - name: lr_degen_thr + category: Linear Response TDDFT + type: Real + description: | + Excited states whose excitation energies lie within this threshold of each other are treated as one degenerate multiplet, and the full gradient matrix $G^{(A\alpha)}_{kl}=\langle X_k|\partial A/\partial R_{A\alpha}|X_l\rangle$ is computed for it in addition to the per-state gradients. Zero (the default) disables this and leaves the per-state gradients as the only output. + + At a $d$-fold degeneracy no single state has a gradient vector: the branch slopes along a displacement $u$ are the eigenvalues of $\sum_{A\alpha}u_{A\alpha}G^{(A\alpha)}$, and the eigenvectors that diagonalise it depend on $u$. The per-state gradients are the diagonal of $G$ in whichever basis the eigensolver happened to return, so only their sum (the trace) is basis-independent, while $G$ itself is the complete first-order information -- it is the linear vibronic coupling Hamiltonian of the multiplet. The extra cost is $d(d-1)/2$ further Z-vector solves per multiplet. + + The threshold proposes candidates; it cannot tell a true degeneracy from an accidental near-degeneracy, where the states have genuinely different excitation energies and the construction does not apply. Each multiplet's actual energy spread and the orthonormality of its eigenvectors are reported in the running log so the distinction can be made there. + + [NOTE] A sensible value is a few times the eigensolver threshold `lr_thr`, so that states split by real physics are not merged. + default_value: "0" + unit: Ry + availability: "" + - name: lr_degen_mode + category: Linear Response TDDFT + type: String + description: | + What `calculation = relax` follows when `lr_target_state` sits inside a degenerate multiplet, as identified by `lr_degen_thr`. It has no effect when the target state is non-degenerate. + + * state: follow the gradient of that one state, as returned by the eigensolver. This is the historical behaviour and is what reproduces earlier results, but inside a multiplet it is not a well-defined quantity: the per-state gradients are the diagonal of the subspace gradient matrix in whichever basis the eigensolver happened to return, so they depend on numerical details of the diagonalisation rather than on physics. + * average: follow the multiplet average $\bar\Omega=\frac{1}{d}\sum_k\Omega_k$, whose gradient is $\operatorname{Tr}G/d$. Unlike the individual states this is a smooth, basis-independent surface, and by symmetry its gradient is totally symmetric, so following it keeps the geometry on the symmetric configuration. Both the reported energy and the reported gradient switch to the average together, which the energy-based optimisers (`cg`, `bfgs`, `lbfgs`) require -- a gradient of one surface line-searched against the energy of another does not converge. This mode deliberately does NOT find the Jahn-Teller distortion, which is orthogonal to the totally symmetric average gradient. + * jt: descend the Jahn-Teller branch. Solves $\min_{\|u\|=1}\lambda_{\min}(\sum_{A\alpha}u_{A\alpha}G^{(A\alpha)})$ -- a joint optimisation over the displacement and the mixing inside the multiplet, since the two are determined together -- and follows the force of the resulting branch. This needs the off-diagonal part of the gradient matrix, so it costs $d(d-1)/2$ further Z-vector solves per step on top of the $d$ diagonal ones. The running log reports the branch's force, its mixing coefficients, and its split into the part common to the multiplet and the part that actually breaks the degeneracy. + + [NOTE] The usual sequence is `average` first, to reach the symmetric stationary point, then `jt` from there: at a stationary point of the average surface the common part vanishes and the whole force is Jahn-Teller. `jt` is self-limiting -- once a step has split the multiplet there is no group left and the ordinary single-state gradient takes over. + + [NOTE] `jt` gives the first-order DIRECTION. The distortion amplitude also needs the harmonic term, and the step norm is Cartesian rather than mass-weighted. A linear molecule has no first-order term at all (the effect is second-order Renner-Teller) and the log says so. + default_value: state + unit: "" + availability: "" + - name: lr_grad_solver + category: Linear Response TDDFT + type: String + description: | + The method to solve the Z-vector (relaxed-density) equation $(A+B)Z=R$ in LR-TDDFT force and relaxation calculations, the linear-equation counterpart of `lr_solver`. Its dimension is $n_k n_{occ} n_{virt}$ summed over spin, where $n_{virt}$ counts every virtual band of the ground state, not only the `nvirt` window of the excitation. + * cg: Solve iteratively with the conjugate-gradient method, applying the orbital Hessian $A+B$ to a vector at each step. The matrix is never built. + * lapack: Construct the full matrix and solve directly with LAPACK (LU). Every MPI process holds the whole matrix and solves the same system. + * scalapack: Construct the matrix distributed over the MPI processes (2D block-cyclic) and solve with ScaLAPACK (LU). + * scalapack_chol: Construct the matrix distributed as for scalapack and solve by a ScaLAPACK Cholesky factorization, about half the flops of the LU and in place. + * elpa: Construct the matrix distributed as for scalapack and solve by an ELPA Cholesky factorization. + + [NOTE] The direct solvers build the matrix column by column, at the cost of one application of $A+B$ per column, which usually dominates the cost of the solve itself; scalapack, scalapack_chol and elpa need an MPI build, elpa also an ELPA build. + + [NOTE] scalapack_chol and elpa require $A+B$ to be positive definite, which holds at a stable ground state. If it is not, scalapack_chol stops with an error, while elpa does so only in a single-process run and hangs in a multi-process one; scalapack (LU) has no such requirement. + default_value: cg + unit: "" + availability: "" + - name: lr_target_spin + category: Linear Response TDDFT + type: String + description: | + Which spin channel `lr_target_state` indexes. + + * singlet / triplet: the two closed-shell channels solved at `nspin = 2`. At `nspin = 1` only `singlet` exists. + * updown: the single spin-conserving channel of an open-shell calculation (`lr_unrestricted`, or a spin-polarised ground state with a non-zero moment). + + An open-shell calculation has only one channel, so any value is accepted there and relaxes that channel; an explicit `triplet` is reported as ignored. A closed-shell calculation rejects `updown`, since singlet and triplet are separate states with separate gradients. + + Ignored outside `calculation = relax`. + default_value: singlet + unit: "" + availability: "" - name: lr_unrestricted category: Linear Response TDDFT type: Boolean diff --git a/source/Makefile.Objects b/source/Makefile.Objects index 62df335c194..0e44f79160c 100644 --- a/source/Makefile.Objects +++ b/source/Makefile.Objects @@ -129,6 +129,7 @@ ${OBJS_DELTASPIN}\ ${OBJS_TENSOR}\ ${OBJS_HSOLVER_PEXSI}\ ${OBJS_LR}\ +${OBJS_LR_GRAD}\ ${OBJS_RDMFT} OBJS_MAIN=main.o\ @@ -437,6 +438,7 @@ OBJS_HAMILT_LCAO=hamilt_lcao.o\ nonlocal_dh.o\ nonlocal_fs.o\ overlap.o\ + ovlp_block.o\ overlap_fs.o\ td_ekinetic_lcao.o\ td_nonlocal_lcao.o\ @@ -846,6 +848,7 @@ OBJS_MODULE_RI=conv_coulomb_pot_k.o\ symm_rot_out.o\ gaussian_abfs.o\ exx_lri_detail.o\ + exx_lr_ws.o\ singular_value.o\ OBJS_BSE=esolver_lr_lcao_bse.o\ @@ -1032,7 +1035,6 @@ OBJS_TENSOR=tensor.o\ refcount.o OBJS_LR=lr_util.o\ - lr_util_hcontainer.o\ utils/lr_io.o\ utils/exciton_plotter.o\ ao_to_mo_parallel.o\ @@ -1048,6 +1050,24 @@ OBJS_TENSOR=tensor.o\ hamilt_casida.o\ esolver_lr_lcao_tddft.o\ +OBJS_LR_GRAD=lr_force.o\ + lr_force_aux.o\ + grad_degen.o\ + root_track.o\ + root_ovlp.o\ + grad_jt.o\ + gradient_output.o\ + lr_grad_cs.o\ + lr_grad_os.o\ + exx_proj.o\ + cvcx_serial.o\ + cvcx_par.o\ + cal_edm.o\ + zeqlin_solv.o\ + pot_grad_xc.o\ + esolver_lr_grad.o\ + esolver_lr_rlx.o\ + OBJS_RDMFT=rdmft.o\ rdmft_tools.o\ rdmft_pot.o\ diff --git a/source/source_base/matrix.cpp b/source/source_base/matrix.cpp index fb036438ce5..7461c8e0787 100644 --- a/source/source_base/matrix.cpp +++ b/source/source_base/matrix.cpp @@ -75,7 +75,12 @@ matrix::matrix( matrix && m_in ) matrix& matrix::operator=( const matrix & m_in ) { this->create( m_in.nr, m_in.nc, false ); - memcpy( c, m_in.c, nr*nc*sizeof(double) ); + // `create` leaves `c` null for an empty matrix, and memcpy's arguments are declared + // non-null even for a zero count -- so assigning an empty matrix is undefined behaviour. + if( nr && nc ) + { + memcpy( c, m_in.c, nr*nc*sizeof(double) ); + } return *this; } @@ -157,6 +162,16 @@ void matrix::create( const int nrow, const int ncol, const bool flag_zero ) } } +/* Unary minus*/ +matrix operator-(const matrix& m1) +{ + matrix tm(m1); + const int size = m1.nr * m1.nc; + for (int i = 0; i < size; i++) + tm.c[i] = -tm.c[i]; + return tm; +} + /* Adding matrices, as a friend */ matrix operator+(const matrix &m1, const matrix &m2) { diff --git a/source/source_base/matrix.h b/source/source_base/matrix.h index 8d2df96c45b..67de0c05876 100644 --- a/source/source_base/matrix.h +++ b/source/source_base/matrix.h @@ -82,7 +82,7 @@ class matrix using type=double; // Peize Lin add 2022.08.08 for template }; - +matrix operator-(const matrix& m1); // unary minus matrix operator+(const matrix &m1, const matrix &m2); matrix operator-(const matrix &m1, const matrix &m2); matrix operator*(const matrix &m1, const matrix &m2); diff --git a/source/source_base/module_external/blacs_connector.h b/source/source_base/module_external/blacs_connector.h index 1aa23a479d9..4a60c94141f 100644 --- a/source/source_base/module_external/blacs_connector.h +++ b/source/source_base/module_external/blacs_connector.h @@ -59,6 +59,9 @@ extern "C" void Czgebs2d(int ConTxt, char *scope, char *top, int m, int n, std::complex *A, int lda); void Czgebr2d(int ConTxt, char *scope, char *top, int m, int n, std::complex *A, int lda, int rsrc, int csrc); + + // element-wise sum over `scope`; rdest = -1 leaves the result on every process + void Cigsum2d(int ConTxt, char *scope, char *top, int m, int n, int *A, int lda, int rdest, int cdest); } // unified interface for broadcast diff --git a/source/source_base/module_external/scalapack_connector.h b/source/source_base/module_external/scalapack_connector.h index 35673385b67..d959219427b 100644 --- a/source/source_base/module_external/scalapack_connector.h +++ b/source/source_base/module_external/scalapack_connector.h @@ -118,6 +118,20 @@ extern "C" int *ipiv, std::complex* B, const int* ib, const int* jb, const int*descb, const int *info ); + void pdgesv_( + const int *n, const int *nrhs, + double *A, const int *ia, const int *ja, const int *desca, + int *ipiv, double* B, const int* ib, const int* jb, const int*descb, int *info + ); + + void pdpotrs_(const char* uplo, const int* n, const int* nrhs, + const double* A, const int* ia, const int* ja, const int* desca, + double* B, const int* ib, const int* jb, const int* descb, int* info); + + void pzpotrs_(const char* uplo, const int* n, const int* nrhs, + const std::complex* A, const int* ia, const int* ja, const int* desca, + std::complex* B, const int* ib, const int* jb, const int* descb, int* info); + void pdsygvx_(const int* itype, const char* jobz, const char* range, const char* uplo, const int* n, double* A, const int* ia, const int* ja, const int*desca, double* B, const int* ib, const int* jb, const int*descb, const double* vl, const double* vu, const int* il, const int* iu, @@ -396,6 +410,33 @@ class ScalapackConnector pzgesv_(&n, &nrhs, A, &ia, &ja, desca, ipiv, B, &ib, &jb, descb, info); } + static inline + void gesv( + const int n, const int nrhs, + double *A, const int ia, const int ja, const int *desca, + int *ipiv, double* B, const int ib, const int jb, const int*descb, int *info) + { + pdgesv_(&n, &nrhs, A, &ia, &ja, desca, ipiv, B, &ib, &jb, descb, info); + } + + static inline + void potrs( + const char uplo, const int n, const int nrhs, + const double* A, const int ia, const int ja, const int* desca, + double* B, const int ib, const int jb, const int* descb, int* info) + { + pdpotrs_(&uplo, &n, &nrhs, A, &ia, &ja, desca, B, &ib, &jb, descb, info); + } + + static inline + void potrs( + const char uplo, const int n, const int nrhs, + const std::complex* A, const int ia, const int ja, const int* desca, + std::complex* B, const int ib, const int jb, const int* descb, int* info) + { + pzpotrs_(&uplo, &n, &nrhs, A, &ia, &ja, desca, B, &ib, &jb, descb, info); + } + static inline void tranu( const int m, const int n, diff --git a/source/source_base/parallel_grid.h b/source/source_base/parallel_grid.h index 8cc7592384c..d74e7a3e70e 100644 --- a/source/source_base/parallel_grid.h +++ b/source/source_base/parallel_grid.h @@ -43,7 +43,7 @@ class Parallel_Grid int get_nz() const { return ncz; } int get_nrxx() const { return nrxx; } - private: + private: void z_distribution(void); diff --git a/source/source_esolver/CMakeLists.txt b/source/source_esolver/CMakeLists.txt index 95488f4be33..84d785eba78 100644 --- a/source/source_esolver/CMakeLists.txt +++ b/source/source_esolver/CMakeLists.txt @@ -21,6 +21,8 @@ if(ENABLE_LCAO) esolver_ks_lcao.cpp esolver_ks_lcao_tddft.cpp esolver_lr_lcao_tddft.cpp + esolver_lr_grad.cpp + esolver_lr_rlx.cpp esolver_gets.cpp lcao_others.cpp esolver_dm2rho.cpp diff --git a/source/source_esolver/esolver_lr_grad.cpp b/source/source_esolver/esolver_lr_grad.cpp new file mode 100644 index 00000000000..a6c581ac88f --- /dev/null +++ b/source/source_esolver/esolver_lr_grad.cpp @@ -0,0 +1,380 @@ +#include "source_esolver/esolver_lr_lcao_tddft.h" +#include "source_lcao/module_lr/zeq_solver.h" +#include "source_lcao/module_lr/cal_edm.h" +#include "source_lcao/module_lr/lr_force.h" +#include "source_lcao/module_lr/gradient_inputs.h" +#include "source_lcao/module_lr/gradient_output.h" +#include "source_lcao/module_lr/lr_amp.h" +#include "source_lcao/module_lr/grad_degen.h" +#include "source_base/parallel_reduce.h" +#include +#include +#include +#include "source_estate/module_dm/dm_from_psi.h" +#include "source_io/module_output/output_log.h" + +using namespace LR; + + +///========================= excited-state geometry relaxation ========================= + +template +void ModuleESolver::ESolver_LR::init_pot_groundstate(const Charge& chg_gs) +{ + ModuleBase::TITLE("ESolver_LR", "init_pot_gs"); + std::vector pot_register; + if (this->inp_->vl_in_h) + { + if (!this->ks_) + { // on the `ks-lr` path `sfac()`/`vloc()` alias the ground-state solver's, which + // `ESolver_FP::before_scf` already refreshed for the current geometry + //! 11) calculate the structure factor + // the has_float_data flag below is read the same way core ABACUS reads it at the + // analogous call site (source_esolver/esolver_fp.cpp's `this->sf.setup(...)`), so this + // mirrors the established convention rather than introducing a new global dependency. + this->sfac().setup(&(*this->ucell_), pgrid(), this->pw_rhod, PARAM.globalv.has_float_data); + this->vloc().init_vloc((*this->ucell_), this->pw_rho); + } + pot_register.push_back("local"); + } + if(this->inp_->vh_in_h) + { + pot_register.push_back("hartree"); + } + pot_register.push_back("xc"); + + // initialize the ground state potential + this->pot_gs = LR_Util::make_unique(this->pw_rhod, this->pw_rho, + &(*this->ucell_), &this->vloc().vloc, &this->sfac(), &this->solvent, + &this->etxc_gs, &this->vtxc_gs); + this->pot_gs.get()->pot_register(pot_register); + XC_Functional::set_xc_type((*this->ucell_).atoms[0].ncpp.xc_func); // set XC type of the ground state + this->pot_gs->init_pot(&chg_gs); // call update_from_charge inside + if (LR_Util::has_local_xc(this->xc_kernel)) + { + XC_Functional::set_xc_type(this->xc_kernel); // recover the excited state xc kernel type + } + if (this->inp_->test_force) + { + this->pot_gs_hartree = LR_Util::make_unique(this->pw_rhod, this->pw_rho, + &(*this->ucell_), &this->vloc().vloc, &this->sfac(), &this->solvent, + &this->etxc_gs, &this->vtxc_gs); + this->pot_gs_hartree->pot_register({ "hartree" }); + this->pot_gs_hartree->init_pot(&chg_gs); // call update_from_charge inside + } +} + +template +ct::Tensor ModuleESolver::ESolver_LR::pad_X_to_z_(const int ispin, const int istate_begin, const int nst) const +{ + return LR::pad_amplitudes(this->X, this->openshell, this->paraX_, this->paraX_z_, + this->nocc, this->nvirt, this->nk, this->nloc_per_state, this->nloc_per_state_z_, + ispin, istate_begin, nst); +} + +template +ct::Tensor ModuleESolver::ESolver_LR::solve_zvector_eqation(const int ispin, const int nst, const ct::Tensor& Xz) +{ + ModuleBase::TITLE("ESolver_LR", "cal_force"); + ModuleBase::timer::start("ESolver_LR", "solve_zvector_eqation"); + // `Z_vector_equation` treats X and Z as `nst` independent blocks of `nloc_per_state_z_`, + // and the blocks need not be eigenvectors of the Casida equation in the order the + // diagonalizer returned them -- `cal_grad_matrix_degenerate` feeds it linear combinations. + // X arrives already widened into the Z window (`Xz`), which spans every virtual band the + // ground state produced rather than the `nvirt` window X was solved in -- the Brillouin + // condition the Z-vector enforces holds in EVERY occupied-virtual rotation. The padded + // entries are zero, so every X-derived quantity (D^X, T) is unchanged; only the space Z is + // solved in grows. Both the closed- and the open-shell path go through here. + ct::Tensor Z = LR_Util::newTensor({ nst, this->nloc_per_state_z_ }); + // construct and solve the Z-vector equation + const LR::GradientInputs inputs = this->gradient_inputs_(); + const T* const x_data = Xz.template data(); + T* const z_data = Z.template data(); + std::weak_ptr pot_weak = inputs.pot[ispin]; + Z_vector_equation(inputs, x_data, z_data, nst, pot_weak, + this->spin_types[ispin], this->in_dir, this->openshell, this->inp_->lr_grad_solver); + ModuleBase::timer::end("ESolver_LR", "solve_zvector_eqation"); + return Z; +} + +template +std::vector ModuleESolver::ESolver_LR::cal_force(const int ispin, const int istate_only) +{ + if (this->inp_->test_force && ispin == 0) { this->test_force(); } + if (this->openshell) { return this->cal_force_openshell(istate_only); } + + const int ist_begin = (istate_only < 0) ? 0 : istate_only; + const int nst = (istate_only < 0) ? this->nstates : 1; + // The whole closed-shell gradient runs in the Z window (every virtual band the ground state + // produced), not the `nvirt` window X was solved in: only there does the Z-vector enforce + // the Brillouin condition in every occupied-virtual rotation. X is zero-padded into it, so + // D^X and T come out bit-identical -- and Omega is untouched, so an existing finite-difference + // reference stays valid. See `fill_z_window_`. + const ct::Tensor Xz = this->pad_X_to_z_(ispin, ist_begin, nst); + std::vector omega(nst); + for (int i = 0; i < nst; ++i) + { + omega[i] = this->pelec->ekb.c[ispin * this->nstates + ist_begin + i]; + } + return this->cal_force_Xz(ispin, Xz, omega, ist_begin); +} + +template +LR::GradientInputs ModuleESolver::ESolver_LR::gradient_inputs_() const +{ + return { *this->ucell_, this->kv, this->gd(), this->orb_cutoff_, this->paraMat_, + this->paraC_z_, this->paraX_z_, *this->psi_ks_z_, this->eig_ks_z_, this->nocc, + this->nvirt_z_, this->nspin, this->nk, this->nbasis, this->nloc_per_state_z_, + this->xc_kernel, this->inp_->dft_functional, this->inp_->ks_solver, + this->inp_->test_force, this->excited_relax_, this->out_dir, this->my_rank_, + this->spin_types, this->pot, this->pot_hxc_gs, this->ofs_running_ +#ifdef __EXX + , this->exx_lri, this->exx_info.info_global.hybrid_alpha +#endif + }; +} + +template +std::vector ModuleESolver::ESolver_LR::cal_force_Xz(const int ispin, + const ct::Tensor& Xz, const std::vector& omega, const int label_begin) +{ + ModuleBase::timer::start("ESolver_LR", "cal_force_Xz"); + const int nst = static_cast(omega.size()); + const int channel = ispin; + const ct::Tensor Z = this->solve_zvector_eqation(channel, nst, Xz); + const auto dm_gs = this->cal_dm_gs(); + const LR::GradientInputs inputs = this->gradient_inputs_(); + LR_Force force_terms(*this->ucell_, this->kv.kvec_d, this->paraMat_, + *this->pw_rhod, *this->pw_rho, this->vloc(), this->sfac(), this->gd(), this->tcb() +#ifdef __EXX + , inputs.exx_lri, inputs.hybrid_alpha +#endif + ); + const auto forces = LR::evaluate_closed_shell_force(inputs, force_terms, dm_gs, Xz, Z, omega, label_begin, ispin); + ModuleBase::timer::end("ESolver_LR", "cal_force_Xz"); + return forces; +} + +template +void ModuleESolver::ESolver_LR::cal_force_and_grad_matrix_(const int ispin, std::ofstream& ofs) +{ + const std::vector forces = this->cal_force(ispin); + if (this->inp_->lr_degen_thr <= 0.0) { return; } + // Rebuilt from scratch for this channel: a relaxation calls this once per ionic step, and a + // stale multiplet from the previous geometry must not survive into the next one. + this->multiplet_lvc_.erase( + std::remove_if(this->multiplet_lvc_.begin(), this->multiplet_lvc_.end(), + [ispin](const MultipletLVC& m) { return m.ispin == ispin; }), + this->multiplet_lvc_.end()); + // The per-state gradients above are the diagonal of the degenerate-subspace gradient matrix in + // whatever basis the eigensolver returned, so inside a multiplet only their trace means + // anything. `lr_degen_thr` asks for the off-diagonal part too, which is the rest of the + // first-order information. + const int ekb_off = this->openshell ? 0 : ispin * this->nstates; + std::vector omega(this->nstates); + for (int ist = 0; ist < this->nstates; ++ist) { omega[ist] = this->pelec->ekb.c[ekb_off + ist]; } + const std::vector> groups + = LR::group_degenerate_states(omega, this->inp_->lr_degen_thr); + for (const std::vector& group : groups) + { + if (group.size() < 2) { continue; } + std::vector diag; + for (const int ist : group) { diag.push_back(forces[ist]); } + MultipletLVC lvc; + lvc.ispin = ispin; + lvc.states = group; + double lo = omega[group.front()]; + double hi = omega[group.front()]; + double sum = 0.0; + for (const int ist : group) + { + lo = std::min(lo, omega[ist]); + hi = std::max(hi, omega[ist]); + sum += omega[ist]; + } + lvc.omega0 = sum / static_cast(group.size()); + lvc.omega_spread = hi - lo; + lvc.g = this->cal_grad_matrix_degenerate(ispin, group, diag, ofs); + this->multiplet_lvc_.push_back(lvc); + } +} + +template +std::vector> +ModuleESolver::ESolver_LR::cal_grad_matrix_degenerate(const int ispin, + const std::vector& group, const std::vector& diag, std::ofstream& ofs) +{ + ModuleBase::TITLE("ESolver_LR", "cal_grad_matrix_degenerate"); + ModuleBase::timer::start("ESolver_LR", "cal_grad_matrix_degenerate"); + const int d = static_cast(group.size()); + assert(d >= 2); + assert(static_cast(diag.size()) == d); + const std::vector> pairs = LR::degenerate_pairs(d); + const int nloc_g = this->nloc_per_state_z_; + const int ekb_off = this->openshell ? 0 : ispin * this->nstates; + + // 1. the multiplet's members, each widened into the Z window + const ct::Tensor Xz = this->pad_group_to_z_(ispin, group); + std::vector omega_member(d); + for (int k = 0; k < d; ++k) { omega_member[k] = this->pelec->ekb.c[ekb_off + group[k]]; } + + // 2. the two preconditions of route (A2), reported rather than enforced: the combinations + // $X_\pm$ are normalized eigenvectors only if the members are orthonormal, and the gradient + // is a single quadratic form only if they share one $\Omega$. An accidental near-degeneracy + // passes the grouping threshold but fails the second, and its `omega_spread` says so. + double max_ovlp_err = 0.0; + for (int k = 0; k < d; ++k) + { + for (int l = k; l < d; ++l) + { + const T* const xk = Xz.template data() + static_cast(k) * nloc_g; + const T* const xl = Xz.template data() + static_cast(l) * nloc_g; + T loc = static_cast(0); + for (int i = 0; i < nloc_g; ++i) { loc += LR_Util::get_conj(xk[i]) * xl[i]; } + Parallel_Reduce::reduce_all(loc); + const double ref = (k == l) ? 1.0 : 0.0; + max_ovlp_err = std::max(max_ovlp_err, std::abs(loc - ref)); + } + } + const double omega_spread = *std::max_element(omega_member.begin(), omega_member.end()) + - *std::min_element(omega_member.begin(), omega_member.end()); + // one $\Omega$ for every combination: the members share it up to `omega_spread`, and the mean + // is the neutral choice for a vector that belongs to no single member + const double omega_mean + = std::accumulate(omega_member.begin(), omega_member.end(), 0.0) / static_cast(d); + ofs << " Degenerate multiplet of " << d << " states, Omega = " << omega_mean + << " Ry, spread = " << omega_spread << " Ry, max | - delta_kl| = " << max_ovlp_err + << std::endl; + if (max_ovlp_err > 1e-6) + { + ofs << " WARNING: this multiplet's eigenvectors are not orthonormal to" + " 1e-6, so (X_k+X_l)/sqrt(2) is not normalized and the assembled gradient matrix is" + " wrong by that much." << std::endl; + } + + // 3. the $d(d-1)/2$ combinations $X_+=(X_k+X_l)/\sqrt2$, all in one Z-vector solve + const int npair = static_cast(pairs.size()); + ct::Tensor Xp = LR_Util::newTensor({ npair, nloc_g }); + Xp.zero(); + for (int ip = 0; ip < npair; ++ip) + { + LR::combine_normalized( + Xz.template data() + static_cast(pairs[ip].first) * nloc_g, + Xz.template data() + static_cast(pairs[ip].second) * nloc_g, + static_cast(nloc_g), + Xp.template data() + static_cast(ip) * nloc_g); + } + const std::vector omega_pair(npair, omega_mean); + const std::vector plus = this->openshell + ? this->cal_force_openshell_Xz(Xp, omega_pair, group[0]) + : this->cal_force_Xz(ispin, Xp, omega_pair, group[0]); + + // 4. $G_{kl}=\mathcal F[X_+]-\tfrac12(G_{kk}+G_{ll})$ + const std::vector> g + = LR::assemble_grad_matrix(diag, plus, pairs); + print_grad_matrix(g, group, (*this->ucell_), ofs); + ModuleBase::timer::end("ESolver_LR", "cal_grad_matrix_degenerate"); + return g; +} + +template +std::vector ModuleESolver::ESolver_LR::cal_force_openshell(const int istate_only) +{ + ModuleBase::TITLE("ESolver_LR", "cal_force_openshell"); + + // Open shell: there is a single eigenproblem whose vector is the concatenation + // [up-block | down-block], and every density matrix has two independent channels. + // The spin-orbital formulas apply verbatim -- unlike the closed-shell singlet/triplet + // algorithm, X here is normalized over BOTH channels, so it carries no implicit sqrt(2) + // and none of the collapsed 2/4 factors are needed. + const int ist_begin_ = (istate_only < 0) ? 0 : istate_only; + const int nst_ = (istate_only < 0) ? this->nstates : 1; + // Like the closed-shell path, the whole gradient runs in the Z window: X is zero-padded + // into it (both channels, each re-based -- see `pad_X_to_z_`), so D^X and T are unchanged + // while the Z-vector gets every occupied-virtual rotation the AO basis supports. + const ct::Tensor Xz = this->pad_X_to_z_(0, ist_begin_, nst_); + std::vector omega_(nst_); + for (int i = 0; i < nst_; ++i) { omega_[i] = this->pelec->ekb.c[ist_begin_ + i]; } + return this->cal_force_openshell_Xz(Xz, omega_, ist_begin_); +} + +template +std::vector ModuleESolver::ESolver_LR::cal_force_openshell_Xz( + const ct::Tensor& Xz, const std::vector& omega, const int label_begin) +{ + ModuleBase::timer::start("ESolver_LR", "cal_force_openshell_Xz"); + const int nst = static_cast(omega.size()); + const int channel = 0; + const ct::Tensor Z = this->solve_zvector_eqation(channel, nst, Xz); + const auto dm_gs = this->cal_dm_gs(); + const LR::GradientInputs inputs = this->gradient_inputs_(); + LR_Force force_terms(*this->ucell_, this->kv.kvec_d, this->paraMat_, + *this->pw_rhod, *this->pw_rho, this->vloc(), this->sfac(), this->gd(), this->tcb() +#ifdef __EXX + , inputs.exx_lri, inputs.hybrid_alpha +#endif + ); + const auto forces = LR::evaluate_open_shell_force(inputs, force_terms, dm_gs, Xz, Z, omega, label_begin); + ModuleBase::timer::end("ESolver_LR", "cal_force_openshell_Xz"); + return forces; +} + +template +module_dm::DensityMatrix ModuleESolver::ESolver_LR::cal_dm_gs() +{ + module_dm::DensityMatrix dm_gs(&this->paraMat_, this->nspin, this->kv.kvec_d, this->nk); + // `psi_ks_all_` is distributed per `this->ks_->pv` on the ks-lr path (it is aliased straight + // from the ground-state solver's own `psi`), but per `this->paraMat_all_` on the + // read-from-file path (where it was allocated with that descriptor's own sizes). Using the + // wrong one here reads `psi_ks_all_`'s local block with the wrong block size/process-grid + // assumption -- invisible at nprocs=1, where every descriptor degenerates to one block, but + // silently wrong at nprocs>1. `refresh_from_ks_`/`fill_z_window_` already make this same + // distinction for their `Cpxgemr2d` source descriptor; this mirrors it. + const Parallel_Orbitals* const pv_all = this->ks_ ? &this->ks_->pv : &this->paraMat_all_; + module_dm::dm_from_psi(pv_all, this->wg_ks_all, *this->psi_ks_all_, dm_gs); // nbands is important here + LR_Util::initialize_DMR(dm_gs, this->paraMat_, (*this->ucell_), this->gd(), this->orb_cutoff_); // nbands is not important here + dm_gs.cal_dmr(-1); + return dm_gs; +} + +template +void ModuleESolver::ESolver_LR::test_force() +{ + LR_Force lr_force((*this->ucell_), this->kv.kvec_d, this->paraMat_, *this->pw_rhod, *this->pw_rho, + this->vloc(), this->sfac(), this->gd(), this->tcb() +#ifdef __EXX + , std::weak_ptr>(this->exx_lri), this->exx_info.info_global.hybrid_alpha +#endif + ); + + const module_dm::DensityMatrix& dm_gs = this->cal_dm_gs(); + // LR_Util::print_DMR(dm_gs, "DM(R) of ground state"); + ///========================== test 1: reproduce the force of ground state ========================= + // energy density matrix of the ground state + module_dm::DensityMatrix edm_gs(&this->paraMat_, this->nspin, this->kv.kvec_d, this->nk); //DX + ModuleBase::matrix wg_ekb_ks_all(nspin, this->inp_->nbands); + std::transform(this->wg_ks_all.c, this->wg_ks_all.c + nspin * this->inp_->nbands, + this->eig_ks_all.c, wg_ekb_ks_all.c, std::multiplies()); + // see `cal_dm_gs()`: `psi_ks_all_` is distributed per `this->ks_->pv` on the ks-lr path, + // not `this->paraMat_all_`. + const Parallel_Orbitals* const pv_all_edm = this->ks_ ? &this->ks_->pv : &this->paraMat_all_; + module_dm::dm_from_psi(pv_all_edm, wg_ekb_ks_all, *this->psi_ks_all_, edm_gs); + LR_Util::initialize_DMR(edm_gs, this->paraMat_, (*this->ucell_), this->gd(), this->orb_cutoff_); + edm_gs.cal_dmr(-1); + // ground-state force + ModuleBase::matrix force_gs = lr_force.reproduce_force_gs(kv, dm_gs, edm_gs); + ModuleIO::print_force(this->ofs_running_, (*this->ucell_), "Ground State FORCE (eV/Angstrom)", force_gs, false); + /// ======================================= END test 1 ========================================= + ///========================== test 2: reproduce the DX Hartree term ========================= + ModuleBase::matrix f_hxc_potgs = lr_force.reproduce_force_gs_loc(dm_gs, *this->pot_gs_hartree); + ModuleIO::print_force(this->ofs_running_, (*this->ucell_), "GS Hartree force calculated by 'cal_pulay_fs' from potential (eV/Angstrom)", f_hxc_potgs, false); + ModuleBase::matrix f_hxc_potlr = lr_force.cal_force_hxc_dmtrans(dm_gs, *this->pot[0]); + ModuleIO::print_force(this->ofs_running_, (*this->ucell_), "GS Hxc force calculated by 'LR_Force' from kernel (eV/Angstrom)", f_hxc_potlr, false); + // `cal_force_hxc_dmtrans` now includes the Pulay -> Pulay+Hellmann-Feynman factor 2 itself, + // so this must match the ground-state Hartree force directly (dm_gs is already symmetric). + /// ======================================= END test 2 ========================================= + +} + +template class ModuleESolver::ESolver_LR; +template class ModuleESolver::ESolver_LR, double>; diff --git a/source/source_esolver/esolver_lr_lcao_bse.cpp b/source/source_esolver/esolver_lr_lcao_bse.cpp index 01506916870..b9e070676e2 100644 --- a/source/source_esolver/esolver_lr_lcao_bse.cpp +++ b/source/source_esolver/esolver_lr_lcao_bse.cpp @@ -22,6 +22,7 @@ void ESolver_BSE::before_all_runners(BaseCell& basecell, const Input_para ModuleBase::TITLE("ESolver_BSE", "before_all_runners"); ModuleBase::timer::start("ESolver_BSE", "before_all_runners"); + this->bind_ground_state_aliases_(); // BSE always owns its grid/orbital objects (no `ks_`) this->ucell_ = &ucell; // xc kernel this->xc_kernel = LR_Util::tolower(inp.xc_kernel); @@ -38,14 +39,14 @@ void ESolver_BSE::before_all_runners(BaseCell& basecell, const Input_para this->parameter_check(); /// read orbitals and build the interpolation table - this->two_center_bundle_.build_orb(ucell.ntype, ucell.orbital_fn.data(), inp.orbital_dir); + this->two_center_bundle_own_.build_orb(ucell.ntype, ucell.orbital_fn.data(), inp.orbital_dir); - this->two_center_bundle_.to_LCAO_Orbitals(this->orb_, inp.lcao_ecut, inp.lcao_dk, inp.lcao_dr, inp.lcao_rmax, + this->two_center_bundle_own_.to_LCAO_Orbitals(this->orb_, inp.lcao_ecut, inp.lcao_dk, inp.lcao_dr, inp.lcao_rmax, inp.out_element_info, inp.cal_force); this->orb_cutoff_ = this->orb_.cutoffs(); if (LR_Util::tolower(this->inp_->abs_gauge) == "velocity") { - this->setup_2center_table(this->two_center_bundle_, this->orb_, ucell); + this->setup_2center_table(this->two_center_bundle_own_, this->orb_, ucell); } this->set_dimension(); @@ -60,11 +61,11 @@ void ESolver_BSE::before_all_runners(BaseCell& basecell, const Input_para this->paraMat_.ncol_bands = this->nbands; #endif - this->psi_ks = new psi::Psi(this->kv.get_nks(), - this->paraMat_.ncol_bands, - this->paraMat_.get_row_size(), - this->kv.ngk, - true); + this->psi_ks.reset(new psi::Psi(this->kv.get_nks(), + this->paraMat_.ncol_bands, + this->paraMat_.get_row_size(), + this->kv.ngk, + true)); this->psi_ks_global = new psi::Psi(this->kv.get_nks(), this->nbands, this->nbasis, @@ -84,7 +85,7 @@ void ESolver_BSE::before_all_runners(BaseCell& basecell, const Input_para #endif ); - this->Pgrid.init(this->pw_rho->nx, + this->pgrid().init(this->pw_rho->nx, this->pw_rho->ny, this->pw_rho->nz, this->pw_rho->nplane, @@ -102,7 +103,7 @@ void ESolver_BSE::before_all_runners(BaseCell& basecell, const Input_para PARAM.globalv.gamma_only_local); atom_arrange::search(PARAM.globalv.search_pbc, GlobalV::ofs_running, - this->gd, + this->gd(), *this->ucell_, search_radius, inp.test_atom_input); @@ -122,7 +123,7 @@ void ESolver_BSE::before_all_runners(BaseCell& basecell, const Input_para this->pw_big->nbzp, this->orb_.Phi, ucell, - this->gd, + this->gd(), inp.nspin, PARAM.globalv.gamma_only_local, PARAM.globalv.domag, @@ -188,7 +189,7 @@ void ESolver_BSE::runner(BaseCell& basecell, const int istep) assert(this->xc_kernel == "bse"); this->lri_init(); BSE::HamiltBSE bse_matrix(this->nspin, this->nbasis, this->nocc, this->nvirt, *this->ucell_, - this->orb_cutoff_, this->gd, *this->psi_ks, *this->psi_ks_global, this->eig_gw, + this->orb_cutoff_, this->gd(), *this->psi_ks, *this->psi_ks_global, this->eig_gw, *this->mo_lri, this->pot[0], this->kv, this->paraX_, this->paraC_, this->paraMat_, this->inp_->bse_spin_types, @@ -397,7 +398,7 @@ void ESolver_BSE::after_all_runners(BaseCell& basecell) std::cout << "plot BSE exciton wavefunction for state: " << this->inp_->plot_istate << ", spin type: " << this->inp_->bse_spin_types[is] << std::endl; LR_Util::ExcitonPlotter eplot(this->nspin, this->nbasis, this->nocc, this->nvirt, *this->psi_ks, - *this->ucell_, this->kv, this->gd, this->orb_cutoff_, this->Pgrid, *this->pw_rho, + *this->ucell_, this->kv, this->gd(), this->orb_cutoff_, this->pgrid(), *this->pw_rho, this->paraX_, this->paraC_, this->paraMat_, output_dir, &this->tda_ene[is * this->nstates], this->X[is].template data(), @@ -478,7 +479,7 @@ void ESolver_BSE::after_all_runners(BaseCell& basecell) if (LR_Util::tolower(this->inp_->abs_gauge) == "velocity" ) { const int nspin_tmp = this->inp_->nspin == 2 ? 2 : 1; - this->velocity_mo = LR_Util::cal_velocity_mo(*this->ucell_, this->gd, this->two_center_bundle_, + this->velocity_mo = LR_Util::cal_velocity_mo(*this->ucell_, this->gd(), this->tcb(), this->paraMat_, this->paraC_, this->kv, *this->psi_ks, this->nk, nspin_tmp, this->nbasis, this->nocc, this->nvirt); } @@ -487,7 +488,7 @@ void ESolver_BSE::after_all_runners(BaseCell& basecell) for (int is = 0; is < this->X.size(); ++is) { LR::LR_Spectrum spectrum(this->nspin, this->nbasis, this->nocc, this->nvirt, *this->pw_rho, *this->psi_ks, - *this->ucell_, this->kv, this->gd, this->orb_cutoff_, this->two_center_bundle_, + *this->ucell_, this->kv, this->gd(), this->orb_cutoff_, this->tcb(), this->paraX_, this->paraC_, this->paraMat_, &this->tda_ene[is * this->nstates], this->eig_ks.c, this->X[is].template data(), this->nstates, false/*openshell*/, @@ -518,7 +519,7 @@ void ESolver_BSE::after_all_runners(BaseCell& basecell) for (int is = 0;is < this->full_X.size();++is) { LR::LR_Spectrum spectrum(this->nspin, this->nbasis, this->nocc, this->nvirt, *this->pw_rho, *this->psi_ks, - *this->ucell_, this->kv, this->gd, this->orb_cutoff_, this->two_center_bundle_, + *this->ucell_, this->kv, this->gd(), this->orb_cutoff_, this->tcb(), this->paraX_, this->paraC_, this->paraMat_, &this->full_ene[is * this->nstates], this->eig_ks.c, this->full_X[is].template data(), this->nstates, false/*openshell*/, @@ -713,12 +714,14 @@ void ESolver_BSE::init_pot(const Charge& chg_gs) { using ST = LR::PotHxcLR::SpinType; case 1: case 2: - this->pot[0] = std::make_shared(this->xc_kernel, *this->pw_rho, *this->ucell_, chg_gs, this->Pgrid, - ST::S1, this->inp_->lr_init_xc_kernel); + this->pot[0] = std::make_shared(this->xc_kernel, *this->pw_rho, *this->ucell_, chg_gs, this->pgrid(), + // BSE multiplies its bare Coulomb matrix by the singlet factor in HamiltBSE. + // S1_gs has weight one; S1 would apply that factor a second time on the grid path. + ST::S1_gs, this->inp_->lr_init_xc_kernel); break; // case 2: - // this->pot[0] = std::make_shared(xc_kernel, *this->pw_rho, ucell, chg_gs, Pgrid, openshell ? ST::S2_updown : ST::S2_singlet, this->inp_->lr_init_xc_kernel); - // this->pot[1] = std::make_shared(xc_kernel, *this->pw_rho, ucell, chg_gs, Pgrid, openshell ? ST::S2_updown : ST::S2_triplet, this->inp_->lr_init_xc_kernel); + // this->pot[0] = std::make_shared(xc_kernel, *this->pw_rho, ucell, chg_gs, pgrid(), openshell ? ST::S2_updown : ST::S2_singlet, this->inp_->lr_init_xc_kernel); + // this->pot[1] = std::make_shared(xc_kernel, *this->pw_rho, ucell, chg_gs, pgrid(), openshell ? ST::S2_updown : ST::S2_triplet, this->inp_->lr_init_xc_kernel); // break; default: throw std::invalid_argument("ESolver_BSE: nspin must be 1 or 2"); diff --git a/source/source_esolver/esolver_lr_lcao_tddft.cpp b/source/source_esolver/esolver_lr_lcao_tddft.cpp index e754ab02041..3b230fed30e 100644 --- a/source/source_esolver/esolver_lr_lcao_tddft.cpp +++ b/source/source_esolver/esolver_lr_lcao_tddft.cpp @@ -1,5 +1,10 @@ #include "esolver_lr_lcao_tddft.h" #include "source_basis/module_pw/pw_basis_big.h" // use PW_Basis_Big +#include +#include "source_base/parallel_reduce.h" + +#include +#include #include "source_lcao/module_lr/utils/lr_io.h" #include "source_lcao/module_lr/utils/lr_util.h" #include "source_lcao/module_lr/hamilt_casida.h" @@ -9,6 +14,7 @@ #include "source_hamilt/module_xc/xc_functional.h" #include "source_lcao/module_lr/hsolver_lrtd.hpp" #include "source_lcao/module_lr/lr_spectrum.h" +#include "source_lcao/module_lr/lr_density.hpp" #include "source_hamilt/module_gint/gint.h" #include #include "source_lcao/hamilt_lcao.h" @@ -27,32 +33,66 @@ #ifdef __EXX #include "source_lcao/module_ri/exx_lri_interface.h" #include "source_hamilt/module_xc/exx_info.h" // for init_exx_info +#include "source_hamilt/module_xc/xc_functional.h" // for set_xc_type +#endif + +// gradient +#include "source_lcao/module_lr/zeq_solver.h" + +#ifdef __EXX +namespace +{ + + /// One `Exx_LRI` carries ONE Coulomb operator, and it may be needed for two different + /// reasons: the LR kernel (when `xc_kernel` is a hybrid) and the ground-state force (when + /// `dft_functional` is a hybrid). That operator is NOT chosen here -- `Exx_LRI` reads + /// `info_ri.coulomb_param`, which `input_conv` builds from `dft_functional` alone. So when + /// the two disagree, the kernel silently gets the ground state's screening, not its own. + /// + /// This used to be decided by a local `exx_ccp_type()` returning Erfc for "hse" and bare + /// Hf otherwise, written onto `info_global.ccp_type`. With general range-separated hybrids + /// that two-way split is not even expressible ($\alpha/r+\beta\,\mathrm{erfc}(\mu r)/r$ + /// is both at once), and the write was in any case dead for the RI path: only `Exx_LRI` + /// runs here, and it never looks at `ccp_type`. + void warn_if_kernel_differs_from_gs(const std::string& xc_kernel, const std::string& dft_functional, + std::ofstream& ofs_running) + { + const bool k = LR::exx_kernel_list().count(xc_kernel) > 0; + const bool g = LR::exx_kernel_list().count(dft_functional) > 0; + if (k && xc_kernel != dft_functional) + { + ofs_running << " WARNING: xc_kernel (" << xc_kernel << ") and dft_functional (" + << dft_functional << ") are not the same functional. The Coulomb operator of" + " Exx_LRI follows dft_functional" << (g ? "" : ", which is not a hybrid at all" + " (no coulomb_param, so the LR exchange kernel vanishes)") << "; set both to the" + " same hybrid to get the kernel you asked for." << std::endl; + } + } +} #endif #ifdef __EXX template<> -void ModuleESolver::ESolver_LR::move_exx_lri(std::shared_ptr>& exx_ks) +void ModuleESolver::ESolver_LR::share_exx_lri(std::shared_ptr>& exx_ks) { - ModuleBase::TITLE("ESolver_LR", "move_exx_lri"); - this->exx_lri = exx_ks; - exx_ks = nullptr; + ModuleBase::TITLE("ESolver_LR", "share_exx_lri"); + this->exx_lri = exx_ks->make_lr_workspace(*this->ucell_, this->kv); } template<> -void ModuleESolver::ESolver_LR>::move_exx_lri(std::shared_ptr>>& exx_ks) +void ModuleESolver::ESolver_LR>::share_exx_lri(std::shared_ptr>>& exx_ks) { - ModuleBase::TITLE("ESolver_LR", "move_exx_lri"); - this->exx_lri = exx_ks; - exx_ks = nullptr; + ModuleBase::TITLE("ESolver_LR", "share_exx_lri"); + this->exx_lri = exx_ks->make_lr_workspace(*this->ucell_, this->kv); } template<> -void ModuleESolver::ESolver_LR>::move_exx_lri(std::shared_ptr>& exx_ks) +void ModuleESolver::ESolver_LR>::share_exx_lri(std::shared_ptr>& exx_ks) { - throw std::runtime_error("ESolver_LR>::move_exx_lri: cannot move double to std::complex"); + throw std::runtime_error("ESolver_LR>::share_exx_lri: cannot share double to std::complex"); } template<> -void ModuleESolver::ESolver_LR::move_exx_lri(std::shared_ptr>>& exx_ks) +void ModuleESolver::ESolver_LR::share_exx_lri(std::shared_ptr>>& exx_ks) { - throw std::runtime_error("ESolver_LR::move_exx_lri: cannot move std::complex to double"); + throw std::runtime_error("ESolver_LR::share_exx_lri: cannot share std::complex to double"); } #endif @@ -63,47 +103,63 @@ int ModuleESolver::ESolver_LR::cal_nupdown_form_occ(const ModuleBase::mat { // only for nspin=2 const int& nk = wg.nr / 2; auto occ_sum_k = [&](const int& is, const int& ib)->double { double o = 0.0; for (int ik = 0;ik < nk;++ik) { o += wg(is * nk + ik, ib); } return o;}; - int nupdown = 0; - for (int ib = 0;ib < wg.nc;++ib) + // Sum the occupations of each channel FIRST and round once, instead of rounding band by band + // and summing the differences. A half-occupied degenerate frontier pair (OH's 2-Pi doublet + // smears its odd electron as 0.5/0.5 over the two pi_down orbitals) otherwise makes the answer + // a coin flip: the stored values are 0.5000000052 and 0.4999999947, so one rounds up and one + // down, and which way they land is pure noise. + double up = 0.0; + double dn = 0.0; + for (int ib = 0;ib < wg.nc;++ib) { up += occ_sum_k(0, ib); dn += occ_sum_k(1, ib); } + // wg is replicated within a pool, but each pool holds different k points. + if (this->kv.para_k.kpar > 1) { - const int nu = static_cast(std::lround(occ_sum_k(0, ib))); - const int nd = static_cast(std::lround(occ_sum_k(1, ib))); - if ((nu + nd) == 0) { break; } - nupdown += nu - nd; + if (this->kv.para_k.rank_in_pool != 0) + { + up = 0.0; + dn = 0.0; + } + Parallel_Reduce::reduce_all(up); + Parallel_Reduce::reduce_all(dn); } - return nupdown; + return static_cast(std::lround(up) - std::lround(dn)); } template void ModuleESolver::ESolver_LR::setup_2center_table(TwoCenterBundle& two_center_bundle, LCAO_Orbitals& orb, UnitCell& ucell) { - // set up 2-center table -#ifdef __FFT_TWO_CENTER - two_center_bundle.tabulate(); -#else - two_center_bundle.tabulate(this->inp_->lcao_ecut, this->inp_->lcao_dk, this->inp_->lcao_dr, this->inp_->lcao_rmax); -#endif if (this->inp_->vnl_in_h) { auto* lcao_nl = new LCAONonlocalInfo(); - lcao_nl->setupNonlocal(ucell.ntype, ucell.atoms, GlobalV::ofs_running, orb, + lcao_nl->setupNonlocal(ucell.ntype, ucell.atoms, this->ofs_running_, orb, this->inp_->basis_type, this->inp_->out_element_info, - this->inp_->lspinorb, this->inp_->nspin, GlobalV::MY_RANK); + this->inp_->lspinorb, this->inp_->nspin, this->my_rank_); ucell.infoNL.reset(lcao_nl); two_center_bundle.build_beta(ucell.ntype, lcao_nl->get_nonlocal().get_Beta_data()); } + // NOTE: tabulate() must be called AFTER build_beta(), otherwise the + // nonlocal (beta) two-center tables are left empty. +#ifdef __FFT_TWO_CENTER + two_center_bundle.tabulate(); +#else + two_center_bundle.tabulate(this->inp_->lcao_ecut, this->inp_->lcao_dk, this->inp_->lcao_dr, this->inp_->lcao_rmax); +#endif } template void ModuleESolver::ESolver_LR::parameter_check()const { const std::set lr_solvers = { "dav", "lapack" , "spectrum", "dav_subspace", "cg", "elpa", "plot" }; - const std::set xc_kernels = { "rpa", "lda", "pwlda", "pbe", "hf", "hse", "bse" }; + // "rpa" and "bse" have no xc kernel at all; everything else is either a (semi)local + // functional or a hybrid, both of which `LR_Util` enumerates. Listing the names a third + // time here is what used to make a newly supported hybrid fail at input parsing. + const std::set kernel_less = { "rpa", "bse" }; const std::set abs_gauge = { "velocity", "length" }; if (lr_solvers.find(this->inp_->lr_solver) == lr_solvers.end()) { throw std::invalid_argument("ESolver_LR: unknown type of lr_solver"); } - if (xc_kernels.find(this->xc_kernel) == xc_kernels.end()) { + if (!kernel_less.count(this->xc_kernel) && !LR_Util::has_local_xc(this->xc_kernel) + && !LR_Util::hybrid_xc_list().count(this->xc_kernel)) { throw std::invalid_argument("ESolver_LR: unknown type of xc_kernel"); } if (this->nspin != 1 && this->nspin != 2) { @@ -112,6 +168,10 @@ void ModuleESolver::ESolver_LR::parameter_check()const if (abs_gauge.find(this->inp_->abs_gauge) == abs_gauge.end()) { throw std::invalid_argument("ESolver_LR: unknown type of abs_gauge"); } + if (this->inp_->cal_force && LR_Util::has_local_xc(this->xc_kernel)) + { + std::cout << "To calculate LR-TDDFT gradients, Libxc should be compiled with kxc, i.e. `-DDISABLE_KXC=OFF` with cmake." << std::endl; + } } template @@ -121,7 +181,29 @@ void ModuleESolver::ESolver_LR::set_dimension() this->nstates = this->inp_->lr_nstates; this->nbasis = PARAM.globalv.nlocal; int ks_nbands = this->inp_->nbands; - this->nocc_max = LR_Util::cal_nocc(LR_Util::cal_nelec(*this->ucell_)); + if (this->nspin == 2) + { + this->nupdown = static_cast(std::lround(this->inp_->nupdown)); + if (this->ks_) + { + this->nupdown = cal_nupdown_form_occ(this->ks_->pelec->wg); + } + else if (this->inp_->ri_hartree_benchmark != "aims" + && this->inp_->ri_hartree_benchmark != "aims-librpa") + { + std::vector populations; + const bool gamma_only = std::is_same::value; + const bool binary = this->inp_->init_wfc_file_format == "binary"; + if (!ModuleIO::read_wfc_nao_spin_populations(this->in_dir, this->kv.get_nkstot(), this->nspin, gamma_only, binary, + this->my_rank_, populations)) + { + ModuleBase::WARNING_QUIT("ESolver_LR", "read complete KS occupations failed"); + } + this->nupdown = static_cast(std::lround(populations[0]) - std::lround(populations[1])); + } + } + // read_pseudo/ParamUpdater has already resolved nelec and applied nelec_delta. + this->nocc_max = LR_Util::cal_nocc(this->inp_->nelec, this->nspin, this->nupdown); if (this->inp_->ri_hartree_benchmark == "aims" || this->inp_->ri_hartree_benchmark == "aims-librpa" && !this->inp_->aims_nbasis.empty()) { @@ -150,9 +232,9 @@ void ModuleESolver::ESolver_LR::set_dimension() } // calculate the number of occupied and unoccupied states // which determines the basis size of the excited states - this->nocc_in = std::max(1, std::min(this->inp_->nocc, this->nocc_max)); + this->nocc_in = LR_Util::cal_nocc_window(this->inp_->nocc, this->nocc_max); this->nvirt_in = ks_nbands - this->nocc_max; //nbands-nocc - if (this->inp_->nvirt > this->nvirt_in) { GlobalV::ofs_running << "ESolver_LR: input nvirt is too large to cover by nbands, set nvirt = nbands - nocc = " << this->nvirt_in << std::endl; } + if (this->inp_->nvirt > this->nvirt_in) { this->ofs_running_ << "ESolver_LR: input nvirt is too large to cover by nbands, set nvirt = nbands - nocc = " << this->nvirt_in << std::endl; } else if (this->inp_->nvirt > 0) { this->nvirt_in = this->inp_->nvirt; } this->nbands = this->nocc_in + this->nvirt_in; this->nk = this->inp_->nspin == 2 ? this->kv.get_nks() / 2 : this->kv.get_nks(); @@ -160,15 +242,18 @@ void ModuleESolver::ESolver_LR::set_dimension() this->nvirt.resize(nspin, nvirt_in); if (this->nstates <= 0) { this->nstates = nk * nocc_in * nvirt_in; - GlobalV::ofs_running << "ESolver_LR: lr_nstates <= 0, set nstates = nk * nocc * nvirt = " << this->nstates << std::endl; + this->ofs_running_ << "ESolver_LR: lr_nstates <= 0, set nstates = nk * nocc * nvirt = " << this->nstates << std::endl; } for (int is = 0;is < nspin;++is) { this->npairs.push_back(nocc[is] * nvirt[is]); } - GlobalV::ofs_running << "Setting LR-TDDFT parameters: " << std::endl; - GlobalV::ofs_running << "number of occupied bands: " << nocc_in << std::endl; - GlobalV::ofs_running << "number of virtual bands: " << nvirt_in << std::endl; - GlobalV::ofs_running << "number of Atom orbitals (LCAO-basis size): " << this->nbasis << std::endl; - GlobalV::ofs_running << "number of KS bands: " << this->eig_ks.nc << std::endl; - GlobalV::ofs_running << "number of excited states to be solved: " << this->nstates << std::endl; + this->ofs_running_ << "Setting LR-TDDFT parameters: " << std::endl; + this->ofs_running_ << "full occupied bands in largest spin channel: " << nocc_max << std::endl; + const int skipped_core_bands = nocc_max - nocc_in; + this->ofs_running_ << "shared core bands omitted from LR window: " << skipped_core_bands << std::endl; + this->ofs_running_ << "number of occupied bands: " << nocc_in << std::endl; + this->ofs_running_ << "number of virtual bands: " << nvirt_in << std::endl; + this->ofs_running_ << "number of Atom orbitals (LCAO-basis size): " << this->nbasis << std::endl; + this->ofs_running_ << "number of KS bands: " << this->eig_ks.nc << std::endl; + this->ofs_running_ << "number of excited states to be solved: " << this->nstates << std::endl; } template @@ -178,14 +263,13 @@ void ModuleESolver::ESolver_LR::reset_dim_spin2() { return; } - if (nupdown == 0) - { - std::cout << " ** Assuming degenerate spin-up and spin-down states **" << std::endl; - } - else + if (nupdown != 0) { this->openshell = true; - nupdown > 0 ? ((nocc[1] -= nupdown) && (nvirt[1] += nupdown)) : ((nocc[0] += nupdown) && (nvirt[0] -= nupdown)); + const int minority_spin = nupdown > 0 ? 1 : 0; + const int polarization = std::abs(nupdown); + nocc[minority_spin] -= polarization; + nvirt[minority_spin] += polarization; npairs = { nocc[0] * nvirt[0], nocc[1] * nvirt[1] }; std::cout << "** Solve the spin-up and spin-down states separately for open-shell system. **" << std::endl; } @@ -205,6 +289,10 @@ void ModuleESolver::ESolver_LR::reset_dim_spin2() { this->openshell = true; } + if (!this->openshell) + { + std::cout << " ** Assuming degenerate spin-up and spin-down states **" << std::endl; + } } template @@ -231,10 +319,13 @@ void ModuleESolver::ESolver_LR::before_all_runners(BaseCell& basecell, co // The embedded KS run happens before Relax_Driver starts its first step. Json::init_output_array_obj(); #endif - ModuleESolver::ESolver_KS_LCAO ks_solver; - ks_solver.before_all_runners(basecell, inp); - ks_solver.runner(basecell, 0); - this->initialize_from_ks_(std::move(ks_solver), ucell, inp); + // the ground-state solver is a member, not a temporary: `runner` re-runs its SCF on every + // ionic step, and the objects aliased from it have to stay alive as long as it does. + // Its SCF is deliberately NOT run here -- `before_all_runners` is called once, outside the + // relaxation loop, so the ground state has to be recomputed from `runner(istep)` instead. + this->ks_ = LR_Util::make_unique>(); + this->ks_->before_all_runners(basecell, inp); + LR_Util::check_force_pp(ucell, inp.cal_force, inp.calculation); } else { @@ -243,190 +334,287 @@ void ModuleESolver::ESolver_LR::before_all_runners(BaseCell& basecell, co } template -void ModuleESolver::ESolver_LR::initialize_from_ks_(ModuleESolver::ESolver_KS_LCAO&& ks_sol, - UnitCell& ucell, - const Input_para& inp) +void ModuleESolver::ESolver_LR::bind_ground_state_aliases_() +{ + if (this->ks_) + { + // `ESolver_FP::before_all_runners` was never called on this object, so its own `pw_rhod`, + // `Pgrid`, `sf` and `locpp` are empty -- the excited-state force needs all four. Point at + // the ground-state solver's, which `before_scf` refreshes for the current geometry every + // ionic step. `pw_rho_flag` stays false so the destructor does not free what it borrowed. + this->pw_rho = this->ks_->pw_rho; + this->pw_rhod = this->ks_->pw_rhod; + this->pw_big = this->ks_->pw_big; + this->pw_rho_flag = false; + this->pgrid_ptr_ = &this->ks_->Pgrid; + this->sf_ptr_ = &this->ks_->sf; + this->locpp_ptr_ = &this->ks_->locpp; + this->gd_ptr_ = &this->ks_->gd; + // the two-center tables depend on the orbitals only, so the ground-state solver's single + // build serves every geometry. This used to be moved over only for `abs_gauge velocity`, + // leaving `LR_Force` and `cal_hs_grad` with an empty bundle in the length gauge. + this->tcb_ptr_ = &this->ks_->two_center_bundle_; + } + else + { + this->pgrid_ptr_ = &this->Pgrid; + this->sf_ptr_ = &this->sf; + this->locpp_ptr_ = &this->locpp; + this->gd_ptr_ = &this->gd_own_; + this->tcb_ptr_ = &this->two_center_bundle_own_; + } +} + +template +void ModuleESolver::ESolver_LR::initialize_from_ks_(UnitCell& ucell, const Input_para& inp) { ModuleBase::TITLE("ESolver_LR", "ESolver_LR(KS)"); + ModuleESolver::ESolver_KS_LCAO& ks_sol = *this->ks_; + this->bind_ground_state_aliases_(); if (this->inp_->lr_solver == "spectrum") { throw std::invalid_argument("when lr_solver==spectrum, esolver_type must be `lr` to skip KS calculation."); } - this->gd = std::move(ks_sol.gd); - // xc kernel this->xc_kernel = LR_Util::tolower(inp.xc_kernel); //kv - this->kv = std::move(ks_sol.kv); + this->kv = ks_sol.kv; // copy: cheap, and the KS solver keeps using its own this->parameter_check(); this->set_dimension(); - // setup_wd_division is not need to be covered in #ifdef __MPI, see its implementation - LR_Util::setup_2d_division(this->paraMat_, 1, this->nbasis, this->nbasis); + // KS and LR must use the same AO distribution, including its BLACS grid. + const int ks_block_size = ks_sol.pv.get_block_size(); + LR_Util::setup_2d_division(this->paraMat_, ks_block_size, this->nbasis, this->nbasis +#ifdef __MPI + , ks_sol.pv.blacs_ctxt +#endif + ); + this->set_parallel_orbitals_band(this->paraMat_, this->nbands); + if (this->inp_->cal_force) + { + LR_Util::setup_2d_division(this->paraMat_all_, ks_block_size, this->nbasis, this->nbasis +#ifdef __MPI + , ks_sol.pv.blacs_ctxt +#endif + ); + this->set_parallel_orbitals_band(this->paraMat_all_, this->inp_->nbands); + } - this->paraMat_.atom_begin_row = std::move(ks_sol.pv.atom_begin_row); - this->paraMat_.atom_begin_col = std::move(ks_sol.pv.atom_begin_col); + this->paraMat_.atom_begin_row = ks_sol.pv.atom_begin_row; + this->paraMat_.atom_begin_col = ks_sol.pv.atom_begin_col; this->paraMat_.iat2iwt_ = ucell.get_iat2iwt(); - LR_Util::setup_2d_division(this->paraC_, 1, this->nbasis, this->nbands + LR_Util::setup_2d_division(this->paraC_, ks_block_size, this->nbasis, this->nbands #ifdef __MPI , this->paraMat_.blacs_ctxt #endif ); - auto move_gs = [&, this]() -> void // move the ground state info - { - this->psi_ks = ks_sol.psi; - ks_sol.psi = nullptr; - //only need the eigenvalues. the 'elecstates' of excited states is different from ground state. - this->eig_ks = std::move(ks_sol.pelec->ekb); - }; + // allocate psi_ks and eig_ks in the [nocc, nvirt] window #ifdef __MPI - if (this->nbands == this->inp_->nbands) - { - move_gs(); - } - else // copy the part of ground state info according to paraC_ - { - this->psi_ks = new psi::Psi(this->kv.get_nks(), - this->paraC_.get_col_size(), - this->paraC_.get_row_size(), - this->kv.ngk, - true); - this->eig_ks.create(this->kv.get_nks(), this->nbands); - const int start_band = this->nocc_max - *std::max_element(nocc.begin(), nocc.end()); - for (int ik = 0;ik < this->kv.get_nks();++ik) - { - Cpxgemr2d(this->nbasis, this->nbands, &(*ks_sol.psi)(ik, 0, 0), 1, start_band + 1, ks_sol.pv.desc_wfc, - &(*this->psi_ks)(ik, 0, 0), 1, 1, this->paraC_.desc, this->paraC_.blacs_ctxt); - for (int ib = 0;ib < this->nbands;++ib) { this->eig_ks(ik, ib) = ks_sol.pelec->ekb(ik, start_band + ib); } - } - } + this->psi_ks.reset(new psi::Psi(this->kv.get_nks(), + this->paraC_.get_col_size(), + this->paraC_.get_row_size(), + this->kv.ngk, + true)); #else - move_gs(); + this->psi_ks.reset(new psi::Psi(this->kv.get_nks(), this->nbands, this->nbasis, this->kv.ngk, true)); #endif - if (nspin == 2) - { - this->nupdown = cal_nupdown_form_occ(ks_sol.pelec->wg); - reset_dim_spin2(); - } - this->gint_info_ = std::move(ks_sol.gint_info_); - // move pw basis - if (this->pw_rho_flag) - { - this->pw_rho_flag = true; - delete this->pw_rho; // newed in ESolver_FP::ESolver_FP - } - this->pw_rho = ks_sol.pw_rho; - ks_sol.pw_rho = nullptr; - //init potential and calculate kernels using ground state charge - init_pot(*ks_sol.pelec->charge); + this->eig_ks.create(this->kv.get_nks(), this->nbands); + this->pelec = new elecstate::ElecStateLCAO(); + orb_cutoff_ = ks_sol.orb_.cutoffs(); #ifdef __EXX - if (xc_kernel == "hf" || xc_kernel == "hse") + // Two independent reasons to need an Exx_LRI: the LR exchange kernel, and the ground-state + // EXX terms of the gradient (the H_gs[T+Z] and W multipliers, gated on gs_is_hybrid). + // `initialize_from_unitcell_` has always covered both; this path used to test only the first, + // so a local kernel on top of a hybrid ground state had no exx_lri at all when asked for forces. + if (exx_kernel_list().count(xc_kernel) || (this->inp_->cal_force && gs_is_hybrid(this->inp_->dft_functional))) { - // if the same kernel is calculated in the esolver_ks, move it std::string dft_functional = LR_Util::tolower(this->inp_->dft_functional); - if (ks_sol.exx_nao.exd && std::is_same::value && xc_kernel == dft_functional) { - this->move_exx_lri(ks_sol.exx_nao.exd->exx_ptr); - } else if (ks_sol.exx_nao.exc && std::is_same>::value && xc_kernel == dft_functional) { - this->move_exx_lri(ks_sol.exx_nao.exc->exx_ptr); - } else // construct C, V from scratch + // Either object would be built from the same `info_ri.coulomb_param`, which `input_conv` + // derives from dft_functional alone -- so whenever the ground-state solver has one of the + // right type its geometry tensors are already up to date. Share those tensors in a + // separate LR electronic workspace, skipping a `cal_exx_ions` per ionic step. + const bool share = (ks_sol.exx_nao.exd && std::is_same::value) + || (ks_sol.exx_nao.exc && std::is_same>::value); + warn_if_kernel_differs_from_gs(xc_kernel, dft_functional, this->ofs_running_); + if (share) { this->exx_owned_ = false; } // `refresh_from_ks_` refreshes the geometry snapshot + else // construct C, V from scratch { - // set ccp_type according to the xc_kernel - if (xc_kernel == "hf") { exx_info.info_global.ccp_type = Conv_Coulomb_Pot_K::Ccp_Type::Hf; } - else if (xc_kernel == "hse") { exx_info.info_global.ccp_type = Conv_Coulomb_Pot_K::Ccp_Type::Erfc; } + // `input_conv` already filled `info_ri.coulomb_param` from INPUT. exx_info.sync_from_global(); // populate ABFs/JLE file lists from UnitCell; keep in sync with Exx_NAO::init exx_info.info_ri.files_abfs = ucell.abfs_orbital_files; exx_info.info_opt_abfs.files_abfs = ucell.abfs_orbital_files; exx_info.info_opt_abfs.files_jles = ucell.jle_orbital_files; this->exx_lri = std::make_shared>(exx_info.info_ri); - this->exx_lri->init(MPI_COMM_WORLD, ucell,this->kv, ks_sol.orb_); - this->exx_lri->cal_exx_ions(ucell,this->inp_->out_ri_cv); + this->exx_lri->init(MPI_COMM_WORLD, ucell, this->kv, ks_sol.orb_); + this->exx_owned_ = true; // the position-dependent `cal_exx_ions` is left to `refresh_from_ks_` } } #endif - this->pelec = new elecstate::ElecStateLCAO(); - orb_cutoff_ = ks_sol.orb_.cutoffs(); - if (LR_Util::tolower(this->inp_->abs_gauge) == "velocity") + + refresh_from_ks_(ucell); + this->ks_initialized_ = true; +} + +template +void ModuleESolver::ESolver_LR::refresh_from_ks_(UnitCell& ucell) +{ + ModuleBase::TITLE("ESolver_LR", "refresh_from_ks_"); + ModuleESolver::ESolver_KS_LCAO& ks_sol = *this->ks_; + // The ground-state solver owns `psi` and reuses it across ionic steps, so it cannot be stolen. + // `eig_ks_all` / `wg_ks_all` are nspin x nbands matrices -- a few kB, copied rather than aliased + // so that they survive the KS solver overwriting `pelec` on the next step. + this->psi_ks_all_ = ks_sol.psi; + this->eig_ks_all = ks_sol.pelec->ekb; + this->wg_ks_all = ks_sol.pelec->wg; + const int start_band = this->nocc_max - *std::max_element(nocc.begin(), nocc.end()); + + for (int ik = 0;ik < this->kv.get_nks();++ik) { - this->two_center_bundle_ = std::move(ks_sol.two_center_bundle_); + // copy the KS orbitals in the [nocc, nvirt] window +#ifdef __MPI + Cpxgemr2d(this->nbasis, this->nbands, &(*this->psi_ks_all_)(ik, 0, 0), 1, start_band + 1, ks_sol.pv.desc_wfc, + &(*this->psi_ks)(ik, 0, 0), 1, 1, this->paraC_.desc, this->paraC_.blacs_ctxt); +#else + // serial: each band is `nbasis` contiguous coefficients, so the window is a plain + // band-by-band copy (this loop used to compute the two pointers and copy nothing, + // leaving `psi_ks` uninitialized in every non-MPI build) + for (int ib = 0;ib < this->nbands;++ib) + { + const auto* start = &(*this->psi_ks_all_)(ik, start_band + ib, 0); + auto* to = &(*this->psi_ks)(ik, ib, 0); + std::copy(start, start + this->nbasis, to); + } +#endif + // copy the KS bands in the [nocc, nvirt] window + for (int ib = 0;ib < this->nbands;++ib) { this->eig_ks(ik, ib) = this->eig_ks_all(ik, start_band + ib); } } + + if (nspin == 2) + { + const int nupdown_now = cal_nupdown_form_occ(ks_sol.pelec->wg); + if (!this->ks_initialized_) + { + this->nupdown = nupdown_now; + reset_dim_spin2(); // shifts nocc/nvirt between the spin channels: must run exactly once + } + else if (nupdown_now != this->nupdown) + { + ModuleBase::WARNING_QUIT("ESolver_LR::refresh_from_ks_", + "the ground-state spin population changed between ionic steps, but nocc/nvirt and" + " the distributions built from them were fixed at the first step."); + } + } + // `gint_info_` stays owned by the KS solver: its `before_scf` rebuilds and re-publishes it + //init potential and calculate kernels using ground state charge + init_pot(*ks_sol.pelec->charge); + +#ifdef __EXX + if (exx_kernel_list().count(xc_kernel) || (this->inp_->cal_force && gs_is_hybrid(this->inp_->dft_functional))) + { + if (this->exx_owned_) + { // Cs/Vs follow the atoms, so they are rebuilt for every geometry + this->exx_lri->cal_exx_ions(ucell, this->inp_->out_ri_cv); + } + else if (ks_sol.exx_nao.exd) { this->share_exx_lri(ks_sol.exx_nao.exd->exx_ptr); } + else if (ks_sol.exx_nao.exc) { this->share_exx_lri(ks_sol.exx_nao.exc->exx_ptr); } + } +#endif + // the grid-integration tables hang off a static pointer that the ground-state solver + // re-publishes in its `before_scf`; make sure it names the object we integrate on + ModuleGint::Gint::set_gint_info(this->ks_->gint_info_.get()); + + // the Z-vector window: after `reset_dim_spin2`, so nocc/nvirt/openshell are final +#ifdef __MPI + this->fill_z_window_(ks_sol.pv.desc_wfc); +#else + this->fill_z_window_(nullptr); +#endif } + template void ModuleESolver::ESolver_LR::initialize_from_unitcell_(UnitCell& ucell, const Input_para& inp) { ModuleBase::TITLE("ESolver_LR", "ESolver_LR(from scratch)"); + this->bind_ground_state_aliases_(); // xc kernel this->xc_kernel = LR_Util::tolower(inp.xc_kernel); // necessary steps in ESolver_FP ESolver_FP::before_all_runners(ucell, inp); + LR_Util::check_force_pp(ucell, inp.cal_force, inp.calculation); this->pelec = new elecstate::ElecStateLCAO(); // necessary steps in ESolver_KS::before_all_runners : symmetry and k-points if (ModuleSymmetry::Symmetry::symm_flag == 1) { const int cal_symm_repr[2] = {this->inp_->cal_symm_repr[0], this->inp_->cal_symm_repr[1]}; - ucell.symm.analy_sys(ucell.lat, ucell.st, ucell.atoms, GlobalV::ofs_running, + ucell.symm.analy_sys(ucell.lat, ucell.st, ucell.atoms, this->ofs_running_, this->inp_->symmetry_prec, this->inp_->nspin, this->inp_->calculation, cal_symm_repr); - ModuleBase::GlobalFunc::DONE(GlobalV::ofs_running, "SYMMETRY"); + ModuleBase::GlobalFunc::DONE(this->ofs_running_, "SYMMETRY"); } const bool use_ibz = false; const bool gamma_only_local = PARAM.globalv.gamma_only_local; const double kspacing[3] = {this->inp_->kspacing[0], this->inp_->kspacing[1], this->inp_->kspacing[2]}; const double koffset[3] = {this->inp_->koffset[0], this->inp_->koffset[1], this->inp_->koffset[2]}; - this->kv.set(ucell, ucell.symm, this->inp_->kpoint_file, this->inp_->nspin, ucell.G, ucell.latvec, GlobalV::ofs_running, GlobalV::ofs_warning, use_ibz, this->out_dir, gamma_only_local, kspacing, this->inp_->kmesh_type, koffset); - ModuleBase::GlobalFunc::DONE(GlobalV::ofs_running, "INIT K-POINTS"); + this->kv.set(ucell, ucell.symm, this->inp_->kpoint_file, this->inp_->nspin, ucell.G, ucell.latvec, this->ofs_running_, this->ofs_warning_, use_ibz, this->out_dir, gamma_only_local, kspacing, this->inp_->kmesh_type, koffset); + ModuleBase::GlobalFunc::DONE(this->ofs_running_, "INIT K-POINTS"); ModuleIO::print_parameters(ucell, this->kv, inp); this->parameter_check(); /// read orbitals and build the interpolation table - two_center_bundle_.build_orb(ucell.ntype, ucell.orbital_fn.data(), inp.orbital_dir); + two_center_bundle_own_.build_orb(ucell.ntype, ucell.orbital_fn.data(), inp.orbital_dir); LCAO_Orbitals orb; - two_center_bundle_.to_LCAO_Orbitals(orb, inp.lcao_ecut, inp.lcao_dk, inp.lcao_dr, inp.lcao_rmax, + two_center_bundle_own_.to_LCAO_Orbitals(orb, inp.lcao_ecut, inp.lcao_dk, inp.lcao_dr, inp.lcao_rmax, inp.out_element_info, inp.cal_force); orb_cutoff_ = orb.cutoffs(); if (LR_Util::tolower(this->inp_->abs_gauge) == "velocity") { - setup_2center_table(this->two_center_bundle_, orb, ucell); + setup_2center_table(this->two_center_bundle_own_, orb, ucell); } this->set_dimension(); // setup 2d-block distribution for AO-matrix and KS wfc LR_Util::setup_2d_division(this->paraMat_, 1, this->nbasis, this->nbasis); -#ifdef __MPI - this->paraMat_.set_desc_wfc_Eij(this->nbasis, this->nbands, paraMat_.get_row_size()); - int err = this->paraMat_.set_nloc_wfc_Eij(this->nbands, GlobalV::ofs_running, GlobalV::ofs_warning); - this->paraMat_.set_atomic_trace(ucell.get_iat2iwt(), ucell.nat, this->nbasis); - if (this->inp_->ri_hartree_benchmark != "aims") { this->paraMat_.set_atomic_trace(ucell.get_iat2iwt(), ucell.nat, this->nbasis); } -#else - this->paraMat_.nrow_bands = this->nbasis; - this->paraMat_.ncol_bands = this->nbands; -#endif + this->set_parallel_orbitals_band(this->paraMat_, this->nbands); + if (this->inp_->cal_force) + { + LR_Util::setup_2d_division(this->paraMat_all_, 1, this->nbasis, this->nbasis); + this->set_parallel_orbitals_band(this->paraMat_all_, this->inp_->nbands); + } // read the ground state info // now ModuleIO::read_wfc_nao needs `Parallel_Orbitals` and can only read all the bands // it need improvement to read only the bands needed - this->psi_ks = new psi::Psi(this->kv.get_nks(), - this->paraMat_.ncol_bands, - this->paraMat_.get_row_size(), - this->kv.ngk, - true); + this->psi_ks.reset(new psi::Psi(this->kv.get_nks(), + this->paraMat_.ncol_bands, + this->paraMat_.get_row_size(), + this->kv.ngk, + true)); this->read_ks_wfc(); + + if (nspin == 2) { - this->nupdown = cal_nupdown_form_occ(this->pelec->wg); + // Complete file populations were read before selecting the occupied window. reset_dim_spin2(); } + // the Z-vector window: after `reset_dim_spin2`, so nocc/nvirt/openshell are final +#ifdef __MPI + this->fill_z_window_(paraMat_all_.desc_wfc); +#else + this->fill_z_window_(nullptr); +#endif LR_Util::setup_2d_division(this->paraC_, 1, this->nbasis, this->nbands #ifdef __MPI @@ -435,10 +623,11 @@ void ModuleESolver::ESolver_LR::initialize_from_unitcell_(UnitCell& ucell ); // clear ks info, new elecstate for excition + delete this->pelec; // the ElecStateLCAO allocated above, only needed while reading the KS data this->pelec = new elecstate::ElecState(); // read the ground state charge density and calculate xc kernel - Pgrid.init(this->pw_rho->nx, + pgrid().init(this->pw_rho->nx, this->pw_rho->ny, this->pw_rho->nz, this->pw_rho->nplane, @@ -452,14 +641,14 @@ void ModuleESolver::ESolver_LR::initialize_from_unitcell_(UnitCell& ucell // search adjacent atoms and init Gint double search_radius = -1.0; - search_radius = atom_arrange::set_sr_NL(GlobalV::ofs_running, + search_radius = atom_arrange::set_sr_NL(this->ofs_running_, this->inp_->out_level, orb.get_rcutmax_Phi(), ucell.infoNL->get_rcutmax_Beta(), PARAM.globalv.gamma_only_local); atom_arrange::search(PARAM.globalv.search_pbc, - GlobalV::ofs_running, - this->gd, + this->ofs_running_, + this->gd(), *this->ucell_, search_radius, this->inp_->test_atom_input); @@ -479,7 +668,7 @@ void ModuleESolver::ESolver_LR::initialize_from_unitcell_(UnitCell& ucell this->pw_big->nbzp, orb.Phi, ucell, - this->gd, + this->gd(), this->inp_->nspin, PARAM.globalv.gamma_only_local, PARAM.globalv.domag, @@ -487,12 +676,15 @@ void ModuleESolver::ESolver_LR::initialize_from_unitcell_(UnitCell& ucell this->inp_->nstream)); ModuleGint::Gint::set_gint_info(gint_info_.get()); // if EXX from scratch, init 2-center integral and calculate Cs, Vs + // when: + // 1. EXX xc_kernel + // 2. cal_force with ground state with EXX functional #ifdef __EXX - if ((xc_kernel == "hf" || xc_kernel == "hse") && this->inp_->lr_solver != "spectrum") + if (((exx_kernel_list().count(xc_kernel)) && this->inp_->lr_solver != "spectrum") + || (this->inp_->cal_force && gs_is_hybrid(this->inp_->dft_functional))) { - // set ccp_type according to the xc_kernel - if (xc_kernel == "hf") { exx_info.info_global.ccp_type = Conv_Coulomb_Pot_K::Ccp_Type::Hf; } - else if (xc_kernel == "hse") { exx_info.info_global.ccp_type = Conv_Coulomb_Pot_K::Ccp_Type::Erfc; } + warn_if_kernel_differs_from_gs(xc_kernel, LR_Util::tolower(this->inp_->dft_functional), this->ofs_running_); + // `input_conv` already filled `info_ri.coulomb_param` from INPUT. exx_info.sync_from_global(); // populate ABFs/JLE file lists from UnitCell; keep in sync with Exx_NAO::init exx_info.info_ri.files_abfs = ucell.abfs_orbital_files; @@ -515,12 +707,37 @@ void ModuleESolver::ESolver_LR::runner(BaseCell& basecell, const int iste ModuleBase::TITLE("ESolver_LR", "runner"); ModuleBase::timer::start("ESolver_LR", "runner"); + + if (this->ks_) + { + // `init_pot_groundstate` leaves the global XC type on the LR kernel, so put it back before + // the ground-state SCF. `ESolver_KS::runner` is re-entrant: its `before_scf` rebuilds the + // neighbour lists, the grid tables and the Hamiltonian for the geometry of this step. + XC_Functional::set_xc_type(ucell.atoms[0].ncpp.xc_func); + this->ks_->runner(ucell, istep); + this->etot_gs_ = this->ks_->cal_energy(); + if (!this->ks_initialized_) + { + this->initialize_from_ks_(ucell, *this->inp_); + this->setup_relax_target_(); // needs the final dimensions, i.e. `openshell` + } + else { this->refresh_from_ks_(ucell); } + } + //allocate 2-particle state and setup 2d division this->setup_eigenvectors_X(); this->pelec->ekb.create(nspin, this->nstates); - auto efile_out = [&](const std::string& label)->std::string {return this->out_dir + "Excitation_Energy_" + label + ".dat";}; - auto vfile_out = [&](const std::string& label)->std::string {return this->out_dir + "Excitation_Amplitude_" + label + "_" + std::to_string(GlobalV::MY_RANK+1) + ".dat";}; + // set once here rather than inside the solver branches: the open-shell `spectrum` branch used + // to leave it empty, and both `after_all_runners` and the gradient index it + this->spin_types = this->openshell ? std::vector({ "updown" }) + : std::vector({ "singlet", "triplet" }); + + // a relaxation writes these once per ionic step; without the suffix every step would overwrite + // the last, and the trajectory would be impossible to inspect afterwards + const std::string step_suffix = this->excited_relax_ ? "_step" + std::to_string(istep) : ""; + auto efile_out = [&](const std::string& label)->std::string {return this->out_dir + "Excitation_Energy_" + label + step_suffix + ".dat";}; + auto vfile_out = [&](const std::string& label)->std::string {return this->out_dir + "Excitation_Amplitude_" + label + step_suffix + "_" + std::to_string(this->my_rank_+1) + ".dat";}; if (this->inp_->lr_solver == "elpa") { ModuleBase::WARNING_QUIT("ESolver_LR", "ESolver_LR doesn't support elpa now."); @@ -530,7 +747,7 @@ void ModuleESolver::ESolver_LR::runner(BaseCell& basecell, const int iste { auto write_states = [&](const std::string& label, const Real* e, const T* v, const int& dim, const int& nst, const int& prec = 8)->void { - if (GlobalV::MY_RANK == 0) { assert(nst == LR_Util::write_value(efile_out(label), prec, e, nst)); } + if (this->my_rank_ == 0) { assert(nst == LR_Util::write_value(efile_out(label), prec, e, nst)); } assert(nst * dim == LR_Util::write_value(vfile_out(label), prec, v, nst, dim)); }; std::vector precondition(this->inp_->lr_solver == "lapack" ? 0 : nloc_per_state, 1.0); @@ -553,7 +770,7 @@ void ModuleESolver::ESolver_LR::runner(BaseCell& basecell, const int iste this->nvirt, *this->ucell_, orb_cutoff_, - this->gd, + this->gd(), *this->psi_ks, this->eig_ks, #ifdef __EXX @@ -577,7 +794,7 @@ void ModuleESolver::ESolver_LR::runner(BaseCell& basecell, const int iste OperatorLRDiag pre_op(this->eig_ks.c, this->paraX_[0], this->nk, this->nocc[0], this->nvirt[0]); pre_op.act(1, nloc_per_state, 1, precondition.data(), precondition.data()); } - auto spin_types = std::vector({ "singlet", "triplet" }); + const std::vector& spin_types = this->spin_types; for (int is = 0;is < nspin;++is) { std::cout << " Calculating " << spin_types[is] << " excitations" << std::endl; @@ -588,7 +805,7 @@ void ModuleESolver::ESolver_LR::runner(BaseCell& basecell, const int iste this->nvirt, *this->ucell_, orb_cutoff_, - this->gd, + this->gd(), *this->psi_ks, this->eig_ks, #ifdef __EXX @@ -618,20 +835,20 @@ void ModuleESolver::ESolver_LR::runner(BaseCell& basecell, const int iste else // lr_solver == "spectrum", read the eigenvalues { auto efile_in = [&](const std::string& label)->std::string {return this->in_dir + "Excitation_Energy_" + label + ".dat";}; - auto vfile_in = [&](const std::string& label)->std::string {return this->in_dir + "Excitation_Amplitude_" + label + "_" + std::to_string(GlobalV::MY_RANK+1) + ".dat";}; + auto vfile_in = [&](const std::string& label)->std::string {return this->in_dir + "Excitation_Amplitude_" + label + "_" + std::to_string(this->my_rank_+1) + ".dat";}; auto read_states = [&](const std::string& label, Real* e, T* v, const int& dim, const int& nst)->void { - if (GlobalV::MY_RANK == 0) { + if (this->my_rank_ == 0) { assert(nst == LR_Util::read_value(efile_in(label), e, nst)); - std::cout <<"Rank "<< GlobalV::MY_RANK << ": finish reading " << efile_in(label) << std::endl; + std::cout <<"Rank "<< this->my_rank_ << ": finish reading " << efile_in(label) << std::endl; } #ifdef __MPI // in velocity gauge, the eigenvalues may be used to calculate the transition dipole, so we'd better broadcast them MPI_Bcast(e, nst, MPI_DOUBLE, 0, MPI_COMM_WORLD); #endif assert(nst * dim == LR_Util::read_value(vfile_in(label), v, nst, dim)); - std::cout <<"Rank "<< GlobalV::MY_RANK << ": finish reading " << vfile_in(label) << std::endl; + std::cout <<"Rank "<< this->my_rank_ << ": finish reading " << vfile_in(label) << std::endl; }; std::cout << "reading the excitation states from file: \n"; if (openshell) @@ -640,10 +857,38 @@ void ModuleESolver::ESolver_LR::runner(BaseCell& basecell, const int iste } else { - auto spin_types = std::vector({ "singlet", "triplet" }); + const std::vector& spin_types = this->spin_types; for (int is = 0;is < nspin;++is) { read_states(spin_types[is], this->pelec->ekb.c + is * nstates, this->X[is].template data(), nloc_per_state, nstates); } } } + if (this->excited_relax_) + { + // Re-select which root to follow BEFORE taking the gradient, so the force belongs to + // the same diabatic state as the previous step's. + this->follow_target_state_(this->ofs_running_); + + // The LR terms only carry the Omega part of the force; the ground-state part is separate + // and comes straight from the KS solver. + this->ks_->cal_force(ucell, this->force_gs_); + this->lr_force_ = this->cal_lr_force_relax_(this->ofs_running_); + + // One line per ionic step with the two halves of the energy and of the gradient. Without + // it the relaxation only reports a force, and whether E_gs + Omega actually goes down -- + // the thing being minimised -- cannot be read off the log at all. + const double omega = this->target_omega_(); + auto max_abs = [](const ModuleBase::matrix& m) -> double + { double v = 0.0; for (int i = 0;i < m.nr * m.nc;++i) { v = std::max(v, std::abs(m.c[i])); } return v; }; + this->ofs_running_ << std::setprecision(8) << std::fixed + << " EXCITED-STATE RELAX step " << istep + << ": E_gs = " << this->etot_gs_ * ModuleBase::Ry_to_eV + << " eV, Omega = " << omega * ModuleBase::Ry_to_eV + << " eV, E_exc = " << (this->etot_gs_ + omega) * ModuleBase::Ry_to_eV << " eV" + << std::setprecision(6) + << " | |F_gs|max = " << max_abs(this->force_gs_) * ModuleBase::Ry_to_eV / ModuleBase::BOHR_TO_A + << ", |F_Omega|max = " << max_abs(this->lr_force_) * ModuleBase::Ry_to_eV / ModuleBase::BOHR_TO_A + << " eV/Angstrom" << std::defaultfloat << std::endl; + } + ModuleBase::timer::end("ESolver_LR", "runner"); return; } @@ -656,11 +901,27 @@ void ModuleESolver::ESolver_LR::after_all_runners(BaseCell& basecell) ModuleBase::TITLE("ESolver_LR", "after_all_runners"); if (this->inp_->ri_hartree_benchmark != "none") { return; } //no need to calculate the spectrum in the benchmark routine + + // cal electron-hole density + if (this->inp_->out_chg[0]) + { + LR_Density lr_density(*this->ucell_, kv, gd(), *psi_ks, orb_cutoff_, pgrid(), + nspin, nocc, nvirt, nbasis, + paraX_, paraC_, paraMat_, openshell); + + if (openshell) + for (int is = 0;is < this->nspin;++is) + lr_density.output_eh_density_all_states(this->X[0].template data(), is, nstates); + else + for (int is = 0;is < this->X.size();++is) + lr_density.output_eh_density_all_states(this->X[is].template data(), is, nstates); + } + //cal spectrum if (LR_Util::tolower(this->inp_->abs_gauge) == "velocity" ) { const int nspin_tmp = this->inp_->nspin == 2 ? 2 : 1; - this->velocity_mo = LR_Util::cal_velocity_mo(*this->ucell_, this->gd, this->two_center_bundle_, + this->velocity_mo = LR_Util::cal_velocity_mo(*this->ucell_, this->gd(), this->tcb(), this->paraMat_, this->paraC_, this->kv, *this->psi_ks, this->nk, nspin_tmp, this->nbasis, this->nocc, this->nvirt); } @@ -673,14 +934,15 @@ void ModuleESolver::ESolver_LR::after_all_runners(BaseCell& basecell) double lambda_diff = std::abs(abs_wavelen_range[1] - abs_wavelen_range[0]); double lambda_min = std::min(abs_wavelen_range[1], abs_wavelen_range[0]); for (int i = 0;i < freq.size();++i) { freq[i] = 91.126664 / (lambda_min + 0.01 * static_cast(i + 1) * lambda_diff); } - auto spin_types = (nspin == 2 && !openshell) ? std::vector({ "singlet", "triplet" }) : std::vector({ "updown" }); + // auto spin_types = (nspin == 2 && !openshell) ? std::vector({ "singlet", "triplet" }) : std::vector({ "updown" }); + // for (int is = 0;is < this->X.size() - 1;++is) for (int is = 0;is < this->X.size();++is) { LR_Spectrum spectrum(nspin, this->nbasis, this->nocc, this->nvirt, *this->pw_rho, *this->psi_ks, - *this->ucell_, this->kv, this->gd, this->orb_cutoff_, this->two_center_bundle_, + *this->ucell_, this->kv, this->gd(), this->orb_cutoff_, this->tcb(), this->paraX_, this->paraC_, this->paraMat_, &this->pelec->ekb.c[is * nstates], this->eig_ks.c, this->X[is].template data(), nstates, openshell, - LR_Util::tolower(this->inp_->abs_gauge), GlobalV::MY_RANK, this->out_dir); + LR_Util::tolower(this->inp_->abs_gauge), this->my_rank_, this->out_dir); if (LR_Util::tolower(this->inp_->abs_gauge) == "velocity" ) {spectrum.set_vmo(this->velocity_mo.data());} spectrum.cal_spectrum(); spectrum.transition_analysis(spin_types[is]+"_tda"); @@ -699,17 +961,41 @@ void ModuleESolver::ESolver_LR::after_all_runners(BaseCell& basecell) // } // =============================================== for test ==================================================== } + if (this->inp_->cal_force && !this->excited_relax_) { this->cal_force_and_grad_matrix_(is, this->ofs_running_); } } } +template +void ModuleESolver::ESolver_LR::set_parallel_orbitals_band(Parallel_Orbitals& pmat, const int nbands_in) +{ +#ifdef __MPI + pmat.set_desc_wfc_Eij(this->nbasis, nbands_in, pmat.get_row_size()); + int err = pmat.set_nloc_wfc_Eij(nbands_in, this->ofs_running_, this->ofs_warning_); + // Skipped for the aims benchmark: with `aims_nbasis` the per-atom orbital counts behind + // `iat2iwt` do not match `nbasis`, so the atomic trace would be wrong. The guard came from + // d4fe3fe84 ("Support different basis number from aims"), and was silently undone by + // 1e4c1c6af (BSE, #7718) re-adding an unconditional call above it. + if (this->inp_->ri_hartree_benchmark != "aims") + { + pmat.set_atomic_trace(this->ucell_->get_iat2iwt(), this->ucell_->nat, this->nbasis); + } +#else + pmat.nrow_bands = this->nbasis; + pmat.ncol_bands = nbands_in; +#endif +} template void ModuleESolver::ESolver_LR::setup_eigenvectors_X() { ModuleBase::TITLE("ESolver_LR", "setup_eigenvectors_X"); + // this function is called once per `runner`, and `paraX_` is only ever appended to, + // so without this reset a second ionic step would double its size + this->paraX_.clear(); + const int block_size = this->paraC_.get_block_size(); for (int is = 0;is < nspin;++is) { Parallel_2D px; - LR_Util::setup_2d_division(px, /*nb2d=*/1, this->nvirt[is], this->nocc[is] + LR_Util::setup_2d_division(px, block_size, this->nvirt[is], this->nocc[is] #ifdef __MPI , this->paraC_.blacs_ctxt #endif @@ -721,7 +1007,6 @@ void ModuleESolver::ESolver_LR::setup_eigenvectors_X() this->X.resize(openshell ? 1 : nspin, LR_Util::newTensor({ nstates, nloc_per_state })); for (auto& x : X) { x.zero(); } - auto spin_types = (nspin == 2 && !openshell) ? std::vector({ "singlet", "triplet" }) : std::vector({ "updown" }); // if spectrum-only, read the LR-eigenstates from file and return if (this->inp_->lr_solver != "spectrum") { set_X_initial_guess(); } } @@ -740,10 +1025,10 @@ void ModuleESolver::ESolver_LR::set_X_initial_guess() // if (E_{lumo}-E_{homo-1} < E_{lumo+1}-E{homo}), mode = 0, else 1(smaller first) bool ix_mode = false; //default if (this->eig_ks.nc > no + 1 && no >= 2 && eig_ks(is, no) - eig_ks(is, no - 2) - 1e-5 > eig_ks(is, no + 1) - eig_ks(is, no - 1)) { ix_mode = true; } - GlobalV::ofs_running << "setting the initial guess of X of spin" << is << std::endl; - if (no >= 2 && eig_ks.nc > no) { GlobalV::ofs_running << "E_{lumo}-E_{homo-1}=" << eig_ks(is, no) - eig_ks(is, no - 2) << std::endl; } - if (no >= 1 && eig_ks.nc > no + 1) { GlobalV::ofs_running << "E_{lumo+1}-E{homo}=" << eig_ks(is, no + 1) - eig_ks(is, no - 1) << std::endl; } - GlobalV::ofs_running << "mode of X-index: " << ix_mode << std::endl; + this->ofs_running_ << "setting the initial guess of X of spin" << is << std::endl; + if (no >= 2 && eig_ks.nc > no) { this->ofs_running_ << "E_{lumo}-E_{homo-1}=" << eig_ks(is, no) - eig_ks(is, no - 2) << std::endl; } + if (no >= 1 && eig_ks.nc > no + 1) { this->ofs_running_ << "E_{lumo+1}-E{homo}=" << eig_ks(is, no + 1) - eig_ks(is, no - 1) << std::endl; } + this->ofs_running_ << "mode of X-index: " << ix_mode << std::endl; /// global index map between (i,c) and ix ModuleBase::matrix ioiv2ix; @@ -776,30 +1061,70 @@ void ModuleESolver::ESolver_LR::set_X_initial_guess() template void ModuleESolver::ESolver_LR::init_pot(const Charge& chg_gs) { - this->pot.resize(nspin, nullptr); + using ST = PotHxcLR::SpinType; + using GX = LR::KernelXC::GxcSpin; + this->pot.assign(nspin, nullptr); // assign, not resize: a re-init must drop the previous geometry's if (this->inp_->ri_hartree_benchmark != "none") { return; } //no need to initialize potential for Hxc kernel in the RI-benchmark routine + + // The singlet and triplet potentials evaluate the *same* kernel arrays and differ only in which + // spin combination of them they read, so they share one `KernelXC` to save memory. + const bool oshell = (nspin == 2) && openshell; + // $g^{xc}$ (third-order) is only ever needed by the LR gradient, and only for the spin + // combinations that are actually going to be requested. + const int gxc_lr = (!this->inp_->cal_force || !LR_Util::has_local_xc(xc_kernel)) ? GX::NoGxc + : ((nspin == 1) ? GX::Singlet : GX::BothSpins); + std::shared_ptr kernel_lr = PotHxcLR::make_kernel( + xc_kernel, *this->pw_rho, *this->ucell_, chg_gs, pgrid(), oshell, gxc_lr, this->inp_->lr_init_xc_kernel); switch (nspin) { - using ST = PotHxcLR::SpinType; case 1: - this->pot[0] = std::make_shared(xc_kernel, *this->pw_rho, *this->ucell_, chg_gs, Pgrid, ST::S1, this->inp_->lr_init_xc_kernel); + this->pot[0] = std::make_shared(kernel_lr, xc_kernel, *this->pw_rho, *this->ucell_, chg_gs.nrxx, ST::S1); break; case 2: - this->pot[0] = std::make_shared(xc_kernel, *this->pw_rho, *this->ucell_, chg_gs, Pgrid, openshell ? ST::S2_updown : ST::S2_singlet, this->inp_->lr_init_xc_kernel); - this->pot[1] = std::make_shared(xc_kernel, *this->pw_rho, *this->ucell_, chg_gs, Pgrid, openshell ? ST::S2_updown : ST::S2_triplet, this->inp_->lr_init_xc_kernel); + this->pot[0] = std::make_shared(kernel_lr, xc_kernel, *this->pw_rho, *this->ucell_, chg_gs.nrxx, oshell ? ST::S2_updown : ST::S2_singlet); + this->pot[1] = std::make_shared(kernel_lr, xc_kernel, *this->pw_rho, *this->ucell_, chg_gs.nrxx, oshell ? ST::S2_updown : ST::S2_triplet); break; default: throw std::invalid_argument("ESolver_LR: nspin must be 1 or 2"); } + // ground-state potentials are needed for calculating the excited state force + if (this->inp_->cal_force) + { + this->init_pot_groundstate(chg_gs); + // `dft_functional == "default"` leaves the raw INPUT string unresolved (it never gets + // overwritten to the actual functional in use); the functional actually read from the + // pseudopotential lives in `ucell.atoms[i].ncpp.xc_func` instead. Comparing against the + // literal "default" string here would always disagree with `xc_kernel`, even when the + // ground state and LR use the same functional. Resolve it before comparing the names + // or constructing a separate GS kernel. + const std::string xc_kernel_gs = (this->inp_->dft_functional == "default") + ? LR_Util::tolower(this->ucell_->atoms[0].ncpp.xc_func) + : LR_Util::tolower(this->inp_->dft_functional); + // `ST::S1` is only correct when nspin=1. `PotHxcLR` builds its `KernelXC` with + // the input `nspin`, so at nspin=2 the kernel arrays carry 3 spin components per grid point + // while the S1 integrand indexes them as if there were 1 -- it does not even read a + // consistent spin combination. Use `ST::S2_gs` there, which is exactly half of S2_singlet, + // matching the `K_Hxc(singlet) = 2 * pot_hxc_gs` convention of the gradient operators. + const ST st_gs = (nspin == 1) ? ST::S1_gs : (oshell ? ST::S2_updown : ST::S2_gs); + // GS response terms retain this kernel. The g^xc terms originating from the LR K + // contribution to W and the force use the corresponding LR potential instead. + const int gxc_gs = (!LR_Util::has_local_xc(xc_kernel) || !LR_Util::has_local_xc(xc_kernel_gs)) ? GX::NoGxc + : ((nspin == 1) ? GX::Singlet : GX::BothSpins); + // When the LR kernel *is* the ground-state functional -- the usual TDDFT case -- the two + // `KernelXC` are bit-for-bit identical, so reuse the one just built. + const bool share_lr = (xc_kernel_gs == xc_kernel) && ((gxc_lr & gxc_gs) == gxc_gs); + std::shared_ptr kernel_gs = share_lr ? kernel_lr + : PotHxcLR::make_kernel(xc_kernel_gs, *this->pw_rho, *this->ucell_, chg_gs, pgrid(), oshell, gxc_gs, this->inp_->lr_init_xc_kernel); + this->pot_hxc_gs = std::make_shared(kernel_gs, xc_kernel_gs, *this->pw_rho, *this->ucell_, chg_gs.nrxx, st_gs); + } } template void ModuleESolver::ESolver_LR::read_ks_wfc() { assert(this->psi_ks != nullptr); - this->pelec->ekb.create(this->kv.get_nks(), this->nbands); - this->pelec->wg.create(this->kv.get_nks(), this->nbands); - + this->eig_ks.create(this->kv.get_nks(), this->nbands); + this->wg_ks.create(this->kv.get_nks(), this->nbands); if (this->inp_->ri_hartree_benchmark == "aims") // for aims benchmark { #ifdef __EXX @@ -808,23 +1133,129 @@ void ModuleESolver::ESolver_LR::read_ks_wfc() std::cout << "ncore=" << ncore << ", nocc=" << nocc_in << ", nvirt=" << nvirt_in << ", nbands=" << this->nbands << std::endl; std::cout << "eig_ks_vec.size()=" << eig_ks_vec.size() << std::endl; if(eig_ks_vec.size() != this->nbands) {ModuleBase::WARNING_QUIT("ESolver_LR", "read_aims_ebands failed.");}; - for (int i = 0;i < nbands;++i) { this->pelec->ekb(0, i) = eig_ks_vec[i]; } + for (int i = 0;i < nbands;++i) { this->eig_ks(0, i) = eig_ks_vec[i]; } RI_Benchmark::read_aims_eigenvectors(*this->psi_ks, this->in_dir + "KS_eigenvectors.out", ncore, nbands, nbasis); #else ModuleBase::WARNING_QUIT("ESolver_LR", "RI benchmark is only supported when compile with LibRI."); #endif } - else if (!ModuleIO::read_wfc_nao(this->in_dir, this->paraMat_, *this->psi_ks, - this->pelec->ekb, - this->pelec->wg, - this->kv.ik2iktot, - this->kv.get_nkstot(), + else if (!ModuleIO::read_wfc_nao(this->in_dir, this->paraMat_, *this->psi_ks, + this->eig_ks, + this->wg_ks, + this->kv.ik2iktot, + this->kv.get_nkstot(), this->inp_->nspin, - this->inp_->init_wfc_file_format == "binary", - /*skip_bands=*/this->nocc_max - this->nocc_in)) { + this->inp_->init_wfc_file_format == "binary", + /*skip_bands=*/this->nocc_max - this->nocc_in)) { ModuleBase::WARNING_QUIT("ESolver_LR", "read ground-state wavefunction failed."); } - this->eig_ks = std::move(this->pelec->ekb); + + if (this->inp_->cal_force) + { // allocate psi_ks_all and eig_ks_all to read all the bands + this->psi_ks_all_own_.reset(new psi::Psi(this->kv.get_nks(), paraMat_all_.ncol_bands, paraMat_all_.get_row_size(), this->kv.ngk, true)); + this->psi_ks_all_ = this->psi_ks_all_own_.get(); + this->eig_ks_all.create(this->kv.get_nks(), this->inp_->nbands); + this->wg_ks_all.create(this->kv.get_nks(), this->inp_->nbands); + if (!ModuleIO::read_wfc_nao(this->in_dir, paraMat_all_, *this->psi_ks_all_, + this->eig_ks_all, + this->wg_ks_all, + this->kv.ik2iktot, + this->kv.get_nkstot(), + this->inp_->nspin, + this->inp_->init_wfc_file_format == "binary", + /*skip_bands=*/0)) + { + ModuleBase::WARNING_QUIT("ESolver_LR", "read all ground-state wavefunctions for force calculation failed."); + } + this->ofs_running_ << " Read in all the KS wavefunctions for force calculation. " << std::endl; + } +} + +template +void ModuleESolver::ESolver_LR::fill_z_window_(const int* desc_src) +{ + ModuleBase::TITLE("ESolver_LR", "fill_z_window_"); + if (!this->inp_->cal_force || this->psi_ks_all_ == nullptr) { return; } + + const int start_band = this->nocc_max - *std::max_element(nocc.begin(), nocc.end()); + // Every band the ground state solved, from the window start upward. `eig_ks_all` is the + // authority on how many there are: on the ks-lr path it is the KS solver's `ekb`, on the + // file path it was read with `skip_bands = 0`. + this->nbands_z_ = this->eig_ks_all.nc - start_band; + if (this->nbands_z_ <= this->nbands) + { // nothing to widen: nbands was already at (or below) the X window + this->nbands_z_ = this->nbands; + } + this->nvirt_z_.assign(this->nspin, 0); + for (int is = 0; is < this->nspin; ++is) { this->nvirt_z_[is] = this->nbands_z_ - this->nocc[is]; } + + const int block_size = this->paraC_.get_block_size(); + LR_Util::setup_2d_division(this->paraC_z_, block_size, this->nbasis, this->nbands_z_ +#ifdef __MPI + , this->paraMat_.blacs_ctxt +#endif + ); + this->paraX_z_.clear(); + for (int is = 0; is < this->nspin; ++is) + { + Parallel_2D px; + LR_Util::setup_2d_division(px, block_size, this->nvirt_z_[is], this->nocc[is] +#ifdef __MPI + , this->paraC_z_.blacs_ctxt +#endif + ); + this->paraX_z_.emplace_back(std::move(px)); + } + // Open shell solves one eigenproblem whose vector is the concatenation [up | down], so a + // state's block is as long as both channels together; the closed-shell singlet/triplet + // algorithm carries a single channel. (Same rule as `nloc_per_state` in + // `setup_eigenvectors_X`, applied to the widened windows.) + this->nloc_per_state_z_ = this->nk * (this->openshell + ? this->paraX_z_[0].get_local_size() + this->paraX_z_[1].get_local_size() + : this->paraX_z_[0].get_local_size()); + +#ifdef __MPI + this->psi_ks_z_.reset(new psi::Psi(this->kv.get_nks(), this->paraC_z_.get_col_size(), + this->paraC_z_.get_row_size(), this->kv.ngk, true)); +#else + this->psi_ks_z_.reset(new psi::Psi(this->kv.get_nks(), this->nbands_z_, this->nbasis, this->kv.ngk, true)); +#endif + this->eig_ks_z_.create(this->kv.get_nks(), this->nbands_z_); + + for (int ik = 0; ik < this->kv.get_nks(); ++ik) + { + // same redistribution `refresh_from_ks_` does for the X window, over more bands +#ifdef __MPI + Cpxgemr2d(this->nbasis, this->nbands_z_, &(*this->psi_ks_all_)(ik, 0, 0), 1, start_band + 1, + const_cast(desc_src), &(*this->psi_ks_z_)(ik, 0, 0), 1, 1, + this->paraC_z_.desc, this->paraC_z_.blacs_ctxt); +#else + for (int ib = 0; ib < this->nbands_z_; ++ib) + { + const auto* start = &(*this->psi_ks_all_)(ik, start_band + ib, 0); + std::copy(start, start + this->nbasis, &(*this->psi_ks_z_)(ik, ib, 0)); + } +#endif + for (int ib = 0; ib < this->nbands_z_; ++ib) + { this->eig_ks_z_(ik, ib) = this->eig_ks_all(ik, start_band + ib); } + } + this->ofs_running_ << "Z-vector window: nbands = " << this->nbands_z_ + << " (X window: " << this->nbands << "), nvirt ="; + for (int is = 0; is < this->nspin; ++is) + { this->ofs_running_ << " " << this->nvirt_z_[is] << "(X window: " << this->nvirt[is] << ")"; } + this->ofs_running_ << std::endl; + if (this->nbands_z_ < this->nbasis) + { + // The Z window can only be as wide as the ground state's band count, so a gradient is + // converged in it only when `nbands` reaches the size of the AO basis. Measured on + // `08_BeH2/rpa_at_lda` (NLOCAL 17), where Omega = eps_a - eps_i makes the finite + // difference exact: nbands 11 -> 18% too high, nbands 17 -> 6 digits. + this->ofs_running_ << " WARNING: the excited-state gradient is not converged with" + " respect to the Z-vector (CPSCF) space: nbands = " << this->nbands_z_ + << " covers only part of the " << this->nbasis << " AO basis functions (NLOCAL)." + " Set nbands = " << this->nbasis << " for a converged gradient; nvirt may stay as" + " it is, since the excitation energies do not depend on this." << std::endl; + } } template @@ -833,19 +1264,20 @@ void ModuleESolver::ESolver_LR::read_ks_chg(Charge& chg_gs) chg_gs.set_rhopw(this->pw_rho); const bool kin_den = XC_Functional::get_ked_flag() || (this->inp_->out_elf[0] > 0); // mohan add 20251202 chg_gs.allocate(this->nspin, kin_den, XC_Functional::get_ked_flag(), this->inp_->test_charge); - GlobalV::ofs_running << " try to read charge from file : "; + this->ofs_running_ << " try to read charge from file : "; for (int is = 0; is < this->nspin; ++is) { std::stringstream ssc; - ssc << this->in_dir << "chgs" << is + 1 << ".cube"; - GlobalV::ofs_running << ssc.str() << std::endl; - ModuleIO::read_vdata_palgrid(Pgrid, - GlobalV::MY_RANK, - GlobalV::ofs_running, + if (this->nspin == 1) { ssc << this->in_dir << "chg.cube"; } + else { ssc << this->in_dir << "chgs" << is + 1 << ".cube"; } + this->ofs_running_ << ssc.str() << std::endl; + if (ModuleIO::read_vdata_palgrid(pgrid(), + this->my_rank_, + this->ofs_running_, ssc.str(), chg_gs.rho[is], - this->ucell_->nat); - GlobalV::ofs_running << " Read in the charge density: " << ssc.str() << std::endl; + this->ucell_->nat)) + this->ofs_running_ << " Read in the charge density: " << ssc.str() << std::endl; } } template class ModuleESolver::ESolver_LR; diff --git a/source/source_esolver/esolver_lr_lcao_tddft.h b/source/source_esolver/esolver_lr_lcao_tddft.h index ddaec8b8358..6789abafeda 100644 --- a/source/source_esolver/esolver_lr_lcao_tddft.h +++ b/source/source_esolver/esolver_lr_lcao_tddft.h @@ -17,12 +17,15 @@ #include "source_estate/module_dm/density_matrix.h" #include "source_lcao/module_lr/potentials/pot_hxc_lrtd.h" #include "source_lcao/module_lr/hamilt_casida.h" +#include "source_lcao/module_lr/root_track.h" #include "source_hamilt/module_gint/gint_info.h" +#include "source_estate/module_pot/potential_new.h" #ifdef __EXX // #include #include "source_lcao/module_ri/exx_lri.h" #include "source_hamilt/module_xc/exx_info.h" // for Exx_Info value member #endif +namespace LR { template struct GradientInputs; } namespace ModuleESolver { ///Excited State Solver: Linear Response TDDFT (Tamm Dancoff Approximation) @@ -31,9 +34,7 @@ namespace ModuleESolver { public: ESolver_LR(const Input_para& inp, const std::string& in_dir, const std::string& out_dir); - ~ESolver_LR() { - delete this->psi_ks; - } + ~ESolver_LR() {} ///input: input, call, basis(LCAO), psi(ground state), elecstate // initialize sth. independent of the ground state @@ -41,25 +42,50 @@ namespace ModuleESolver virtual void runner(BaseCell& basecell, int istep) override; virtual void after_all_runners(BaseCell& basecell) override; - virtual double cal_energy() override { return 0.0; }; - virtual void cal_force(BaseCell& basecell, ModuleBase::matrix& force) override - { - static_cast(force); - basecell.require_kind(BaseCell::Kind::unitcell, __FUNCTION__); - }; - virtual void cal_stress(BaseCell& basecell, ModuleBase::matrix& stress) override - { - static_cast(stress); - basecell.require_kind(BaseCell::Kind::unitcell, __FUNCTION__); - }; + /// Total energy of the excited state being relaxed: E_gs + Omega. Zero outside a + /// relaxation, where nothing consumes it and the old behaviour is kept. + virtual double cal_energy() override; + /// Force on the atoms in the relaxed excited state, F = -d(E_gs + Omega)/dR (Ry/Bohr). + virtual void cal_force(BaseCell& basecell, ModuleBase::matrix& force) override; + /// Not implemented: there is no excited-state stress. `cell-relax` is rejected at input. + virtual void cal_stress(BaseCell& basecell, ModuleBase::matrix& stress) override; protected: const std::string in_dir; const std::string out_dir; const UnitCell* ucell_ = nullptr; - Grid_Driver gd; std::vector orb_cutoff_; + /// cached aliases, read once here instead of at every log/print call site below + std::ofstream& ofs_running_ = GlobalV::ofs_running; + std::ofstream& ofs_warning_ = GlobalV::ofs_warning; + const int my_rank_ = GlobalV::MY_RANK; + + /// @brief the ground-state solver, kept alive across ionic steps (esolver_type = "ks-lr"). + /// Null on the `lr` path, where the ground state comes from files instead. + std::unique_ptr> ks_; + + // Geometry-dependent objects that the ground-state solver already builds for the current + // structure. They are aliased rather than rebuilt: `ESolver_KS_LCAO::before_scf` refreshes + // its own copies every ionic step, so recomputing them here would be both wasted work and + // a chance for the two to disagree. On the `lr` path the pointers are bound to this + // object's own members below, which `initialize_from_unitcell_` fills from file. + Grid_Driver gd_own_; ///< only used when `ks_` is null + TwoCenterBundle two_center_bundle_own_; ///< only used when `ks_` is null + Grid_Driver* gd_ptr_ = nullptr; + const TwoCenterBundle* tcb_ptr_ = nullptr; + Parallel_Grid* pgrid_ptr_ = nullptr; + Structure_Factor* sf_ptr_ = nullptr; + pseudopot_cell_vl* locpp_ptr_ = nullptr; + + Grid_Driver& gd() const { return *this->gd_ptr_; } + const TwoCenterBundle& tcb() const { return *this->tcb_ptr_; } + Parallel_Grid& pgrid() const { return *this->pgrid_ptr_; } + Structure_Factor& sfac() const { return *this->sf_ptr_; } + pseudopot_cell_vl& vloc() const { return *this->locpp_ptr_; } + /// bind the aliases above; `ks_` must already be set (or null for the `lr` path) + void bind_ground_state_aliases_(); + // not to use ElecState because 2-particle state is quite different from 1-particle state. // implement a independent one (ExcitedState) to pack physical properties if needed. // put the components of ElecState here: @@ -68,10 +94,27 @@ namespace ModuleESolver // ground state info /// @brief ground state wave function - psi::Psi* psi_ks = nullptr; + std::unique_ptr> psi_ks; ///< KS orbitals used in the [nocc+nvirt] window + /// @brief all KS orbitals. On the `ks-lr` path this aliases the ground-state solver's + /// `psi` (nk x nbands x nbasis -- far too big to copy every ionic step, and only read + /// here); on the `lr` path it points at `psi_ks_all_own_`, filled from file. + psi::Psi* psi_ks_all_ = nullptr; + std::unique_ptr> psi_ks_all_own_; /// @brief ground state bands, read from the file, or moved from ESolver_FP::pelec.ekb - ModuleBase::matrix eig_ks;///< energy of ground state + ModuleBase::matrix eig_ks;///< ground state eigenvalues in the [nocc+nvirt] window + ModuleBase::matrix eig_ks_all; ///< all eigenvalues of ground state, read from the file, or moved from ESolver_FP::pelec.ekb + ModuleBase::matrix wg_ks; /// occupation numbers of ground state in the [nocc+nvirt] window + ModuleBase::matrix wg_ks_all; /// occupation number of all bands of ground state + + + // @brief only needed for force calculation + std::unique_ptr pot_gs; + std::unique_ptr pot_gs_hartree; /// ground-state Hartree potential, only used for test_force + double etxc_gs = 0.; + double vtxc_gs = 0.; + + std::shared_ptr pot_hxc_gs; /// used in lr-grad, in the ground-state Hxc gradient term coming from dF/dC /// @brief Excited state wavefunction (locc, lvirt are local size of nocc and nvirt in each process) /// size of X: [neq][{nstate, nloc_per_state}], namely: @@ -82,7 +125,7 @@ namespace ModuleESolver std::vector nocc; ///< number of occupied orbitals for each spin used in the calculation int nocc_in = 1; ///< nocc read from input (adjusted by nelec): max(spin-up, spindown) - int nocc_max = 1; ///< nelec/2 + int nocc_max = 1; ///< full occupied count in the largest spin channel std::vector nvirt; ///< number of virtual orbitals for each spin used in the calculation int nvirt_in = 1; ///< nvirt read from input (adjusted by nelec): min(spin-up, spindown) int nbands = 2; @@ -98,9 +141,46 @@ namespace ModuleESolver std::string xc_kernel; void initialize_from_unitcell_(UnitCell& ucell, const Input_para& inp); - void initialize_from_ks_(ModuleESolver::ESolver_KS_LCAO&& ks_sol, - UnitCell& ucell, - const Input_para& inp); + /// one-time setup from the ground-state solver; ends by calling `refresh_from_ks_` + void initialize_from_ks_(UnitCell& ucell, const Input_para& inp); + /// re-read everything that depends on the atomic positions, once per ionic step + void refresh_from_ks_(UnitCell& ucell); + bool ks_initialized_ = false; ///< whether `initialize_from_ks_` has already run + bool exx_owned_ = false; ///< `exx_lri` was built here (so its Cs/Vs are ours to refresh) + + // ---------- geometry relaxation on an excited state ---------- + /// resolve and validate `lr_target_state` / `lr_target_spin`; call once the dimensions + /// (and therefore `openshell`) are final + void setup_relax_target_(); + bool excited_relax_ = false; ///< driving `calculation = relax` from an excited state + int target_is_ = 0; ///< spin block of the relaxed state in `X` and `pelec->ekb` + double etot_gs_ = 0.0; ///< ground-state total energy of the current step (Ry) + ModuleBase::matrix force_gs_; ///< ground-state force of the current step (Ry/Bohr, F = -dE/dR) + /// The LR part of the excited-state force, -d(Omega)/dR (Ry/Bohr). + ModuleBase::matrix lr_force_; + /// The state currently being followed, as an index into `X` / `pelec->ekb`. + /// Seeded from `lr_target_state` on the first ionic step, then re-chosen at every + /// later step by maximum overlap with the previous step's amplitude (below). + /// Following a fixed INDEX instead is what makes a relaxation fail near a + /// degeneracy: the index always names the n-th lowest root, so as soon as two + /// surfaces cross, "the target" jumps to a different diabatic state and the force + /// is discontinuous. CG assumes a conservative field and cannot recover from that. + int target_state_ = -1; + /// Previous ionic step's amplitude for the followed state (local part), the + /// reference the overlap is taken against. Empty on the first step. + std::vector target_X_prev_; + LR::RootBasis target_basis_prev_; + /// Re-select `target_state_` as argmax_j || and refresh the reference. + /// `ofs` receives the note when the followed root changes index, and the warning when + /// no current root resembles the previous one. + void follow_target_state_(std::ofstream& ofs); + + /// index of the relaxed state inside `pelec->ekb` + int target_ekb_offset_() const + { return this->openshell ? this->target_state_ + : this->target_is_ * this->nstates + this->target_state_; } + + std::vector spin_types; std::unique_ptr gint_info_ = nullptr; void set_gint(); @@ -109,10 +189,33 @@ namespace ModuleESolver Parallel_2D paraC_; /// @brief variables for parallel distribution of excited states std::vector paraX_; + + // ---------------- the Z-vector (CPSCF) window ---------------- + // The Z-vector equation enforces the Brillouin condition in EVERY occupied-virtual + // rotation, so it must not be confined to the `nvirt` window X lives in. + // It should be the whole AO virtual space. + // + // These mirror `psi_ks` / `eig_ks` / `paraC_` / `paraX_` / `nvirt` / `nloc_per_state` + // but span every virtual band the ground state produced (the input `nbands`), so the + // window is widened by raising *nbands*, not `nvirt`. X keeps its own window, so Omega + // -- and with it any finite-difference reference -- is untouched. + std::unique_ptr> psi_ks_z_; + ModuleBase::matrix eig_ks_z_; + Parallel_2D paraC_z_; + std::vector paraX_z_; + std::vector nvirt_z_; + int nbands_z_ = 0; + int nloc_per_state_z_ = 0; + /// (re)build the Z window from `psi_ks_all_` / `eig_ks_all`. `desc_src` describes the + /// source wavefunction's 2D layout; unused (and may be null) in a serial build. + void fill_z_window_(const int* desc_src); + /// Widen `nst` X blocks starting at `istate_begin` from the X window into the Z window, + /// zero-filling the virtual rows X does not have. + ct::Tensor pad_X_to_z_(const int ispin, const int istate_begin, const int nst) const; /// @brief variables for parallel distribution of matrix in AO representation Parallel_Orbitals paraMat_; + Parallel_Orbitals paraMat_all_; // for the parallelized size of the KS orbitals - TwoCenterBundle two_center_bundle_; LCAO_Orbitals orb_; ///< numerical atomic orbital data for single-point evaluation std::vector> velocity_mo; ///< store the velocity matrix elements in MO basis @@ -137,11 +240,120 @@ namespace ModuleESolver /// reset nocc, nvirt, npairs after read ground-state wavefunction when nspin=2 void reset_dim_spin2(); + /// setup Parallel_Orbitals info. beyond Parallel_2D + void set_parallel_orbitals_band(Parallel_Orbitals& p, const int nbands_in); + + ///========================== for gradient calculation ========================= + void init_pot_groundstate(const Charge& chg_gs); + /// Solve the Z-vector equation for `nst` independent blocks of `Xz`. + ct::Tensor solve_zvector_eqation(const int ispin, const int nst, const ct::Tensor& Xz); + /// Excited-state gradients d(Omega)/dR, one matrix per state solved. `istate_only >= 0` + /// restricts it to that one state: geometry relaxation follows a single state, and the + /// Z-vector solve dominates the cost. -1 does all `nstates`. + std::vector cal_force(const int ispin, const int istate_only = -1); + /// open-shell (spin-unrestricted) excited-state force: X holds [up | down] and every + /// density matrix has two independent channels + std::vector cal_force_openshell(const int istate_only = -1); + /// @brief Gradients for excitation vectors supplied by the caller, already widened into + /// the Z window -- the two functions above are thin wrappers that widen the stored + /// eigenvectors and look up their `omega`. + /// + /// The blocks of `Xz` need not be the eigenvectors the Casida diagonalizer returned. Any + /// normalized vector inside a degenerate multiplet is an eigenvector with the same + /// `omega`, so passing a linear combination is what turns the per-state gradient into the + /// full degenerate-subspace gradient matrix; see `cal_grad_matrix_degenerate` and + /// `grad_degen.h`. + /// + /// @param omega excitation energy of each block (Ry); its size sets the block count + /// @param label_begin state index the first block is reported under (labels only) + std::vector cal_force_Xz(const int ispin, const ct::Tensor& Xz, + const std::vector& omega, const int label_begin); + /// open-shell counterpart of `cal_force_Xz` + std::vector cal_force_openshell_Xz(const ct::Tensor& Xz, + const std::vector& omega, const int label_begin); + /// @brief The linear vibronic coupling (LVC) data of one degenerate multiplet. + /// + /// $H(\delta R)=\Omega_0\mathbb{1}+\sum_{A\alpha}\delta R_{A\alpha}G^{(A\alpha)}$ is the + /// complete first-order description of a degeneracy, and $G$ is its parameter set. Kept as + /// one object rather than as loose force matrices because the excited-state relaxation and + /// (later) non-adiabatic dynamics need exactly the same data: near a degeneracy the correct + /// propagation is on this coupled $d\times d$ model, not on an adiabatic gradient. + struct MultipletLVC + { + int ispin = 0; ///< which spin channel (index into `spin_types`) + std::vector states; ///< the multiplet's state indices, ascending + double omega0 = 0.0; ///< the common excitation energy (Ry) + double omega_spread = 0.0; ///< max - min over the members (Ry); 0 if exact + /// G[k][l], symmetric, each a (nat, 3) force matrix -- i.e. the 3N matrices of d x d + std::vector> g; + int dim() const { return static_cast(states.size()); } + /// $\operatorname{Tr}G/d$: the one smooth, basis-independent 3N vector field the + /// multiplet has. Following it preserves the symmetric configuration. + ModuleBase::matrix average_force() const; + }; + /// LVC data of every multiplet found at this geometry, rebuilt on each call of + /// `cal_force_and_grad_matrix_`. Empty unless `lr_degen_thr > 0`. + std::vector multiplet_lvc_; + /// The multiplet `lr_target_state` belongs to, or empty when the target is non-degenerate + /// or `lr_degen_mode = state`. Refreshed every ionic step by + /// `resolve_target_multiplet_`, and it is what makes `cal_energy` and the reported gradient + /// describe the same surface. + std::vector target_group_; + /// Fill `target_group_` from the excitation energies of the current geometry. + void resolve_target_multiplet_(); + /// The excitation energy the relaxation is minimising: the target state's own, or the + /// multiplet average when `target_group_` is set. + double target_omega_() const; + /// The LR half of the force for the current geometry, following whichever surface + /// `lr_degen_mode` selects. `ofs` receives the note when that is not a single state. + ModuleBase::matrix cal_lr_force_relax_(std::ofstream& ofs); + /// @brief The force of the steepest-descending branch of the target multiplet, i.e. + /// `lr_degen_mode = jt`. + /// + /// Assembles the off-diagonal part of the gradient matrix (which `average` does not need), + /// solves the joint direction/mixing optimization in `LR::find_jt_direction`, and returns + /// that branch's own force. Once a step has split the multiplet there is no group left and + /// the ordinary single-state path takes over, so the mode is self-limiting. + /// + /// @param diag the multiplet's per-state forces, already computed + ModuleBase::matrix cal_jt_force_(const std::vector& diag, + std::ofstream& ofs); + LR::GradientInputs gradient_inputs_() const; + /// Widen a multiplet's eigenvectors into the Z window, one block each. Members need not be + /// contiguous, so they are padded one at a time. + ct::Tensor pad_group_to_z_(const int ispin, const std::vector& group) const; + /// @brief Per-state gradients of every state, plus the gradient matrix of each degenerate + /// multiplet when `lr_degen_thr` asks for it. The single-point entry point. + /// + /// Fills `multiplet_lvc_`. + void cal_force_and_grad_matrix_(const int ispin, std::ofstream& ofs); + /// @brief The gradient matrix of one degenerate multiplet, + /// $G^{(A\alpha)}_{kl}=\langle X_k|\partial A/\partial R_{A\alpha}|X_l\rangle$. + /// + /// At a degeneracy no single state has a gradient vector -- the branch slopes along a + /// displacement $u$ are the eigenvalues of $\sum_{A\alpha}u_{A\alpha}G^{(A\alpha)}$, whose + /// eigenvectors depend on $u$ -- so this whole matrix, not its diagonal, is the first-order + /// information. It is obtained from the polarization identity + /// $G_{kl}=\mathcal F[(X_k{+}X_l)/\sqrt2]-\tfrac12(G_{kk}+G_{ll})$, which needs no new + /// physics. `grad_degen.h` derives why that is exact. + /// + /// @param group state indices of the multiplet, from `LR::group_degenerate_states` + /// @param diag their per-state gradients, i.e. $G_{kk}$, already computed by `cal_force` + /// @param ofs log stream the matrix and the precondition diagnostics are written to + /// @return G[k][l], symmetric, each entry a (nat, 3) force matrix + std::vector> cal_grad_matrix_degenerate(const int ispin, + const std::vector& group, const std::vector& diag, + std::ofstream& ofs); + void test_force(); // test: reproduce the force of ground state + module_dm::DensityMatrix cal_dm_gs(); ///< ground-state density matrix + #ifdef __EXX /// Tdata of Exx_LRI is same as T, for the reason, see operator_lr_exx.h std::shared_ptr> exx_lri = nullptr; - void move_exx_lri(std::shared_ptr>&); - void move_exx_lri(std::shared_ptr>>&); + /// share the ground-state solver's Exx_LRI. It is shared, not stolen: the KS solver + /// keeps using it on the next ionic step. + void share_exx_lri(std::shared_ptr>&); + void share_exx_lri(std::shared_ptr>>&); Exx_Info exx_info; #endif }; diff --git a/source/source_esolver/esolver_lr_rlx.cpp b/source/source_esolver/esolver_lr_rlx.cpp new file mode 100644 index 00000000000..b2302893b9e --- /dev/null +++ b/source/source_esolver/esolver_lr_rlx.cpp @@ -0,0 +1,325 @@ +#include "source_esolver/esolver_lr_lcao_tddft.h" +#include "source_lcao/module_lr/zeq_solver.h" +#include "source_lcao/module_lr/cal_edm.h" +#include "source_lcao/module_lr/lr_force.h" +#include "source_lcao/module_lr/gradient_inputs.h" +#include "source_lcao/module_lr/gradient_output.h" +#include "source_lcao/module_lr/lr_amp.h" +#include "source_lcao/module_lr/grad_degen.h" +#include "source_base/parallel_reduce.h" +#include +#include +#include +#include "source_estate/module_dm/dm_from_psi.h" +#include "source_io/module_output/output_log.h" + +using namespace LR; + + +template +void ModuleESolver::ESolver_LR::setup_relax_target_() +{ + this->excited_relax_ = (this->inp_->calculation == "relax"); + if (!this->excited_relax_) { return; } + + // The Z-vector equation has no complex solver (see zeq_solver.hpp), so an + // excited-state gradient only exists at gamma. Fail here rather than after the SCF. + if (!std::is_same::value) + { + ModuleBase::WARNING_QUIT("ESolver_LR", + "excited-state relaxation currently requires gamma-only sampling: the complex " + "Z-vector solver is not implemented."); + } + if (this->inp_->lr_target_state >= this->nstates) + { + ModuleBase::WARNING_QUIT("ESolver_LR", + "lr_target_state is beyond the states actually solved (lr_nstates <= 0 expands to " + "all particle-hole pairs, which may be fewer than requested)."); + } + + // `openshell` is only settled after the ground-state occupations have been read, which is why + // this cannot live in the input-file checks + const std::string& spin = this->inp_->lr_target_spin; + if (this->openshell) + { + // An open-shell calculation solves one spin-conserving channel, so there is nothing to + // choose: whatever lr_target_spin says, this is the state that gets relaxed. Only an + // explicit `triplet` is worth mentioning -- `singlet` is the default and expresses no + // intent, and `updown` is already the right name for this channel. + if (spin == "triplet") + { + this->ofs_running_ << " WARNING: lr_target_spin=triplet is ignored. This is an" + " open-shell calculation with a single spin-conserving channel (updown), which is" + " what the relaxation will follow." << std::endl; + } + this->target_is_ = 0; + } + else + { + // Closed shell is the opposite case: singlet and triplet are genuinely different states + // with different gradients, so `updown` here is ambiguous rather than redundant -- it + // usually means lr_unrestricted was meant to be set. + if (spin == "updown") + { + ModuleBase::WARNING_QUIT("ESolver_LR", + "lr_target_spin=updown, but this is a closed-shell calculation, where singlet and " + "triplet are separate states with separate gradients. Pick one of them, or set " + "lr_unrestricted to run spin-unrestricted."); + } + this->target_is_ = (spin == "triplet") ? 1 : 0; + if (this->target_is_ >= this->nspin) + { + ModuleBase::WARNING_QUIT("ESolver_LR", + "lr_target_spin=triplet requires nspin=2: the triplet channel is not built here."); + } + } + this->force_gs_.create(this->ucell_->nat, 3); + this->lr_force_.create(this->ucell_->nat, 3); + this->target_state_ = this->inp_->lr_target_state; // seed; overlap takes over from step 2 + this->ofs_running_ << " Excited-state relaxation follows state " << this->inp_->lr_target_state + << " of the " << (this->openshell ? "updown" : (this->target_is_ == 1 ? "triplet" : "singlet")) + << " channel, tracked by cross-geometry orbital and amplitude overlap." << std::endl; +} + +/// Compare roots in a common electron-hole basis using exact cross-geometry AO overlap. +template +void ModuleESolver::ESolver_LR::follow_target_state_(std::ofstream& ofs) +{ + if (!this->excited_relax_) { return; } + const int channel = this->openshell ? 0 : this->target_is_; + const T* const amplitudes = this->X[channel].template data(); + const TwoCenterIntegrator& overlap_integrator = *this->tcb().overlap_orb; + const LR::RootInputs inputs{*this->ucell_, this->orb_cutoff_, overlap_integrator, + *this->psi_ks, this->paraC_, this->paraMat_, this->paraX_, this->nocc, this->nvirt, this->openshell, this->target_is_}; + LR::follow_cross_root(inputs, amplitudes, this->nloc_per_state, this->nstates, + this->inp_->lr_target_state, this->target_state_, this->target_X_prev_, this->target_basis_prev_, ofs); +} + +template +double ModuleESolver::ESolver_LR::cal_energy() +{ + // Outside a relaxation nothing consumes this, and returning a non-zero value would change + // what the existing single-point outputs report. + if (!this->excited_relax_) { return 0.0; } + // `target_omega_()` is the multiplet average under `lr_degen_mode = average` and the + // single state otherwise, matching whatever `cal_lr_force_relax_` produced the gradient of. + // The energy-based optimisers line-search on this, so the two must not describe different + // surfaces. + return this->etot_gs_ + this->target_omega_(); +} + +template +void ModuleESolver::ESolver_LR::cal_force(BaseCell& basecell, ModuleBase::matrix& force) +{ + basecell.require_kind(BaseCell::Kind::unitcell, __FUNCTION__); + const UnitCell& ucell = static_cast(basecell); + if (!this->excited_relax_) + { // single-point runs print the gradients of every state from `after_all_runners` instead + return; + } + if (this->lr_force_.nr != ucell.nat) + { + ModuleBase::WARNING_QUIT("ESolver_LR::cal_force", + "the excited-state gradient has not been computed for this geometry."); + } + // Both halves already follow the ABACUS force convention F = -dE/dR (Ry/Bohr), so they add. + // The "Gradients of each excited state" heading that `cal_force(int)` prints under is a + // misnomer: the finite-difference reference it was validated against (`abacus-fd lr-custom`) + // computes (E(-h) - E(+h))/h, which is -d(Omega)/dR, and the two agree in sign. + force.create(ucell.nat, 3); + force = this->force_gs_ + this->lr_force_; + ModuleIO::print_force(this->ofs_running_, ucell, "EXCITED-STATE TOTAL-FORCE (eV/Angstrom)", force, false); +} + +template +void ModuleESolver::ESolver_LR::cal_stress(BaseCell& basecell, ModuleBase::matrix& stress) +{ + static_cast(stress); + basecell.require_kind(BaseCell::Kind::unitcell, __FUNCTION__); + ModuleBase::WARNING_QUIT("ESolver_LR::cal_stress", + "the excited-state stress is not implemented (every LR gradient stress term is a " + "dummy passed with isstress=false)."); +} + +template +ct::Tensor ModuleESolver::ESolver_LR::pad_group_to_z_(const int ispin, + const std::vector& group) const +{ + const int d = static_cast(group.size()); + const int nloc_g = this->nloc_per_state_z_; + ct::Tensor Xz = LR_Util::newTensor({ d, nloc_g }); + Xz.zero(); + // One at a time rather than as a range: `group` is sorted, but nothing guarantees its members + // are contiguous in the state list. + for (int k = 0; k < d; ++k) + { + const ct::Tensor one = this->pad_X_to_z_(ispin, group[k], 1); + std::copy(one.template data(), one.template data() + nloc_g, + Xz.template data() + static_cast(k) * nloc_g); + } + return Xz; +} + +template +void ModuleESolver::ESolver_LR::resolve_target_multiplet_() +{ + // Keyed on `target_state_`, not on `lr_target_state`: the overlap tracking re-chooses the + // followed root at every ionic step, and the multiplet has to be the one containing the root + // actually being followed. They differ as soon as two surfaces have crossed. + this->target_group_.clear(); + if (LR_Util::tolower(this->inp_->lr_degen_mode) == "state") { return; } + const int ekb_off = this->openshell ? 0 : this->target_is_ * this->nstates; + std::vector omega(this->nstates); + for (int ist = 0; ist < this->nstates; ++ist) { omega[ist] = this->pelec->ekb.c[ekb_off + ist]; } + const std::vector> groups + = LR::group_degenerate_states(omega, this->inp_->lr_degen_thr); + for (const std::vector& g : groups) + { + if (std::find(g.begin(), g.end(), this->target_state_) == g.end()) { continue; } + // a one-member group means the target is not degenerate here, and averaging over it would + // be the single-state path with extra steps + if (g.size() > 1) { this->target_group_ = g; } + break; + } +} + +template +double ModuleESolver::ESolver_LR::target_omega_() const +{ + if (this->target_group_.empty()) { return this->pelec->ekb.c[this->target_ekb_offset_()]; } + const int ekb_off = this->openshell ? 0 : this->target_is_ * this->nstates; + double sum = 0.0; + for (const int ist : this->target_group_) { sum += this->pelec->ekb.c[ekb_off + ist]; } + return sum / static_cast(this->target_group_.size()); +} + +template +ModuleBase::matrix ModuleESolver::ESolver_LR::cal_lr_force_relax_(std::ofstream& ofs) +{ + this->resolve_target_multiplet_(); + if (this->target_group_.empty()) + { + return this->cal_force(this->target_is_, this->target_state_)[0]; + } + const bool jt_mode = (LR_Util::tolower(this->inp_->lr_degen_mode) == "jt"); + // `average` needs only the DIAGONAL of the gradient matrix: the average is basis-independent by + // construction, so the off-diagonal part (and the extra d(d-1)/2 solves it costs) is not + // involved. The Jahn-Teller direction is orthogonal to the average and does need them. + const int d = static_cast(this->target_group_.size()); + const int ekb_off = this->openshell ? 0 : this->target_is_ * this->nstates; + const ct::Tensor Xz = this->pad_group_to_z_(this->target_is_, this->target_group_); + std::vector omega(d); + for (int k = 0; k < d; ++k) { omega[k] = this->pelec->ekb.c[ekb_off + this->target_group_[k]]; } + const std::vector forces = this->openshell + ? this->cal_force_openshell_Xz(Xz, omega, this->target_group_.front()) + : this->cal_force_Xz(this->target_is_, Xz, omega, this->target_group_.front()); + ofs << " Followed state " << this->target_state_ << " is degenerate with"; + for (const int ist : this->target_group_) + { + if (ist != this->target_state_) { ofs << " " << ist; } + } + ofs << " (Omega_bar = " << this->target_omega_() << " Ry)." << std::endl; + if (!jt_mode) + { + ofs << " lr_degen_mode=average: following the multiplet average, which keeps the" + " geometry on the symmetric configuration." << std::endl; + return average_forces(forces); + } + return this->cal_jt_force_(forces, ofs); +} + +template +ModuleBase::matrix ModuleESolver::ESolver_LR::cal_jt_force_( + const std::vector& diag, std::ofstream& ofs) +{ + // Regime (c): descend the Jahn-Teller branch. This needs the whole gradient matrix, so the + // off-diagonal elements are assembled here -- d(d-1)/2 further Z-vector solves on top of the + // d diagonal ones already in `diag`. + const int d = static_cast(this->target_group_.size()); + const std::vector> g + = this->cal_grad_matrix_degenerate(this->target_is_, this->target_group_, diag, ofs); + const int nat = diag[0].nr; + const int ncoord = nat * 3; + // flatten to the layout `find_jt_direction` takes: one d x d block per nuclear coordinate. + // FORCES go in, not gradients, so the direction that comes back points downhill. + std::vector gflat(static_cast(ncoord) * d * d, 0.0); + for (int iat = 0; iat < nat; ++iat) + { + for (int ixyz = 0; ixyz < 3; ++ixyz) + { + const int a = iat * 3 + ixyz; + for (int k = 0; k < d; ++k) + { + for (int l = 0; l < d; ++l) + { + gflat[static_cast(a) * d * d + k * d + l] = g[k][l](iat, ixyz); + } + } + } + } + const LR::JTDirection jt = LR::find_jt_direction(gflat, ncoord, d); + const int channel = this->openshell ? 0 : this->target_is_; + const T* const amplitudes = this->X[channel].template data(); + LR::save_mixed_root(amplitudes, this->nloc_per_state, this->target_group_, jt.mixing, this->target_X_prev_); + std::vector jt_part; + const std::vector sym = LR::split_symmetric_part(gflat, ncoord, d, jt.mixing, jt_part); + + const double fac = ModuleBase::Ry_to_eV / ModuleBase::BOHR_TO_A; + ofs << " lr_degen_mode=jt: descending the steepest branch of the multiplet." << std::endl + << " |F| of that branch = " << jt.slope * fac << " eV/Angstrom" << std::endl + << " mixing v ="; + for (int k = 0; k < d; ++k) { ofs << " " << jt.mixing[k]; } + ofs << std::endl + << " starts agreeing = " << jt.restarts_agreeing << " of " + << (d + (1 << (d - 1))) << " (iterations " << jt.iterations << ")" << std::endl; + // The two halves matter separately: the symmetric part is common to the whole multiplet and + // only relaxes the geometry, while the remainder is what actually breaks the degeneracy. At a + // stationary point of the average surface the first is zero and the whole force is Jahn-Teller. + double nsym = 0.0; + double njt = 0.0; + for (int a = 0; a < ncoord; ++a) + { + nsym += sym[a] * sym[a]; + njt += jt_part[a] * jt_part[a]; + } + ofs << " |symmetric part| = " << std::sqrt(nsym) * fac << " eV/Angstrom (common to the" + " multiplet; relaxes the geometry without splitting it)" << std::endl + << " |Jahn-Teller part| = " << std::sqrt(njt) * fac << " eV/Angstrom (the symmetry-" + "breaking remainder)" << std::endl; + if (std::sqrt(njt) < 1e-8) + { + ofs << " WARNING: the symmetry-breaking part vanishes -- every branch of this multiplet has" + " the same gradient, so there is no Jahn-Teller direction to descend here. This is the" + " expected outcome for a linear molecule, where the effect is second order" + " (Renner-Teller) and no first-order term exists." << std::endl; + } + ofs << " NOTE: this is the first-order DIRECTION only. The distortion amplitude also needs the" + " harmonic term, and the step norm is Cartesian rather than mass-weighted." << std::endl; + + // The force handed back is that of the descending branch, q(v) with v the optimal mixing -- + // i.e. the branch's own gradient, which is what a relaxation must follow. Once a step has + // split the multiplet, `resolve_target_multiplet_` finds no group and the ordinary + // single-state path takes over, so this mode is self-limiting by construction. + ModuleBase::matrix f(nat, 3); + for (int iat = 0; iat < nat; ++iat) + { + for (int ixyz = 0; ixyz < 3; ++ixyz) + { + const int a = iat * 3 + ixyz; + f(iat, ixyz) = sym[a] + jt_part[a]; + } + } + return f; +} + +template +ModuleBase::matrix ModuleESolver::ESolver_LR::MultipletLVC::average_force() const +{ + assert(!this->g.empty()); + std::vector diag; + for (int k = 0; k < this->dim(); ++k) { diag.push_back(this->g[k][k]); } + return average_forces(diag); +} + +template class ModuleESolver::ESolver_LR; +template class ModuleESolver::ESolver_LR, double>; diff --git a/source/source_estate/CMakeLists.txt b/source/source_estate/CMakeLists.txt index b5dd498251c..51640e345c3 100644 --- a/source/source_estate/CMakeLists.txt +++ b/source/source_estate/CMakeLists.txt @@ -98,8 +98,6 @@ endif() if(ENABLE_LCAO) add_subdirectory(module_dm) if(BUILD_TESTING) - if(ENABLE_MPI) - add_subdirectory(module_dm/unittests) - endif() + add_subdirectory(module_dm/unittests) endif() endif() diff --git a/source/source_estate/module_dm/density_matrix.h b/source/source_estate/module_dm/density_matrix.h index 811d85349d9..c382d30fb2e 100644 --- a/source/source_estate/module_dm/density_matrix.h +++ b/source/source_estate/module_dm/density_matrix.h @@ -45,7 +45,7 @@ struct ShiftRealComplex> // DensityMatrix,TR>::cal_dmr() is illegal in C++, so module_dm is used instead. template extern void cal_dmr( - DensityMatrix &dm, + const DensityMatrix &dm, std::vector*> &dmR_out, const int ik_in); @@ -70,7 +70,7 @@ struct ShiftRealComplex> */ template extern void accumulate_dmr( - DensityMatrix &dm, + const DensityMatrix &dm, std::vector*> &dmR_out, const std::map, std::complex>& phase_hybrid, const int ik_in, @@ -340,6 +340,10 @@ class DensityMatrix * please make sure the size of TK* is correct */ void set_dmk_ptr(const int ik, TK* DMK_in); + void set_DMK_vector(const int ik, const std::vector& v) { this->dmk[ik] = v; } + + /// number of k-slots stored in `dmk` (spin_mult * _nk, flattened) + int get_DMK_nks() const { return static_cast(this->dmk.size()); } /** * @brief calculate density matrix DMR from dm(k) using blas::axpy @@ -459,7 +463,7 @@ class DensityMatrix std::vector dmr_tmp; friend void module_dm::cal_dmr( - DensityMatrix& dm, + const DensityMatrix& dm, std::vector*>& dmR_out, const int ik_in); friend void module_dm::cal_dmr_td( @@ -473,7 +477,7 @@ class DensityMatrix hamilt::HContainer>* dmR_out, const int ik_in); friend void module_dm::accumulate_dmr( - DensityMatrix& dm, + const DensityMatrix& dm, std::vector*>& dmR_out, const std::map, std::complex>& phase_hybrid, const int ik_in, diff --git a/source/source_estate/module_dm/dmr_k.cpp b/source/source_estate/module_dm/dmr_k.cpp index 9fbbcbf8d22..a22fc6e6379 100644 --- a/source/source_estate/module_dm/dmr_k.cpp +++ b/source/source_estate/module_dm/dmr_k.cpp @@ -13,7 +13,7 @@ namespace module_dm // shared inner loop of cal_dmr / cal_dmr_td: accumulate kphase * DMK into DMR blocks template void accumulate_dmr( - DensityMatrix& dm, + const DensityMatrix& dm, std::vector*>& dmR_out, const std::map, std::complex>& phase_hybrid, const int ik_in, @@ -78,7 +78,7 @@ void accumulate_dmr( // calculate DMR from DMK using blas for multi-k calculation template void cal_dmr( - DensityMatrix& dm, + const DensityMatrix& dm, std::vector*>& dmR_out, const int ik_in) { @@ -101,7 +101,6 @@ void cal_dmr( const std::map, std::complex> no_hybrid_phase; accumulate_dmr(dm, dmR_out, no_hybrid_phase, ik_in, "module_dm::cal_dmr"); - dm._dmr_ready = true; ModuleBase::timer::end("DensityMatrix", "cal_dmr"); } @@ -109,26 +108,28 @@ template <> void DensityMatrix, double>::cal_dmr(const int ik_in) { module_dm::cal_dmr(*this, this->dmr, ik_in); + this->_dmr_ready = true; } template <> void DensityMatrix, std::complex>::cal_dmr(const int ik_in) { module_dm::cal_dmr(*this, this->dmr, ik_in); + this->_dmr_ready = true; } // explicit instantiations for accumulate_dmr (used by both cal_dmr here and // cal_dmr_td in dmr_td.cpp; without these the TD instantiations are missing // at link time) template void accumulate_dmr, double, double>( - DensityMatrix, double>&, + const DensityMatrix, double>&, std::vector*>&, const std::map, std::complex>&, const int, const char*); template void accumulate_dmr, std::complex, std::complex>( - DensityMatrix, std::complex>&, + const DensityMatrix, std::complex>&, std::vector>*>&, const std::map, std::complex>&, const int, diff --git a/source/source_estate/module_dm/unittests/CMakeLists.txt b/source/source_estate/module_dm/unittests/CMakeLists.txt index 3535cc7fe7a..1064492f639 100644 --- a/source/source_estate/module_dm/unittests/CMakeLists.txt +++ b/source/source_estate/module_dm/unittests/CMakeLists.txt @@ -4,6 +4,7 @@ abacus_disable_feature_definitions(__ROCM) install(DIRECTORY support DESTINATION ${CMAKE_CURRENT_BINARY_DIR}) +if(ENABLE_MPI) AddTest( TARGET MODULE_ESTATE_dm_constructor_test LIBS parameter base device symmetry @@ -30,10 +31,14 @@ AddTest( ${ABACUS_SOURCE_DIR}/source_cell/reciprocal_grid.cpp ) +endif() + +# The calculation fixture initializes both serial and distributed orbital layouts. AddTest( TARGET MODULE_ESTATE_dm_cal_DMR_test - LIBS parameter base device symmetry + LIBS parameter base device symmetry container SOURCES test_cal_dm_r.cpp ../density_matrix.cpp ../dmr_gamma.cpp ../dmr_init.cpp ../dm_setter.cpp ../dm_getter.cpp ../dm_tools.cpp ../dmr_k.cpp ../dmr_td.cpp ../dmr_full.cpp tmp_mocks.cpp + ${ABACUS_SOURCE_DIR}/source_lcao/module_lr/utils/lr_util.cpp ${ABACUS_SOURCE_DIR}/source_hamilt/module_hcontainer/base_matrix.cpp ${ABACUS_SOURCE_DIR}/source_hamilt/module_hcontainer/hcontainer.cpp ${ABACUS_SOURCE_DIR}/source_hamilt/module_hcontainer/atom_pair.cpp @@ -43,6 +48,7 @@ AddTest( ${ABACUS_SOURCE_DIR}/source_cell/reciprocal_grid.cpp ) +if(ENABLE_MPI) AddTest( TARGET MODULE_ESTATE_dm_soc_magnetization_roundtrip_test LIBS parameter base device @@ -52,3 +58,4 @@ AddTest( ${ABACUS_SOURCE_DIR}/source_hamilt/module_hcontainer/atom_pair.cpp ${ABACUS_SOURCE_DIR}/source_basis/module_ao/parallel_orbitals.cpp ) +endif() diff --git a/source/source_estate/module_dm/unittests/test_cal_dm_r.cpp b/source/source_estate/module_dm/unittests/test_cal_dm_r.cpp index dbad57eb1de..3c3873918de 100644 --- a/source/source_estate/module_dm/unittests/test_cal_dm_r.cpp +++ b/source/source_estate/module_dm/unittests/test_cal_dm_r.cpp @@ -5,6 +5,7 @@ #include "source_estate/module_dm/density_matrix.h" #include "source_hamilt/module_hcontainer/hcontainer.h" #include "source_cell/klist.h" +#include "source_lcao/module_lr/utils/lr_util_hcontainer.h" /************************************************ * unit test of DensityMatrix constructor @@ -61,7 +62,6 @@ class DMTest : public testing::Test ucell.atoms[0].iw2n[iw] = 0; } ucell.set_iat2iwt(1); - init_parav(); // set paraV init_parav(); @@ -87,6 +87,10 @@ class DMTest : public testing::Test #else void init_parav() { + const int n = test_size * test_nw; + paraV = new Parallel_Orbitals(); + paraV->set_serial(n, n); + paraV->set_atomic_trace(ucell.get_iat2iwt(), test_size, n); } #endif }; @@ -205,7 +209,9 @@ TEST_F(DMTest, cal_DMR_blas_double) } // calculate this->dmr std::chrono::high_resolution_clock::time_point start_time = std::chrono::high_resolution_clock::now(); + EXPECT_FALSE(DM.is_dmr_ready()); DM.cal_dmr(-1); + EXPECT_TRUE(DM.is_dmr_ready()); std::chrono::high_resolution_clock::time_point end_time = std::chrono::high_resolution_clock::now(); std::chrono::duration elapsed_time = std::chrono::duration_cast>(end_time - start_time); @@ -271,7 +277,9 @@ TEST_F(DMTest, cal_DMR_blas_complex) DM.init_dmr(&gd, &ucell); // calculate this->dmr std::chrono::high_resolution_clock::time_point start_time = std::chrono::high_resolution_clock::now(); + EXPECT_FALSE(DM.is_dmr_ready()); DM.cal_dmr(-1); + EXPECT_TRUE(DM.is_dmr_ready()); std::chrono::high_resolution_clock::time_point end_time = std::chrono::high_resolution_clock::now(); std::chrono::duration elapsed_time = std::chrono::duration_cast>(end_time - start_time); @@ -411,7 +419,9 @@ TEST_F(DMTest, cal_DMR_soc_pauli_branch) DM.init_dmr(&gd, &ucell); // Gamma-only: reduce R vectors to (0, 0, 0), as cal_DMR_blas_double does DM.get_dmr_ptr(1)->fix_gamma(); + EXPECT_FALSE(DM.is_dmr_ready()); DM.cal_dmr(-1); + EXPECT_TRUE(DM.is_dmr_ready()); // check the Gamma (R = 0) block: rho_0 must be 2a (Pauli), NOT a (real projection); // rho_x = rho_y = rho_z = 0 for uu == dd and zero spin off-diagonals. @@ -445,6 +455,129 @@ TEST_F(DMTest, cal_DMR_soc_pauli_branch) #endif } + +TEST_F(DMTest, LRDotRCountsEachImageOnce) +{ + hamilt::HContainer h1(paraV); + hamilt::HContainer h2(paraV); + for (int r = -1; r <= 1; ++r) + { + hamilt::AtomPair a(0, 1, r, 0, 0, paraV); + const int opposite_r = -r; + hamilt::AtomPair b(0, 1, opposite_r, 0, 0, paraV); + h1.insert_pair(a); + h2.insert_pair(b); + } + h1.allocate(nullptr, true); + h2.allocate(nullptr, true); + double expected = 0.0; + for (int r = -1; r <= 1; ++r) + { + auto& a = h1.find_pair(0, 1)->get_HR_values(r, 0, 0); + auto& b = h2.find_pair(0, 1)->get_HR_values(r, 0, 0); + const int size = a.get_row_size() * a.get_col_size(); + for (int i = 0; i < size; ++i) + { + const double x = r + 2.0; + const double y = i + 1.0; + a.get_pointer()[i] = x; + b.get_pointer()[i] = y; + expected += x * y; + } + } + EXPECT_DOUBLE_EQ(LR_Util::dot_R_matrix(h1, h2), expected); +} + +TEST_F(DMTest, LRTransposePreservesAtomIndices) +{ + const int n = test_size * test_nw; + const std::vector> kvec(1); + module_dm::DensityMatrix dm(paraV, 1, kvec, 1); + hamilt::HContainer hr(paraV); + for (int ia = 0; ia < test_size; ++ia) + { + for (int ja = 0; ja < test_size; ++ja) + { + hamilt::AtomPair pair(ia, ja, paraV); + hr.insert_pair(pair); + } + } + hr.allocate(nullptr, true); + hr.fix_gamma(); + dm.init_dmr(hr); + auto& dk = dm.get_dmk_vec()[0]; + for (int j = 0; j < paraV->ncol; ++j) + { + for (int i = 0; i < paraV->nrow; ++i) + { + const int mu = paraV->local2global_row(i); + const int nu = paraV->local2global_col(j); + dk[j * paraV->nrow + i] = mu * n + nu; + } + } + dm.cal_dmr(-1); + const int passes = 2; + for (int pass = 0; pass < passes; ++pass) + { + for (int ip = 0; ip < dm.get_dmr_ptr(1)->size_atom_pairs(); ++ip) + { + const auto& pair = dm.get_dmr_ptr(1)->get_atom_pair(ip); + const auto* values = pair.get_HR_values(0).get_pointer(); + for (int i = 0; i < pair.get_row_size(); ++i) + { + for (int j = 0; j < pair.get_col_size(); ++j) + { + const int mu = paraV->local2global_row(pair.get_begin_row() + i); + const int nu = paraV->local2global_col(pair.get_begin_col() + j); + const double expected = pass == 0 ? mu * n + nu : nu * n + mu; + EXPECT_DOUBLE_EQ(values[i * pair.get_col_size() + j], expected); + } + } + } + if (pass == 0) { LR_Util::transpose_DMR(dm, *paraV); } + } +} + +#ifdef __EXX +namespace RI_2D_Comm +{ +// Observe the adapter's spin/k input before the expensive RI redistribution. +template <> +std::vector>>> +split_m2D_ktoR>( + const UnitCell&, const K_Vectors&, const std::vector*>& mks, + const Parallel_2D&, const int nspin, const bool) +{ + EXPECT_EQ(nspin, 2); + EXPECT_EQ(mks.size(), 4U); + for (std::size_t ik = 0; ik < mks.size(); ++ik) + { + EXPECT_DOUBLE_EQ(mks[ik]->front(), ik + 1.0); + } + return std::vector>>>(nspin); +} +} + +TEST_F(DMTest, LRExxAdapterIncludesBothSpinChannels) +{ + K_Vectors kv; + kv.set_nks(4); + kv.kvec_d.resize(4); + module_dm::DensityMatrix dm(paraV, 2, kv.kvec_d, 2); + hamilt::HContainer hr(paraV); + hr.allocate(nullptr, true); + dm.init_dmr(hr); + auto& dmk = dm.get_dmk_vec(); + ASSERT_EQ(dmk.size(), 4U); + for (std::size_t ik = 0; ik < dmk.size(); ++ik) + { + std::fill(dmk[ik].begin(), dmk[ik].end(), ik + 1.0); + } + const auto ds = LR_Util::get_exx_Ds_gs(dm, ucell, kv, *paraV); + EXPECT_EQ(ds.size(), 2U); +} +#endif + int main(int argc, char** argv) { #ifdef __MPI diff --git a/source/source_hamilt/module_hcontainer/test/test_func_folding.cpp b/source/source_hamilt/module_hcontainer/test/test_func_folding.cpp index 12f5c9b987e..67c7d34ca6a 100644 --- a/source/source_hamilt/module_hcontainer/test/test_func_folding.cpp +++ b/source/source_hamilt/module_hcontainer/test/test_func_folding.cpp @@ -210,4 +210,42 @@ TEST_F(FoldingTest, folding_HR_d2d) std::cout << "HR init time: " << elapsed_time0.count()<<" fix_gamma time: "< overlap(&distribution); + const ModuleBase::Vector3 minus(-1, 0, 0); + const ModuleBase::Vector3 center(0, 0, 0); + const ModuleBase::Vector3 plus(1, 0, 0); + const hamilt::AtomPair forward_minus(0, 1, minus, &distribution); + const hamilt::AtomPair forward_center(0, 1, center, &distribution); + const hamilt::AtomPair forward_plus(0, 1, plus, &distribution); + const hamilt::AtomPair reverse(1, 0, center, &distribution); + overlap.insert_pair(forward_minus); + overlap.insert_pair(forward_center); + overlap.insert_pair(forward_plus); + overlap.insert_pair(reverse); + overlap.allocate(nullptr, true); + auto* forward = overlap.find_pair(0, 1); + forward->get_HR_values(-1, 0, 0).get_pointer()[0] = 2.0; + forward->get_HR_values(0, 0, 0).get_pointer()[0] = 3.0; + forward->get_HR_values(1, 0, 0).get_pointer()[0] = 5.0; + overlap.find_pair(1, 0)->get_pointer(0)[0] = 7.0; + const std::complex zero(0.0, 0.0); + std::vector> gamma_matrix(4, zero); + const ModuleBase::Vector3 gamma(0.0, 0.0, 0.0); + hamilt::folding_HR(overlap, gamma_matrix.data(), gamma, 2, 0); + EXPECT_EQ(gamma_matrix[1], std::complex(10.0, 0.0)); + EXPECT_EQ(gamma_matrix[2], std::complex(7.0, 0.0)); + std::vector> k_matrix(4, zero); + const ModuleBase::Vector3 k(0.25, 0.0, 0.0); + hamilt::folding_HR(overlap, k_matrix.data(), k, 2, 0); + EXPECT_NEAR(k_matrix[1].real(), 3.0, 1e-14); + EXPECT_NEAR(k_matrix[1].imag(), 3.0, 1e-14); + EXPECT_EQ(k_matrix[2], std::complex(7.0, 0.0)); +} diff --git a/source/source_hamilt/operator.h b/source/source_hamilt/operator.h index 8848fcc7bac..370d0e53056 100644 --- a/source/source_hamilt/operator.h +++ b/source/source_hamilt/operator.h @@ -26,6 +26,11 @@ enum class calculation_type lcao_dftu, lcao_sc_lambda, lcao_tddft_periodic, + lr_dmtrans_hxc, + lr_dmtrans_gxc, + lr_dmdiff_hxc, + lr_dmtrans_exx, + lr_dmdiff_exx }; // Basic class for operator module, diff --git a/source/source_hsolver/diago_dav_subspace.cpp b/source/source_hsolver/diago_dav_subspace.cpp index 331eb9857e4..189206abbaa 100644 --- a/source/source_hsolver/diago_dav_subspace.cpp +++ b/source/source_hsolver/diago_dav_subspace.cpp @@ -20,6 +20,7 @@ #ifdef __MPI #include #include "source_base/parallel_comm.h" +#include "source_base/parallel_device.h" #endif using namespace hsolver; @@ -714,9 +715,10 @@ void Diago_DavSubspace::diag_zhegvx(const int& nbase, // vcc: nbase * nband for (int i = 0; i < nband; i++) { - MPI_Bcast(&vcc[i * this->nbase_x], nbase, MPI_DOUBLE_COMPLEX, 0, this->diag_comm.comm); + // fix a bug here when vcc is real: bcast according to the true type + Parallel_Common::bcast_data(&vcc[i * this->nbase_x], nbase, this->diag_comm.comm, 0); } - MPI_Bcast((*eigenvalue_iter).data(), nband, MPI_DOUBLE, 0, this->diag_comm.comm); + Parallel_Common::bcast_data((*eigenvalue_iter).data(), nband, this->diag_comm.comm, 0); } #endif diff --git a/source/source_io/module_parameter/input_parameter.h b/source/source_io/module_parameter/input_parameter.h index 7deb3732dbf..03c63d28162 100644 --- a/source/source_io/module_parameter/input_parameter.h +++ b/source/source_io/module_parameter/input_parameter.h @@ -389,6 +389,11 @@ struct Input_para // ============== #Parameters (10.lr-tddft) =========================== int lr_nstates = 1; ///< the number of 2-particle states to be solved + int lr_target_state = 0; ///< which excited state the geometry relaxation follows (0-based) + std::string lr_target_spin = "singlet"; ///< spin channel of that state: singlet / triplet / updown + double lr_degen_thr = 0.0; ///< max excitation-energy spread of a degenerate multiplet whose gradient matrix is computed (Ry); 0 disables + std::string lr_degen_mode = "state"; ///< what a relaxation follows when the target state sits in a degenerate multiplet: state / average + std::string lr_grad_solver = "cg"; ///< the linear solver of the Z-vector equation for LR-TDDFT gradients: cg / lapack / scalapack / scalapack_chol / elpa std::vector lr_init_xc_kernel = {}; ///< The method to initalize the xc kernel int nocc = -1; ///< the number of occupied orbitals to form the 2-particle basis int nvirt = 1; ///< the number of virtual orbitals to form the 2-particle basis (nocc + nvirt <= nbands) diff --git a/source/source_io/module_parameter/read_inp_sys.cpp b/source/source_io/module_parameter/read_inp_sys.cpp index 97a40123eb8..78de1666024 100644 --- a/source/source_io/module_parameter/read_inp_sys.cpp +++ b/source/source_io/module_parameter/read_inp_sys.cpp @@ -246,7 +246,9 @@ Socket mode always computes energy. Force and stress extraction follows cal_forc * nep: Neuroevolution Potential * ks-lr: Kohn-Sham density functional theory + LR-TDDFT (Under Development Feature) * lr: LR-TDDFT with given KS orbitals (Under Development Feature) -* dfpt: density functional perturbation theory (Under Development Feature))"; +* dfpt: density functional perturbation theory (Under Development Feature) + +[NOTE] Excited-state forces (`cal_force = 1`) and atomic relaxation with `ks-lr` or `lr` currently require `gamma_only = 1` and pseudopotentials without nonlinear core correction (NLCC). NLCC core-density response-force and XC-kernel derivative terms are not implemented; these force requests are rejected after reading the pseudopotentials, before force evaluation. Multi-k spectra and spectra with NLCC pseudopotentials remain available with `cal_force = 0`. Ground-state forces are unaffected by this LR restriction.)"; item.default_value = "ksdft"; read_sync_string(input.esolver_type); item.check_value = [](const Input_Item& item, const Parameter& para) { @@ -271,6 +273,42 @@ Socket mode always computes energy. Force and stress extraction follows cal_forc "esolver_type=lr requires calculation=nscf (it reads the ground state " "wave function computed by a separate SCF run); please set calculation=nscf."); } + const bool is_lr = (para.input.esolver_type == "lr" || para.input.esolver_type == "ks-lr"); + const bool requests_lr_force = para.input.cal_force || para.input.calculation == "relax"; + if (is_lr && requests_lr_force && !para.input.gamma_only) + { + ModuleBase::WARNING_QUIT("ReadInput", + "LR-TDDFT excited-state forces and relaxation currently require gamma_only=1; " + "complex/multi-k gradients are not implemented. Use cal_force=0 for multi-k spectra."); + } + if (is_lr && para.input.calculation == "cell-relax") + { + ModuleBase::WARNING_QUIT("ReadInput", + "LR-TDDFT has no excited-state stress, so calculation=cell-relax cannot be driven by it. " + "Use calculation=relax to relax the atomic positions at fixed cell."); + } + if (is_lr && para.input.calculation == "md") + { + ModuleBase::WARNING_QUIT("ReadInput", + "excited-state MD is not supported: the non-adiabatic couplings between excited states " + "are not implemented, so a trajectory cannot switch surfaces at a crossing, and a " + "single-surface run would silently follow a fixed state index straight through one. " + "Use calculation=relax instead."); + } + if (para.input.esolver_type == "lr" && para.input.calculation == "relax") + { + ModuleBase::WARNING_QUIT("ReadInput", + "esolver_type=lr reads a ground state from disk that belongs to one fixed geometry, " + "so it cannot follow moving ions. Use esolver_type=ks-lr, which runs the SCF itself " + "at every ionic step."); + } + if (para.input.esolver_type == "ks-lr" && para.input.calculation == "relax" + && para.input.lr_solver == "spectrum") + { + ModuleBase::WARNING_QUIT("ReadInput", + "lr_solver=spectrum only reads previously written excitation amplitudes; it solves " + "nothing, so it cannot produce gradients for a relaxation."); + } }; this->add_item(item); } diff --git a/source/source_io/module_parameter/read_inp_tddft.cpp b/source/source_io/module_parameter/read_inp_tddft.cpp index 4436947b3d9..7ba7ec9eee4 100644 --- a/source/source_io/module_parameter/read_inp_tddft.cpp +++ b/source/source_io/module_parameter/read_inp_tddft.cpp @@ -996,7 +996,7 @@ void ReadInput::item_lr_tddft() item.annotation = "exchange correlation (XC) kernel for LR-TDDFT"; item.category = "Linear Response TDDFT"; item.type = "String"; - item.description = "The exchange-correlation kernel used in the calculation. Currently supported: RPA, LDA, PBE, HSE, HF."; + item.description = "The exchange-correlation kernel used in the calculation. Currently supported: RPA, LDA, PWLDA, PBE, and the hybrids HF, PBE0, HSE, B3LYP, CAM_PBEH, LC_PBE, LC_WPBE, LRC_WPBE, LRC_WPBEH. A hybrid kernel needs the ground state to use the same functional: the exact-exchange operator $[\\alpha+\\beta\\,\\mathrm{erfc}(\\mu r)]/r$ is built from exx_fock_alpha ($\\alpha$), exx_erfc_alpha ($\\beta$) and exx_erfc_omega ($\\mu$), which are keyed off dft_functional, not off this parameter."; item.default_value = "LDA"; item.unit = ""; read_sync_string(input.xc_kernel); @@ -1051,18 +1051,16 @@ void ReadInput::item_lr_tddft() } { Input_Item item("nocc"); - item.annotation = "the number of occupied orbitals to form the 2-particle basis ( <= nelec/2)"; + item.annotation = "the occupied orbital window ending at HOMO for LR-TDDFT"; item.category = "Linear Response TDDFT"; item.type = "Integer"; - item.description = R"(The number of occupied orbitals (up to HOMO) used in the LR-TDDFT calculation. -* Note: If the value is illegal ( > nelec/2 or <= 0), it will be autoset to nelec/2.)"; - item.default_value = "nband"; + item.description = R"(The number of occupied orbitals (up to HOMO) retained in the majority-spin LR-TDDFT window. A positive value selects a shared core prefix to discard from both spin channels; it does not change the ground-state occupations. +* If omitted, non-positive, or larger than the occupied majority-spin channel, all occupied orbitals are used. +* The full occupied window is determined by the effective electron number (including nelec_delta once) and the ground-state spin populations. For nspin=2, the minority-spin window has abs(N_up-N_down) fewer occupied orbitals.)"; + item.default_value = "all occupied orbitals"; item.unit = ""; read_sync_int(input.nocc); - item.reset_value = [](const Input_Item& item, Parameter& para) { - const int nocc_default = std::max(static_cast(para.input.nelec + 1) / 2, para.input.nbands); - if (para.input.nocc <= 0 || para.input.nocc > nocc_default) { para.input.nocc = nocc_default; } - }; + // Resolve the occupied window only after effective nelec and KS spin populations are known. this->add_item(item); } { @@ -1088,6 +1086,166 @@ void ReadInput::item_lr_tddft() read_sync_int(input.lr_nstates); this->add_item(item); } + { + Input_Item item("lr_target_state"); + item.annotation = "the initial excited state for geometry relaxation (0-based)"; + item.category = "Linear Response TDDFT"; + item.type = "Integer"; + item.description = R"(Initial excited-state index for `calculation = relax`, counted from 0 within the spin channel selected by `lr_target_spin`. + +The gradient of the followed state (or its degenerate multiplet selected by `lr_degen_mode`) is computed, since solving the Z-vector equation dominates the cost of an excited-state gradient. It also selects the state whose excitation energy is added to the ground-state total energy, which is the quantity the energy-based relaxation algorithms (`cg`, `bfgs`, `lbfgs`) line-search on. + +Ignored outside `calculation = relax`: a single-point run solves and reports the gradients of every state. + +[NOTE] This index seeds the first ionic step. Subsequent steps compute cross-geometry AO overlaps and transform them with the saved and current KS orbitals to compare excitation amplitudes in a common occupied/virtual basis. Orbital sign changes and rotations within those subspaces do not change the overlap criterion. Index changes and low overlaps are reported. In JT mode the selected normalized multiplet mixture and its orbital basis become the reference. The projection is not renormalized, so window leakage remains visible. Tracking currently requires a fixed cell and unchanged orbital windows; degeneracy and a state leaving the solved window still limit state identification. Increase `lr_nstates` or reduce the ionic step when the overlap is low.)"; + item.default_value = "0"; + item.unit = ""; + item.check_value = [](const Input_Item& item, const Parameter& para) { + // Both parameters only steer a relaxation. A single-point run solves and reports every + // state, so they are dead there and must not be able to abort it. + if (para.input.calculation != "relax") { return; } + if (para.input.lr_target_state < 0) + { + ModuleBase::WARNING_QUIT("ReadInput", "lr_target_state must be >= 0"); + } + // lr_nstates <= 0 means "all particle-hole pairs"; that count is only known once the + // ground state has been read, so ESolver_LR::setup_relax_target_() re-checks there + if (para.input.lr_nstates > 0 && para.input.lr_target_state >= para.input.lr_nstates) + { + ModuleBase::WARNING_QUIT("ReadInput", "lr_target_state must be < lr_nstates"); + } + const std::vector spins = { "singlet", "triplet", "updown" }; + if (std::find(spins.begin(), spins.end(), para.input.lr_target_spin) == spins.end()) + { + ModuleBase::WARNING_QUIT("ReadInput", "lr_target_spin must be singlet, triplet or updown"); + } + if (para.input.lr_target_spin == "triplet" && para.input.nspin == 1) + { + ModuleBase::WARNING_QUIT("ReadInput", + "lr_target_spin=triplet requires nspin=2: only the singlet channel is built at nspin=1"); + } + }; + read_sync_int(input.lr_target_state); + this->add_item(item); + } + { + Input_Item item("lr_degen_thr"); + item.annotation = "max excitation-energy spread of a degenerate multiplet whose gradient matrix is computed (Ry); 0 disables"; + item.category = "Linear Response TDDFT"; + item.type = "Real"; + item.description = R"(Excited states whose excitation energies lie within this threshold of each other are treated as one degenerate multiplet, and the full gradient matrix $G^{(A\alpha)}_{kl}=\langle X_k|\partial A/\partial R_{A\alpha}|X_l\rangle$ is computed for it in addition to the per-state gradients. Zero (the default) disables this and leaves the per-state gradients as the only output. + +At a $d$-fold degeneracy no single state has a gradient vector: the branch slopes along a displacement $u$ are the eigenvalues of $\sum_{A\alpha}u_{A\alpha}G^{(A\alpha)}$, and the eigenvectors that diagonalise it depend on $u$. The per-state gradients are the diagonal of $G$ in whichever basis the eigensolver happened to return, so only their sum (the trace) is basis-independent, while $G$ itself is the complete first-order information -- it is the linear vibronic coupling Hamiltonian of the multiplet. The extra cost is $d(d-1)/2$ further Z-vector solves per multiplet. + +The threshold proposes candidates; it cannot tell a true degeneracy from an accidental near-degeneracy, where the states have genuinely different excitation energies and the construction does not apply. Each multiplet's actual energy spread and the orthonormality of its eigenvectors are reported in the running log so the distinction can be made there. + +[NOTE] A sensible value is a few times the eigensolver threshold `lr_thr`, so that states split by real physics are not merged.)"; + item.default_value = "0"; + item.unit = "Ry"; + item.check_value = [](const Input_Item& item, const Parameter& para) { + if (para.input.lr_degen_thr < 0.0) + { + ModuleBase::WARNING_QUIT("ReadInput", "lr_degen_thr must be >= 0"); + } + }; + read_sync_double(input.lr_degen_thr); + this->add_item(item); + } + { + Input_Item item("lr_degen_mode"); + item.annotation = "what a relaxation follows when the target state is degenerate: state or average"; + item.category = "Linear Response TDDFT"; + item.type = "String"; + item.description = R"(What `calculation = relax` follows when `lr_target_state` sits inside a degenerate multiplet, as identified by `lr_degen_thr`. It has no effect when the target state is non-degenerate. + +* state: follow the gradient of that one state, as returned by the eigensolver. This is the historical behaviour and is what reproduces earlier results, but inside a multiplet it is not a well-defined quantity: the per-state gradients are the diagonal of the subspace gradient matrix in whichever basis the eigensolver happened to return, so they depend on numerical details of the diagonalisation rather than on physics. +* average: follow the multiplet average $\bar\Omega=\frac{1}{d}\sum_k\Omega_k$, whose gradient is $\operatorname{Tr}G/d$. Unlike the individual states this is a smooth, basis-independent surface, and by symmetry its gradient is totally symmetric, so following it keeps the geometry on the symmetric configuration. Both the reported energy and the reported gradient switch to the average together, which the energy-based optimisers (`cg`, `bfgs`, `lbfgs`) require -- a gradient of one surface line-searched against the energy of another does not converge. This mode deliberately does NOT find the Jahn-Teller distortion, which is orthogonal to the totally symmetric average gradient. +* jt: descend the Jahn-Teller branch. Solves $\min_{\|u\|=1}\lambda_{\min}(\sum_{A\alpha}u_{A\alpha}G^{(A\alpha)})$ -- a joint optimisation over the displacement and the mixing inside the multiplet, since the two are determined together -- and follows the force of the resulting branch. This needs the off-diagonal part of the gradient matrix, so it costs $d(d-1)/2$ further Z-vector solves per step on top of the $d$ diagonal ones. The running log reports the branch's force, its mixing coefficients, and its split into the part common to the multiplet and the part that actually breaks the degeneracy. + +[NOTE] The usual sequence is `average` first, to reach the symmetric stationary point, then `jt` from there: at a stationary point of the average surface the common part vanishes and the whole force is Jahn-Teller. `jt` is self-limiting -- once a step has split the multiplet there is no group left and the ordinary single-state gradient takes over. + +[NOTE] `jt` gives the first-order DIRECTION. The distortion amplitude also needs the harmonic term, and the step norm is Cartesian rather than mass-weighted. A linear molecule has no first-order term at all (the effect is second-order Renner-Teller) and the log says so.)"; + item.default_value = "state"; + item.unit = ""; + item.check_value = [](const Input_Item& item, const Parameter& para) { + const std::vector modes = { "state", "average", "jt" }; + if (std::find(modes.begin(), modes.end(), para.input.lr_degen_mode) == modes.end()) + { + ModuleBase::WARNING_QUIT("ReadInput", + "lr_degen_mode must be state, average or jt"); + } + // Both non-default modes need to know which states form the multiplet, and that + // grouping is what lr_degen_thr defines; without it there is nothing to act on. + if (para.input.lr_degen_mode != "state" && para.input.lr_degen_thr <= 0.0) + { + ModuleBase::WARNING_QUIT("ReadInput", + "lr_degen_mode=" + para.input.lr_degen_mode + + " requires lr_degen_thr > 0 to define the multiplet"); + } + }; + read_sync_string(input.lr_degen_mode); + this->add_item(item); + } + { + Input_Item item("lr_grad_solver"); + item.annotation = "the linear solver of the Z-vector equation for LR-TDDFT gradients"; + item.category = "Linear Response TDDFT"; + item.type = "String"; + item.description = R"(The method to solve the Z-vector (relaxed-density) equation $(A+B)Z=R$ in LR-TDDFT force and relaxation calculations, the linear-equation counterpart of `lr_solver`. Its dimension is $n_k n_{occ} n_{virt}$ summed over spin, where $n_{virt}$ counts every virtual band of the ground state, not only the `nvirt` window of the excitation. +* cg: Solve iteratively with the conjugate-gradient method, applying the orbital Hessian $A+B$ to a vector at each step. The matrix is never built. +* lapack: Construct the full matrix and solve directly with LAPACK (LU). Every MPI process holds the whole matrix and solves the same system. +* scalapack: Construct the matrix distributed over the MPI processes (2D block-cyclic) and solve with ScaLAPACK (LU). +* scalapack_chol: Construct the matrix distributed as for scalapack and solve by a ScaLAPACK Cholesky factorization, about half the flops of the LU and in place. +* elpa: Construct the matrix distributed as for scalapack and solve by an ELPA Cholesky factorization. + +[NOTE] The direct solvers build the matrix column by column, at the cost of one application of $A+B$ per column, which usually dominates the cost of the solve itself; scalapack, scalapack_chol and elpa need an MPI build, elpa also an ELPA build. + +[NOTE] scalapack_chol and elpa require $A+B$ to be positive definite, which holds at a stable ground state. If it is not, scalapack_chol stops with an error, while elpa does so only in a single-process run and hangs in a multi-process one; scalapack (LU) has no such requirement.)"; + item.default_value = "cg"; + item.unit = ""; + item.check_value = [](const Input_Item& item, const Parameter& para) { + const std::string& solver = para.input.lr_grad_solver; + const std::vector solvers = { "cg", "lapack", "scalapack", "scalapack_chol", "elpa" }; + if (std::find(solvers.begin(), solvers.end(), solver) == solvers.end()) + { + ModuleBase::WARNING_QUIT("ReadInput", "lr_grad_solver must be cg, lapack, scalapack, scalapack_chol or elpa"); + } +#ifndef __MPI + if (solver == "scalapack" || solver == "scalapack_chol" || solver == "elpa") + { + ModuleBase::WARNING_QUIT("ReadInput", + "lr_grad_solver = " + solver + " needs an MPI build; use cg or lapack"); + } +#endif +#ifndef __ELPA + if (solver == "elpa") + { + ModuleBase::WARNING_QUIT("ReadInput", + "lr_grad_solver = elpa needs ABACUS compiled with ELPA; use scalapack"); + } +#endif + }; + read_sync_string(input.lr_grad_solver); + this->add_item(item); + } + { + Input_Item item("lr_target_spin"); + item.annotation = "spin channel of lr_target_state: singlet, triplet or updown"; + item.category = "Linear Response TDDFT"; + item.type = "String"; + item.description = R"(Which spin channel `lr_target_state` indexes. + +* singlet / triplet: the two closed-shell channels solved at `nspin = 2`. At `nspin = 1` only `singlet` exists. +* updown: the single spin-conserving channel of an open-shell calculation (`lr_unrestricted`, or a spin-polarised ground state with a non-zero moment). + +An open-shell calculation has only one channel, so any value is accepted there and relaxes that channel; an explicit `triplet` is reported as ignored. A closed-shell calculation rejects `updown`, since singlet and triplet are separate states with separate gradients. + +Ignored outside `calculation = relax`.)"; + item.default_value = "singlet"; + item.unit = ""; + read_sync_string(input.lr_target_spin); + this->add_item(item); + } { Input_Item item("lr_unrestricted"); item.annotation = "Whether to use unrestricted construction for LR-TDDFT"; diff --git a/source/source_io/module_wf/read_wfc_nao.cpp b/source/source_io/module_wf/read_wfc_nao.cpp index 3148e2c0e2f..e7cd0d9a7e7 100644 --- a/source/source_io/module_wf/read_wfc_nao.cpp +++ b/source/source_io/module_wf/read_wfc_nao.cpp @@ -72,6 +72,88 @@ bool read_binary_wfc_data(std::ifstream& ifs, std::complex& data) } // namespace +bool ModuleIO::read_wfc_nao_spin_populations(const std::string& readin_dir, + int nkstot, + int nspin, + bool gamma_only, + bool binary, + int my_rank, + std::vector& populations) +{ + populations.assign(nspin, 0.0); + bool success = (nspin == 1 || nspin == 2) && nkstot > 0 && nkstot % nspin == 0; + std::vector ik2iktot; + for (int ik = 0; ik < nkstot; ++ik) { ik2iktot.push_back(ik); } + if (gamma_only && nkstot != nspin) { success = false; } + if (my_rank == 0 && success) + { + const int nk = ik2iktot.size() / nspin; + const int read_type = binary ? 2 : 1; + for (int ik = 0; ik < static_cast(ik2iktot.size()) && success; ++ik) + { + const std::string file = ModuleIO::filename_output(readin_dir, "wf", "nao", + ik, ik2iktot, nspin, nkstot, read_type, false, gamma_only, -1); + const std::ios_base::openmode mode = binary ? std::ios::in | std::ios::binary : std::ios::in; + std::ifstream ifs(file.c_str(), mode); + if (!gamma_only) + { + int ik_file = 0; + double k = 0.0; + success = read_record_value(ifs, ik_file, binary); + if (binary) + { + success = success && read_binary_value(ifs, k) + && read_binary_value(ifs, k) && read_binary_value(ifs, k); + } + else + { + ifs >> k >> k >> k; + success = success && static_cast(ifs); + } + success = success && ik_file == ik + 1; + } + int nbands = 0; + int nbasis = 0; + success = success && read_record_value(ifs, nbands, binary) + && read_record_value(ifs, nbasis, binary) && nbands > 0 && nbasis > 0; + for (int ib = 0; ib < nbands && success; ++ib) + { + int ib_file = 0; + double energy = 0.0; + double occupation = 0.0; + success = read_record_value(ifs, ib_file, binary) + && read_record_value(ifs, energy, binary) + && read_record_value(ifs, occupation, binary) && ib_file == ib + 1; + populations[ik / nk] += occupation; + const int ncoeff = gamma_only ? nbasis : 2 * nbasis; + if (binary) + { + const std::streamoff bytes = ncoeff * sizeof(double); + ifs.ignore(bytes); + success = success && ifs.gcount() == bytes; + } + else + { + double coefficient = 0.0; + for (int i = 0; i < ncoeff && success; ++i) + { + ifs >> coefficient; + success = static_cast(ifs); + } + } + } + } + } +#ifdef __MPI + Parallel_Common::bcast_bool(success); + if (success) + { + Parallel_Common::bcast_double(populations.data(), nspin); + } +#endif + return success; +} + // mohan add 2025-10-19 void ModuleIO::read_wfc_nao_one_data(std::ifstream& ifs, float& data) { diff --git a/source/source_io/module_wf/read_wfc_nao.h b/source/source_io/module_wf/read_wfc_nao.h index 0b3b23e11e2..865669e1468 100644 --- a/source/source_io/module_wf/read_wfc_nao.h +++ b/source/source_io/module_wf/read_wfc_nao.h @@ -8,6 +8,17 @@ // mohan add 2010-09-09 namespace ModuleIO { +/** Read spin populations from all bands, without allocating wavefunction matrices. + * File occupations are already weighted by their k-point weights. + */ +bool read_wfc_nao_spin_populations(const std::string& readin_dir, + int nkstot, + int nspin, + bool gamma_only, + bool binary, + int my_rank, + std::vector& populations); + /** * @brief Reads a single data value from an input file stream. * diff --git a/source/source_io/test/read_input_ptest.cpp b/source/source_io/test/read_input_ptest.cpp index 4784aa3f643..19b32ff768f 100644 --- a/source/source_io/test/read_input_ptest.cpp +++ b/source/source_io/test/read_input_ptest.cpp @@ -452,12 +452,15 @@ TEST_F(InputParaTest, ParaRead) EXPECT_EQ(param.inp.sc_scf_thr, 1e-3); EXPECT_EQ(param.inp.sc_drop_thr, 1e-3); EXPECT_EQ(param.inp.lr_nstates, 1); - EXPECT_EQ(param.inp.nocc, param.inp.nbands); + // nocc is resolved from ground-state spin populations after SCF (see esolver_lr_lcao_tddft.cpp), + // not at INPUT-parse time, so a run with no explicit `nocc` keeps its raw unresolved default here. + EXPECT_EQ(param.inp.nocc, -1); EXPECT_EQ(param.inp.nvirt, 1); EXPECT_EQ(param.inp.xc_kernel, "LDA"); EXPECT_EQ(param.inp.lr_init_xc_kernel[0], "default"); EXPECT_EQ(param.inp.lr_solver, "dav"); EXPECT_DOUBLE_EQ(param.inp.lr_thr, 1e-2); + EXPECT_EQ(param.inp.lr_grad_solver, "cg"); EXPECT_FALSE(param.inp.lr_unrestricted); EXPECT_FALSE(param.inp.out_wfc_lr); EXPECT_EQ(param.inp.abs_wavelen_range.size(), 2); diff --git a/source/source_io/test/read_wfc_nao_test.cpp b/source/source_io/test/read_wfc_nao_test.cpp index 6107db23036..7ecc84447e4 100644 --- a/source/source_io/test/read_wfc_nao_test.cpp +++ b/source/source_io/test/read_wfc_nao_test.cpp @@ -453,6 +453,99 @@ TEST_F(ReadWfcNaoTest, RejectTruncatedBinary) +TEST_F(ReadWfcNaoTest, CompleteSpinPopulations) +{ + const int nbands = 4; + const int nlocal = 2; + const int nspin = 2; + const std::vector ik2iktot = {0, 1}; + ModuleBase::matrix energies(2, nbands); + ModuleBase::matrix occupations(2, nbands); + const std::vector real_coefficients(nbands * nlocal, 0.25); + const std::vector> complex_coefficients(nbands * nlocal, {0.25, 0.125}); + for (int ib = 0; ib < nbands; ++ib) + { + occupations(0, ib) = ib < 3 ? 1.0 : 0.0; + occupations(1, ib) = ib < 1 ? 1.0 : 0.0; + } + for (bool gamma_only : {true, false}) + { + for (bool binary : {false, true}) + { + const int file_type = binary ? 2 : 1; + std::vector files; + for (int ik = 0; ik < 2; ++ik) + { + const std::string file = ModuleIO::filename_output(binary_test_dir, "wf", "nao", + ik, ik2iktot, nspin, 2, file_type, false, gamma_only, -1); + files.push_back(file); + if (my_rank == 0) + { + if (gamma_only) + { + ModuleIO::wfc_nao_write2file(file, real_coefficients.data(), nlocal, + ik, energies, occupations, binary, false); + } + else + { + const ModuleBase::Vector3 kpoint(0.25, 0.0, 0.0); + ModuleIO::wfc_nao_write2file_complex(file, complex_coefficients.data(), nlocal, + ik, kpoint, energies, occupations, binary, false); + } + } + } + std::vector populations; + ASSERT_TRUE(ModuleIO::read_wfc_nao_spin_populations(binary_test_dir, 2, nspin, gamma_only, binary, my_rank, populations)); + EXPECT_DOUBLE_EQ(populations[0], 3.0); + EXPECT_DOUBLE_EQ(populations[1], 1.0); + if (my_rank == 0) + { + for (const auto& file : files) { std::remove(file.c_str()); } + } + } + } +} + +TEST_F(ReadWfcNaoTest, CompleteGlobalKPointPopulations) +{ + const int nks = 4; // two k points per spin, independent of the reader's pool + const int nbands = 4; + const int nlocal = 2; + const std::vector global_indices = {0, 1, 2, 3}; + ModuleBase::matrix energies(nks, nbands); + ModuleBase::matrix occupations(nks, nbands); + const std::vector> coefficients(nbands * nlocal, {0.25, 0.125}); + std::vector files; + for (int ik = 0; ik < nks; ++ik) + { + for (int ib = 0; ib < nbands; ++ib) + { + const int occupied = ik < 2 ? 3 : 1; + occupations(ik, ib) = ib < occupied ? 0.5 : 0.0; + } + const std::string file = ModuleIO::filename_output(binary_test_dir, "wf", "nao", + ik, global_indices, 2, nks, 1, false, false, -1); + files.push_back(file); + if (my_rank == 0) + { + const ModuleBase::Vector3 kpoint(0.25 * ik, 0.0, 0.0); + ModuleIO::wfc_nao_write2file_complex(file, coefficients.data(), nlocal, + ik, kpoint, energies, occupations, false, false); + } + } + std::vector populations; + ASSERT_TRUE(ModuleIO::read_wfc_nao_spin_populations(binary_test_dir, nks, + 2, false, false, my_rank, populations)); + EXPECT_DOUBLE_EQ(populations[0], 3.0); + EXPECT_DOUBLE_EQ(populations[1], 1.0); + EXPECT_FALSE(ModuleIO::read_wfc_nao_spin_populations(binary_test_dir, 0, + 2, false, false, my_rank, populations)); + if (my_rank == 0) + { + for (const auto& file : files) { std::remove(file.c_str()); } + } +} + #ifdef __MPI int main(int argc, char** argv) { diff --git a/source/source_io/test_serial/read_input_item_test.cpp b/source/source_io/test_serial/read_input_item_test.cpp index ffe37e6ba80..8f351fa72c9 100644 --- a/source/source_io/test_serial/read_input_item_test.cpp +++ b/source/source_io/test_serial/read_input_item_test.cpp @@ -2239,16 +2239,42 @@ TEST_F(InputTest, Item_test2) output = testing::internal::GetCapturedStdout(); EXPECT_THAT(output, testing::HasSubstr("NOTICE")); } - { // nocc - auto it = find_label("nocc", readinput.input_lists); - TestParameters::input(param).nocc = 5; - TestParameters::input(param).nbands = 4; - TestParameters::input(param).nelec = 0.0; - it->second.reset_value(it->second, param); - EXPECT_EQ(TestParameters::input(param).nocc, 4); - TestParameters::input(param).nocc = 0; - it->second.reset_value(it->second, param); - EXPECT_EQ(TestParameters::input(param).nocc, 4); + // nocc no longer has a reset_value here: it's resolved from ground-state spin populations + // after SCF (see esolver_lr_lcao_tddft.cpp), not at INPUT-parse time. + { // lr_grad_solver + auto it = find_label("lr_grad_solver", readinput.input_lists); + for (const std::string solver : { "cg", "lapack" }) + { + TestParameters::input(param).lr_grad_solver = solver; + it->second.check_value(it->second, param); // accepted in every build + } + TestParameters::input(param).lr_grad_solver = "gmres"; + testing::internal::CaptureStdout(); + EXPECT_EXIT(it->second.check_value(it->second, param), ::testing::ExitedWithCode(1), ""); + output = testing::internal::GetCapturedStdout(); + EXPECT_THAT(output, testing::HasSubstr("lr_grad_solver must be cg, lapack, scalapack, scalapack_chol or elpa")); + for (const std::string solver : { "scalapack", "scalapack_chol" }) + { + TestParameters::input(param).lr_grad_solver = solver; +#ifdef __MPI + it->second.check_value(it->second, param); +#else + testing::internal::CaptureStdout(); + EXPECT_EXIT(it->second.check_value(it->second, param), ::testing::ExitedWithCode(1), ""); + output = testing::internal::GetCapturedStdout(); + EXPECT_THAT(output, testing::HasSubstr("needs an MPI build")); +#endif + } +#if defined(__MPI) && defined(__ELPA) + TestParameters::input(param).lr_grad_solver = "elpa"; + it->second.check_value(it->second, param); +#else // rejected by the MPI or the ELPA check; both messages name the value + TestParameters::input(param).lr_grad_solver = "elpa"; + testing::internal::CaptureStdout(); + EXPECT_EXIT(it->second.check_value(it->second, param), ::testing::ExitedWithCode(1), ""); + output = testing::internal::GetCapturedStdout(); + EXPECT_THAT(output, testing::HasSubstr("lr_grad_solver = elpa")); +#endif } } diff --git a/source/source_io/test_serial/read_input_test.cpp b/source/source_io/test_serial/read_input_test.cpp index 236b8e8f87d..6070e36910e 100644 --- a/source/source_io/test_serial/read_input_test.cpp +++ b/source/source_io/test_serial/read_input_test.cpp @@ -325,6 +325,24 @@ TEST_F(InputTest, ValidateLrRequiresNscf) EXPECT_EQ(valid_param.inp.calculation, "nscf"); } +TEST_F(InputTest, RejectComplexLrForcesBeforeSCF) +{ + const std::string base = "basis_type lcao\nesolver_type ks-lr\ngamma_only 0\n"; + const std::string reason = "currently require gamma_only=1"; + const std::string cg = base + "cal_force 1\nlr_grad_solver cg\n"; + const std::string lapack = base + "cal_force 1\nlr_grad_solver lapack\n"; + const std::string relax = base + "calculation relax\ncal_force 0\n"; + expect_invalid_input("complex_force_cg_INPUT", cg, reason); + expect_invalid_input("complex_force_lapack_INPUT", lapack, reason); + expect_invalid_input("complex_relax_INPUT", relax, reason); + + Parameter valid; + const std::string spectra = base + "cal_force 0\n"; + EXPECT_NO_THROW(read_parameters("complex_spectra_INPUT", spectra, valid)); + EXPECT_FALSE(valid.inp.gamma_only); + EXPECT_FALSE(valid.inp.cal_force); +} + TEST_F(InputTest, ValidateDeepksOutputFrequency) { Parameter default_param; @@ -397,3 +415,12 @@ TEST_F(InputTest, Check) EXPECT_THAT(output, testing::HasSubstr("INPUT parameters have been successfully checked!")); EXPECT_TRUE(std::remove("./INPUT.ref") == 0); } + +TEST_F(InputTest, ExplicitNoccSurvivesUnresolvedElectronNumber) +{ + Parameter parameters; + read_parameters("nocc_window_INPUT", "basis_type lcao\nesolver_type ks-lr\nnocc 4\n", parameters); + EXPECT_EQ(parameters.inp.nelec, 0.0); + EXPECT_EQ(parameters.inp.nbands, 0); + EXPECT_EQ(parameters.inp.nocc, 4); +} diff --git a/source/source_lcao/CMakeLists.txt b/source/source_lcao/CMakeLists.txt index 41f0085a014..5e79e0f7c85 100644 --- a/source/source_lcao/CMakeLists.txt +++ b/source/source_lcao/CMakeLists.txt @@ -19,6 +19,7 @@ if(ENABLE_LCAO) module_operator_lcao/deepks_lcao.cpp module_operator_lcao/op_exx_lcao.cpp module_operator_lcao/overlap.cpp + module_operator_lcao/ovlp_block.cpp module_operator_lcao/overlap_fs.cpp module_operator_lcao/ekinetic.cpp module_operator_lcao/ekinetic_fs.cpp diff --git a/source/source_lcao/force_stress_lcao.cpp b/source/source_lcao/force_stress_lcao.cpp index aab0dc61166..691db97de2c 100644 --- a/source/source_lcao/force_stress_lcao.cpp +++ b/source/source_lcao/force_stress_lcao.cpp @@ -303,7 +303,7 @@ void Force_Stress_LCAO::cal_operator_fs(UnitCell& ucell, // Calculate local potential force/stress (vl_dphi) // This uses grid integration, not operator-based method edm_cal.ParaV = &pv; - PulayForceStress::cal_pulay_fs(parts.fvl_dphi, sparts.svl_dphi, *dmat.dm, ucell, pelec->pot, + PulayForceStress::cal_pulay_fs(cfg.nspin, parts.fvl_dphi, sparts.svl_dphi, *dmat.dm, ucell, pelec->pot, isforce, isstress, false /*reset dm to gint*/); } else if (cfg.nspin == 4) @@ -340,7 +340,7 @@ void Force_Stress_LCAO::cal_operator_fs(UnitCell& ucell, // Local-potential (vl_dphi) Pulay term via grid integration edm_cal.ParaV = &pv; - PulayForceStress::cal_pulay_fs(parts.fvl_dphi, sparts.svl_dphi, *dmat.dm, ucell, pelec->pot, + PulayForceStress::cal_pulay_fs(cfg.nspin, parts.fvl_dphi, sparts.svl_dphi, *dmat.dm, ucell, pelec->pot, isforce, isstress, false); } diff --git a/source/source_lcao/module_bse/hamilt_bse.cpp b/source/source_lcao/module_bse/hamilt_bse.cpp index a162fe3ffe8..f331e2464a8 100644 --- a/source/source_lcao/module_bse/hamilt_bse.cpp +++ b/source/source_lcao/module_bse/hamilt_bse.cpp @@ -582,7 +582,7 @@ void HamiltBSE>::grid_calculation(hamilt::HContainer void { - LR_Util::get_DMR_real_imag_part(*this->DM_trans, DM_trans_real_imag, ucell.nat, type); + LR_Util::get_DMR_real_imag_part(*this->DM_trans, DM_trans_real_imag, type); // if (this->first_print)LR_Util::print_DMR(DM_trans_real_imag, ucell.nat, "DMR(2d, real)"); // 4.1. transition density rho on grid @@ -603,7 +603,7 @@ void HamiltBSE>::grid_calculation(hamilt::HContainerucell.nat, "VR(real, 2d)"); - LR_Util::set_HR_real_imag_part(HR_real_imag, VR, ucell.nat, type); + LR_Util::set_HR_real_imag_part(HR_real_imag, VR, type); }; VR.set_zero(); dmR_to_hR('R'); //real diff --git a/source/source_lcao/module_lr/CMakeLists.txt b/source/source_lcao/module_lr/CMakeLists.txt index d57d02afc4f..121653cd2e2 100644 --- a/source/source_lcao/module_lr/CMakeLists.txt +++ b/source/source_lcao/module_lr/CMakeLists.txt @@ -3,31 +3,44 @@ if(ENABLE_LCAO) add_subdirectory(ao_to_mo_transformer) add_subdirectory(dm_trans) add_subdirectory(ri_benchmark) + if(BUILD_TESTING) + add_subdirectory(test) + endif() list(APPEND objects utils/lr_util.cpp - utils/lr_util_hcontainer.cpp utils/lr_io.cpp utils/exciton_plotter.cpp ao_to_mo_transformer/ao_to_mo_parallel.cpp ao_to_mo_transformer/ao_to_mo_serial.cpp + ao_to_mo_transformer/cvcx_par.cpp + ao_to_mo_transformer/cvcx_serial.cpp dm_trans/dm_trans_parallel.cpp dm_trans/dm_trans_serial.cpp operator_casida/operator_lr_hxc.cpp operator_casida/operator_lr_exx.cpp potentials/pot_hxc_lrtd.cpp + potentials/pot_grad_xc.cpp lr_spectrum.cpp lr_spectrum_velocity.cpp hamilt_casida.cpp - potentials/xc_kernel.cpp) + potentials/xc_kernel.cpp + gradient_output.cpp + lr_grad_cs.cpp + lr_grad_os.cpp + lr_force.cpp + exx_proj.cpp + lr_force_aux.cpp + grad_degen.cpp + root_track.cpp + root_ovlp.cpp + grad_jt.cpp + cal_edm.cpp + zeqlin_solv.cpp) - # BSE-related code: only compiled with LibRI (__EXX), all its - # consumers (ESolver_BSE, RI benchmark in hamilt_casida.h) are - # already guarded by __EXX + # The RI k-point list reader requires LibRI headers. if(ENABLE_LIBRI) - list(APPEND objects - utils/lr_io_krlist.cpp - ) + list(APPEND objects utils/lr_io_krlist.cpp) endif() add_library( diff --git a/source/source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo.h b/source/source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo.h index cc812f4b1d9..5c21c3fbf64 100644 --- a/source/source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo.h +++ b/source/source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo.h @@ -18,7 +18,8 @@ namespace LR const int& nocc, const int& nvirt, T* const mat_mo, - const LR_Util::MO_TYPE type = LR_Util::VO); + const LR_Util::MO_TYPE type = LR_Util::VO, + const T factor = static_cast(1.0)); template void ao_to_mo_blas( const std::vector& mat_ao, @@ -27,7 +28,8 @@ namespace LR const int& nvirt, T* const mat_mo, const bool add_on = true, - const LR_Util::MO_TYPE type = LR_Util::VO); + const LR_Util::MO_TYPE type = LR_Util::VO, + const T factor = static_cast(1.0)); #ifdef __MPI template void ao_to_mo_pblas( @@ -41,7 +43,8 @@ namespace LR const Parallel_2D& pmat_mo, T* const mat_mo, const bool add_on = true, - const LR_Util::MO_TYPE type = LR_Util::VO); + const LR_Util::MO_TYPE type = LR_Util::VO, + const T factor = static_cast(1.0)); #endif } diff --git a/source/source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo_parallel.cpp b/source/source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo_parallel.cpp index f4d1bb79dea..e48d8fe1ba7 100644 --- a/source/source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo_parallel.cpp +++ b/source/source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo_parallel.cpp @@ -21,7 +21,8 @@ namespace LR const Parallel_2D& pmat_mo, double* mat_mo, const bool add_on, - const LR_Util::MO_TYPE type) + const LR_Util::MO_TYPE type, + const double factor) { ModuleBase::TITLE("LR", "ao_to_mo_pblas"); assert(pmat_ao.comm() == pcoeff.comm() && pmat_ao.comm() == pmat_mo.comm()); @@ -61,7 +62,7 @@ namespace LR // mat_mo = c ^ TVc // descC puts M(nvirt) to row ScalapackConnector::gemm(transa, transb, nmo2, nmo1, naos, - alpha, coeff.get_pointer(), i1, imo2, pcoeff.desc, + factor, coeff.get_pointer(), i1, imo2, pcoeff.desc, Vc.data(), i1, i1, pVc.desc, beta, mat_mo + start, i1, i1, pmat_mo.desc); @@ -80,7 +81,8 @@ namespace LR const Parallel_2D& pmat_mo, std::complex* const mat_mo, const bool add_on, - const LR_Util::MO_TYPE type) + const LR_Util::MO_TYPE type, + const std::complex factor) { ModuleBase::TITLE("LR", "ao_to_mo_pblas"); assert(pmat_ao.comm() == pcoeff.comm() && pmat_ao.comm() == pmat_mo.comm()); @@ -120,7 +122,7 @@ namespace LR // mat_mo = c ^ TVc // descC puts M(nvirt) to row ScalapackConnector::gemm(transa, transb, nmo2, nmo1, naos, - alpha, coeff.get_pointer(), i1, imo2, pcoeff.desc, + factor, coeff.get_pointer(), i1, imo2, pcoeff.desc, Vc.data>(), i1, i1, pVc.desc, beta, mat_mo + start, i1, i1, pmat_mo.desc); } diff --git a/source/source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo_serial.cpp b/source/source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo_serial.cpp index 74afc6e8033..3c191077633 100644 --- a/source/source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo_serial.cpp +++ b/source/source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo_serial.cpp @@ -11,7 +11,8 @@ namespace LR const int& nocc, const int& nvirt, double* mat_mo, - const LR_Util::MO_TYPE type) + const LR_Util::MO_TYPE type, + const double factor) { ModuleBase::TITLE("LR", "ao_to_mo_forloop_serial"); const int nks = mat_ao.size(); @@ -33,7 +34,7 @@ namespace LR { for (int mu = 0;mu < naos;++mu) { - mat_mo[start + p * nmo2 + q] += coeff(imo2 + q, mu) * mat_ao[isk].data()[nu * naos + mu] * coeff(imo1 + p, nu); + mat_mo[start + p * nmo2 + q] += coeff(imo2 + q, mu) * mat_ao[isk].data()[nu * naos + mu] * coeff(imo1 + p, nu) * factor; } } } @@ -47,7 +48,8 @@ namespace LR const int& nocc, const int& nvirt, std::complex* const mat_mo, - const LR_Util::MO_TYPE type) + const LR_Util::MO_TYPE type, + const std::complex factor) { ModuleBase::TITLE("LR", "ao_to_mo_forloop_serial"); const int nks = mat_ao.size(); @@ -69,7 +71,7 @@ namespace LR { for (int mu = 0;mu < naos;++mu) { - mat_mo[start + p * nmo2 + q] += std::conj(coeff(imo2 + q, mu)) * mat_ao[isk].data>()[nu * naos + mu] * coeff(imo1 + p, nu); + mat_mo[start + p * nmo2 + q] += std::conj(coeff(imo2 + q, mu)) * mat_ao[isk].data>()[nu * naos + mu] * coeff(imo1 + p, nu) * factor; } } } @@ -84,7 +86,8 @@ namespace LR const int& nvirt, double* mat_mo, const bool add_on, - const LR_Util::MO_TYPE type) + const LR_Util::MO_TYPE type, + const double factor) { ModuleBase::TITLE("LR", "ao_to_mo_blas"); const int nks = mat_ao.size(); @@ -110,7 +113,7 @@ namespace LR transa = 'T'; //mat_mo=coeff^TVc (nvirt major) - BlasConnector::gemm(transb, transa, nmo1, nmo2, naos, alpha, + BlasConnector::gemm(transb, transa, nmo1, nmo2, naos, factor, Vc.data(), naos, coeff.get_pointer(imo2), naos, beta, mat_mo + start, nmo2); } @@ -123,7 +126,8 @@ namespace LR const int& nvirt, std::complex* const mat_mo, const bool add_on, - const LR_Util::MO_TYPE type) + const LR_Util::MO_TYPE type, + const std::complex factor) { ModuleBase::TITLE("LR", "ao_to_mo_blas"); const int nks = mat_ao.size(); @@ -149,7 +153,7 @@ namespace LR transa = 'C'; //mat_mo=coeff^\dagger Vc (nvirt major) - BlasConnector::gemm(transb, transa, nmo1, nmo2, naos, alpha, + BlasConnector::gemm(transb, transa, nmo1, nmo2, naos, factor, Vc.data>(), naos, coeff.get_pointer(imo2), naos, beta, mat_mo + start, nmo2); } diff --git a/source/source_lcao/module_lr/ao_to_mo_transformer/cvcx.h b/source/source_lcao/module_lr/ao_to_mo_transformer/cvcx.h new file mode 100644 index 00000000000..10021879017 --- /dev/null +++ b/source/source_lcao/module_lr/ao_to_mo_transformer/cvcx.h @@ -0,0 +1,88 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_AO_TO_MO_TRANSFORMER_CVCX_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_AO_TO_MO_TRANSFORMER_CVCX_H +#include +#include "source_psi/psi.h" +#include +#ifdef __MPI +#include "source_base/parallel_2d.h" +#endif +namespace LR +{ + // occ + /// $\sum_{k\mu\nu}C^*_{\mu i}K_{\mu\nu}C_{\nu k}X_{ak}^*$ + template + void CVCX_occ_forloop_serial( + const std::vector& V_istate, + const psi::Psi& c, + const T* const X_istate, + const int& naos, + const int& nocc, + const int& nvirt, + T* const AX_istate); + template + void CVCX_occ_blas( + const std::vector& V_istate, + const psi::Psi& c, + const T* const X_istate, + const int& naos, + const int& nocc, + const int& nvirt, + T* const AX_istate, + const bool add_on = true, + const T factor = (T)1.0); +#ifdef __MPI + template + void CVCX_occ_pblas( + const std::vector& V_istate, + const Parallel_2D& pmat, + const psi::Psi& c, + const Parallel_2D& pc, + const T* const X_istate, + const Parallel_2D& px, + const int& naos, + const int& nocc, + const int& nvirt, + T* const AX_istate, + const bool add_on = true, + const T factor = (T)1.0); +#endif + // virt + /// $\sum_{b\mu\nu}X^*_{bi}C^*_{\mu b}K_{\mu\nu}C_{\nu a}$ + template + void CVCX_virt_forloop_serial( + const std::vector& V_istate, + const psi::Psi& c, + const T* const X_istate, + const int& naos, + const int& nocc, + const int& nvirt, + T* const AX_istate); + template + void CVCX_virt_blas( + const std::vector& V_istate, + const psi::Psi& c, + const T* const X_istate, + const int& naos, + const int& nocc, + const int& nvirt, + T* const AX_istate, + const bool add_on = true, + const T factor = (T)1.0); +#ifdef __MPI + template + void CVCX_virt_pblas( + const std::vector& V_istate, + const Parallel_2D& pmat, + const psi::Psi& c, + const Parallel_2D& pc, + const T* const X_istate, + const Parallel_2D& px, + const int& naos, + const int& nocc, + const int& nvirt, + T* const AX_istate, + const bool add_on = true, + const T factor = (T)1.0); +#endif +} +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_AO_TO_MO_TRANSFORMER_CVCX_H diff --git a/source/source_lcao/module_lr/ao_to_mo_transformer/cvcx_par.cpp b/source/source_lcao/module_lr/ao_to_mo_transformer/cvcx_par.cpp new file mode 100644 index 00000000000..afb6dd37355 --- /dev/null +++ b/source/source_lcao/module_lr/ao_to_mo_transformer/cvcx_par.cpp @@ -0,0 +1,264 @@ +#ifdef __MPI +#include "cvcx.h" +#include "source_base/module_external/scalapack_connector.h" +#include "source_base/tool_title.h" +#include "source_lcao/module_lr/utils/lr_util.h" +#include "source_lcao/module_lr/utils/lr_util_print.h" +namespace LR +{ + template <> + void CVCX_occ_pblas( + const std::vector& V_istate, + const Parallel_2D& pmat, + const psi::Psi& c, + const Parallel_2D& pc, + const double* const X_istate, + const Parallel_2D& px, + const int& naos, + const int& nocc, + const int& nvirt, + double* const AX_istate, + const bool add_on, + const double factor) + { + ModuleBase::TITLE("hamilt_lrtd", "CVCX_occ_pblas"); + assert(pmat.comm() == pc.comm()); + assert(pmat.comm() == px.comm()); + assert(pmat.blacs_ctxt == pc.blacs_ctxt); + assert(pmat.blacs_ctxt == px.blacs_ctxt); + assert(px.get_local_size() > 0); + + const int nks = c.get_nk(); + assert(V_istate.size() == nks); + + Parallel_2D pcv; + LR_Util::setup_2d_division(pcv, pmat.get_block_size(), nocc, naos, pmat.blacs_ctxt); + Parallel_2D pcx; + LR_Util::setup_2d_division(pcx, pmat.get_block_size(), naos, nvirt, pmat.blacs_ctxt); + for (int isk = 0;isk < nks;++isk) + { + c.fix_k(isk); + const int start = isk * px.get_local_size(); + + const int i1 = 1; + const int ivirt = nocc + 1; + const char trans = 'T'; + const char notrans = 'N'; //c is col major + const double one = 1.0; + const double zero = 0.0; + + // c^TV[nocc*naos] + container::Tensor cv(DAT::DT_DOUBLE, DEV::CpuDevice, { pcv.get_col_size(), pcv.get_row_size() }); + pdgemm_(&trans, ¬rans, &nocc, &naos, &naos, + &one, c.get_pointer(), &i1, &i1, pc.desc, + V_istate[isk].data(), &i1, &i1, pmat.desc, + &zero, cv.data(), &i1, &i1, pcv.desc); + + // cX^T[naos*nvirt] + container::Tensor cx(DAT::DT_DOUBLE, DEV::CpuDevice, { pcx.get_col_size(), pcx.get_row_size() }); + pdgemm_(¬rans, &trans, &naos, &nvirt, &nocc, + &one, c.get_pointer(), &i1, &i1, pc.desc, + X_istate + start, &i1, &i1, px.desc, + &zero, cx.data(), &i1, &i1, pcx.desc); + + //AX_istate=[cX^T]^T[c^TV]^T (nvirt major) + pdgemm_(&trans, &trans, &nvirt, &nocc, &naos, + &factor, cx.data(), &i1, &i1, pcx.desc, + cv.data(), &i1, &i1, pcv.desc, + add_on ? &one : &zero, AX_istate + start, &i1, &i1, px.desc); + } + } + + template <> + void CVCX_occ_pblas( + const std::vector& V_istate, + const Parallel_2D& pmat, + const psi::Psi, base_device::DEVICE_CPU>& c, + const Parallel_2D& pc, + const std::complex* const X_istate, + const Parallel_2D& px, + const int& naos, + const int& nocc, + const int& nvirt, + std::complex* const AX_istate, + const bool add_on, + const std::complex factor) + { + ModuleBase::TITLE("hamilt_lrtd", "CVCX_occ_pblas"); + assert(pmat.comm() == pc.comm()); + assert(pmat.comm() == px.comm()); + assert(pmat.blacs_ctxt == pc.blacs_ctxt); + assert(pmat.blacs_ctxt == px.blacs_ctxt); + assert(px.get_local_size() > 0); + + int nks = c.get_nk(); + assert(V_istate.size() == nks); + + Parallel_2D pcv; + LR_Util::setup_2d_division(pcv, pmat.get_block_size(), nocc, naos, pmat.blacs_ctxt); + Parallel_2D pcx; + LR_Util::setup_2d_division(pcx, pmat.get_block_size(), naos, nvirt, pmat.blacs_ctxt); + for (int isk = 0;isk < nks;++isk) + { + c.fix_k(isk); + const int start = isk * px.get_local_size(); + + const int i1 = 1; + const int ivirt = nocc + 1; + const char trans = 'T'; + const char dagger = 'C'; + const char notrans = 'N'; //c is col major + const std::complex one(1.0, 0.0); + const std::complex zero(0.0, 0.0); + + // c^TV[nocc*naos] + container::Tensor cv(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { pcv.get_col_size(), pcv.get_row_size() }); + pzgemm_(&dagger, ¬rans, &nocc, &naos, &naos, + &one, c.get_pointer(), &i1, &i1, pc.desc, + V_istate[isk].data>(), &i1, &i1, pmat.desc, + &zero, cv.data>(), &i1, &i1, pcv.desc); + + // cX^T[naos*nvirt] + container::Tensor cx(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { pcx.get_col_size(), pcx.get_row_size() }); + pzgemm_(¬rans, &dagger, &naos, &nvirt, &nocc, + &one, c.get_pointer(), &i1, &i1, pc.desc, + X_istate + start, &i1, &i1, px.desc, + &zero, cx.data>(), &i1, &i1, pcx.desc); + + //AX_istate=[cX^T]^T[c^TV]^T (nvirt major) + pzgemm_(&trans, &trans, &nvirt, &nocc, &naos, + &factor, cx.data>(), &i1, &i1, pcx.desc, + cv.data>(), &i1, &i1, pcv.desc, + add_on ? &one : &zero, AX_istate + start, &i1, &i1, px.desc); + } + } + + template <> + void CVCX_virt_pblas( + const std::vector& V_istate, + const Parallel_2D& pmat, + const psi::Psi& c, + const Parallel_2D& pc, + const double* const X_istate, + const Parallel_2D& px, + const int& naos, + const int& nocc, + const int& nvirt, + double* const AX_istate, + const bool add_on, + const double factor) + { + ModuleBase::TITLE("hamilt_lrtd", "CVCX_virt_pblas"); + assert(pmat.comm() == pc.comm()); + assert(pmat.comm() == px.comm()); + assert(pmat.blacs_ctxt == pc.blacs_ctxt); + assert(pmat.blacs_ctxt == px.blacs_ctxt); + assert(px.get_local_size() > 0); + + const int nks = c.get_nk(); + assert(V_istate.size() == nks); + + Parallel_2D pcv; + LR_Util::setup_2d_division(pcv, pmat.get_block_size(), naos, nvirt, pmat.blacs_ctxt); + Parallel_2D pcx; + LR_Util::setup_2d_division(pcx, pmat.get_block_size(), nocc, naos, pmat.blacs_ctxt); + for (int isk = 0;isk < nks;++isk) + { + c.fix_k(isk); + const int start = isk * px.get_local_size(); + + const int i1 = 1; + const int ivirt = nocc + 1; + const char trans = 'T'; + const char dagger = 'C'; + const char notrans = 'N'; //c is col major + const double one = 1.0; + const double zero = 0.0; + + // VC[naos*nvirt] + container::Tensor cv(DAT::DT_DOUBLE, DEV::CpuDevice, { pcv.get_col_size(), pcv.get_row_size() }); + pdgemm_(¬rans, ¬rans, &naos, &nvirt, &naos, + &one, V_istate[isk].data(), &i1, &i1, pmat.desc, + c.get_pointer(), &i1, &ivirt, pc.desc, + &zero, cv.data(), &i1, &i1, pcv.desc); + + // X^TC^T[nocc*naos] + container::Tensor cx(DAT::DT_DOUBLE, DEV::CpuDevice, { pcx.get_col_size(), pcx.get_row_size() }); + pdgemm_(&dagger, &dagger, &nocc, &naos, &nvirt, + &one, X_istate + start, &i1, &i1, px.desc, + c.get_pointer(), &i1, &ivirt, pc.desc, + &zero, cx.data(), &i1, &i1, pcx.desc); + + //AX_istate=[VC]^T[X^TC^T]^T (nvirt major) + pdgemm_(&trans, &trans, &nvirt, &nocc, &naos, + &factor, cv.data(), &i1, &i1, pcv.desc, + cx.data(), &i1, &i1, pcx.desc, + add_on ? &one : &zero, AX_istate + start, &i1, &i1, px.desc); + } + } + + template <> + void CVCX_virt_pblas( + const std::vector& V_istate, + const Parallel_2D& pmat, + const psi::Psi, base_device::DEVICE_CPU>& c, + const Parallel_2D& pc, + const std::complex* const X_istate, + const Parallel_2D& px, + const int& naos, + const int& nocc, + const int& nvirt, + std::complex* const AX_istate, + const bool add_on, + const std::complex factor) + { + ModuleBase::TITLE("hamilt_lrtd", "CVCX_virt_pblas"); + assert(pmat.comm() == pc.comm()); + assert(pmat.comm() == px.comm()); + assert(pmat.blacs_ctxt == pc.blacs_ctxt); + assert(pmat.blacs_ctxt == px.blacs_ctxt); + assert(px.get_local_size() > 0); + + const int nks = c.get_nk(); + assert(V_istate.size() == nks); + + Parallel_2D pcv; + LR_Util::setup_2d_division(pcv, pmat.get_block_size(), naos, nvirt, pmat.blacs_ctxt); + Parallel_2D pcx; + LR_Util::setup_2d_division(pcx, pmat.get_block_size(), nocc, naos, pmat.blacs_ctxt); + for (int isk = 0;isk < nks;++isk) + { + c.fix_k(isk); + const int start = isk * px.get_local_size(); + + const int i1 = 1; + const int ivirt = nocc + 1; + const char trans = 'T'; + const char dagger = 'C'; + const char notrans = 'N'; //c is col major + const std::complex one(1.0, 0.0); + const std::complex zero(0.0, 0.0); + + // VC[naos*nvirt] + container::Tensor cv(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { pcv.get_col_size(), pcv.get_row_size() }); + pzgemm_(¬rans, ¬rans, &naos, &nvirt, &naos, + &one, V_istate[isk].data>(), &i1, &i1, pmat.desc, + c.get_pointer(), &i1, &ivirt, pc.desc, + &zero, cv.data>(), &i1, &i1, pcv.desc); + + // X^TC^T[nocc*naos] + container::Tensor cx(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { pcx.get_col_size(), pcx.get_row_size() }); + pzgemm_(&dagger, &dagger, &nocc, &naos, &nvirt, + &one, X_istate + start, &i1, &i1, px.desc, + c.get_pointer(), &i1, &ivirt, pc.desc, + &zero, cx.data>(), &i1, &i1, pcx.desc); + + //AX_istate=[VC]^T[X^TC^T]^T (nvirt major) + pzgemm_(&trans, &trans, &nvirt, &nocc, &naos, + &factor, cv.data>(), &i1, &i1, pcv.desc, + cx.data>(), &i1, &i1, pcx.desc, + add_on ? &one : &zero, AX_istate + start, &i1, &i1, px.desc); + } + } +} +#endif \ No newline at end of file diff --git a/source/source_lcao/module_lr/ao_to_mo_transformer/cvcx_serial.cpp b/source/source_lcao/module_lr/ao_to_mo_transformer/cvcx_serial.cpp new file mode 100644 index 00000000000..87aa95012ff --- /dev/null +++ b/source/source_lcao/module_lr/ao_to_mo_transformer/cvcx_serial.cpp @@ -0,0 +1,301 @@ +#include "cvcx.h" +#include "source_base/module_external/blas_connector.h" +#include "source_base/tool_title.h" +#include "source_lcao/module_lr/utils/lr_util.h" +namespace LR +{ + //=====================occ======================== + template <> + void CVCX_occ_forloop_serial( + const std::vector& V_istate, + const psi::Psi& c, + const double* const X_istate, + const int& naos, + const int& nocc, + const int& nvirt, + double* const AX_istate) + { + ModuleBase::TITLE("hamilt_lrtd", "CVCX_occ_forloop_serial"); + int nks = c.get_nk(); + assert(V_istate.size() == nks); + assert(naos == c.get_nbasis()); + ModuleBase::GlobalFunc::ZEROS(AX_istate, nks * nocc * nvirt); + for (int isk = 0;isk < nks;++isk) + { + c.fix_k(isk); + const int start = isk * nocc * nvirt; + for (int i = 0;i < nocc;++i) + for (int a = 0;a < nvirt;++a) + for (int nu = 0;nu < naos;++nu) + for (int mu = 0;mu < naos;++mu) + for (int j = 0;j < nocc;++j) + AX_istate[start + i * nvirt + a] += X_istate[start + j * nvirt + a] * c(i, mu) * V_istate[isk].data()[nu * naos + mu] * c(j, nu); + } + } + template <> + void CVCX_occ_forloop_serial( + const std::vector& V_istate, + const psi::Psi, base_device::DEVICE_CPU>& c, + const std::complex* const X_istate, + const int& naos, + const int& nocc, + const int& nvirt, + std::complex* const AX_istate) + { + ModuleBase::TITLE("hamilt_lrtd", "CVCX_occ_forloop_serial"); + int nks = c.get_nk(); + assert(V_istate.size() == nks); + assert(naos == c.get_nbasis()); + ModuleBase::GlobalFunc::ZEROS(AX_istate, nks * nocc * nvirt); + for (int isk = 0;isk < nks;++isk) + { + c.fix_k(isk); + const int start = isk * nocc * nvirt; + for (int i = 0;i < nocc;++i) + for (int a = 0;a < nvirt;++a) + for (int nu = 0;nu < naos;++nu) + for (int mu = 0;mu < naos;++mu) + for (int j = 0;j < nocc;++j) + AX_istate[start + i * nvirt + a] += std::conj(X_istate[start + j * nvirt + a] * c(i, mu)) * V_istate[isk].data>()[nu * naos + mu] * c(j, nu); + } + } + + template <> + void CVCX_occ_blas( + const std::vector& V_istate, + const psi::Psi& c, + const double* const X_istate, + const int& naos, + const int& nocc, + const int& nvirt, + double* const AX_istate, + const bool add_on, + const double factor) + { + ModuleBase::TITLE("hamilt_lrtd", "CVCX_occ_AX_blas"); + int nks = c.get_nk(); + assert(V_istate.size() == nks); + assert(naos == c.get_nbasis()); + + for (int isk = 0;isk < nks;++isk) + { + c.fix_k(isk); + const int start = isk * nocc * nvirt; + const char trans = 'T'; + const char notrans = 'N'; //c is col major + const double one = 1.0; + const double zero = 0.0; + + // c^TV[nocc*naos] + container::Tensor cv(DAT::DT_DOUBLE, DEV::CpuDevice, { naos, nocc }); + dgemm_(&trans, ¬rans, &nocc, &naos, &naos, &one, + c.get_pointer(), &naos, V_istate[isk].data(), &naos, &zero, + cv.data(), &nocc); + + // cX^T[naos*nvirt] + container::Tensor cx(DAT::DT_DOUBLE, DEV::CpuDevice, { nvirt, naos }); + dgemm_(¬rans, &trans, &naos, &nvirt, &nocc, &one, + c.get_pointer(), &naos, X_istate + start, &nvirt, &zero, + cx.data(), &naos); + + //AX_istate=[cX^T]^T[c^TV]^T (nvirt major) + dgemm_(&trans, &trans, &nvirt, &nocc, &naos, &factor, + cx.data(), &naos, cv.data(), &nocc, add_on ? &one : &zero, + AX_istate + start, &nvirt); + } + } + + template <> + void CVCX_occ_blas( + const std::vector& V_istate, + const psi::Psi, base_device::DEVICE_CPU>& c, + const std::complex* const X_istate, + const int& naos, + const int& nocc, + const int& nvirt, + std::complex* const AX_istate, + const bool add_on, + const std::complex factor) + { + ModuleBase::TITLE("hamilt_lrtd", "CVCX_occ_AX_blas"); + int nks = c.get_nk(); + assert(V_istate.size() == nks); + assert(naos == c.get_nbasis()); + + for (int isk = 0;isk < nks;++isk) + { + c.fix_k(isk); + const int start = isk * nocc * nvirt; + const char trans = 'T'; + const char notrans = 'N'; //c is col major + const char dagger = 'C'; + const std::complex one(1.0, 0.0); + const std::complex zero(0.0, 0.0); + + // c^TV[nocc*naos] + container::Tensor cv(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { naos, nocc }); + zgemm_(&dagger, ¬rans, &nocc, &naos, &naos, &one, + c.get_pointer(), &naos, V_istate[isk].data>(), &naos, &zero, + cv.data>(), &nocc); + + // cX^T[naos*nvirt] + container::Tensor cx(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { nvirt, naos }); + zgemm_(¬rans, &dagger, &naos, &nvirt, &nocc, &one, + c.get_pointer(), &naos, X_istate + start, &nvirt, &zero, + cx.data>(), &naos); + + //AX_istate=[cX^T]^T[c^TV]^T (nvirt major) + zgemm_(&trans, &trans, &nvirt, &nocc, &naos, &factor, + cx.data>(), &naos, cv.data>(), &nocc, add_on ? &one : &zero, + AX_istate + start, &nvirt); + } + } + + + //=====================virt======================== + template <> + void CVCX_virt_forloop_serial( + const std::vector& V_istate, + const psi::Psi& c, + const double* const X_istate, + const int& naos, + const int& nocc, + const int& nvirt, + double* const AX_istate) + { + ModuleBase::TITLE("hamilt_lrtd", "CVCX_virt_forloop_serial"); + int nks = c.get_nk(); + assert(V_istate.size() == nks); + assert(naos == c.get_nbasis()); + ModuleBase::GlobalFunc::ZEROS(AX_istate, nks * nocc * nvirt); + for (int isk = 0;isk < nks;++isk) + { + c.fix_k(isk); + const int start = isk * nocc * nvirt; + for (int i = 0;i < nocc;++i) + for (int a = 0;a < nvirt;++a) + for (int nu = 0;nu < naos;++nu) + for (int mu = 0;mu < naos;++mu) + for (int b = 0;b < nvirt;++b) + AX_istate[start + i * nvirt + a] += X_istate[start + i * nvirt + b] * c(nocc + b, mu) * V_istate[isk].data()[nu * naos + mu] * c(nocc + a, nu); + } + } + template <> + void CVCX_virt_forloop_serial( + const std::vector& V_istate, + const psi::Psi, base_device::DEVICE_CPU>& c, + const std::complex* const X_istate, + const int& naos, + const int& nocc, + const int& nvirt, + std::complex* const AX_istate) + { + ModuleBase::TITLE("hamilt_lrtd", "CVCX_virt_forloop_serial"); + int nks = c.get_nk(); + assert(V_istate.size() == nks); + assert(naos == c.get_nbasis()); + ModuleBase::GlobalFunc::ZEROS(AX_istate, nks * nocc * nvirt); + for (int isk = 0;isk < nks;++isk) + { + c.fix_k(isk); + const int start = isk * nocc * nvirt; + for (int i = 0;i < nocc;++i) + for (int a = 0;a < nvirt;++a) + for (int nu = 0;nu < naos;++nu) + for (int mu = 0;mu < naos;++mu) + for (int b = 0;b < nvirt;++b) + AX_istate[start + i * nvirt + a] += std::conj(X_istate[start + i * nvirt + b] * c(nocc + b, mu)) * V_istate[isk].data>()[nu * naos + mu] * c(nocc + a, nu); + } + } + + template <> + void CVCX_virt_blas( + const std::vector& V_istate, + const psi::Psi& c, + const double* const X_istate, + const int& naos, + const int& nocc, + const int& nvirt, + double* const AX_istate, + const bool add_on, + const double factor) + { + ModuleBase::TITLE("hamilt_lrtd", "CVCX_virt_AX_blas"); + const int nks = c.get_nk(); + assert(V_istate.size() == nks); + assert(naos == c.get_nbasis()); + + for (int isk = 0;isk < nks;++isk) + { + c.fix_k(isk); + const int start = isk * nocc * nvirt; + const char trans = 'T'; + const char notrans = 'N'; //c is col major + const double one = 1.0; + const double zero = 0.0; + + // VC[naos*nvirt] + container::Tensor cv(DAT::DT_DOUBLE, DEV::CpuDevice, { nvirt, naos }); + dgemm_(¬rans, ¬rans, &naos, &nvirt, &naos, &one, + V_istate[isk].data(), &naos, c.get_pointer(nocc), &naos, &zero, + cv.data(), &naos); + + // X^TC^T[nocc*naos] + container::Tensor cx(DAT::DT_DOUBLE, DEV::CpuDevice, { naos, nocc }); + dgemm_(&trans, &trans, &nocc, &naos, &nvirt, &one, + X_istate + start, &nvirt, c.get_pointer(nocc), &naos, &zero, + cx.data(), &nocc); + + //AX_istate=[VC]^T[X^TC^T]^T (nvirt major) + dgemm_(&trans, &trans, &nvirt, &nocc, &naos, &factor, + cv.data(), &naos, cx.data(), &nocc, add_on ? &one : &zero, + AX_istate + start, &nvirt); + } + } + + template <> + void CVCX_virt_blas( + const std::vector& V_istate, + const psi::Psi, base_device::DEVICE_CPU>& c, + const std::complex* const X_istate, + const int& naos, + const int& nocc, + const int& nvirt, + std::complex* const AX_istate, + const bool add_on, + const std::complex factor) + { + ModuleBase::TITLE("hamilt_lrtd", "CVCX_virt_AX_blas"); + int nks = c.get_nk(); + assert(V_istate.size() == nks); + assert(naos == c.get_nbasis()); + + for (int isk = 0;isk < nks;++isk) + { + c.fix_k(isk); + const int start = isk * nocc * nvirt; + const char trans = 'T'; + const char notrans = 'N'; //c is col major + const char dagger = 'C'; + const std::complex one(1.0, 0.0); + const std::complex zero(0.0, 0.0); + + // VC[naos*nvirt] + container::Tensor cv(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { nvirt, naos }); + zgemm_(¬rans, ¬rans, &naos, &nvirt, &naos, &one, + V_istate[isk].data>(), &naos, c.get_pointer(nocc), &naos, &zero, + cv.data>(), &naos); + + // X^TC^T[nocc*naos] + container::Tensor cx(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { naos, nocc }); + zgemm_(&dagger, &dagger, &nocc, &naos, &nvirt, &one, + X_istate+start, &nvirt, c.get_pointer(nocc), &naos, &zero, + cx.data>(), &nocc); + + //AX_istate=[VC]^T[X^TC^T]^T (nvirt major) + zgemm_(&trans, &trans, &nvirt, &nocc, &naos, &factor, + cv.data>(), &naos, cx.data>(), &nocc, add_on ? &one : &zero, + AX_istate+start, &nvirt); + } + } +} \ No newline at end of file diff --git a/source/source_lcao/module_lr/ao_to_mo_transformer/test/CMakeLists.txt b/source/source_lcao/module_lr/ao_to_mo_transformer/test/CMakeLists.txt index de891024277..c8baf3f3e76 100644 --- a/source/source_lcao/module_lr/ao_to_mo_transformer/test/CMakeLists.txt +++ b/source/source_lcao/module_lr/ao_to_mo_transformer/test/CMakeLists.txt @@ -3,4 +3,10 @@ AddTest( TARGET MODULE_LR_ao_to_mo_test LIBS parameter base container device psi SOURCES ao_to_mo_test.cpp ../../utils/lr_util.cpp ../ao_to_mo_parallel.cpp ../ao_to_mo_serial.cpp +) + +AddTest( + TARGET MODULE_LR_CVCX_test + LIBS base parameter ${math_libs} container device psi + SOURCES test_cvcx.cpp ../../utils/lr_util.cpp ../cvcx_par.cpp ../cvcx_serial.cpp ) \ No newline at end of file diff --git a/source/source_lcao/module_lr/ao_to_mo_transformer/test/test_cvcx.cpp b/source/source_lcao/module_lr/ao_to_mo_transformer/test/test_cvcx.cpp new file mode 100644 index 00000000000..dfa10633898 --- /dev/null +++ b/source/source_lcao/module_lr/ao_to_mo_transformer/test/test_cvcx.cpp @@ -0,0 +1,355 @@ +#include +#include "mpi.h" +#include "../cvcx.h" + +#include "source_lcao/module_lr/utils/lr_util.h" + +struct matsize +{ + int nks = 1; + int naos; + int nocc; + int nvirt; + int nb = 1; + matsize(int nks, int naos, int nocc, int nvirt, int nb = 1) + :nks(nks), naos(naos), nocc(nocc), nvirt(nvirt), nb(nb) { + assert(nocc + nvirt <= naos); + }; +}; + +class AXTest : public testing::Test +{ +public: + std::vector sizes{ + // {2, 3, 2, 1} + {2, 13, 7, 4}, + {2, 14, 8, 5} + }; + int nstate = 2; + std::ofstream ofs_running; + int my_rank; +#ifdef __MPI + void SetUp() override + { + MPI_Comm_rank(MPI_COMM_WORLD, &my_rank); + this->ofs_running.open("log" + std::to_string(my_rank) + ".txt"); + ofs_running << "my_rank = " << my_rank << std::endl; + } + void TearDown() override + { + ofs_running.close(); + } +#endif + + void set_ones(double* data, int size) { for (int i = 0;i < size;++i) data[i] = 1.0; }; + void set_int(double* data, int size) { for (int i = 0;i < size;++i) data[i] = static_cast(i + 1); }; + void set_int(std::complex* data, int size) { for (int i = 0;i < size;++i) data[i] = std::complex(i + 1, -i - 1); }; + void set_rand(double* data, int size) { for (int i = 0;i < size;++i) data[i] = double(rand()) / double(RAND_MAX) * 10.0 - 5.0; }; + void set_rand(std::complex* data, int size) { for (int i = 0;i < size;++i) data[i] = std::complex(rand(), rand()) / double(RAND_MAX) * 10.0 - 5.0; }; + void check_eq(double* data1, double* data2, int size) { for (int i = 0;i < size;++i) EXPECT_NEAR(data1[i], data2[i], 1e-8); }; + void check_eq(std::complex* data1, std::complex* data2, int size) + { + for (int i = 0;i < size;++i) + { + EXPECT_NEAR(data1[i].real(), data2[i].real(), 1e-8); + EXPECT_NEAR(data1[i].imag(), data2[i].imag(), 1e-8); + } + }; +}; + +TEST_F(AXTest, DoubleSerial) +{ + for (auto s : this->sizes) + { + psi::Psi X(s.nks, nstate, s.nocc * s.nvirt, {}, false); + psi::Psi AX_for(s.nks, nstate, s.nocc * s.nvirt, {}, false); + psi::Psi AX_blas(s.nks, nstate, s.nocc * s.nvirt, {}, false); + const int size_x = nstate * s.nks * s.nocc * s.nvirt; + set_rand(X.get_pointer(), size_x); + + const int size_c = s.nks * (s.nocc + s.nvirt) * s.naos; + const int size_v = s.naos * s.naos; + for (int istate = 0;istate < nstate;++istate) + { + psi::Psi c(s.nks, s.nocc + s.nvirt, s.naos, {}, true); + std::vector V(s.nks, container::Tensor(DAT::DT_DOUBLE, DEV::CpuDevice, { s.naos, s.naos })); + set_rand(c.get_pointer(), size_c); + for (auto& v : V)set_rand(v.data(), size_v); + X.fix_b(istate); + AX_for.fix_b(istate); + AX_blas.fix_b(istate); + // occ + LR::CVCX_occ_forloop_serial(V, c, X.get_pointer(), s.naos, s.nocc, s.nvirt, AX_for.get_pointer()); + LR::CVCX_occ_blas(V, c, X.get_pointer(), s.naos, s.nocc, s.nvirt, AX_blas.get_pointer(), false); + AX_for.fix_k(0); + AX_blas.fix_k(0); + check_eq(AX_for.get_pointer(), AX_blas.get_pointer(), s.nks * s.nocc * s.nvirt); + // virt + LR::CVCX_virt_forloop_serial(V, c, X.get_pointer(), s.naos, s.nocc, s.nvirt, AX_for.get_pointer()); + LR::CVCX_virt_blas(V, c, X.get_pointer(), s.naos, s.nocc, s.nvirt, AX_blas.get_pointer(), false); + AX_for.fix_k(0); + AX_blas.fix_k(0); + check_eq(AX_for.get_pointer(), AX_blas.get_pointer(), s.nks * s.nocc * s.nvirt); + } + } +} + +TEST_F(AXTest, ComplexSerial) +{ + for (auto s : this->sizes) + { + psi::Psi, base_device::DEVICE_CPU> X(s.nks, nstate, s.nocc * s.nvirt, {}, false); + psi::Psi, base_device::DEVICE_CPU> AX_for(s.nks, nstate, s.nocc * s.nvirt, {}, false); + psi::Psi, base_device::DEVICE_CPU> AX_blas(s.nks, nstate, s.nocc * s.nvirt, {}, false); + const int size_x = nstate * s.nks * s.nocc * s.nvirt; + set_rand(X.get_pointer(), size_x); + + int size_c = s.nks * (s.nocc + s.nvirt) * s.naos; + int size_v = s.naos * s.naos; + for (int istate = 0;istate < nstate;++istate) + { + psi::Psi, base_device::DEVICE_CPU> c(s.nks, s.nocc + s.nvirt, s.naos, {}, true); + std::vector V(s.nks, container::Tensor(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { s.naos, s.naos })); + set_rand(c.get_pointer(), size_c); + for (auto& v : V)set_rand(v.data>(), size_v); + X.fix_b(istate); + AX_for.fix_b(istate); + AX_blas.fix_b(istate); + // occ + LR::CVCX_occ_forloop_serial(V, c, X.get_pointer(), s.naos, s.nocc, s.nvirt, AX_for.get_pointer()); + LR::CVCX_occ_blas(V, c, X.get_pointer(), s.naos, s.nocc, s.nvirt, AX_blas.get_pointer(), false); + AX_for.fix_k(0); + AX_blas.fix_k(0); + check_eq(AX_for.get_pointer(), AX_blas.get_pointer(), s.nks * s.nocc * s.nvirt); + // virt + LR::CVCX_virt_forloop_serial(V, c, X.get_pointer(), s.naos, s.nocc, s.nvirt, AX_for.get_pointer()); + LR::CVCX_virt_blas(V, c, X.get_pointer(), s.naos, s.nocc, s.nvirt, AX_blas.get_pointer(), false); + AX_for.fix_k(0); + AX_blas.fix_k(0); + check_eq(AX_for.get_pointer(), AX_blas.get_pointer(), s.nks * s.nocc * s.nvirt); + } + } +} +#ifdef __MPI +TEST_F(AXTest, DoubleParallel) +{ + for (auto s : this->sizes) + { + // c: nao*nbands in para2d, nbands*nao in psi (row-para and constructed: nao) + // X: nvirt*nocc in para2d, nocc*nvirt in psi (row-para and constructed: nvirt) + Parallel_2D pV; + LR_Util::setup_2d_division(pV, s.nb, s.naos, s.naos); + std::vector V(s.nks, container::Tensor(DAT::DT_DOUBLE, DEV::CpuDevice, { pV.get_col_size(), pV.get_row_size() })); + Parallel_2D pc; + const int nmo = s.nocc + s.nvirt; + LR_Util::setup_2d_division(pc, s.nb, s.naos, nmo, pV.blacs_ctxt); + psi::Psi c(s.nks, pc.get_col_size(), pc.get_row_size(), {}, true); + Parallel_2D px; + LR_Util::setup_2d_division(px, s.nb, s.nvirt, s.nocc, pV.blacs_ctxt); + + EXPECT_EQ(pV.dim0, pc.dim0); + EXPECT_EQ(pV.dim1, pc.dim1); + EXPECT_GE(s.nvirt, px.dim0); + EXPECT_GE(s.nocc, px.dim1); + EXPECT_GE(s.naos, pc.dim0); + + psi::Psi AX_pblas_loc(s.nks, nstate, px.get_local_size(), {}, false); + psi::Psi AX_gather(s.nks, nstate, s.nocc * s.nvirt, {}, false); + + //set X and X_full + psi::Psi X(s.nks, nstate, px.get_local_size(), {}, false); + set_rand(X.get_pointer(), nstate * s.nks * px.get_local_size()); + psi::Psi X_full(s.nks, nstate, s.nocc * s.nvirt, {}, false); // allocate X_full + X_full.zero_out(); + for (int istate = 0;istate < nstate;++istate) + { + X.fix_b(istate); + X_full.fix_b(istate); + for (int isk = 0;isk < s.nks;++isk) + { + X.fix_k(isk); + X_full.fix_k(isk); + LR_Util::gather_2d_to_full(px, X.get_pointer(), X_full.get_pointer(), false, s.nvirt, s.nocc); + } + } + + for (int istate = 0;istate < nstate;++istate) + { + for (int isk = 0;isk < s.nks;++isk) + { + set_rand(V.at(isk).data(), pV.get_local_size()); + c.fix_k(isk); + set_rand(c.get_pointer(), pc.get_local_size()); + } + X.fix_b(istate); + X_full.fix_b(istate); + AX_pblas_loc.fix_b(istate); + AX_gather.fix_b(istate); + LR::CVCX_occ_pblas(V, pV, c, pc, X.get_pointer(), px, s.naos, s.nocc, s.nvirt, AX_pblas_loc.get_pointer(), false); + AX_gather.zero_out(); + // gather AX and output + for (int isk = 0;isk < s.nks;++isk) + { + AX_pblas_loc.fix_k(isk); + AX_gather.fix_k(isk); + LR_Util::gather_2d_to_full(px, AX_pblas_loc.get_pointer(), AX_gather.get_pointer(), false, s.nvirt, s.nocc); + } + // compare to global AX + std::vector V_full(s.nks, container::Tensor(DAT::DT_DOUBLE, DEV::CpuDevice, { s.naos, s.naos })); + psi::Psi c_full(s.nks, s.nocc + s.nvirt, s.naos, {}, true); + c_full.zero_out(); + for (int isk = 0;isk < s.nks;++isk) + { + V_full.at(isk).zero(); + LR_Util::gather_2d_to_full(pV, V.at(isk).data(), V_full.at(isk).data(), false, s.naos, s.naos); + c.fix_k(isk); + c_full.fix_k(isk); + LR_Util::gather_2d_to_full(pc, c.get_pointer(), c_full.get_pointer(), false, s.naos, nmo); + } + if (my_rank == 0) + { + psi::Psi AX_full_istate(s.nks, 1, s.nocc * s.nvirt, {}, true); + LR::CVCX_occ_blas(V_full, c_full, X_full.get_pointer(), s.naos, s.nocc, s.nvirt, AX_full_istate.get_pointer(), false); + AX_full_istate.fix_b(0); + AX_gather.fix_b(istate); + check_eq(AX_full_istate.get_pointer(), AX_gather.get_pointer(), s.nks * s.nocc * s.nvirt); + } + + // //============ the same for virtual ========== + X.fix_b(istate); + X_full.fix_b(istate); + AX_pblas_loc.fix_b(istate); + AX_gather.fix_b(istate); + LR::CVCX_virt_pblas(V, pV, c, pc, X.get_pointer(), px, s.naos, s.nocc, s.nvirt, AX_pblas_loc.get_pointer(), false); + AX_gather.zero_out(); + for (int isk = 0;isk < s.nks;++isk) + { + AX_pblas_loc.fix_k(isk); + AX_gather.fix_k(isk); + LR_Util::gather_2d_to_full(px, AX_pblas_loc.get_pointer(), AX_gather.get_pointer(), false, s.nvirt, s.nocc); + } + if (my_rank == 0) + { + psi::Psi AX_full_istate(s.nks, 1, s.nocc * s.nvirt, {}, true); + LR::CVCX_virt_blas(V_full, c_full, X_full.get_pointer(), s.naos, s.nocc, s.nvirt, AX_full_istate.get_pointer(), false); + AX_full_istate.fix_b(0); + AX_gather.fix_b(istate); + check_eq(AX_full_istate.get_pointer(), AX_gather.get_pointer(), s.nks * s.nocc * s.nvirt); + } + } + } +} +TEST_F(AXTest, ComplexParallel) +{ + for (auto s : this->sizes) + { + // c: nao*nbands in para2d, nbands*nao in psi (row-para and constructed: nao) + // X: nvirt*nocc in para2d, nocc*nvirt in psi (row-para and constructed: nvirt) + Parallel_2D pV; + LR_Util::setup_2d_division(pV, s.nb, s.naos, s.naos); + std::vector V(s.nks, container::Tensor(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { pV.get_col_size(), pV.get_row_size() })); + Parallel_2D pc; + const int nmo = s.nocc + s.nvirt; + LR_Util::setup_2d_division(pc, s.nb, s.naos, nmo, pV.blacs_ctxt); + psi::Psi, base_device::DEVICE_CPU> c(s.nks, pc.get_col_size(), pc.get_row_size(), {}, true); + Parallel_2D px; + LR_Util::setup_2d_division(px, s.nb, s.nvirt, s.nocc, pV.blacs_ctxt); + + psi::Psi, base_device::DEVICE_CPU> AX_pblas_loc(s.nks, nstate, px.get_local_size(), {}, false); + psi::Psi, base_device::DEVICE_CPU> AX_gather(s.nks, nstate, s.nocc * s.nvirt, {}, false); + + //set X and X_full + psi::Psi, base_device::DEVICE_CPU> X(s.nks, nstate, px.get_local_size(), {}, false); + set_rand(X.get_pointer(), nstate * s.nks * px.get_local_size()); + psi::Psi, base_device::DEVICE_CPU> X_full(s.nks, nstate, s.nocc * s.nvirt, {}, false); // allocate X_full + X_full.zero_out(); + for (int istate = 0;istate < nstate;++istate) + { + X.fix_b(istate); + X_full.fix_b(istate); + for (int isk = 0;isk < s.nks;++isk) + { + X.fix_k(isk); + X_full.fix_k(isk); + LR_Util::gather_2d_to_full(px, X.get_pointer(), X_full.get_pointer(), false, s.nvirt, s.nocc); + } + } + + for (int istate = 0;istate < nstate;++istate) + { + for (int isk = 0;isk < s.nks;++isk) + { + set_rand(V.at(isk).data>(), pV.get_local_size()); + c.fix_k(isk); + set_rand(c.get_pointer(), pc.get_local_size()); + } + X.fix_b(istate); + X_full.fix_b(istate); + AX_pblas_loc.fix_b(istate); + AX_gather.fix_b(istate); + LR::CVCX_occ_pblas(V, pV, c, pc, X.get_pointer(), px, s.naos, s.nocc, s.nvirt, AX_pblas_loc.get_pointer(), false); + AX_gather.zero_out(); + + // gather AX and output + for (int isk = 0;isk < s.nks;++isk) + { + AX_pblas_loc.fix_k(isk); + AX_gather.fix_k(isk); + LR_Util::gather_2d_to_full(px, AX_pblas_loc.get_pointer(), AX_gather.get_pointer(), false, s.nvirt, s.nocc); + } + // compare to global AX + std::vector V_full(s.nks, container::Tensor(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { s.naos, s.naos })); + psi::Psi, base_device::DEVICE_CPU> c_full(s.nks, s.nocc + s.nvirt, s.naos, {}, true); + c_full.zero_out(); + for (int isk = 0;isk < s.nks;++isk) + { + V_full.at(isk).zero(); + LR_Util::gather_2d_to_full(pV, V.at(isk).data>(), V_full.at(isk).data>(), false, s.naos, s.naos); + c.fix_k(isk); + c_full.fix_k(isk); + LR_Util::gather_2d_to_full(pc, c.get_pointer(), c_full.get_pointer(), false, s.naos, nmo); + } + if (my_rank == 0) + { + psi::Psi, base_device::DEVICE_CPU> AX_full_istate(s.nks, 1, s.nocc * s.nvirt, {}, false); + LR::CVCX_occ_blas(V_full, c_full, X_full.get_pointer(), s.naos, s.nocc, s.nvirt, AX_full_istate.get_pointer(), false); + AX_full_istate.fix_b(0); + AX_gather.fix_b(istate); + check_eq(AX_full_istate.get_pointer(), AX_gather.get_pointer(), s.nks * s.nocc * s.nvirt); + } + // //============ the same for virtual ========== + X.fix_b(istate); + X_full.fix_b(istate); + AX_pblas_loc.fix_b(istate); + AX_gather.fix_b(istate); + LR::CVCX_virt_pblas(V, pV, c, pc, X.get_pointer(), px, s.naos, s.nocc, s.nvirt, AX_pblas_loc.get_pointer(), false); + AX_gather.zero_out(); + for (int isk = 0;isk < s.nks;++isk) + { + AX_pblas_loc.fix_k(isk); + AX_gather.fix_k(isk); + LR_Util::gather_2d_to_full(px, AX_pblas_loc.get_pointer(), AX_gather.get_pointer(), false, s.nvirt, s.nocc); + } + if (my_rank == 0) + { + psi::Psi, base_device::DEVICE_CPU> AX_full_istate(s.nks, 1, s.nocc * s.nvirt, {}, false); + LR::CVCX_virt_blas(V_full, c_full, X_full.get_pointer(), s.naos, s.nocc, s.nvirt, AX_full_istate.get_pointer(), false); + AX_full_istate.fix_b(0); + AX_gather.fix_b(istate); + check_eq(AX_full_istate.get_pointer(), AX_gather.get_pointer(), s.nks * s.nocc * s.nvirt); + } + } + } +} +#endif + + +int main(int argc, char** argv) +{ + srand(time(NULL)); // for random number generator + MPI_Init(&argc, &argv); + testing::InitGoogleTest(&argc, argv); + int result = RUN_ALL_TESTS(); + MPI_Finalize(); + return result; +} \ No newline at end of file diff --git a/source/source_lcao/module_lr/cal_edm.cpp b/source/source_lcao/module_lr/cal_edm.cpp new file mode 100644 index 00000000000..07b23acb6c4 --- /dev/null +++ b/source/source_lcao/module_lr/cal_edm.cpp @@ -0,0 +1,102 @@ +#include "cal_edm.h" +#include "source_base/module_external/scalapack_connector.h" +#include "source_base/module_external/blas_connector.h" +namespace LR +{ + // $X_{\mu i}=\sum_a c_{\mu a} X_{ai}$ + template<> + void cal_X_ao_occ(const double* const X, const Parallel_2D& px, + const double* const c, const Parallel_2D& pc, + double* const X_ao_occ, const Parallel_2D& px_ao_occ) + { + const int nocc = px.get_global_col_size(); + const int nvirt = px.get_global_row_size(); + const int naos = pc.get_global_row_size(); + const double alpha = 1.0; + const double beta = 0.0; + const char transa = 'N', transb = 'N'; +#ifdef __MPI + const int i1 = 1; + const int ivirt = nocc + 1; + pdgemm_(&transa, &transb, &naos, &nocc, &nvirt, + &alpha, c, &i1, &ivirt, pc.desc, + X, &i1, &i1, px.desc, + &beta, X_ao_occ, &i1, &i1, px_ao_occ.desc); +#else + dgemm_(&transa, &transb, &naos, &nocc, &nvirt, + &alpha, c + nocc * naos, &naos, + X, &nvirt, + &beta, X_ao_occ, &naos); +#endif + } + template<> + void cal_X_ao_occ(const std::complex* const X, const Parallel_2D& px, + const std::complex* const c, const Parallel_2D& pc, + std::complex* const X_ao_occ, const Parallel_2D& px_ao_occ) + { + const int nocc = px.get_global_col_size(); + const int nvirt = px.get_global_row_size(); + const int naos = pc.get_global_row_size(); + const std::complex alpha(1.0, 0.0), beta(0.0, 0.0); + const char transa = 'N', transb = 'N'; +#ifdef __MPI + const int i1 = 1; + const int ivirt = nocc + 1; + pzgemm_(&transa, &transb, &naos, &nocc, &nvirt, + &alpha, c, &i1, &ivirt, pc.desc, + X, &i1, &i1, px.desc, + &beta, X_ao_occ, &i1, &i1, px_ao_occ.desc); +#else + zgemm_(&transa, &transb, &naos, &nocc, &nvirt, + &alpha, c + nocc * naos, &naos, + X, &nvirt, + &beta, X_ao_occ, &naos); +#endif + } + // D=X1*X2^T + // $D_{\mu\nu} = \sum_i X1_{\mu i}X2_{\nu i}$ + template<> + void matdot(const double* const vec1, const double* const vec2, const Parallel_2D& pvec, + double* const dm, const Parallel_2D& pmat) + { + const int nocc = pvec.get_global_col_size(); + const int naos = pvec.get_global_row_size(); + const double alpha = 1.0; + const double beta = 0.0; + const char transa = 'N', transb = 'T'; +#ifdef __MPI + const int i1 = 1; + pdgemm_(&transa, &transb, &naos, &naos, &nocc, + &alpha, vec1, &i1, &i1, pvec.desc, + vec2, &i1, &i1, pvec.desc, + &beta, dm, &i1, &i1, pmat.desc); +#else + dgemm_(&transa, &transb, &naos, &naos, &nocc, + &alpha, vec1, &naos, + vec2, &naos, + &beta, dm, &naos); +#endif + } + + template<> + void matdot(const std::complex* const vec1, const std::complex* const vec2, const Parallel_2D& pvec, + std::complex* const dm, const Parallel_2D& pmat) + { + const int nocc = pvec.get_global_col_size(); + const int naos = pvec.get_global_row_size(); + const std::complex alpha(1.0, 0.0), beta(0.0, 0.0); + const char transa = 'N', transb = 'C'; +#ifdef __MPI + const int i1 = 1; + pzgemm_(&transa, &transb, &naos, &naos, &nocc, + &alpha, vec1, &i1, &i1, pvec.desc, + vec2, &i1, &i1, pvec.desc, + &beta, dm, &i1, &i1, pmat.desc); +#else + zgemm_(&transa, &transb, &naos, &naos, &nocc, + &alpha, vec1, &naos, + vec2, &naos, + &beta, dm, &naos); +#endif + } +} // namespace LR diff --git a/source/source_lcao/module_lr/cal_edm.h b/source/source_lcao/module_lr/cal_edm.h new file mode 100644 index 00000000000..e601241c76e --- /dev/null +++ b/source/source_lcao/module_lr/cal_edm.h @@ -0,0 +1,371 @@ +#ifndef ABACUS_LR_CAL_EDM_H +#define ABACUS_LR_CAL_EDM_H +#include "source_basis/module_ao/parallel_orbitals.h" +#include "source_lcao/module_lr/dm_trans/dm_trans.h" +#include "source_lcao/module_lr/utils/lr_util.h" +#include "source_lcao/module_lr/utils/lr_util_print.h" +#include "cal_w_from_z.h" +#include "gradient_inputs.h" +#include +#ifdef __EXX +#include "source_lcao/module_lr/operator_casida/operator_lr_exx.h" +#endif +namespace LR +{ + + // $X_{\mu i}=\sum_a c_{\mu a} X_{ai}$ + template + void cal_X_ao_occ(const T* const X, const Parallel_2D& px, + const T* const c, const Parallel_2D& pc, + T* const X_ao_occ, const Parallel_2D& px_ao_occ); + + // D=X1*X2^T + // $D_{\mu\nu} = \sum_i X1_{\mu i}X2_{\nu i}$ + template + void matdot(const T* const vec1, const T* const vec2, const Parallel_2D& pvec, + T* const dm, const Parallel_2D& pmat); + + template + void multiply_eig_onto_vec(const T* const vec, const double* const eig, const Parallel_2D& pvec, T* const evec) + { + for (int i = 0;i < pvec.get_col_size();++i) + { + const int gi = pvec.local2global_col(i); + for (int j = 0;j < pvec.get_row_size();++j) + { + const int idx = i * pvec.get_row_size() + j; + evec[idx] = vec[idx] * eig[gi]; + } + } + } + + template + ct::Tensor cal_edm_single_kpoint(const T* const vec, const Parallel_2D& pvec, const double* eig, const Parallel_2D& pmat) + { + const int nocc = pvec.get_global_col_size(); + const int naos = pvec.get_global_row_size(); + std::vector eig_times_vec(pvec.get_local_size()); + multiply_eig_onto_vec(vec, eig, pvec, eig_times_vec.data()); + ct::Tensor edm_result = LR_Util::newTensor({ pmat.get_col_size(), pmat.get_row_size() }); + matdot(eig_times_vec.data(), vec, pvec, edm_result.data(), pmat); + return edm_result; + } + + template + std::vector cal_edm_term4(const T* X, + const double eig_ext_istate, //1, the excitation energy of one state + const double* const eig_ks, // gocc+gvirt + const psi::Psi& c, + const Parallel_2D& px, + const Parallel_2D& pc, + const Parallel_Orbitals& pmat) + { + const int& naos = pmat.get_global_row_size(); + const int& nocc = px.get_global_col_size(); + const int& nvirt = px.get_global_row_size(); + // 4. $\sum_i (\Omega + \epsilon_i) \sum_{ab} C_{\mu a} X_{ia} C_{\nu b} X_{ib}$ + std::vector edm(c.get_nk()); + Parallel_2D px_ao_occ; + LR_Util::setup_2d_division(px_ao_occ, px.get_block_size(), naos, nocc +#ifdef __MPI + , px.blacs_ctxt +#endif + ); + for (int ik = 0;ik < c.get_nk();++ik) + { + const int idx_X = ik * px.get_local_size(); + std::vector X_ao_occ(px_ao_occ.get_local_size()); + cal_X_ao_occ(X + idx_X, px, &c(ik, 0, 0), pc, X_ao_occ.data(), px_ao_occ); + std::vector eig_ks_plus_ext(nocc, 0.0); + const int idx_eig_ks = ik * (nocc + nvirt); + std::transform(eig_ks + idx_eig_ks, eig_ks + idx_eig_ks + nocc, eig_ks_plus_ext.begin(), [eig_ext_istate](double x) {return x + eig_ext_istate;}); + edm[ik] = cal_edm_single_kpoint(X_ao_occ.data(), px_ao_occ, eig_ks_plus_ext.data(), pmat); + } + return edm; + } + + // calculate the excited state energy density matrix (for multiplying the overlap gradient in the gradient of lagrangian) + // multi-k has not been supported yet + template + std::vector cal_edm_terms_from_XZWK( + const T* const X, //lvirt*locc + const T* const Z, //lvirt*locc + const T* const W, //locc*locc + const T* const K_cvcx, //lvirt*locc + const double eig_ext_istate, //1, the excitation energy of one state + const double* const eig_ks, // gocc+gvirt + const psi::Psi& c, + const int nspin, + const bool test_force, + const Parallel_2D& p_occ_occ, + const Parallel_2D& p_virt_occ, + const Parallel_2D& pc, + const Parallel_Orbitals& pmat) + { + const int& naos = pmat.get_global_col_size(); + const int& nocc = p_occ_occ.get_global_col_size(); + const int& nvirt = p_virt_occ.get_global_row_size(); + const Parallel_2D& px = p_virt_occ; + + // 1. c * W * c +#ifdef __MPI + const std::vector cWc = cal_dm_trans_pblas(W, p_occ_occ, c, pc, naos, nocc, nvirt, pmat, (T)1., LR_Util::MO_TYPE::OO); +#else + const std::vector cWc = cal_dm_trans_blas(W, c, nocc, nvirt, (T)1., LR_Util::MO_TYPE::OO); +#endif + + // 2. edm of Z : $\sum_i \sum_a c_{\mu a} \epsilon_i Z_{ai} c_{\nu i}$ + std::vector epsi_Z(px.get_local_size() * c.get_nk()); + for (int ik = 0;ik < c.get_nk();++ik) + { + multiply_eig_onto_vec(Z + ik * px.get_local_size(), eig_ks + ik * (nocc + nvirt), + px, epsi_Z.data() + ik * px.get_local_size()); + } +#ifdef __MPI + std::vector cZc = cal_dm_trans_pblas(epsi_Z.data(), px, c, pc, naos, nocc, nvirt, pmat); + std::for_each(cZc.begin(), cZc.end(), [&](ct::Tensor& s) { LR_Util::matsym(s.data(), naos, pmat); }); +#else + std::vector cZc = cal_dm_trans_blas(epsi_Z.data(), c, nocc, nvirt); + std::for_each(cZc.begin(), cZc.end(), [&](ct::Tensor& s) { LR_Util::matsym(s.data(), naos); }); +#endif + + //3. c * K_cvcx * c + // $\sum_{kl}K_{kl}[D^X](c_{\kappa k}X_{\lambda l}+X_{\kappa k}c_{\lambda l})$. + // `K_cvcx` already carries the factor 2 of $W^X_{ij}=2K_{ij}[D^X]$ (see `op_K_cvcx` above), + // `matsym` then supplies the 1/2 that turns $2\,X_\kappa K c_\lambda$ into the symmetric pair above. +#ifdef __MPI + std::vector cKc = cal_dm_trans_pblas(K_cvcx, px, c, pc, naos, nocc, nvirt, pmat, (T)1.0); + std::for_each(cKc.begin(), cKc.end(), [&](ct::Tensor& s) { LR_Util::matsym(s.data(), naos, pmat); }); +#else + std::vector cKc = cal_dm_trans_blas(K_cvcx, c, nocc, nvirt, (T)1.0); + std::for_each(cKc.begin(), cKc.end(), [&](ct::Tensor& s) { LR_Util::matsym(s.data(), naos); }); +#endif + + // 4. $\sum_i (\Omega + \epsilon_i) \sum_{ab} C_{\mu a} X_{ia} C_{\nu b} X_{ib}$ + const std::vector edm = cal_edm_term4(X, eig_ext_istate, eig_ks, c, px, pc, pmat); + + if (test_force) + { + std::cout << "cWc: " << std::endl; + LR_Util::print_value(cWc[0].data(), pmat.get_col_size(), pmat.get_row_size()); + std::cout << "cZc: " << std::endl; + LR_Util::print_value(cZc[0].data(), pmat.get_col_size(), pmat.get_row_size()); + std::cout << "cKc: " << std::endl; + LR_Util::print_value(cKc[0].data(), pmat.get_col_size(), pmat.get_row_size()); + std::cout << "edm term 4: " << std::endl; + LR_Util::print_value(edm[0].data(), pmat.get_col_size(), pmat.get_row_size()); + } + return edm + cWc + cZc + cKc; + } + + template + std::vector cal_edm_from_XZ_istate( // for one excited state + const GradientInputs& inputs, + const T* const X, + const T* const Z, + const double eig_ext_istate, + const double* const eig_ks, + module_dm::DensityMatrix& dm_trans, + const psi::Psi& c, // preserve the caller's single-spin view + std::weak_ptr pot, + const std::string& spin_type = "singlet") + { + const int& nspin = inputs.nspin; + const int& naos = inputs.nbasis; + const std::vector& nocc = inputs.nocc; + const std::vector& nvirt = inputs.nvirt; + const UnitCell& ucell = inputs.ucell; + const std::vector& orb_cutoff = inputs.orb_cutoff; + const Grid_Driver& gd = inputs.gd; + const K_Vectors& kv = inputs.kv; + const std::vector& px = inputs.px; + const Parallel_2D& pc = inputs.pc; + const Parallel_Orbitals& pmat = inputs.pmat; +#ifdef __EXX + std::weak_ptr> exx_lri = inputs.exx_lri; + const double& exx_alpha = inputs.hybrid_alpha; +#endif + const bool test_force = inputs.test_force; + const std::string xc_kernel = inputs.xc_kernel; + const int nk = kv.get_nks() / nspin; + // 1. calculate W multiplier + std::vector p_occ_occ(nspin); + for (int is = 0;is < nspin;++is) + { + const int block_size = px[is].get_block_size(); + LR_Util::setup_2d_division(p_occ_occ[is], block_size, nocc[is], nocc[is] +#ifdef __MPI + , px[is].blacs_ctxt +#endif + ); + } + std::vector W(p_occ_occ[0].get_local_size() * nk, 0.0); + cal_W_from_Z(inputs, W.data(), Z, X, eig_ext_istate, eig_ks, c, p_occ_occ, spin_type); + // std::cout << "W: " << std::endl; + // LR_Util::print_value(W.data(), nk, p_occ_occ[0].get_col_size(), p_occ_occ[0].get_row_size()); + + // 2. build K_cvcx (nvirt*nocc) = \sum_i X_{ia} K_{ij} = \sum_i X_{ia} \sum_{\mu\nu} c_{\mu i} c_{\nu j} K_{\mu\nu}[D^X] + // $2\sum_i X_{ai} K_{ij}[D_X]$ (D_X is symmetrized) + OperatorLRHxc op_K_cvcx(nspin, naos, nocc, nvirt, c, + dm_trans, pot, ucell, orb_cutoff, gd, kv, px, pc, pmat, + { 0 }, T(2.0), OperatorLRHxc::MO_TO_AO_TYPE::CXC_o); +#ifdef __EXX + // this EDM term only runs on the force-calculation path, so cal_force is always true here. + OperatorLREXX op_K_exx(nspin, naos, nocc[0], nvirt[0], ucell, c, + dm_trans, exx_lri, kv, px[0], pc, pmat, + /*cal_force=*/true, 2.0 * exx_alpha, OperatorLREXX::MO_TO_AO_TYPE::CXC_o); +#endif + const int ld_vo = nk * px[0].get_local_size(); + std::vector K_cvcx(ld_vo, 0.0); + op_K_cvcx.act(/*nbands=*/1, ld_vo, /*npol=*/1, X, K_cvcx.data()); +#ifdef __EXX + if (LR::exx_kernel_list().count(xc_kernel)) + op_K_exx.act(/*nbands=*/1, ld_vo, /*npol=*/1, X, K_cvcx.data()); +#endif + + return cal_edm_terms_from_XZWK(X, Z, W.data(), K_cvcx.data(), eig_ext_istate, eig_ks, c, nspin, test_force, p_occ_occ[0], px[0], pc, pmat); + } + + /// @brief Open-shell (spin-unrestricted) counterpart of `cal_edm_from_XZ_istate`. + /// Returns the energy-weighted density matrix of each spin channel: `[is][ik]`. + /// + /// The four EDM terms are all spin-diagonal AO outer products; the spin coupling only + /// enters when building the two multipliers, $W^c$ (see `cal_W_from_Z_openshell`) and + /// $W^X_{ki\sigma}=2K_{ki\sigma}[D^X]$ (the `op_K_cvcx` blocks below). + template + std::vector> cal_edm_from_XZ_istate_openshell( + const GradientInputs& inputs, + const T* const X, + const T* const Z, + const double eig_ext_istate, + const double* const eig_ks, + const module_dm::DensityMatrix& dm_trans, // retained for signature symmetry + std::weak_ptr pot) + { + const int& nspin = inputs.nspin; + const int& naos = inputs.nbasis; + const std::vector& nocc = inputs.nocc; + const std::vector& nvirt = inputs.nvirt; + const UnitCell& ucell = inputs.ucell; + const std::vector& orb_cutoff = inputs.orb_cutoff; + const Grid_Driver& gd = inputs.gd; + const K_Vectors& kv = inputs.kv; + const std::vector& px = inputs.px; + const Parallel_2D& pc = inputs.pc; + const Parallel_Orbitals& pmat = inputs.pmat; +#ifdef __EXX + std::weak_ptr> exx_lri = inputs.exx_lri; + const double& exx_alpha = inputs.hybrid_alpha; +#endif + const psi::Psi& psi_ks = inputs.psi_ks; + const bool test_force = inputs.test_force; + const std::string xc_kernel = inputs.xc_kernel; + using ATYPE = typename OperatorLRHxc::MO_TO_AO_TYPE; +#ifdef __EXX + using ATYPE_EXX = typename OperatorLREXX::MO_TO_AO_TYPE; +#endif + const int nk = kv.get_nks() / nspin; + const std::vector ld_x = { static_cast(nk * px[0].get_local_size()), static_cast(nk * px[1].get_local_size()) }; + const std::vector off_x = { 0, ld_x[0] }; + const int nband_window = nocc[0] + nvirt[0]; + + // 1. the W^c multiplier, one occ-occ block per spin + std::vector p_occ_occ(2); + for (int is : {0, 1}) + { + const int block_size = px[is].get_block_size(); + LR_Util::setup_2d_division(p_occ_occ[is], block_size, nocc[is], nocc[is] +#ifdef __MPI + , px[is].blacs_ctxt +#endif + ); + } + std::vector> W; + cal_W_from_Z_openshell(inputs, W, Z, X, eig_ext_istate, eig_ks, p_occ_occ); + + // 2. $W^X_{ai\sigma}=2\sum_j X_{aj\sigma}K_{ji\sigma}[D^X]$. + // The free spin sits on X (hence `psi_in = X + off_x[sl]`, laid out over `px[sl]`), + // the summed spin sits on $D^X$. + std::vector> K_cvcx(2); + for (int is : {0, 1}) { K_cvcx[is].assign(ld_x[is], T(0.0)); } + + module_dm::DensityMatrix DM_trans(&pmat, 1, kv.kvec_d, nk); + LR_Util::initialize_DMR(DM_trans, pmat, ucell, gd, orb_cutoff); + std::vector>> op_K(4); + for (int sl : {0, 1}) + { + for (int sr : {0, 1}) + { + op_K[(sl << 1) + sr] = LR_Util::make_unique>(nspin, naos, nocc, nvirt, psi_ks, + DM_trans, pot, ucell, orb_cutoff, gd, kv, px, pc, pmat, + std::vector({ sl, sr }), T(2.0), ATYPE::CXC_o); + } + } + std::vector> psi_ks_spin; + for (int is : {0, 1}) { psi_ks_spin.push_back(LR_Util::get_psi_spin(psi_ks, is, nk)); } +#ifdef __EXX + std::vector>> op_K_exx(2); + const bool with_exx_lr = LR::exx_kernel_list().count(xc_kernel) > 0; + if (with_exx_lr) + { + for (int is : {0, 1}) + { + op_K_exx[is] = LR_Util::make_unique>(nspin, naos, nocc[is], nvirt[is], + ucell, psi_ks_spin[is], DM_trans, exx_lri, kv, px[is], pc, pmat, + /*cal_force=*/true, 2.0 * exx_alpha, ATYPE_EXX::CXC_o); + } + } +#endif + // $D^X$ is fed to the CXC_o operators TRANSPOSED, exactly as the closed-shell + // `cal_force` does (it hands `cal_edm_from_XZ_istate` a `transpose_DMR`-ed $D^X$): + // `CVCX_occ` produces the kernel matrix with its two MO indices in the opposite order + // to what $W^X_{ai\sigma}=2\sum_jX_{aj\sigma}K_{ji\sigma}[D^X]$ needs, and since + // $(K[D])^T=K[D^T]$, transposing on the way in restores it. + std::vector dmx_buf; + auto set_dm_trans = [&](const int is)->void + { +#ifdef __MPI + dmx_buf = cal_dm_trans_pblas(X + off_x[is], px[is], psi_ks_spin[is], pc, naos, nocc[is], nvirt[is], pmat); + for (auto& t : dmx_buf) { LR_Util::mattrans(t.data(), naos, pmat); } +#else + dmx_buf = cal_dm_trans_blas(X + off_x[is], psi_ks_spin[is], nocc[is], nvirt[is]); + for (auto& t : dmx_buf) + { + T* d = t.data(); + for (int u = 0;u < naos;++u) + { + for (int v = u + 1;v < naos;++v) { std::swap(d[u * naos + v], d[v * naos + u]); } + } + } +#endif + for (int ik = 0;ik < nk;++ik) { DM_trans.set_dmk_ptr(ik, dmx_buf[ik].data()); } + }; + for (int sr : {0, 1}) + { + set_dm_trans(sr); + for (int sl : {0, 1}) + { + op_K[(sl << 1) + sr]->act(/*nbands=*/1, ld_x[sl], /*npol=*/1, + X + off_x[sl], K_cvcx[sl].data()); + } +#ifdef __EXX + if (with_exx_lr) + { + op_K_exx[sr]->act(/*nbands=*/1, ld_x[sr], /*npol=*/1, X + off_x[sr], K_cvcx[sr].data()); + } +#endif + } + + // 3. assemble the four EDM terms, per spin channel + std::vector> edm(2); + for (int is : {0, 1}) + { + edm[is] = cal_edm_terms_from_XZWK(X + off_x[is], Z + off_x[is], W[is].data(), K_cvcx[is].data(), + eig_ext_istate, eig_ks + is * nk * nband_window, psi_ks_spin[is], nspin, test_force, + p_occ_occ[is], px[is], pc, pmat); + } + return edm; + } +} + +#endif // ABACUS_LR_CAL_EDM_H diff --git a/source/source_lcao/module_lr/cal_hs_grad.h b/source/source_lcao/module_lr/cal_hs_grad.h new file mode 100644 index 00000000000..054683b10dc --- /dev/null +++ b/source/source_lcao/module_lr/cal_hs_grad.h @@ -0,0 +1,152 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_CAL_HS_GRAD_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_CAL_HS_GRAD_H +#include +#include "source_cell/module_neighbor/sltk_grid_driver.h" +#include "source_cell/unitcell.h" +#include "source_basis/module_ao/parallel_orbitals.h" +#include "source_basis/module_nao/two_center_bundle.h" +#include "source_hamilt/module_hcontainer/hcontainer.h" + +inline void filter_adjs_by_rcut(const UnitCell& ucell, + const int iat0, + AdjacentAtomInfo& adjs) +{ + std::vector is_adj(adjs.adj_num + 1, false); + for (int ad = 0;ad < adjs.adj_num + 1;++ad) + { + const int it0 = ucell.iat2it[iat0]; + const int it1 = adjs.ntype[ad]; + const int ia1 = adjs.natom[ad]; + const int iat1 = ucell.itia2iat(it1,ia1); + const ModuleBase::Vector3& R_index1 = adjs.box[ad]; + if(ucell.cal_dtau(iat0, iat1, R_index1).norm() * ucell.lat0 < ucell.atoms[it0].Rcut + ucell.atoms[it1].Rcut - 1e-15) + { + is_adj[ad] = true; + } + } + filter_adjs(is_adj, adjs); +} + +/// @brief Return a local Hamiltonian operator in type HContainer. +/// It can replace OverlapNew::initialize_SR() and EkineticNew::initialize_HR(). +template +hamilt::HContainer build_hcontainer_local_op(const UnitCell& ucell, const Grid_Driver& gd, const Parallel_Orbitals& pv) +{ + hamilt::HContainer hcontainer(&pv); + for (int iat0 = 0;iat0 < ucell.nat; ++iat0) + { + const int it0 = ucell.iat2it[iat0]; + const int ia0 = ucell.iat2ia[iat0]; + AdjacentAtomInfo adjs; + gd.Find_atom(ucell, ucell.get_tau(iat0), it0, ia0, &adjs); + filter_adjs_by_rcut(ucell, iat0, adjs); + + for (int ad = 0;ad < adjs.adj_num + 1;++ad) + { + const int it1 = adjs.ntype[ad]; + const int ia1 = adjs.natom[ad]; + const int iat1 = ucell.itia2iat(it1, ia1); + if (pv.get_nrow_atom(iat0) * pv.get_ncol_atom(iat1)) + { + hamilt::AtomPair ap(iat0, iat1, adjs.box[ad], &pv); + hcontainer.insert_pair(ap); + } + } + } + hcontainer.allocate(nullptr, true); + return hcontainer; +} + +/// @brief Calculate or by 2-center integration +inline std::vector> cal_hs_grad(const char job, + const UnitCell& ucell, + const Parallel_Orbitals& pv, + const Grid_Driver& gd, + const TwoCenterBundle& two_center_bundle) +{ + if(job != 'S' && job != 'T') + { + throw std::invalid_argument("job must be 'S' or 'T'"); + } + // allocate dHS (better to use HContainer> for access continuity) + // hamilt::HContainer> dHS(&pv); + std::vector> dHS(3, build_hcontainer_local_op(ucell, gd, pv)); + // std::vector> dHS(3, build_hcontainer_local_op(ucell, gd, pv)); + std::vector tmp_deriv(3); + const int npol = ucell.get_npol(); + + for (int ixyz = 0;ixyz < 3;++ixyz) + { + // traverse ijR to calculate dHS + std::vector ijr_info = dHS.at(ixyz).get_ijr_info(); + // std::vector ijr_info = dHS.at(0).get_ijr_info(); + std::vector::iterator it = ijr_info.begin(); + const int npairs = *it++; + int npairs_count = 0; + while (it != ijr_info.end()) + { + ++npairs_count; + const int iat0 = *it++; + const int it0 = ucell.iat2it[iat0]; + const Atom& atom0 = ucell.atoms[it0]; + const ModuleBase::Vector3 tau0 = ucell.get_tau(iat0); + auto row_indexes = pv.get_indexes_row(iat0); + + const int iat1 = *it++; + const int it1 = ucell.iat2it[iat1]; + const Atom& atom1 = ucell.atoms[it1]; + const ModuleBase::Vector3 tau1 = ucell.get_tau(iat1); + auto col_indexes = pv.get_indexes_col(iat1); + + const int nR = *it++; + + for (int iR = 0;iR < nR;++iR) + { + // Read the three components in separate statements. The order in which function + // arguments are evaluated is UNSPECIFIED in C++, so `Vector3(*it++, *it++, *it++)` + // may store the R triple permuted and silently corrupt periodic systems. + const int Rx = *it++; + const int Ry = *it++; + const int Rz = *it++; + const ModuleBase::Vector3 R(Rx, Ry, Rz); // int to double + ModuleBase::Vector3 relative_position = (tau1 - tau0 + R * ucell.latvec) * ucell.lat0; + hamilt::BaseMatrix* dHS_block = dHS[ixyz].find_matrix(iat0, iat1, R.x, R.y, R.z); + // `ijr_info` came from this very container, so a miss means the indices are wrong. + assert(dHS_block != nullptr); + + // OMP can be used here + for (int lw0 = 0;lw0 < row_indexes.size();lw0 += npol) // spin 1-3 of dHS is not needed at nspin=4 + { + const int gw0 = row_indexes[lw0] / npol; + const int l0 = atom0.iw2l[gw0]; + const int n0 = atom0.iw2n[gw0]; + const int m0 = atom0.iw2m[gw0]; + const int M0 = (m0 % 2 == 0) ? -m0 / 2 : (m0 + 1) / 2; // convert m (0,1,...2l) to M (-l, -l+1, ..., l-1, l) + for (int lw1 = 0;lw1 < col_indexes.size();lw1 += npol) + { + const int& gw1 = col_indexes[lw1] / npol; + const int l1 = atom1.iw2l[gw1]; + const int n1 = atom1.iw2n[gw1]; + const int m1 = atom1.iw2m[gw1]; + const int M1 = (m1 % 2 == 0) ? -m1 / 2 : (m1 + 1) / 2; + switch (job) + { + case 'S': + two_center_bundle.overlap_orb->calculate(it0, l0, n0, M0, it1, l1, n1, M1, + relative_position, nullptr, tmp_deriv.data()); + break; + case 'T': + two_center_bundle.kinetic_orb->calculate(it0, l0, n0, M0, it1, l1, n1, M1, + relative_position, nullptr, tmp_deriv.data()); + break; + } + dHS_block->get_value(lw0, lw1) = tmp_deriv[ixyz]; + } + } + } + } + assert(npairs == npairs_count); + } + return dHS; +} +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_CAL_HS_GRAD_H diff --git a/source/source_lcao/module_lr/cal_w_from_z.h b/source/source_lcao/module_lr/cal_w_from_z.h new file mode 100644 index 00000000000..76196ea93ee --- /dev/null +++ b/source/source_lcao/module_lr/cal_w_from_z.h @@ -0,0 +1,348 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_CAL_W_FROM_Z_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_CAL_W_FROM_Z_H +#include "gradient_inputs.h" +#include "source_hamilt/hamilt.h" +#include "source_lcao/module_lr/dm_trans/dm_diff.h" +#include "source_estate/module_dm/density_matrix.h" +#include "source_lcao/module_lr/potentials/pot_grad_xc.h" +#include "source_lcao/module_lr/operator_casida/op_gxc_ulr.h" +#include "source_lcao/module_lr/potentials/pot_hxc_lrtd.h" +#include "source_lcao/module_lr/operator_casida/operator_lr_hxc.h" +#include "source_basis/module_ao/parallel_orbitals.h" +#include +#ifdef __EXX +#include "source_lcao/module_lr/operator_casida/operator_lr_exx.h" +#endif +#include "source_base/module_external/scalapack_connector.h" +#include "source_base/module_external/blas_connector.h" + +namespace LR +{ + // $\sum_a X^*_{ai} X_{aj}(\Omega - \epsilon_a)$ + template + void add_ediff_term(T* const inout, + const T* const X, + const double eig_ext_istate, + const double* const eig_ks, + const int nk, + const int nocc, + const int nvirt, + const Parallel_2D& px, + const Parallel_2D& p_occ_occ) + { + for (int ik = 0;ik < nk;++ik) + { + // 1. calculate (\Omega-\epsilon_a)X_{aj} + std::vector wX(px.get_local_size()); + const int eigks_start_k = ik * (nocc + nvirt); + const int x_start_k = ik * px.get_local_size(); + const int inout_start_k = ik * p_occ_occ.get_local_size(); + for (int la = 0;la < px.get_row_size();++la) + { + const int ga = px.local2global_row(la); + const double weight = eig_ext_istate - eig_ks[eigks_start_k + nocc + ga]; + for (int li = 0;li < px.get_col_size();++li) + { + const int idx = li * px.get_row_size() + la; + wX[idx] = weight * X[x_start_k + idx]; + } + } + // 2. matrix multiplication (parallel) +#ifdef __MPI + const int i1 = 1; + ScalapackConnector::gemm('C', 'N', nocc, nocc, nvirt, + T(1.0), X + x_start_k, i1, i1, px.desc, + wX.data(), i1, i1, px.desc, + T(1.0)/*add-on*/, inout + inout_start_k, i1, i1, p_occ_occ.desc); +#else + BlasConnector::gemm_cm('C', 'N', nocc, nocc, nvirt, + T(1.0), X + x_start_k, nvirt, + wX.data(), nvirt, + T(1.0)/*add-on*/, inout + inout_start_k, nocc); +#endif + } + } + + template + void cal_W_from_Z(const GradientInputs& inputs, + T* const W, + const T* const Z, + const T* const X, + const double eig, + const double* const eig_ks, + const psi::Psi& psi_ks, // the closed-shell caller supplies a single-spin view + const std::vector& p_occ_occ, + const std::string& spin_type = "singlet") + { + const int& nspin = inputs.nspin; + const int& naos = inputs.nbasis; + const std::vector& nocc = inputs.nocc; + const std::vector& nvirt = inputs.nvirt; + const UnitCell& ucell = inputs.ucell; + const std::vector& orb_cutoff = inputs.orb_cutoff; + const Grid_Driver& gd = inputs.gd; + const K_Vectors& kv = inputs.kv; + const std::vector& px = inputs.px; + const Parallel_2D& pc = inputs.pc; + const Parallel_Orbitals& pmat = inputs.pmat; + const std::string& dft_functional = inputs.dft_functional; + std::weak_ptr pot_hxc_gs = inputs.pot_hxc_gs; +#ifdef __EXX + std::weak_ptr> exx_lri = inputs.exx_lri; + const double& exx_alpha = inputs.hybrid_alpha; +#endif + const std::string xc_kernel = inputs.xc_kernel; + ModuleBase::TITLE("cal_W_from_Z", "cal_W_from_Z"); + using ATYPE = typename OperatorLRHxc::MO_TO_AO_TYPE; +#ifdef __EXX + using ATYPE_EXX = typename OperatorLREXX::MO_TO_AO_TYPE; +#endif + const int nk = kv.get_nks() / nspin; + // allocate memory for DMs + module_dm::DensityMatrix DM_trans(&pmat, 1, kv.kvec_d, nk); //DX + LR_Util::initialize_DMR(DM_trans, pmat, ucell, gd, orb_cutoff); + module_dm::DensityMatrix DM_diff_relaxed(&pmat, 1, kv.kvec_d, nk); //T+DZ + LR_Util::initialize_DMR(DM_diff_relaxed, pmat, ucell, gd, orb_cutoff); + /// operators + // 1. 0.5$H_ij[T+Z]$, equals to $K_ij[T+Z]$ when $(T+Z)$ is symmetrized + // Note that K_Hxc(singlet) = 2* pot_hxc_gs, that's why there's factor 2 here. + OperatorLRHxc op_ht(nspin, naos, nocc, nvirt, psi_ks, + DM_diff_relaxed, pot_hxc_gs, ucell, orb_cutoff, gd, kv, p_occ_occ, pc, pmat, + { 0 }, T(2.0), ATYPE::CC_oo); +#ifdef __EXX + // cal_W_from_Z only runs on the force-calculation path, so cal_force is always true here. + OperatorLREXX op_ht_exx(nspin, naos, nocc[0], nvirt[0], ucell, psi_ks, + DM_diff_relaxed, exx_lri, kv, p_occ_occ[0], pc, pmat, + /*cal_force=*/true, exx_alpha, ATYPE_EXX::CC_oo); +#endif + // 2. $2\sum_{jb,kc} g^{xc}_{ia, jb, kc}X_{jb}X_{kc}$ + // use pointer here for polymorphism + // but `weak_ptr = make_shared()` will cause a segment fault because the shared_ptr is a temporary object + // correct way is to use `shared_ptr = make_shared()` and then assign it to weak_ptr + // `weak_ptr=shared_ptr` is automatically called in the constructor of OperatorLRHxc, so we don't need to do it manually + // if `pot_grad` is passed into a function rather than a class, we need to write `weak_ptr=shared_ptr` explicitly + // This quadratic term differentiates the LR kernel K, unlike H[T+DZ] above. + const bool triplet = spin_type == "triplet"; + const int ispin = triplet ? 1 : 0; + const std::shared_ptr& pot_lr = inputs.pot[ispin]; + std::shared_ptr pot_grad = + std::make_shared(pot_lr->xc_kernel_components(), pot_lr->get_rho_basis(), + ucell, pot_lr->nrxx, triplet); + OperatorLRHxc op_gxc(nspin, naos, nocc, nvirt, psi_ks, + DM_trans, pot_grad, ucell, orb_cutoff, gd, kv, p_occ_occ, pc, pmat, + // Factor 1.0 according to the $W^c$ formula (`pot_grad` carries $2*g^{xc}$: uu+ud or uu-ud). + // NOTE this factor is NOT shared with the Z-vector RHS `op_gxc` in `hamilt_zeq_right.h`: + // that one is a different object and its original -2.0 is correct, as the formula and H2-DZP scan confirms. + { 0 }, T(1.0), ATYPE::CC_oo); + + std::vector dm_trans_2d, dm_diff_2d; + auto cal_dm_trans = [&](const int is, const T* const x_ptr)->void //DX + { + const auto psi_ks_is = LR_Util::get_psi_spin(psi_ks, is, nk); +#ifdef __MPI + dm_trans_2d = cal_dm_trans_pblas(x_ptr, px[is], psi_ks_is, pc, naos, nocc[is], nvirt[is], pmat); + for (auto& t : dm_trans_2d) LR_Util::matsym(t.data(), naos, pmat); +#else + dm_trans_2d = cal_dm_trans_blas(x_ptr, psi_ks_is, nocc[is], nvirt[is]); + for (auto& t : dm_trans_2d) LR_Util::matsym(t.data(), naos); +#endif + for (int ik = 0;ik < nk;++ik) { DM_trans.set_dmk_ptr(ik, dm_trans_2d[ik].data()); } + }; + auto cal_dm_diff_relaxed = [&](const int& is, const T* const x_ptr, const T* const z_ptr)->void // T+DZ + { + const auto psi_ks_is = LR_Util::get_psi_spin(psi_ks, is, nk); +#ifdef __MPI + std::vector z_2d = cal_dm_trans_pblas(z_ptr, px[is], psi_ks_is, pc, naos, nocc[is], nvirt[is], pmat); + for (auto& t : z_2d) LR_Util::matsym(t.data(), naos, pmat); +#else + std::vector z_2d = cal_dm_trans_blas(z_ptr, psi_ks_is, nocc[is], nvirt[is]); + for (auto& t : z_2d) LR_Util::matsym(t.data(), naos); +#endif + +#ifdef __MPI + dm_diff_2d = cal_dm_diff_pblas(x_ptr, px[is], psi_ks_is, pc, naos, nocc[is], nvirt[is], pmat); +#else + dm_diff_2d = cal_dm_diff_blas(x_ptr, psi_ks_is, naos, nocc[is], nvirt[is]); +#endif + for (int ik = 0;ik < nk;++ik) + { + dm_diff_2d[ik] = dm_diff_2d[ik] + z_2d[ik]; + DM_diff_relaxed.set_dmk_ptr(ik, dm_diff_2d[ik].data()); + } + }; + + // act the operators onto current state + const int ld_vo = nk * px[0].get_local_size(); + const int ld_oo = nk * p_occ_occ[0].get_local_size(); + cal_dm_trans(0, X); // transition density matrix DX + cal_dm_diff_relaxed(0, X, Z); // relaxed difference density matrix T+DZ + // the 3 terms + op_ht.act(/*nband=*/1, ld_oo, /*npol=*/1, X, W); //comment out this line to test H[T+Z]=0 + // std::cout << "W (H[T+Z])) local terms: " << std::endl; + // LR_Util::print_value(W, nk, p_occ_occ[0].get_col_size(), p_occ_occ[0].get_row_size()); +#ifdef __EXX + if (LR::gs_is_hybrid(dft_functional)) // H[T+Z] term depends on ground-state kernel + op_ht_exx.act(/*nband=*/1, ld_oo, /*npol=*/1, X, W); +#endif + // std::cout << "W (H[T+Z])) local +exx terms: " << std::endl; + // LR_Util::print_value(W, nk, p_occ_occ[0].get_col_size(), p_occ_occ[0].get_row_size()); + // Not singlet-only: $K^T_{xc}=f_{uu}-f_{ud}\ne0$ for a local functional, so $W^{c,T}$ has a + // $g^{xc}$ term as well, built from the "-" spin combination (the `triplet` flag above). + if (LR_Util::has_local_xc(xc_kernel)) + op_gxc.act(/*nband=*/1, ld_oo, /*npol=*/1, X, W); + + add_ediff_term(W, X, eig, eig_ks, nk, nocc[0], nvirt[0], px[0], p_occ_occ[0]); + } + + /// @brief Open-shell (spin-unrestricted) counterpart of `cal_W_from_Z`. + /// + /// $$W^c_{ij\sigma}=\tfrac12 H_{ij\sigma}[T+D^Z] + /// +\sum_a X_{ai\sigma}(\Omega-\epsilon_{a\sigma})X_{aj\sigma} + /// +\sum_{\kappa\lambda\sigma'}\sum_{\alpha\beta\sigma''}D^X D^X g^{xc}_{\dots,ij\sigma}$$ + /// + /// `W[is]` is the occ-occ block of spin channel `is`, distributed over `p_occ_occ[is]`. + /// Factor 1.0 on the $H[T+D^Z]$ operator (the closed-shell version uses 2.0): there + /// $\tfrac12 H^S = K^S = 2\cdot$`pot_hxc_gs` because `pot_hxc_gs` is the halved `S2_gs`; + /// here it is `S2_updown`, one $K_{\sigma\sigma'}$ component, and the $\sum_{\sigma'}$ + /// is done by the block loop below. + template + void cal_W_from_Z_openshell(const GradientInputs& inputs, + std::vector>& W, + const T* const Z, + const T* const X, + const double eig, + const double* const eig_ks, + const std::vector& p_occ_occ) + { + const int& nspin = inputs.nspin; + const int& naos = inputs.nbasis; + const std::vector& nocc = inputs.nocc; + const std::vector& nvirt = inputs.nvirt; + const UnitCell& ucell = inputs.ucell; + const std::vector& orb_cutoff = inputs.orb_cutoff; + const Grid_Driver& gd = inputs.gd; + const K_Vectors& kv = inputs.kv; + const std::vector& px = inputs.px; + const Parallel_2D& pc = inputs.pc; + const Parallel_Orbitals& pmat = inputs.pmat; + const std::string& dft_functional = inputs.dft_functional; + std::weak_ptr pot_hxc_gs = inputs.pot_hxc_gs; +#ifdef __EXX + std::weak_ptr> exx_lri = inputs.exx_lri; + const double& exx_alpha = inputs.hybrid_alpha; +#endif + const psi::Psi& psi_ks = inputs.psi_ks; + const std::string xc_kernel = inputs.xc_kernel; + const std::string& ks_solver = inputs.ks_solver; + ModuleBase::TITLE("cal_W_from_Z_openshell", "cal_W_from_Z_openshell"); + using ATYPE = typename OperatorLRHxc::MO_TO_AO_TYPE; +#ifdef __EXX + using ATYPE_EXX = typename OperatorLREXX::MO_TO_AO_TYPE; +#endif + const int nk = kv.get_nks() / nspin; + const std::vector ld_x = { static_cast(nk * px[0].get_local_size()), static_cast(nk * px[1].get_local_size()) }; + const std::vector ld_oo = { static_cast(nk * p_occ_occ[0].get_local_size()), static_cast(nk * p_occ_occ[1].get_local_size()) }; + const std::vector off_x = { 0, ld_x[0] }; + const int nband_window = nocc[0] + nvirt[0]; // common KS window, see `set_dimension` + + W.assign(2, {}); + for (int is : {0, 1}) { W[is].assign(ld_oo[is], T(0.0)); } + + module_dm::DensityMatrix DM_diff_relaxed(&pmat, 1, kv.kvec_d, nk); // T+D^Z of one channel + LR_Util::initialize_DMR(DM_diff_relaxed, pmat, ucell, gd, orb_cutoff); + + // $\tfrac12 H_{ij\sigma}[T+D^Z]=K_{ij\sigma}[T+D^Z]$, one operator per (out, in) spin pair + std::vector>> op_ht(4); + for (int sl : {0, 1}) + { + for (int sr : {0, 1}) + { + op_ht[(sl << 1) + sr] = LR_Util::make_unique>(nspin, naos, nocc, nvirt, psi_ks, + DM_diff_relaxed, pot_hxc_gs, ucell, orb_cutoff, gd, kv, p_occ_occ, pc, pmat, + std::vector({ sl, sr }), T(1.0), ATYPE::CC_oo); + } + } +#ifdef __EXX + std::vector> psi_ks_spin; + for (int is : {0, 1}) { psi_ks_spin.push_back(LR_Util::get_psi_spin(psi_ks, is, nk)); } + std::vector>> op_ht_exx(2); + const bool with_exx = LR::gs_is_hybrid(dft_functional); + if (with_exx) + { // exchange is spin-diagonal + for (int is : {0, 1}) + { + op_ht_exx[is] = LR_Util::make_unique>(nspin, naos, nocc[is], nvirt[is], + ucell, psi_ks_spin[is], DM_diff_relaxed, exx_lri, kv, p_occ_occ[is], pc, pmat, + /*cal_force=*/true, exx_alpha, ATYPE_EXX::CC_oo); + } + } +#endif + // the relaxed difference density matrix of one spin channel + std::vector dm_buf; + auto set_dm_diff_relaxed = [&](const int is)->void + { + const auto psi_ks_is = LR_Util::get_psi_spin(psi_ks, is, nk); + const T* const x_ptr = X + off_x[is]; + const T* const z_ptr = Z + off_x[is]; +#ifdef __MPI + std::vector z_2d = cal_dm_trans_pblas(z_ptr, px[is], psi_ks_is, pc, naos, nocc[is], nvirt[is], pmat); + for (auto& t : z_2d) { LR_Util::matsym(t.data(), naos, pmat); } + dm_buf = cal_dm_diff_pblas(x_ptr, px[is], psi_ks_is, pc, naos, nocc[is], nvirt[is], pmat); +#else + std::vector z_2d = cal_dm_trans_blas(z_ptr, psi_ks_is, nocc[is], nvirt[is]); + for (auto& t : z_2d) { LR_Util::matsym(t.data(), naos); } + dm_buf = cal_dm_diff_blas(x_ptr, psi_ks_is, naos, nocc[is], nvirt[is]); +#endif + for (int ik = 0;ik < nk;++ik) + { + dm_buf[ik] = dm_buf[ik] + z_2d[ik]; + DM_diff_relaxed.set_dmk_ptr(ik, dm_buf[ik].data()); + } + }; + + for (int is_in : {0, 1}) + { + set_dm_diff_relaxed(is_in); + for (int is_out : {0, 1}) + { + op_ht[(is_out << 1) + is_in]->act(/*nband=*/1, ld_oo[is_out], /*npol=*/1, + X + off_x[is_in], W[is_out].data()); + } +#ifdef __EXX + if (with_exx) + { // $\delta_{\sigma\sigma'}$: only the diagonal block contributes + op_ht_exx[is_in]->act(/*nband=*/1, ld_oo[is_in], /*npol=*/1, + X + off_x[is_in], W[is_in].data()); + } +#endif + } + + // $\sum_{\sigma'\sigma''}D^X_{\sigma'}D^X_{\sigma''}g^{xc}_{\dots,ij\tau}$, coefficient 1 + // in the spin-orbital formula. Quadratic in $D^X$, so it does not fit the block loop above. + // Use the LR kernel's density derivative, consistently with the Z-vector RHS. + // The separate H[T+DZ] operators above retain the GS kernel. + if (LR_Util::has_local_xc(xc_kernel)) + { + const std::shared_ptr& pot_lr = inputs.pot[0]; + OperatorGxcULR gxc(pot_lr->xc_kernel_components(), pot_lr->get_rho_basis(), + ucell, orb_cutoff, gd, kv, pmat, pc, psi_ks, nocc, nvirt, naos, + px, p_occ_occ, LR_Util::MO_TYPE::OO, T(1.0), nspin, ks_solver); + std::vector w_flat(ld_oo[0] + ld_oo[1], T(0.0)); + gxc.act(X, w_flat.data()); + for (int is : {0, 1}) + { + const int off = is * ld_oo[0]; + for (int i = 0;i < ld_oo[is];++i) { W[is][i] += w_flat[off + i]; } + } + } + + // $\sum_a X_{ai\sigma}(\Omega-\epsilon_{a\sigma})X_{aj\sigma}$ -- spin-diagonal + for (int is : {0, 1}) + { + add_ediff_term(W[is].data(), X + off_x[is], eig, eig_ks + is * nk * nband_window, + nk, nocc[is], nvirt[is], px[is], p_occ_occ[is]); + } + } +} + +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_CAL_W_FROM_Z_H diff --git a/source/source_lcao/module_lr/dm_trans/dm_diff.h b/source/source_lcao/module_lr/dm_trans/dm_diff.h new file mode 100644 index 00000000000..5c269b8125b --- /dev/null +++ b/source/source_lcao/module_lr/dm_trans/dm_diff.h @@ -0,0 +1,53 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_DM_TRANS_DM_DIFF_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_DM_TRANS_DM_DIFF_H +#include +#include "source_psi/psi.h" +#include +#include "source_lcao/module_lr/utils/lr_util.h" +namespace LR +{ + // use templates in the future. +#ifdef __MPI +/// @brief calculate the 2d-block difference density matrix in AO basis using p?gemm +/// \f[ T=(C_v*X) * (C_v * X)^\dagger + (C_o*X^T) * (C_o*X^T)^\dagger \f] + template + std::vector cal_dm_diff_pblas( + const T* const X_istate, + const Parallel_2D& px, + const psi::Psi& c, + const Parallel_2D& pc, + const int& naos, + const int& nocc, + const int& nvirt, + const Parallel_2D& pmat, + const bool renorm_k = true, + const int nspin = 1); +#endif + + /// @brief calculate the 2d-block transition density matrix in AO basis using ?gemm + template + std::vector cal_dm_diff_blas( + const T* const X_istate, + const psi::Psi& c, + const int& naos, + const int& nocc, + const int& nvirt, + const bool renorm_k = true, + const int nspin = 1); + + // for test + /// @brief calculate the 2d-block transition density matrix in AO basis using for loop (for test) + template + std::vector cal_dm_diff_forloop( + const T* const X_istate, + const psi::Psi& c, + const int& naos, + const int& nocc, + const int& nvirt, + const bool renorm_k = true, + const int nspin = 1); +} + +#include "dm_diff_ser.hpp" +#include "dm_diff_par.hpp" +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_DM_TRANS_DM_DIFF_H diff --git a/source/source_lcao/module_lr/dm_trans/dm_diff_par.hpp b/source/source_lcao/module_lr/dm_trans/dm_diff_par.hpp new file mode 100644 index 00000000000..967490bd919 --- /dev/null +++ b/source/source_lcao/module_lr/dm_trans/dm_diff_par.hpp @@ -0,0 +1,194 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_DM_TRANS_DM_DIFF_PAR_HPP +#define ABACUS_SOURCE_LCAO_MODULE_LR_DM_TRANS_DM_DIFF_PAR_HPP +#ifdef __MPI +// #include +#include "source_base/module_container/ATen/core/tensor_types.h" +#include "source_base/module_external/scalapack_connector.h" +#include "source_base/tool_title.h" +#include "source_lcao/module_lr/utils/lr_util.h" +#include "dm_diff.h" +namespace LR +{ + inline void CvX( + const double* C, + const Parallel_2D& pc, + const double* X, + const Parallel_2D& px, + const int& naos, + const int& nocc, + const int& nvirt, + double* CvX, + const Parallel_2D& pcx) + { + const int i1 = 1; + const int ivirt = nocc + 1; + char transa = 'N'; + char transb = 'N'; + const double alpha = 1.0; + const double beta = 0; + pdgemm_(&transa, &transb, &naos, &nocc, &nvirt, + &alpha, C, &i1, &ivirt, pc.desc, + X, &i1, &i1, px.desc, + &beta, CvX, &i1, &i1, pcx.desc); + } + inline void CvX( + const std::complex* C, + const Parallel_2D& pc, + const std::complex* X, + const Parallel_2D& px, + const int& naos, + const int& nocc, + const int& nvirt, + std::complex* CvX, + const Parallel_2D& pcx) + { + const int i1 = 1; + const int ivirt = nocc + 1; + char transa = 'N'; + char transb = 'N'; + const std::complex alpha(1.0, 0.0); + const std::complex beta(0.0, 0.0); + pzgemm_(&transa, &transb, &naos, &nocc, &nvirt, + &alpha, C, &i1, &ivirt, pc.desc, + X, &i1, &i1, px.desc, + &beta, CvX, &i1, &i1, pcx.desc); + } + + inline void CoXT( + const double* C, + const Parallel_2D& pc, + const double* X, + const Parallel_2D& px, + const int& naos, + const int& nocc, + const int& nvirt, + double* CoXT, + const Parallel_2D& pcxt) + { + const int i1 = 1; + char transa = 'N'; + char transb = 'T'; + const double alpha = 1.0; + const double beta = 0; + pdgemm_(&transa, &transb, &naos, &nvirt, &nocc, + &alpha, C, &i1, &i1, pc.desc, + X, &i1, &i1, px.desc, + &beta, CoXT, &i1, &i1, pcxt.desc); + } + inline void CoXT( + const std::complex* C, + const Parallel_2D& pc, + const std::complex* X, + const Parallel_2D& px, + const int& naos, + const int& nocc, + const int& nvirt, + std::complex* CoXT, + const Parallel_2D& pcxt) + { + const int i1 = 1; + char transa = 'N'; + char transb = 'T'; + const std::complex alpha(1.0, 0.0); + const std::complex beta(0.0, 0.0); + pzgemm_(&transa, &transb, &naos, &nvirt, &nocc, + &alpha, C, &i1, &i1, pc.desc, + X, &i1, &i1, px.desc, + &beta, CoXT, &i1, &i1, pcxt.desc); + } + + inline void AAT( + const double* A, + const Parallel_2D& pa, + const int& nrow, + const int& ncol, + double* C, + const Parallel_2D& pc, + const bool add = false, + const double& alpha = 1.0) + { + const int i1 = 1; + char transa = 'N'; + char transb = 'T'; + const double beta = add ? 1.0 : 0.0; + pdgemm_(&transa, &transb, &nrow, &nrow, &ncol, + &alpha, A, &i1, &i1, pa.desc, + A, &i1, &i1, pa.desc, + &beta, C, &i1, &i1, pc.desc); + } + inline void AAT( + const std::complex* A, + const Parallel_2D& pa, + const int& nrow, + const int& ncol, + std::complex* C, + const Parallel_2D& pc, + const bool add = false, + const std::complex& alpha = std::complex(1.0, 0.0)) + { + const int i1 = 1; + char transa = 'N'; + char transb = 'T'; + const std::complex beta(add ? 1.0 : 0.0, 0.0); + // take conjugate of A + std::vector> A_conj(pa.get_local_size()); + for (int i = 0;i < pa.get_local_size();++i) A_conj[i] = std::conj(A[i]); + pzgemm_(&transa, &transb, &nrow, &nrow, &ncol, + &alpha, A_conj.data(), &i1, &i1, pa.desc, + A, &i1, &i1, pa.desc, + &beta, C, &i1, &i1, pc.desc); + } + + //output: col first, consistent with blas + // c: nao*nbands in para2d, nbands*nao in psi (row-para and constructed: nao) + // X: nvirt*nocc in para2d, nocc*nvirt in psi (row-para and constructed: nvirt) + template + std::vector cal_dm_diff_pblas( + const T* const X_istate, + const Parallel_2D& px, + const psi::Psi& c, + const Parallel_2D& pc, + const int& naos, + const int& nocc, + const int& nvirt, + const Parallel_2D& pmat, + const bool renorm_k, + const int nspin) + { + ModuleBase::TITLE("hamilt_lrtd", "cal_dm_diff_pblas"); + assert(px.comm() == pc.comm() && px.comm() == pmat.comm()); + assert(px.blacs_ctxt == pc.blacs_ctxt && px.blacs_ctxt == pmat.blacs_ctxt); + const int nks = c.get_nk(); + const int nk = nks / nspin; + + Parallel_2D pcx; + LR_Util::setup_2d_division(pcx, px.get_block_size(), naos, nocc, px.blacs_ctxt); + ct::Tensor cvx(ct::DataTypeToEnum::value, DEV::CpuDevice, { pcx.get_col_size(), pcx.get_row_size() }); + Parallel_2D pcxt; + LR_Util::setup_2d_division(pcxt, px.get_block_size(), naos, nvirt, px.blacs_ctxt); + ct::Tensor coxt(ct::DataTypeToEnum::value, DEV::CpuDevice, { pcxt.get_col_size(), pcxt.get_row_size() }); + std::vector dm_diff(nks, ct::Tensor(ct::DataTypeToEnum::value, DEV::CpuDevice, { pmat.get_col_size(), pmat.get_row_size() })); + for (int iks = 0;iks < nks;++iks) + { + c.fix_k(iks); + const int start = iks * px.get_local_size(); + // 1. C_virt * X + CvX(c.get_pointer(), pc, X_istate + start, px, naos, nocc, nvirt, cvx.data(), pcx); + // 2. C_occ * X^T + CoXT(c.get_pointer(), pc, X_istate + start, px, naos, nocc, nvirt, coxt.data(), pcxt); + // print_colfirst(c.get_pointer(), "c_pblas", naos, nocc + nvirt); + // print_colfirst(X_istate.get_pointer(), "X_pblas", nvirt, nocc); + // print_colfirst(cvx.data(), "cvx_pblas", naos, nocc); + // print_colfirst(coxt.data(), "coxt_pblas", naos, nvirt); + // 3. cvx*cvx^T - coxt*coxt^T + AAT(cvx.data(), pcx, naos, nocc, dm_diff[iks].data(), pmat, false, renorm_k ? (T)(1.0 / (double)nk) : (T)1.0); + // print_colfirst(dm_diff[iks].data(), "dm_diff_1_pblas", naos, naos); + AAT(coxt.data(), pcxt, naos, nvirt, dm_diff[iks].data(), pmat, true, renorm_k ? (T)(-1.0 / (double)nk) : (T)(-1.0)); + // print_colfirst(dm_diff[iks].data(), "dm_diff_2_pblas", naos, naos); + } + return dm_diff; + } +} +#endif + +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_DM_TRANS_DM_DIFF_PAR_HPP diff --git a/source/source_lcao/module_lr/dm_trans/dm_diff_ser.hpp b/source/source_lcao/module_lr/dm_trans/dm_diff_ser.hpp new file mode 100644 index 00000000000..f4c6cbc198a --- /dev/null +++ b/source/source_lcao/module_lr/dm_trans/dm_diff_ser.hpp @@ -0,0 +1,207 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_DM_TRANS_DM_DIFF_SER_HPP +#define ABACUS_SOURCE_LCAO_MODULE_LR_DM_TRANS_DM_DIFF_SER_HPP +#include "source_base/module_container/ATen/core/tensor_types.h" +#include "source_base/module_external/blas_connector.h" +#include "source_base/tool_title.h" +#include "source_lcao/module_lr/utils/lr_util.h" +#include "dm_diff.h" +namespace LR +{ + template + inline void print_colfirst(const T* ptr, const std::string& name, const int& nrow, const int& ncol) + { + std::cout << name << std::endl; + for (int i = 0;i < nrow;++i) + { + for (int j = 0;j < ncol;++j) + std::cout << ptr[j * nrow + i] << " "; + std::cout << std::endl; + } + } + inline void CvX( + const double* C, + const double* X, + const int& naos, + const int& nocc, + const int& nvirt, + double* CvX) + { + char transa = 'N'; + char transb = 'N'; + const double alpha = 1.0; + const double beta = 0; + dgemm_(&transa, &transb, &naos, &nocc, &nvirt, + &alpha, C + nocc * naos, &naos, + X, &nvirt, + &beta, CvX, &naos); + } + inline void CvX( + const std::complex* C, + const std::complex* X, + const int& naos, + const int& nocc, + const int& nvirt, + std::complex* CvX) + { + char transa = 'N'; + char transb = 'N'; + const std::complex alpha(1.0, 0.0); + const std::complex beta(0.0, 0.0); + zgemm_(&transa, &transb, &naos, &nocc, &nvirt, + &alpha, C + nocc * naos, &naos, + X, &nvirt, + &beta, CvX, &naos); + } + + inline void CoXT( + const double* C, + const double* X, + const int& naos, + const int& nocc, + const int& nvirt, + double* CoXT) + { + char transa = 'N'; + char transb = 'T'; + const double alpha = 1.0; + const double beta = 0; + dgemm_(&transa, &transb, &naos, &nvirt, &nocc, + &alpha, C, &naos, + X, &nvirt, + &beta, CoXT, &naos); + } + inline void CoXT( + const std::complex* C, + const std::complex* X, + const int& naos, + const int& nocc, + const int& nvirt, + std::complex* CoXT) + { + char transa = 'N'; + char transb = 'T'; + const std::complex alpha(1.0, 0.0); + const std::complex beta(0.0, 0.0); + zgemm_(&transa, &transb, &naos, &nvirt, &nocc, + &alpha, C, &naos, + X, &nvirt, + &beta, CoXT, &naos); + } + + inline void AAT( + const double* A, + const int& nrow, + const int& ncol, + double* C, + const bool add = false, + const double& alpha = 1.0) + { + char transa = 'N'; + char transb = 'T'; + const double beta = add ? 1.0 : 0.0; + dgemm_(&transa, &transb, &nrow, &nrow, &ncol, + &alpha, A, &nrow, + A, &nrow, + &beta, C, &nrow); + } + inline void AAT( + const std::complex* A, + const int& nrow, + const int& ncol, + std::complex* C, + const bool add = false, + const std::complex& alpha = std::complex(1.0, 0.0)) + { + char transa = 'N'; + char transb = 'T'; + const std::complex beta(add ? 1.0 : 0.0, 0.0); + // take conjugate of A + std::vector> A_conj(nrow * ncol); + for (int i = 0;i < A_conj.size();++i) A_conj[i] = std::conj(A[i]); + zgemm_(&transa, &transb, &nrow, &nrow, &ncol, + &alpha, A_conj.data(), &nrow, + A, &nrow, + &beta, C, &nrow); + } + + //output: col first, consistent with blas + // c: nao*nbands in para2d, nbands*nao in psi (row-para and constructed: nao) + // X: nvirt*nocc in para2d, nocc*nvirt in psi (row-para and constructed: nvirt) + template + std::vector cal_dm_diff_blas( + const T* const X_istate, + const psi::Psi& c, + const int& naos, + const int& nocc, + const int& nvirt, + const bool renorm_k, + const int nspin) + { + ModuleBase::TITLE("hamilt_lrtd", "cal_dm_diff_blas"); + const int nks = c.get_nk(); + const int nk = nks / nspin; + + ct::Tensor cvx(ct::DataTypeToEnum::value, DEV::CpuDevice, { nocc, naos }); + ct::Tensor coxt(ct::DataTypeToEnum::value, DEV::CpuDevice, { nvirt, naos }); + std::vector dm_diff(nks, ct::Tensor(ct::DataTypeToEnum::value, DEV::CpuDevice, { naos, naos })); + for (int iks = 0;iks < nks;++iks) + { + c.fix_k(iks); + const int start = iks * nocc * nvirt; + // 1. C_virt * X + CvX(c.get_pointer(), X_istate + start, naos, nocc, nvirt, cvx.data()); + // 2. C_occ * X^T + CoXT(c.get_pointer(), X_istate + start, naos, nocc, nvirt, coxt.data()); + // 3. cvx*cvx^T + coxt*coxt^T + AAT(cvx.data(), naos, nocc, dm_diff[iks].data(), false, renorm_k ? (T)(1.0 / (double)nk) : (T)1.0); + AAT(coxt.data(), naos, nvirt, dm_diff[iks].data(), true, renorm_k ? (T)(-1.0 / (double)nk) : (T)(-1.0)); + } + return dm_diff; + } + + inline double get_conj(const double& x) { return x; } + inline std::complex get_conj(const std::complex& x) { return std::conj(x); } + + template + std::vector cal_dm_diff_forloop( + const T* const X_istate, + const psi::Psi& c, + const int& naos, + const int& nocc, + const int& nvirt, + const bool renorm_k, + const int nspin) + { + ModuleBase::TITLE("hamilt_lrtd", "cal_dm_diff_forloop"); + const int nks = c.get_nk(); + const int nk = nks / nspin; + + std::vector dm_diff(nks, ct::Tensor(ct::DataTypeToEnum::value, DEV::CpuDevice, { naos, naos })); + for (int iks = 0;iks < nks;++iks) + { + dm_diff[iks].zero(); + c.fix_k(iks); + const int start = iks * nocc * nvirt; + for (int nu = 0;nu < naos;++nu)//col + for (int mu = 0;mu < naos;++mu)//row + { + for (int i = 0;i < nocc;++i) + for (int a = 0;a < nvirt;++a) + { + for (int b = 0;b < nvirt;++b) + dm_diff[iks].data()[nu * naos + mu] + += get_conj(c.get_pointer()[(nocc + a) * naos + mu] * X_istate[start + i * nvirt + a]) + * c.get_pointer()[(nocc + b) * naos + nu] * X_istate[start + i * nvirt + b]; + for (int j = 0;j < nocc;++j) + dm_diff[iks].data()[nu * naos + mu] + -= get_conj(c.get_pointer()[i * naos + mu] * X_istate[start + i * nvirt + a]) + * c.get_pointer()[j * naos + nu] * X_istate[start + j * nvirt + a]; + } + if (renorm_k) + dm_diff[iks].data()[nu * naos + mu] /= (double)nk; + } + } + return dm_diff; + } +} +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_DM_TRANS_DM_DIFF_SER_HPP diff --git a/source/source_lcao/module_lr/dm_trans/dm_trans_parallel.cpp b/source/source_lcao/module_lr/dm_trans/dm_trans_parallel.cpp index f8198afcafe..745f39b7550 100644 --- a/source/source_lcao/module_lr/dm_trans/dm_trans_parallel.cpp +++ b/source/source_lcao/module_lr/dm_trans/dm_trans_parallel.cpp @@ -104,7 +104,7 @@ std::vector cal_dm_trans_pblas(const std::complex* co // char transb = 'C'; // // 1. [X*C_occ^\dagger]^\dagger=C_occ*X^\dagger // Parallel_2D pXc; - // LR_Util::setup_2d_division(pXc, px.get_block_size(), naos, nvirt, px.comm_2D, px.blacs_ctxt); + // LR_Util::setup_2d_division(pXc, px.get_block_size(), naos, nvirt,px.blacs_ctxt); // container::Tensor Xc(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { pXc.get_col_size(), pXc.get_row_size() // });//row is "inside"(memory contiguity) for pblas Xc.zero(); const std::complex alpha(1.0, 0.0); // const std::complex beta(0.0, 0.0); diff --git a/source/source_lcao/module_lr/dm_trans/test/CMakeLists.txt b/source/source_lcao/module_lr/dm_trans/test/CMakeLists.txt index 6894e57ab2c..11507970b06 100644 --- a/source/source_lcao/module_lr/dm_trans/test/CMakeLists.txt +++ b/source/source_lcao/module_lr/dm_trans/test/CMakeLists.txt @@ -11,4 +11,10 @@ AddTest( # ../../../source_base/module_container/base/core/cpu_allocator.cpp # ../../../source_base/module_container/base/core/refcount.cpp # ../../../source_base/module_container/ATen/kernels/memory_impl.cpp +) + +AddTest( + TARGET MODULE_LR_dm_diff_test + LIBS psi base parameter ${math_libs} device container + SOURCES dm_diff_tst.cpp ../../utils/lr_util.cpp ) \ No newline at end of file diff --git a/source/source_lcao/module_lr/dm_trans/test/dm_diff_tst.cpp b/source/source_lcao/module_lr/dm_trans/test/dm_diff_tst.cpp new file mode 100644 index 00000000000..ec09cff8b31 --- /dev/null +++ b/source/source_lcao/module_lr/dm_trans/test/dm_diff_tst.cpp @@ -0,0 +1,246 @@ +#include +#include "mpi.h" +#include +#include "source_lcao/module_lr/utils/lr_util.h" +#include "../dm_diff.h" + +struct matsize +{ + int nks = 1; + int naos; + int nocc; + int nvirt; + int nb = 1; + matsize(int nks, int naos, int nocc, int nvirt, int nb = 1) + :nks(nks), naos(naos), nocc(nocc), nvirt(nvirt), nb(nb) { + assert(nocc + nvirt <= naos); + }; +}; + +class DMDiffTest : public testing::Test +{ +public: + std::vector sizes{ + // {1, 3, 2, 1}, + { 2, 14, 9, 4 }, + {2, 20, 10, 7} + }; + int nstate = 2; + std::ofstream ofs_running; + int my_rank; +#ifdef __MPI + void SetUp() override + { + MPI_Comm_rank(MPI_COMM_WORLD, &my_rank); + this->ofs_running.open("log" + std::to_string(my_rank) + ".txt"); + ofs_running << "my_rank = " << my_rank << std::endl; + } + void TearDown() override + { + ofs_running.close(); + } +#endif + + void set_ones(double* data, int size) { for (int i = 0;i < size;++i) data[i] = 1.0; }; + void set_int(double* data, int size) { for (int i = 0;i < size;++i) data[i] = static_cast(i + 1); }; + void set_int(std::complex* data, int size) { for (int i = 0;i < size;++i) data[i] = std::complex(i + 1, -i - 1); }; + void set_rand(double* data, int size) { for (int i = 0;i < size;++i) data[i] = double(rand()) / double(RAND_MAX) * 10.0 - 5.0; }; + void set_rand(std::complex* data, int size) { for (int i = 0;i < size;++i) data[i] = std::complex(rand(), rand()) / double(RAND_MAX) * 10.0 - 5.0; }; + void check_eq(double* data1, double* data2, int size) { for (int i = 0;i < size;++i) EXPECT_NEAR(data1[i], data2[i], 1e-8); }; + void check_eq(std::complex* data1, std::complex* data2, int size) + { + for (int i = 0;i < size;++i) + { + EXPECT_NEAR(data1[i].real(), data2[i].real(), 1e-8); + EXPECT_NEAR(data1[i].imag(), data2[i].imag(), 1e-8); + } + }; +}; + +TEST_F(DMDiffTest, DoubleSerial) +{ + for (auto s : this->sizes) + { + psi::Psi X(s.nks, nstate, s.nocc * s.nvirt, {}, false); + set_rand(X.get_pointer(), nstate * s.nks * s.nocc * s.nvirt); + for (int istate = 0;istate < nstate;++istate) + { + int size_c = s.nks * (s.nocc + s.nvirt) * s.naos; + psi::Psi c(s.nks, s.nocc + s.nvirt, s.naos, {}, true); + set_rand(c.get_pointer(), size_c); + X.fix_b(istate); + const std::vector& dm_for = LR::cal_dm_diff_forloop(X.get_pointer(), c, s.naos, s.nocc, s.nvirt); + const std::vector& dm_blas = LR::cal_dm_diff_blas(X.get_pointer(), c, s.naos, s.nocc, s.nvirt); + for (int isk = 0;isk < s.nks;++isk) check_eq(dm_for[isk].data(), dm_blas[isk].data(), s.naos * s.naos); + } + + } +} +TEST_F(DMDiffTest, ComplexSerial) +{ + for (auto s : this->sizes) + { + psi::Psi, base_device::DEVICE_CPU> X(s.nks, nstate, s.nocc * s.nvirt, {}, false); + set_rand(X.get_pointer(), nstate * s.nks * s.nocc * s.nvirt); + for (int istate = 0;istate < nstate;++istate) + { + int size_c = s.nks * (s.nocc + s.nvirt) * s.naos; + psi::Psi, base_device::DEVICE_CPU> c(s.nks, s.nocc + s.nvirt, s.naos, {}, true); + set_rand(c.get_pointer(), size_c); + X.fix_b(istate); + const std::vector& dm_for = LR::cal_dm_diff_forloop(X.get_pointer(), c, s.naos, s.nocc, s.nvirt); + const std::vector& dm_blas = LR::cal_dm_diff_blas(X.get_pointer(), c, s.naos, s.nocc, s.nvirt); + for (int isk = 0;isk < s.nks;++isk) check_eq(dm_for[isk].data>(), dm_blas[isk].data>(), s.naos * s.naos); + } + } +} + +#ifdef __MPI +TEST_F(DMDiffTest, DoubleParallel) +{ + for (auto s : this->sizes) + { + // c: nao*nbands in para2d, nbands*nao in psi (row-para and constructed: nao) + // X: nvirt*nocc in para2d, nocc*nvirt in psi (row-para and constructed: nvirt) + Parallel_2D px; + LR_Util::setup_2d_division(px, s.nb, s.nvirt, s.nocc); + psi::Psi X(s.nks, nstate, px.get_local_size(), {}, false); + Parallel_2D pc; + const int nmo = s.nocc + s.nvirt; + LR_Util::setup_2d_division(pc, s.nb, s.naos, nmo, px.blacs_ctxt); + psi::Psi c(s.nks, pc.get_col_size(), pc.get_row_size(), {}, true); + Parallel_2D pmat; + LR_Util::setup_2d_division(pmat, s.nb, s.naos, s.naos, px.blacs_ctxt); + + EXPECT_EQ(px.dim0, pc.dim0); + EXPECT_EQ(px.dim1, pc.dim1); + EXPECT_GE(s.nvirt, px.dim0); + EXPECT_GE(s.nocc, px.dim1); + EXPECT_GE(s.naos, pc.dim0); + + set_rand(X.get_pointer(), nstate * s.nks * px.get_local_size()); //set X and X_full + psi::Psi X_full(s.nks, nstate, s.nocc * s.nvirt, {}, false); // allocate X_full + X_full.zero_out(); + for (int istate = 0;istate < nstate;++istate) + { + X.fix_b(istate); + X_full.fix_b(istate); + for (int isk = 0;isk < s.nks;++isk) + { + X.fix_k(isk); + X_full.fix_k(isk); + LR_Util::gather_2d_to_full(px, X.get_pointer(), X_full.get_pointer(), false, s.nvirt, s.nocc); + } + } + for (int istate = 0;istate < nstate;++istate) + { + c.fix_k(0); + set_rand(c.get_pointer(), s.nks * pc.get_local_size()); // set c + + X.fix_b(istate); + X_full.fix_b(istate); + + std::vector dm_pblas_loc = LR::cal_dm_diff_pblas(X.get_pointer(), px, c, pc, s.naos, s.nocc, s.nvirt, pmat); + + // gather dm and output + std::vector dm_gather(s.nks, container::Tensor(DAT::DT_DOUBLE, DEV::CpuDevice, { s.naos, s.naos })); + for (int isk = 0;isk < s.nks;++isk) + { + dm_gather[isk].zero(); + LR_Util::gather_2d_to_full(pmat, dm_pblas_loc[isk].data(), dm_gather[isk].data(), false, s.naos, s.naos); + } + + // compare to global matrix + psi::Psi c_full(s.nks, s.nocc + s.nvirt, s.naos, {}, true); + c_full.zero_out(); + for (int isk = 0;isk < s.nks;++isk) + { + c.fix_k(isk); + c_full.fix_k(isk); + LR_Util::gather_2d_to_full(pc, c.get_pointer(), c_full.get_pointer(), false, s.naos, nmo); + } + if (my_rank == 0) + { + const std::vector& dm_full = LR::cal_dm_diff_blas(X_full.get_pointer(), c_full, s.naos, s.nocc, s.nvirt); + for (int isk = 0;isk < s.nks;++isk) check_eq(dm_full[isk].data(), dm_gather[isk].data(), s.naos * s.naos); + } + } + } +} +TEST_F(DMDiffTest, ComplexParallel) +{ + for (auto s : this->sizes) + { + // c: nao*nbands in para2d, nbands*nao in psi (row-para and constructed: nao) + // X: nvirt*nocc in para2d, nocc*nvirt in psi (row-para and constructed: nvirt) + Parallel_2D px; + LR_Util::setup_2d_division(px, s.nb, s.nvirt, s.nocc); + psi::Psi, base_device::DEVICE_CPU> X(s.nks, nstate, px.get_local_size(), {}, false); + Parallel_2D pc; + const int nmo = s.nocc + s.nvirt; + LR_Util::setup_2d_division(pc, s.nb, s.naos, nmo, px.blacs_ctxt); + psi::Psi, base_device::DEVICE_CPU> c(s.nks, pc.get_col_size(), pc.get_row_size(), {}, true); + Parallel_2D pmat; + LR_Util::setup_2d_division(pmat, s.nb, s.naos, s.naos, px.blacs_ctxt); + + set_rand(X.get_pointer(), nstate * s.nks * px.get_local_size()); //set X and X_full + psi::Psi, base_device::DEVICE_CPU> X_full(s.nks, nstate, s.nocc * s.nvirt, {}, false); // allocate X_full + X_full.zero_out(); + for (int istate = 0;istate < nstate;++istate) + { + X.fix_b(istate); + X_full.fix_b(istate); + for (int isk = 0;isk < s.nks;++isk) + { + X.fix_k(isk); + X_full.fix_k(isk); + LR_Util::gather_2d_to_full(px, X.get_pointer(), X_full.get_pointer(), false, s.nvirt, s.nocc); + } + } + for (int istate = 0;istate < nstate;++istate) + { + c.fix_k(0); + set_rand(c.get_pointer(), s.nks * pc.get_local_size()); // set c + + X.fix_b(istate); + X_full.fix_b(istate); + + std::vector dm_pblas_loc = LR::cal_dm_diff_pblas(X.get_pointer(), px, c, pc, s.naos, s.nocc, s.nvirt, pmat); + + // gather dm and output + std::vector dm_gather(s.nks, container::Tensor(DAT::DT_COMPLEX_DOUBLE, DEV::CpuDevice, { s.naos, s.naos })); + for (int isk = 0;isk < s.nks;++isk) + { + dm_gather[isk].zero(); + LR_Util::gather_2d_to_full(pmat, dm_pblas_loc[isk].data>(), dm_gather[isk].data>(), false, s.naos, s.naos); + } + + // compare to global matrix + psi::Psi, base_device::DEVICE_CPU> c_full(s.nks, s.nocc + s.nvirt, s.naos, {}, true); + c_full.zero_out(); + for (int isk = 0;isk < s.nks;++isk) + { + c.fix_k(isk); + c_full.fix_k(isk); + LR_Util::gather_2d_to_full(pc, c.get_pointer(), c_full.get_pointer(), false, s.naos, nmo); + } + if (my_rank == 0) + { + std::vector dm_full = LR::cal_dm_diff_blas(X_full.get_pointer(), c_full, s.naos, s.nocc, s.nvirt); + for (int isk = 0;isk < s.nks;++isk) check_eq(dm_full[isk].data>(), dm_gather[isk].data>(), s.naos * s.naos); + } + } + } +} +#endif + + +int main(int argc, char** argv) +{ + srand(time(NULL)); // for random number generator + MPI_Init(&argc, &argv); + testing::InitGoogleTest(&argc, argv); + int result = RUN_ALL_TESTS(); + MPI_Finalize(); + return result; +} \ No newline at end of file diff --git a/source/source_lcao/module_lr/exx_force_dm.h b/source/source_lcao/module_lr/exx_force_dm.h new file mode 100644 index 00000000000..d132f2c72d7 --- /dev/null +++ b/source/source_lcao/module_lr/exx_force_dm.h @@ -0,0 +1,56 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_EXX_FORCE_DM_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_EXX_FORCE_DM_H +#ifdef __EXX +#include + +#include +#include +#include +#include + +namespace LR +{ + /// @brief `RI::Exx` plus a `cal_force` overload taking two *different* density matrices. + /// + /// LibRI used to ship this as `RI::LR` (a subclass of `RI::Exx` in + /// `RI/physics/LR.h`). LibRI commit 5c6c262 repurposed the name `RI::LR` for an + /// unrelated k-space CVCX/Hartree helper built on `LRI_k`, so the piece the + /// LR-TDDFT analytical gradient needs is kept here instead. + /// + /// `RI::Exx::cal_force` contracts the 3-center derivative dH with whatever sits in + /// `post_2D.saves["Ds_"+suffix]`, while the density matrix handed to `set_Ds` is the + /// one consumed inside the loop-3 contraction. Overwriting the former therefore lets + /// the two sides of Tr[D_IJ * dH[D_KL]] differ, which is what the + /// Pulay / Hellmann-Feynman split of the EXX gradient requires. + template + class ExxForceTwoDM : public RI::Exx + { + using Base = RI::Exx; + + public: + using TC = std::array; + using TAC = std::pair; + + ExxForceTwoDM() = default; + /// take over an existing Exx kernel (Cs/Vs/dCs/dVs already set up) + ExxForceTwoDM(Base&& exx) : Base(std::move(exx)) {} + + using Base::cal_force; + + /// @param Ds_left the D_IJ contracted with dH; if empty, behaves as `Exx::cal_force` + /// @param save_names_suffix "Cs", "Vs", "Ds", "dCs", "dVs" + void cal_force(const std::map>>& Ds_left, + const std::array& save_names_suffix = { "","","","","" }) + { + if (!Ds_left.empty()) + { + this->post_2D.saves["Ds_" + save_names_suffix[2]] + = this->post_2D.set_tensors_map2(Ds_left); + } + this->cal_force(save_names_suffix); + } + }; +} +#endif + +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_EXX_FORCE_DM_H diff --git a/source/source_lcao/module_lr/exx_proj.cpp b/source/source_lcao/module_lr/exx_proj.cpp new file mode 100644 index 00000000000..c790d7ebffc --- /dev/null +++ b/source/source_lcao/module_lr/exx_proj.cpp @@ -0,0 +1,46 @@ +#include "exx_proj.h" + +#include "source_base/module_external/blas_connector.h" + +namespace LR +{ +void project_exx(const double* h, + const double* left, + const double* right, + const int naos, + const int nleft, + const int nright, + const double factor, + double* scratch, + double* result) +{ + if (nleft == 0 || nright == 0) { return; } + const double one = 1.0; + const double zero = 0.0; + // Transpose, rather than symmetrize: K[D^X] need not be symmetric. + BlasConnector::gemm_cm('T', 'N', naos, nleft, naos, one, + h, naos, left, naos, zero, scratch, naos); + BlasConnector::gemm_cm('T', 'N', nright, nleft, naos, factor, + right, naos, scratch, naos, one, result, nright); +} + +void project_exx(const std::complex* h, + const std::complex* left, + const std::complex* right, + const int naos, + const int nleft, + const int nright, + const double factor, + std::complex* scratch, + std::complex* result) +{ + if (nleft == 0 || nright == 0) { return; } + const std::complex one(1.0, 0.0); + const std::complex zero(0.0, 0.0); + const std::complex scale(factor, 0.0); + BlasConnector::gemm_cm('N', 'N', naos, nleft, naos, one, + h, naos, left, naos, zero, scratch, naos); + BlasConnector::gemm_cm('C', 'N', nright, nleft, naos, scale, + right, naos, scratch, naos, one, result, nright); +} +} diff --git a/source/source_lcao/module_lr/exx_proj.h b/source/source_lcao/module_lr/exx_proj.h new file mode 100644 index 00000000000..9ddc72d3416 --- /dev/null +++ b/source/source_lcao/module_lr/exx_proj.h @@ -0,0 +1,35 @@ +#ifndef ABACUS_LR_EXX_PROJECTION_H +#define ABACUS_LR_EXX_PROJECTION_H + +#include + +namespace LR +{ +/// Project a real, generally nonsymmetric AO matrix against two sets of columns. +/// All matrices are column major. result(right,left) += factor * left^T H right. +/// The caller owns the naos*nleft scratch buffer; inputs and outputs cannot alias. +void project_exx(const double* h, + const double* left, + const double* right, + int naos, + int nleft, + int nright, + double factor, + double* scratch, + double* result); + +/// Complex Bloch projection: result(right,left) += factor * right^H H(k) left. +/// H(k) must contain the positive Bloch phase, conjugate to the probe density. +/// All matrices are column major; scratch has naos*nleft elements. +void project_exx(const std::complex* h, + const std::complex* left, + const std::complex* right, + int naos, + int nleft, + int nright, + double factor, + std::complex* scratch, + std::complex* result); +} + +#endif diff --git a/source/source_lcao/module_lr/force_funcs.h b/source/source_lcao/module_lr/force_funcs.h new file mode 100644 index 00000000000..64be8d0eac0 --- /dev/null +++ b/source/source_lcao/module_lr/force_funcs.h @@ -0,0 +1,140 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_FORCE_FUNCS_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_FORCE_FUNCS_H +#include "source_pw/module_pwdft/force_pw.h" +#include "source_lcao/module_operator_lcao/nonlocal.h" +#include "source_estate/module_dm/density_matrix.h" +#include "source_basis/module_nao/two_center_bundle.h" +#include "source_io/module_output/output_log.h" + +// This file extracts out some terms from the ground state force code +// to calculate the excited state force with a different density matrix. +template +class ForcePWTerms +{ +public: + ModuleBase::matrix operator()(const UnitCell& ucell, + const Charge& chr, + const ModulePW::PW_Basis& rhopw, + const pseudopot_cell_vl& locpp, + const Structure_Factor& sf, + const int nspin, + const bool test_force, + std::ofstream& ofs_running, + const bool with_ewald = true, + const elecstate::ElecState* pelec = nullptr) + { + if (nspin == 4) { throw std::runtime_error("ForcePWTerms: nspin=4 is not supported."); } + ModuleBase::TITLE("Force_Stress_LCAO", "cal_force_pw"); + Forces f_pw(ucell.nat); + ModuleBase::matrix fvl_dvl(ucell.nat, 3), fewalds(ucell.nat, 3), fcc(ucell.nat, 3), fscc(ucell.nat, 3); + //-------------------------------------------------------- + // local pseudopotential force: + // use charge density; plane wave; local pseudopotential; + //-------------------------------------------------------- + f_pw.cal_force_loc(ucell, fvl_dvl, &rhopw, locpp.vloc, &chr); + //-------------------------------------------------------- + // ewald force: use plane wave only. + //-------------------------------------------------------- + if (with_ewald) + { + f_pw.cal_force_ew(ucell, fewalds, &rhopw, &sf); + } + //-------------------------------------------------------- + // force due to core correlation. + //-------------------------------------------------------- + UnitCell& ucell_noconst = const_cast(ucell); + // domag/domag_z/gga_grad only matter for the noncollinear (nspin=4) case, which the + // throw-guard above already excludes, so they are passed as fixed, inert defaults + // here instead of reading the corresponding global flags. + f_pw.cal_force_cc(fcc, &rhopw, &chr, locpp.numeric, ucell_noconst, + nspin, /*domag=*/false, /*domag_z=*/false, /*gga_grad=*/0); + //-------------------------------------------------------- + // force due to self-consistent charge (invalid in from-scratch LR case) + //-------------------------------------------------------- + if (pelec) + { + f_pw.cal_force_scc(fscc, &rhopw, pelec->vnew, pelec->vnew_exist, locpp.numeric, ucell); + } + if (test_force) + { + ModuleIO::print_force(ofs_running, ucell, "VL_dVL FORCE (eV/Angstrom)", fvl_dvl, false); + ModuleIO::print_force(ofs_running, ucell, "EWALD FORCE (eV/Angstrom)", fewalds, false); + ModuleIO::print_force(ofs_running, ucell, "NLCC FORCE (eV/Angstrom)", fcc, false); + ModuleIO::print_force(ofs_running, ucell, "SCC FORCE (eV/Angstrom)", fscc, false); + } + return fvl_dvl + fewalds + fcc + fscc; + } +}; + +// calculate force and stress for Nonlocal part +// nspin = 2 or 4 will not be used in the current excited state force calculation +// for nspin = 1 or 2 +template +ModuleBase::matrix cal_force_nonlocal( + const UnitCell& ucell, + const std::vector>& kvec_d, + const Grid_Driver& gd, + const TwoCenterBundle& two_center_bundle, + const module_dm::DensityMatrix& dm ) +{ + ModuleBase::TITLE("Force_Stress_LCAO", "cal_force_nonlocal_dvnl"); + std::vector orb_cutoffs(ucell.ntype); + for (int it = 0;it < ucell.ntype;++it) { orb_cutoffs[it] = ucell.atoms[it].Rcut; } + + hamilt::Nonlocal> tmp_nonlocal(nullptr, + kvec_d, + nullptr, + &ucell, + orb_cutoffs, + &gd, + two_center_bundle.overlap_orb_beta.get()); + const int nspin = dm.get_dmr_vec().size(); + if(nspin==2) + { + const_cast*>(&dm)->switch_dmr(1); //spin-up + spin-down + } + const hamilt::HContainer* dmr = dm.get_dmr_ptr(1); + ModuleBase::matrix fvnl(ucell.nat, 3); + ModuleBase::matrix svnl; // no use now, only for passing into interfaces + tmp_nonlocal.cal_force_stress(/*force*/true, /*stress*/false, dmr, fvnl, svnl); + if (nspin == 2) + { + const_cast*>(&dm)->switch_dmr(0); + } + return fvnl; +} +// calculate force and stress for Nonlocal part (Hellmann-Feynman term) +// for nspin = 4 +template +ModuleBase::matrix cal_force_nonlocal_dvnl( + UnitCell ucell, + const std::vector>& kvec_d, + const Grid_Driver& gd, + const TwoCenterBundle& two_center_bundle, + const module_dm::DensityMatrix>& dm) +{ + ModuleBase::TITLE("Force_Stress_LCAO", "cal_force_nonlocal_dvnl"); + + std::vector orb_cutoffs(ucell.ntype); + for (int it = 0;it < ucell.ntype;++it) { orb_cutoffs[it] = ucell.atoms[it].Rcut; } + + hamilt::Nonlocal, std::complex>> tmp_nonlocal( + nullptr, + kvec_d, + nullptr, + &ucell, + orb_cutoffs, + &gd, + two_center_bundle.overlap_orb_beta.get()); + + hamilt::HContainer> tmp_dmr(dm.get_dmr_ptr(1)->get_paraV()); + std::vector ijrs = dm.get_dmr_ptr(1)->get_ijr_info(); + tmp_dmr.insert_ijrs(&ijrs); + tmp_dmr.allocate(); + dm.cal_DMR_full(&tmp_dmr); + ModuleBase::matrix fvnl(ucell.nat, 3); + ModuleBase::matrix svnl; // no use now, only for passing into interfaces + tmp_nonlocal.cal_force_stress(/*force*/true, /*stress*/false, &tmp_dmr, fvnl, svnl); + return fvnl; +} +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_FORCE_FUNCS_H diff --git a/source/source_lcao/module_lr/grad_degen.cpp b/source/source_lcao/module_lr/grad_degen.cpp new file mode 100644 index 00000000000..31e7bf052dd --- /dev/null +++ b/source/source_lcao/module_lr/grad_degen.cpp @@ -0,0 +1,62 @@ +#include "grad_degen.h" + +#include +#include +#include +#include + +namespace LR +{ + std::vector> group_degenerate_states(const std::vector& omega, + const double thr) + { + const int nst = static_cast(omega.size()); + std::vector> groups; + if (nst == 0) { return groups; } + + // Casida output happens to come out ascending, but nothing here should depend on that. + std::vector order(nst); + std::iota(order.begin(), order.end(), 0); + std::stable_sort(order.begin(), order.end(), + [&omega](const int a, const int b) { return omega[a] < omega[b]; }); + + if (thr <= 0.0) + { + for (int i = 0; i < nst; ++i) { groups.push_back(std::vector(1, order[i])); } + return groups; + } + + double anchor = omega[order[0]]; + groups.push_back(std::vector(1, order[0])); + for (int i = 1; i < nst; ++i) + { + const int ist = order[i]; + // measured against the group's first (lowest) member, so the spread inside a group is + // bounded by `thr` however many members it collects + if (omega[ist] - anchor < thr) + { + groups.back().push_back(ist); + } + else + { + groups.push_back(std::vector(1, ist)); + anchor = omega[ist]; + } + } + for (std::vector& g : groups) { std::sort(g.begin(), g.end()); } + return groups; + } + + std::vector> degenerate_pairs(const int d) + { + std::vector> pairs; + if (d < 2) { return pairs; } + pairs.reserve(static_cast(d) * (d - 1) / 2); + for (int k = 0; k < d; ++k) + { + for (int l = k + 1; l < d; ++l) { pairs.push_back(std::make_pair(k, l)); } + } + return pairs; + } + +} diff --git a/source/source_lcao/module_lr/grad_degen.h b/source/source_lcao/module_lr/grad_degen.h new file mode 100644 index 00000000000..3dba8634051 --- /dev/null +++ b/source/source_lcao/module_lr/grad_degen.h @@ -0,0 +1,189 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_GRAD_DEGEN_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_GRAD_DEGEN_H +#include +#include +#include + +/// @file +/// Algebra of the degenerate-subspace gradient matrix +/// $G^{(A\alpha)}_{kl}=\langle X_k|\partial A/\partial R_{A\alpha}|X_l\rangle$. +/// +/// At a $d$-fold degeneracy no single state has a gradient vector: the branch slopes along a +/// displacement $u$ are the eigenvalues of $M(u)=\sum_{A\alpha}u_{A\alpha}G^{(A\alpha)}$, whose +/// eigenvectors depend on $u$. The full first-order object is $3N$ matrices of size $d\times d$, +/// and only $\operatorname{Tr}M$ is basis-invariant. +/// +/// The per-state analytic gradient implemented in `esolver_lr_grad.cpp` is, restricted to one +/// multiplet, a *quadratic form* in the excitation vector: every ingredient ($D^X$, $T$, $R$, $Z$, +/// $D^Z$, $W^X$, $\Lambda$) is built from $X\otimes X$, the CPSCF operator $L$ depends only on the +/// ground state, and $\Omega$ is a constant inside the multiplet. Write it $\mathcal F[X]$. Because +/// every $X$ in the degenerate subspace is itself a legitimate eigenvector with the same $\Omega$, +/// $\mathcal F[X]=G[X,X]$ holds on the whole diagonal -- and a symmetric bilinear form is +/// determined by its diagonal. So the existing code already contains all of $G$; it has only ever +/// been evaluated on the basis the diagonalizer happened to return. +/// +/// Recovering the off-diagonal elements is then the polarization identity. With +/// $X_\pm=(X_k\pm X_l)/\sqrt2$ and $\mathcal F[\alpha X]=\alpha^2\mathcal F[X]$, +/// (A1) $G_{kl}=\tfrac12(\mathcal F[X_+]-\mathcal F[X_-])$, +/// (A2) $G_{kl}=\mathcal F[X_+]-\tfrac12(G_{kk}+G_{ll})$, +/// and (A2) reuses the $d$ diagonal gradients the code already computes, so the whole matrix costs +/// $d(d+1)/2$ evaluations -- exactly its number of independent components. +/// +/// The construction admits two checks that need no finite differences and are worth running on +/// any new case: $\mathcal F[2X]=4\mathcal F[X]$ (it is a quadratic form at all), and +/// covariance under a rotation of the subspace basis, $\mathcal F[X'_k]=(U^\top GU)_{kk}$ with +/// $X'_k=\sum_lU_{lk}X_l$. Both are exercised in `test/test_grad_degen.cpp`. +/// +/// The alternative -- deriving an explicitly bilinear Z-vector right-hand side that takes two +/// different $X$ -- yields the same $G$, but it has to re-derive every factor and hand-polarize +/// the $g^{xc}$ potential, so it is the more expensive and more error-prone route. +/// +/// This header holds only the basis-independent bookkeeping, so that it is unit-testable without a +/// ground state: grouping states into multiplets, enumerating the pairs, forming the normalized +/// combination, and assembling $G$. Everything needing the parallel layout or the Z-vector solver +/// stays in `esolver_lr_grad.cpp`. + +namespace LR +{ + /// @brief Split states into multiplets of (near-)degenerate excitation energies. + /// + /// A state joins the group it is within `thr` of *the group's first member*, not of its + /// predecessor: chaining on consecutive gaps would let a run of small steps span a spread far + /// larger than `thr`, and "degenerate" has to mean a bounded total spread. + /// + /// The threshold alone does NOT decide whether route (A2) applies: an accidental near-degeneracy + /// falls inside any loose threshold, yet there $\Omega$ differs between the states, so they are + /// not one quadratic form and the combinations $X_\pm$ are not eigenvectors. The discriminator + /// is whether both the analytic and the finite-difference side split by the same amount (see + /// section 2(B) of the document above); this function only proposes the candidates. + /// + /// @param omega excitation energies, any order (Ry) + /// @param thr maximum spread inside one multiplet (Ry); <= 0 puts every state alone + /// @return groups of indices into `omega`, each group ascending, groups ordered by their + /// lowest energy + std::vector> group_degenerate_states(const std::vector& omega, + const double thr); + + /// @brief The $(k,l)$, $k> degenerate_pairs(const int d); + + /// @brief $X_+=(X_k+X_l)/\sqrt2$, normalized when $X_k$ and $X_l$ are orthonormal. + /// + /// Using the normalized combination rather than $X_k+X_l$ is what lets the gradient be + /// evaluated by the untouched per-state path: $X_+$ is then a genuine normalized eigenvector, + /// so every normalization, $\Omega$ and $W^X$ convention inside that path still holds. + template + void combine_normalized(const T* const Xk, const T* const Xl, const size_t nloc, T* const Xplus) + { + // 1/sqrt(2) as a literal: `std::sqrt` is not constexpr under the C++11 baseline. + const T inv_sqrt2 = static_cast(0.70710678118654752440); + for (size_t i = 0; i < nloc; ++i) { Xplus[i] = inv_sqrt2 * (Xk[i] + Xl[i]); } + } + + /// @brief Assemble $G_{kl}$ of one multiplet from the diagonal gradients and the combinations, + /// i.e. route (A2) above. + /// + /// @param diag $G_{kk}=\mathcal F[X_k]$, `d` entries + /// @param plus $\mathcal F[(X_k+X_l)/\sqrt2]$, one per entry of `pairs`, in that order + /// @param pairs as returned by `degenerate_pairs(d)` + /// @return G[k][l], symmetric by construction + /// + /// `TMat` needs `operator+`, `operator-` and `operator*(double)`; `ModuleBase::matrix` and + /// plain `double` both qualify, which is what keeps this testable. + template + std::vector> assemble_grad_matrix(const std::vector& diag, + const std::vector& plus, + const std::vector>& pairs) + { + const int d = static_cast(diag.size()); + std::vector> g(d, std::vector(d)); + for (int k = 0; k < d; ++k) { g[k][k] = diag[k]; } + for (size_t ip = 0; ip < pairs.size(); ++ip) + { + const int k = pairs[ip].first; + const int l = pairs[ip].second; + // $G_{kl}=\mathcal F[X_+]-\tfrac12(G_{kk}+G_{ll})$ + const TMat off = plus[ip] - (diag[k] + diag[l]) * 0.5; + g[k][l] = off; + g[l][k] = off; + } + return g; + } + + /// @brief The outcome of the Jahn-Teller search: the displacement direction that lowers one + /// branch of the multiplet fastest, and the electronic state that follows it. + struct JTDirection + { + std::vector displacement; ///< the $3N$ direction, normalized, laid out (atom, xyz) + std::vector mixing; ///< $v$, that branch's $d$ coefficients in the multiplet basis + double slope = 0.0; ///< $\|q(v)\|$, the branch's steepest descent rate + int restarts_agreeing = 0; ///< how many starting points reached this same optimum + int iterations = 0; ///< iterations used by the winning start + }; + + /// @brief Find the Jahn-Teller direction: the displacement that splits the multiplet and lowers + /// one branch as fast as possible. + /// + /// The Jahn-Teller theorem says a degenerate electronic state of a non-linear molecule is + /// unstable against some symmetry-lowering displacement. Finding it is a JOINT optimization over + /// the displacement and the mixing inside the subspace -- the two are determined together, which + /// is why it cannot be done one Cartesian axis at a time: + /// + /// $\min_{\|u\|=1}\lambda_{\min}\big(\sum_a u_a G^{(a)}\big)$. + /// + /// Written that way it looks like a non-convex problem on the unit sphere in $3N$ dimensions. It + /// is not: since $\lambda_{\min}(M)=\min_{\|v\|=1}v^\top Mv$, the two minimizations can be + /// swapped, and the inner one over $u$ has a closed form, + /// + /// $\min_{\|u\|=1}\ u\cdot q(v)=-\|q(v)\|$, where $q(v)_a=v^\top G^{(a)}v$, + /// + /// leaving + /// + /// $\min_{\|u\|=1}\lambda_{\min}\big(M(u)\big)=-\max_{\|v\|=1}\|q(v)\|$, + /// with the optimal direction $u^*=-q(v^*)/\|q(v^*)\|$. + /// + /// So the search runs over the $d$-dimensional subspace, not over $3N$ coordinates, and it has a + /// plain physical reading: $q(v)$ is the gradient of the mixed state + /// $|v\rangle=\sum_kv_k|X_k\rangle$, so the Jahn-Teller direction is the steepest-descent + /// direction of whichever state in the multiplet has the largest gradient. + /// + /// Solved by alternating exact minimization -- $u\leftarrow-q(v)/\|q(v)\|$, then $v\leftarrow$ + /// the $\lambda_{\min}$ eigenvector of $M(u)$ -- which decreases the joint objective + /// monotonically. The objective is homogeneous of degree one and only piecewise smooth, so + /// several deterministic starting points are tried and the best kept; `restarts_agreeing` + /// reports how many landed on it, which is the practical signal that the optimum is global. + /// + /// This is first order only. It gives the DIRECTION; the actual distortion amplitude needs the + /// harmonic term as well ($Q\approx-g/k$), and a norm that means anything physically should be + /// taken in mass-weighted coordinates rather than plain Cartesian ones. + /// + /// @param gflat the gradient matrix, flattened as `gflat[a * d * d + k * d + l]`, symmetric in + /// (k, l). The sign convention is the caller's: feeding FORCES + /// ($-\partial\Omega/\partial R$) makes `displacement` point downhill directly, + /// while feeding gradients makes it point uphill. + /// @param ncoord $3N$ + /// @param d the multiplet's dimension, must be >= 2 + JTDirection find_jt_direction(const std::vector& gflat, const int ncoord, const int d); + + /// @brief Split a branch gradient into the part shared by the whole multiplet and the part that + /// actually breaks the degeneracy. + /// + /// $q(v)=\bar q+\big(q(v)-\bar q\big)$ with $\bar q_a=\operatorname{Tr}G^{(a)}/d$. The first term + /// is the multiplet average: totally symmetric, common to every branch, and it only relaxes the + /// geometry without splitting anything. The second is the Jahn-Teller part. The distinction + /// matters when reading the result -- at a stationary point of the average surface, which is + /// where an `lr_degen_mode = average` relaxation ends up, the first term vanishes and the + /// whole direction is Jahn-Teller. + /// + /// @return the symmetric part $\bar q$; `jt_part` receives $q(v)-\bar q$ + std::vector split_symmetric_part(const std::vector& gflat, + const int ncoord, + const int d, + const std::vector& mixing, + std::vector& jt_part); +} + +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_GRAD_DEGEN_H diff --git a/source/source_lcao/module_lr/grad_jt.cpp b/source/source_lcao/module_lr/grad_jt.cpp new file mode 100644 index 00000000000..87d6fd9e69a --- /dev/null +++ b/source/source_lcao/module_lr/grad_jt.cpp @@ -0,0 +1,208 @@ +#include "grad_degen.h" + +#include +#include +#include +#include + +namespace LR +{ + namespace + { + /// $q(v)_a=v^\top G^{(a)}v$: the gradient of the mixed state $\sum_kv_k|X_k\rangle$. + std::vector branch_gradient(const std::vector& gflat, + const int ncoord, + const int d, + const std::vector& v) + { + std::vector q(ncoord, 0.0); + for (int a = 0; a < ncoord; ++a) + { + const double* const block = gflat.data() + static_cast(a) * d * d; + double s = 0.0; + for (int k = 0; k < d; ++k) + { + for (int l = 0; l < d; ++l) { s += v[k] * block[k * d + l] * v[l]; } + } + q[a] = s; + } + return q; + } + + double norm2(const std::vector& x) + { + double s = 0.0; + for (size_t i = 0; i < x.size(); ++i) { s += x[i] * x[i]; } + return std::sqrt(s); + } + + /// $M(u)=\sum_a u_aG^{(a)}$, as a dense $d\times d$ row-major block. + std::vector contract_direction(const std::vector& gflat, + const int ncoord, + const int d, + const std::vector& u) + { + std::vector m(static_cast(d) * d, 0.0); + for (int a = 0; a < ncoord; ++a) + { + const double* const block = gflat.data() + static_cast(a) * d * d; + for (int i = 0; i < d * d; ++i) { m[i] += u[a] * block[i]; } + } + return m; + } + + /// The eigenvector of the smallest eigenvalue of a small symmetric matrix, by Jacobi + /// rotations. Written out rather than taken from LAPACK so that this file stays free of + /// the parallel/linear-algebra layer and can be unit-tested on its own; $d$ is the + /// dimension of an electronic multiplet, so 2 or 3 in practice and never large. + std::vector min_eigenvector(std::vector m, const int d) + { + std::vector ev(static_cast(d) * d, 0.0); + for (int i = 0; i < d; ++i) { ev[i * d + i] = 1.0; } + for (int sweep = 0; sweep < 100; ++sweep) + { + double off = 0.0; + for (int i = 0; i < d; ++i) + { + for (int j = i + 1; j < d; ++j) { off += m[i * d + j] * m[i * d + j]; } + } + if (off < 1e-30) { break; } + for (int i = 0; i < d; ++i) + { + for (int j = i + 1; j < d; ++j) + { + const double aij = m[i * d + j]; + if (std::abs(aij) < 1e-300) { continue; } + const double theta = 0.5 * (m[j * d + j] - m[i * d + i]) / aij; + const double t = (theta >= 0.0 ? 1.0 : -1.0) + / (std::abs(theta) + std::sqrt(theta * theta + 1.0)); + const double c = 1.0 / std::sqrt(t * t + 1.0); + const double s = t * c; + for (int k = 0; k < d; ++k) + { + const double mik = m[i * d + k]; + const double mjk = m[j * d + k]; + m[i * d + k] = c * mik - s * mjk; + m[j * d + k] = s * mik + c * mjk; + } + for (int k = 0; k < d; ++k) + { + const double mki = m[k * d + i]; + const double mkj = m[k * d + j]; + m[k * d + i] = c * mki - s * mkj; + m[k * d + j] = s * mki + c * mkj; + const double eki = ev[k * d + i]; + const double ekj = ev[k * d + j]; + ev[k * d + i] = c * eki - s * ekj; + ev[k * d + j] = s * eki + c * ekj; + } + } + } + } + int best = 0; + for (int i = 1; i < d; ++i) + { + if (m[i * d + i] < m[best * d + best]) { best = i; } + } + std::vector v(d, 0.0); + for (int k = 0; k < d; ++k) { v[k] = ev[k * d + best]; } + const double n = norm2(v); + for (int k = 0; k < d; ++k) { v[k] /= n; } + return v; + } + + /// Deterministic starting points: the unit vectors, then the normalized all-ones-with-signs + /// patterns. Deterministic because a relaxation must give the same answer twice. + std::vector> jt_start_points(const int d) + { + std::vector> starts; + for (int k = 0; k < d; ++k) + { + std::vector v(d, 0.0); + v[k] = 1.0; + starts.push_back(v); + } + const int nsign = 1 << (d - 1); // fix the first sign: v and -v give the same q(v) + for (int mask = 0; mask < nsign; ++mask) + { + std::vector v(d, 1.0 / std::sqrt(static_cast(d))); + for (int k = 1; k < d; ++k) + { + if ((mask >> (k - 1)) & 1) { v[k] = -v[k]; } + } + starts.push_back(v); + } + return starts; + } + } + + JTDirection find_jt_direction(const std::vector& gflat, const int ncoord, const int d) + { + assert(d >= 2); + assert(gflat.size() == static_cast(ncoord) * d * d); + JTDirection best; + // A stationary multiplet still needs a valid branch for the force decomposition and + // next-step tracking. No displacement is selected when every branch gradient vanishes. + best.mixing.assign(d, 0.0); + best.mixing[0] = 1.0; + best.displacement.assign(ncoord, 0.0); + const std::vector> starts = jt_start_points(d); + for (size_t is = 0; is < starts.size(); ++is) + { + std::vector v = starts[is]; + double obj = -1.0; + int it = 0; + std::vector q; + for (; it < 200; ++it) + { + q = branch_gradient(gflat, ncoord, d, v); + const double nq = norm2(q); + // A vanishing gradient means this branch is already stationary: there is no + // direction to report from this start, so leave it to the others. + if (nq < 1e-14) { break; } + if (nq - obj < 1e-12 * std::max(1.0, nq)) { obj = nq; break; } + obj = nq; + // u minimizes u.q(v) at fixed v; v then minimizes v^T M(u) v at fixed u. Both are + // exact, so the joint objective -\|q\| decreases monotonically. + std::vector u(ncoord); + for (int a = 0; a < ncoord; ++a) { u[a] = -q[a] / nq; } + v = min_eigenvector(contract_direction(gflat, ncoord, d, u), d); + } + if (obj <= 0.0) { continue; } + if (obj > best.slope * (1.0 + 1e-9)) + { + best.slope = obj; + best.mixing = v; + best.displacement.assign(ncoord, 0.0); + for (int a = 0; a < ncoord; ++a) { best.displacement[a] = q[a] / obj; } + best.iterations = it; + best.restarts_agreeing = 1; + } + else if (obj > best.slope * (1.0 - 1e-6)) + { + ++best.restarts_agreeing; + } + } + return best; + } + + std::vector split_symmetric_part(const std::vector& gflat, + const int ncoord, + const int d, + const std::vector& mixing, + std::vector& jt_part) + { + const std::vector q = branch_gradient(gflat, ncoord, d, mixing); + std::vector sym(ncoord, 0.0); + jt_part.assign(ncoord, 0.0); + for (int a = 0; a < ncoord; ++a) + { + const double* const block = gflat.data() + static_cast(a) * d * d; + double tr = 0.0; + for (int k = 0; k < d; ++k) { tr += block[k * d + k]; } + sym[a] = tr / static_cast(d); + jt_part[a] = q[a] - sym[a]; + } + return sym; + } +} diff --git a/source/source_lcao/module_lr/gradient_checks.h b/source/source_lcao/module_lr/gradient_checks.h new file mode 100644 index 00000000000..6c1c800cfff --- /dev/null +++ b/source/source_lcao/module_lr/gradient_checks.h @@ -0,0 +1,41 @@ +#ifndef ABACUS_LR_GRADIENT_CHECKS_H +#define ABACUS_LR_GRADIENT_CHECKS_H +#include "utils/lr_util_print.h" +namespace LR +{ +// check C_uaC_va-C_uiC_vi of lumo-homo, nocc=1, nk=1 +template +inline void test_dm_diff_H2(const T* dm, const psi::Psi& c, const int nbasis) +{ + std::cout << "difference dm cal: " << std::endl; + LR_Util::print_value(dm, nbasis, nbasis); + std::cout << "difference dm ref: " << std::endl; + for (int i = 0;i < nbasis;++i) + { + for (int j = 0;j < nbasis;++j) + { + std::cout << c(0, 1, i) * c(0, 1, j) - c(0, 0, i) * c(0, 0, j) << " "; + } + std::cout << std::endl; + } +} + +// check e_aC_uaC_va-e_iC_uiC_vi of lumo-homo, nocc=1, nk=1 +template +inline void test_edm_H2(const T* const edm, const double* const eig_ks, const psi::Psi& c, const int nbasis) +{ + std::cout << "edm cal: " << std::endl; + LR_Util::print_value(edm, nbasis, nbasis); + std::cout << "edm ref: " << std::endl; + for (int i = 0;i < nbasis;++i) + { + for (int j = 0;j < nbasis;++j) + { + std::cout << eig_ks[1] * c(0, 1, i) * c(0, 1, j) - eig_ks[0] * c(0, 0, i) * c(0, 0, j) << " "; + } + std::cout << std::endl; + } +} + +} +#endif diff --git a/source/source_lcao/module_lr/gradient_inputs.h b/source/source_lcao/module_lr/gradient_inputs.h new file mode 100644 index 00000000000..35494f51e19 --- /dev/null +++ b/source/source_lcao/module_lr/gradient_inputs.h @@ -0,0 +1,55 @@ +#ifndef ABACUS_LR_GRADIENT_INPUTS_H +#define ABACUS_LR_GRADIENT_INPUTS_H +#include "lr_force.h" +#include "source_lcao/module_lr/potentials/pot_hxc_lrtd.h" +namespace LR +{ +// Explicit, borrowed inputs for one gradient evaluation. No solver ownership or caches. +template +struct GradientInputs +{ + const UnitCell& ucell; + const K_Vectors& kv; + const Grid_Driver& gd; + const std::vector& orb_cutoff; + const Parallel_Orbitals& pmat; + const Parallel_2D& pc; + const std::vector& px; + const psi::Psi& psi_ks; + const ModuleBase::matrix& eig_ks; + const std::vector& nocc; + const std::vector& nvirt; + int nspin; + int nk; + int nbasis; + int nloc; + const std::string& xc_kernel; + const std::string& dft_functional; + const std::string& ks_solver; + bool test_force; + bool excited_relax; + const std::string& out_dir; + int my_rank; + const std::vector& spin_types; + const std::vector>& pot; + const std::shared_ptr& pot_hxc_gs; + std::ofstream& ofs; +#ifdef __EXX + const std::shared_ptr>& exx_lri; + double hybrid_alpha; +#endif +}; +template +std::vector evaluate_closed_shell_force( + const GradientInputs& inputs, LR_Force& force_terms, + const module_dm::DensityMatrix& dm_gs, + const ct::Tensor& Xz, const ct::Tensor& Z, + const std::vector& omega, int label_begin, int ispin); +template +std::vector evaluate_open_shell_force( + const GradientInputs& inputs, LR_Force& force_terms, + const module_dm::DensityMatrix& dm_gs, + const ct::Tensor& Xz, const ct::Tensor& Z, + const std::vector& omega, int label_begin); +} +#endif diff --git a/source/source_lcao/module_lr/gradient_output.cpp b/source/source_lcao/module_lr/gradient_output.cpp new file mode 100644 index 00000000000..8b86f214310 --- /dev/null +++ b/source/source_lcao/module_lr/gradient_output.cpp @@ -0,0 +1,108 @@ +#include "gradient_output.h" +#include "grad_degen.h" +#include "utils/lr_util.h" +#include "source_io/module_output/output_log.h" +#include "source_base/constants.h" +#include "source_cell/unitcell.h" +#include +#include +namespace LR +{ +void print_lr_force(const std::vector& force, std::ostream& ofs, const int istate_begin) +{ + const std::ios::fmtflags old_flags = ofs.flags(); + const std::streamsize old_precision = ofs.precision(); + const int nstate = force.size(); + ofs << "Forces (-gradients) of each excited state: (eV/Angstrom)" << std::endl; + ofs << std::fixed << std::setprecision(10) << std::setw(6) << "state" << std::setw(6) << "atom" + << std::setw(20) << "x" << std::setw(20) << "y" << std::setw(20) << "z" << std::endl; + const double fac = ModuleBase::Ry_to_eV / ModuleBase::BOHR_TO_A; + for (int i = 0;i < nstate;++i) + { + for (int iat = 0;iat < force[i].nr;++iat) + { + std::string istate = iat == 0 ? std::to_string(istate_begin + i) : " "; + ofs << std::setw(6) << istate << std::setw(6) << iat << std::setw(6) << "force"; + for (int ixyz = 0;ixyz < 3;++ixyz) { ofs << std::setw(20) << force[i](iat, ixyz) * fac; } + ofs << std::endl; + } + } + ofs.flags(old_flags); + ofs.precision(old_precision); +} + +/// @brief The mean of a set of force matrices. +/// +/// Used for the multiplet average $\operatorname{Tr}G/d$, the one smooth, basis-independent $3N$ +/// vector field a degenerate multiplet has. Factored out because three callers need it from +/// different inputs: the printout and the stored LVC hold the whole gradient matrix, while the +/// relaxation only ever computes its diagonal. +ModuleBase::matrix average_forces(const std::vector& f) +{ + assert(!f.empty()); + ModuleBase::matrix avg(f[0].nr, f[0].nc); + for (size_t k = 0; k < f.size(); ++k) { avg += f[k]; } + avg *= 1.0 / static_cast(f.size()); + return avg; +} + +/// @brief Print the degenerate-subspace gradient matrix $G_{kl}$, one $d\times d$ block per +/// nuclear coordinate, plus the multiplet average on its diagonal. +/// +/// The individual diagonal entries are basis-dependent: only the eigenvalues of +/// $M(u)=\sum_a u_aG^{(a)}$ are branch slopes, and only $\operatorname{Tr}G$ is invariant. The +/// average $\operatorname{Tr}G/d$ is printed because it IS a smooth, basis-independent $3N$ vector +/// field -- the one a symmetry-constrained relaxation can follow. +void print_grad_matrix(const std::vector>& g, + const std::vector& group, const UnitCell& ucell, std::ofstream& ofs) +{ + const double fac = ModuleBase::Ry_to_eV / ModuleBase::BOHR_TO_A; + const int d = static_cast(g.size()); + const int nat = g[0][0].nr; + ofs << std::endl << " DEGENERATE-SUBSPACE GRADIENT MATRIX G_kl (eV/Angstrom), states"; + for (int k = 0; k < d; ++k) { ofs << " " << group[k]; } + ofs << std::endl + << " Branch slopes along a displacement u are the EIGENVALUES of sum_a u_a G^(a); the" + << std::endl + << " diagonal entries alone are basis-dependent and only their trace is invariant." + << std::endl; + ofs << " For each coordinate the matrix is followed by its eigenvalues, which ARE the branch" + << std::endl + << " slopes for displacing that one atom along that one axis. They must NOT be combined" + << std::endl + << " across axes: G^(x), G^(y), G^(z) do not commute in general, so eigenvalues are not" + << std::endl + << " additive and picking one per axis describes no adiabatic state at all." << std::endl; + ofs << std::setprecision(6); + for (int iat = 0; iat < nat; ++iat) + { + for (int ixyz = 0; ixyz < 3; ++ixyz) + { + ofs << " atom " << std::setw(5) << iat << " dir " << std::setw(2) << ixyz << std::endl; + std::vector block(static_cast(d) * d); + for (int k = 0; k < d; ++k) + { + ofs << " "; + for (int l = 0; l < d; ++l) + { + const double v = g[k][l](iat, ixyz) * fac; + ofs << std::setw(15) << v; + block[static_cast(k) * d + l] = v; + } + ofs << std::endl; + } + // `diag_lapack` overwrites its input with the eigenvectors, hence the scratch copy + std::vector eig(d, 0.0); + LR_Util::diag_lapack(d, block.data(), eig.data()); + ofs << " eig"; + for (int k = 0; k < d; ++k) { ofs << std::setw(15) << eig[k]; } + ofs << std::endl; + } + } + std::vector diag; + for (int k = 0; k < d; ++k) { diag.push_back(g[k][k]); } + const ModuleBase::matrix average = average_forces(diag); + ModuleIO::print_force(ofs, ucell, "MULTIPLET-AVERAGE FORCE Tr(G)/d (eV/Angstrom)", average, false); +} + +} diff --git a/source/source_lcao/module_lr/gradient_output.h b/source/source_lcao/module_lr/gradient_output.h new file mode 100644 index 00000000000..a11d0e462ca --- /dev/null +++ b/source/source_lcao/module_lr/gradient_output.h @@ -0,0 +1,14 @@ +#ifndef ABACUS_LR_GRADIENT_OUTPUT_H +#define ABACUS_LR_GRADIENT_OUTPUT_H +#include "source_base/matrix.h" +#include +#include +class UnitCell; +namespace LR +{ +void print_lr_force(const std::vector& force, std::ostream& ofs, int istate_begin); +ModuleBase::matrix average_forces(const std::vector& forces); +void print_grad_matrix(const std::vector>& gradient, + const std::vector& group, const UnitCell& ucell, std::ofstream& ofs); +} +#endif diff --git a/source/source_lcao/module_lr/hamilt_casida.cpp b/source/source_lcao/module_lr/hamilt_casida.cpp index e51ddee7933..7ce77052580 100644 --- a/source/source_lcao/module_lr/hamilt_casida.cpp +++ b/source/source_lcao/module_lr/hamilt_casida.cpp @@ -1,4 +1,6 @@ #include "hamilt_casida.h" +#include +#include #include "source_lcao/module_lr/utils/lr_util_print.h" namespace LR { @@ -56,9 +58,6 @@ namespace LR } } } - // output Amat - std::cout << "Full A matrix: (elements < 1e-10 is set to 0)" << std::endl; - LR_Util::print_value(Amat_full.data(), nk * npairs, nk * npairs); return Amat_full; } diff --git a/source/source_lcao/module_lr/hamilt_casida.h b/source/source_lcao/module_lr/hamilt_casida.h index b17a13fdc45..218d04f5f8a 100644 --- a/source/source_lcao/module_lr/hamilt_casida.h +++ b/source/source_lcao/module_lr/hamilt_casida.h @@ -20,7 +20,7 @@ namespace LR class HamiltLR { public: - HamiltLR(std::string& xc_kernel, + HamiltLR(const std::string& xc_kernel, const int& nspin, const int& naos, const std::vector& nocc, @@ -31,10 +31,10 @@ namespace LR const psi::Psi& psi_ks_in, const ModuleBase::matrix& eig_ks, #ifdef __EXX - std::weak_ptr> exx_lri_in, - const double& exx_alpha, + std::weak_ptr> exx_lri_in, + const double& exx_alpha, #endif - std::weak_ptr pot_in, + std::weak_ptr pot_in, const K_Vectors& kv_in, const std::vector& pX_in, const Parallel_2D& pc_in, @@ -111,22 +111,27 @@ namespace LR else #endif { - OperatorLRHxc* lr_hxc = new OperatorLRHxc(nspin, naos, nocc, nvirt, psi_ks_in, - this->DM_trans, pot_in, ucell_in, orb_cutoff, gd_in, kv_in, pX_in, pc_in, pmat_in); + hamilt::Operator* lr_hxc = new OperatorLRHxc(nspin, naos, nocc, nvirt, psi_ks_in, + *this->DM_trans, pot_in, ucell_in, orb_cutoff, gd_in, kv_in, pX_in, pc_in, pmat_in); this->ops->add(lr_hxc); } -#ifdef __EXX// 3.add Exx operator - if (xc_kernel == "hf" || xc_kernel == "hse") - { +#ifdef __EXX + if (exx_kernel_list().count(xc_kernel) ) + { //add Exx operator if (ri_hartree_benchmark != "none" && spin_type == "singlet") { exx_lri_in.lock()->reset_Cs(Cs_read); exx_lri_in.lock()->reset_Vs(Vs_read); } // std::cout << "exx_alpha=" << exx_alpha << std::endl; // the default value of exx_alpha is 0.25 when dft_functional is pbe or hse + // this ground-state Casida operator is always MO_TO_AO_TYPE::CC_vo, which never + // reads the force-only coxt_full/cvx_full buffers, so cal_force is always false. hamilt::Operator* lr_exx = new OperatorLREXX(nspin, naos, nocc[0], nvirt[0], ucell_in, psi_ks_in, - this->DM_trans, exx_lri_in, kv_in, pX_in[0], pc_in, pmat_in, - (xc_kernel == "hf") ? 1.0 : exx_alpha); + *this->DM_trans, exx_lri_in, kv_in, pX_in[0], pc_in, pmat_in, + /*cal_force=*/false, + (xc_kernel == "hf" ? 1.0 : exx_alpha), //alpha + OperatorLREXX::MO_TO_AO_TYPE::CC_vo, + aims_nbasis); this->ops->add(lr_exx); } #endif @@ -150,7 +155,7 @@ namespace LR std::vector matrix()const; - void hPsi(const T* const psi_in, T* const hpsi, const int ld_psi, const int& nband) const + virtual void hPsi(const T* const psi_in, T* const hpsi, const int ld_psi, const int nband) const { assert(ld_psi == nk * pX[0].get_local_size()); for (int ib = 0;ib < nband;++ib) @@ -190,13 +195,14 @@ namespace LR // } // } - private: + // const references const std::vector& nocc; const std::vector& nvirt; const int nspin = 1; const int nk = 1; - const bool tdm_sym = false; ///< whether to symmetrize the transition density matrix const std::vector& pX; + protected: + const bool tdm_sym = false; ///< whether to symmetrize the transition density matrix T one()const; /// transition density matrix in AO representation /// calculate on the same address for each bands, and commonly used by all the operators diff --git a/source/source_lcao/module_lr/hamilt_ulr.hpp b/source/source_lcao/module_lr/hamilt_ulr.hpp index b9b88cd42d6..cfd69e61bea 100644 --- a/source/source_lcao/module_lr/hamilt_ulr.hpp +++ b/source/source_lcao/module_lr/hamilt_ulr.hpp @@ -17,7 +17,7 @@ namespace LR class HamiltULR { public: - HamiltULR(std::string& xc_kernel, + HamiltULR(const std::string& xc_kernel, const int& nspin, const int& naos, const std::vector& nocc, ///< {up, down} @@ -45,25 +45,28 @@ namespace LR // this->DM_trans->init_dmr(&gd_in, &ucell_in); // too large due to not restricted by orb_cutoff this->ops.resize(4); - this->ops[0] = new OperatorLRDiag(eig_ks.c, pX_in[0], nk, nocc[0], nvirt[0]); - this->ops[3] = new OperatorLRDiag(eig_ks.c + nk * (nocc[0] + nvirt[0]), pX_in[1], nk, nocc[1], nvirt[1]); + this->ops[0] = new OperatorLRDiag(eig_ks.c, pX_in[0], nk, nocc[0], nvirt[0], /*add_on=*/true); + this->ops[3] = new OperatorLRDiag(eig_ks.c + nk * (nocc[0] + nvirt[0]), pX_in[1], nk, nocc[1], nvirt[1], /*add_on=*/true); auto newHxc = [&](const int& sl, const int& sr) { return new OperatorLRHxc(nspin, naos, nocc, nvirt, psi_ks_in, - this->DM_trans, pot_in[sl], ucell_in, orb_cutoff, gd_in, kv_in, pX_in, pc_in, pmat_in, { sl,sr }); }; + *this->DM_trans, pot_in[sl], ucell_in, orb_cutoff, gd_in, kv_in, pX_in, pc_in, pmat_in, { sl,sr }); }; this->ops[0]->add(newHxc(0, 0)); this->ops[1] = newHxc(0, 1); this->ops[2] = newHxc(1, 0); this->ops[3]->add(newHxc(1, 1)); #ifdef __EXX - if (xc_kernel == "hf" || xc_kernel == "hse") + if (exx_kernel_list().count(xc_kernel) ) { std::vector> psi_ks_spin = { LR_Util::get_psi_spin(psi_ks_in, 0, nk), LR_Util::get_psi_spin(psi_ks_in, 1, nk) }; for (int is : {0, 1}) { + // this ground-state Casida operator defaults to MO_TO_AO_TYPE::CC_vo, which + // never reads the force-only coxt_full/cvx_full buffers. this->ops[(is << 1) + is]->add(new OperatorLREXX(nspin, naos, nocc[is], nvirt[is], ucell_in, psi_ks_spin[is], - this->DM_trans, exx_lri_in, kv_in, pX_in[is], pc_in, pmat_in, - xc_kernel == "hf" ? 1.0 : exx_alpha)); + *this->DM_trans, exx_lri_in, kv_in, pX_in[is], pc_in, pmat_in, + /*cal_force=*/false, + (xc_kernel == "hf" ? 1.0 : exx_alpha))); } } #endif @@ -97,17 +100,23 @@ namespace LR for (int ib = 0;ib < nband;++ib) { const int offset_band = ib * ld_psi; + // every one of the four spin blocks accumulates into this band's output, and the + // diagonal `OperatorLRDiag`s now add rather than assign, so clear it first + std::fill(hpsi + offset_band, hpsi + offset_band + ld_psi, T(0)); for (int is_bj : {0, 1}) { const int offset_bj = offset_band + is_bj * xdim_is[0]; cal_dm_trans(is_bj, psi_in + offset_bj); // calculate transition density matrix here + typename OperatorLRHxc::TransitionDensityCache density_cache; for (int is_ai : {0, 1}) { const int offset_ai = offset_band + is_ai * xdim_is[0]; hamilt::Operator* node(this->ops[(is_ai << 1) + is_bj]); while (node != nullptr) { - node->act(/*nband=*/1, xdim_is[is_bj], /*npol=*/1, psi_in + offset_bj, hpsi + offset_ai); + const T* input = psi_in + offset_bj; + T* output = hpsi + offset_ai; + act_with_shared_density(node, xdim_is[is_bj], input, output, density_cache); node = (hamilt::Operator*)(node->next_op); } } @@ -156,12 +165,18 @@ namespace LR #ifdef __MPI for (int ik_ai = 0;ik_ai < this->nk;++ik_ai) { + // The block being gathered is the OUT channel's, so its global + // shape is (nvirt[is_ai], nocc[is_ai]). Passing the IN channel's + // `nv`/`no` copies the wrong number of elements as soon as the + // two channels differ in size. LR_Util::gather_2d_to_full(pax, Aloc_col.data() + loffset_ai + ik_ai * pax.get_local_size(), Amat_full.data() + gcol * gdim /*col, bj*/ + goffset_ai + ik_ai * npairs[is_ai]/*row, ai*/, - false, nv, no); + false, this->nvirt[is_ai], this->nocc[is_ai]); } #else - std::memcpy(Amat_full.data() + gcol * gdim + goffset_ai, Aloc_col.data() + goffset_ai, gdim_is[is_ai] * sizeof(T)); + // `Aloc_col` is the LOCAL vector, so it is indexed by `loffset_ai` + // (the two coincide only on one process). + std::memcpy(Amat_full.data() + gcol * gdim + goffset_ai, Aloc_col.data() + loffset_ai, gdim_is[is_ai] * sizeof(T)); #endif } } diff --git a/source/source_lcao/module_lr/hamilt_zeq_l.h b/source/source_lcao/module_lr/hamilt_zeq_l.h new file mode 100644 index 00000000000..594b81be8c4 --- /dev/null +++ b/source/source_lcao/module_lr/hamilt_zeq_l.h @@ -0,0 +1,211 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_HAMILT_ZEQ_L_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_HAMILT_ZEQ_L_H +#include "source_lcao/module_lr/hamilt_casida.h" +#include "hamilt_zequlr.h" +#include "source_estate/module_dm/density_matrix.h" +#include "source_lcao/module_lr/potentials/pot_grad_xc.h" +#include "source_lcao/module_lr/operator_casida/operator_lr_hxc.h" +#include "source_lcao/module_lr/dm_trans/dm_diff.h" +#include "source_basis/module_ao/parallel_orbitals.h" +#ifdef __EXX +#include "source_lcao/module_lr/operator_casida/operator_lr_exx.h" +#endif +namespace LR +{ + template + class Z_vector_L : public HamiltLR + { + using ATYPE = typename OperatorLRHxc::MO_TO_AO_TYPE; +#ifdef __EXX + using ATYPE_EXX = typename OperatorLREXX::MO_TO_AO_TYPE; +#endif + public: + Z_vector_L(const std::string& xc_kernel, + const int& nspin, + const int& naos, + const std::vector& nocc, + const std::vector& nvirt, + const UnitCell& ucell, + const std::vector& orb_cutoff, + const Grid_Driver& gd, + const psi::Psi& psi_ks, + const ModuleBase::matrix& eig_ks, +#ifdef __EXX + std::weak_ptr> exx_lri, + const double& exx_alpha, +#endif + std::weak_ptr pot_hxc_gs, + const K_Vectors& kv, + const std::vector& pX, + const Parallel_2D& pc, + const Parallel_Orbitals& pmat, + const std::string& spin_type, + const std::string& in_dir, + const std::string& out_dir, + const std::string& dft_functional) + : HamiltLR(xc_kernel, nspin, naos, nocc, nvirt, ucell, orb_cutoff, gd, psi_ks, eig_ks, +#ifdef __EXX + exx_lri, exx_alpha, +#endif + pot_hxc_gs, kv, pX, pc, pmat, spin_type, in_dir, out_dir) + { + ModuleBase::TITLE("Z_vector_L", "Z_vector_L"); + // Destroy the base chain while its borrowed density matrix is still alive. + delete this->ops; + this->ops = nullptr; + this->DM_trans = LR_Util::make_unique>(&pmat, 1, kv.kvec_d, this->nk); + LR_Util::initialize_DMR(*this->DM_trans, pmat, ucell, gd, orb_cutoff); + // Hessian (A+B) with GS XC kernel + // 1. diag term in A + this->ops = new OperatorLRDiag(eig_ks.c, pX[0], kv.get_nks() / nspin, nocc[0], nvirt[0]); + // 2. $H_{ia}[D^Z]$, equals to $2K_{ab}[D^Z]$ when $D^Z$ is symmetrized. + // Factor 4 (not 2): the singlet kernel is $K^S_\text{Hxc}=2$`pot_hxc_gs` (not doubled), + // while `pot` is the already-doubled singlet potential, so $H^S=2K^S$ here needs 4. + // The EXX line below is already $2\alpha$ and is consistent. + hamilt::Operator* op_hz = new OperatorLRHxc(nspin, naos, nocc, nvirt, psi_ks, + *this->DM_trans, pot_hxc_gs, ucell, orb_cutoff, gd, kv, pX, pc, pmat, + { 0 }, 4.0, ATYPE::CC_vo); + this->ops->add(op_hz); +#ifdef __EXX + if (gs_is_hybrid(dft_functional)) + { + // Z_vector_L only exists on the force-calculation path, so cal_force is always true here. + hamilt::Operator* op_hz_exx = new OperatorLREXX(nspin, naos, nocc[0], nvirt[0], ucell, psi_ks, + *this->DM_trans, exx_lri, kv, pX[0], pc, pmat, + /*cal_force=*/true, + 2.0 * exx_alpha, //alpha; H=2K when D is symmetrized + ATYPE_EXX::CC_vo); + this->ops->add(op_hz_exx); + } +#endif + this->cal_dm_trans = [&, this](const int& is, const T* X)->void + { + const auto psi_ks_is = LR_Util::get_psi_spin(psi_ks, is, this->nk); +#ifdef __MPI + std::vector dm_trans_2d = cal_dm_trans_pblas(X, this->pX[is], psi_ks_is, pc, naos, nocc[is], nvirt[is], pmat); + for (auto& t : dm_trans_2d) LR_Util::matsym(t.data(), naos, pmat); +#else + std::vector dm_trans_2d = cal_dm_trans_blas(X, psi_ks_is, nocc[is], nvirt[is]); + for (auto& t : dm_trans_2d) LR_Util::matsym(t.data(), naos); +#endif + // LR_Util::print_tensor(dm_trans_2d[0], "dm_trans_2d[0]", &pmat); + // tensor to vector, then set DMK + for (int ik = 0;ik < this->nk;++ik) { this->DM_trans->set_dmk_ptr(ik, dm_trans_2d[ik].data()); } + }; + } + }; + + /// @brief Open-shell (spin-unrestricted) counterpart of `Z_vector_L`. + /// + /// Left-hand side of the Z-vector equation: the orbital Hessian built with the + /// GROUND-STATE kernel, + /// $H_{ai\sigma}[D^Z]+Z_{ai\sigma}(\epsilon_{a\sigma}-\epsilon_{i\sigma})$. + /// + /// Factor 2.0 on the Hxc blocks (the closed-shell `Z_vector_L` uses 4.0): there + /// `pot_hxc_gs` is `S2_gs = S2_singlet/2`, so recovering $H^S=2K^S=4\cdot$`pot` needs 4. + /// Here `pot_hxc_gs` is `S2_updown`, which IS one $K_{\sigma\sigma'}$ component with no + /// halving, so only the $H=2K$ factor remains. The spin sum $\sum_{\sigma'}$ that the + /// singlet combination had baked in is now done by the block loop in `ZeqULR::hPsi`. + template + class Z_vector_UL : public ZeqULR + { + using ATYPE = typename OperatorLRHxc::MO_TO_AO_TYPE; +#ifdef __EXX + using ATYPE_EXX = typename OperatorLREXX::MO_TO_AO_TYPE; +#endif + public: + Z_vector_UL(const std::string& xc_kernel, + const int& nspin, + const int& naos, + const std::vector& nocc, + const std::vector& nvirt, + const UnitCell& ucell, + const std::vector& orb_cutoff, + const Grid_Driver& gd, + const psi::Psi& psi_ks, + const ModuleBase::matrix& eig_ks, +#ifdef __EXX + std::weak_ptr> exx_lri, + const double& exx_alpha, +#endif + std::weak_ptr pot_hxc_gs, + const K_Vectors& kv, + const std::vector& pX, + const Parallel_2D& pc, + const Parallel_Orbitals& pmat, + const std::string& dft_functional) + : ZeqULR(nocc, nvirt, pX, kv.get_nks() / nspin), + naos_(naos), pc_(pc), pmat_(pmat), psi_ks_(psi_ks) + { + ModuleBase::TITLE("Z_vector_UL", "Z_vector_UL"); + this->DM_trans = LR_Util::make_unique>(&pmat, 1, kv.kvec_d, this->nk); + LR_Util::initialize_DMR(*this->DM_trans, pmat, ucell, gd, orb_cutoff); + + // 1. the orbital-energy difference, diagonal blocks only + this->ops[0] = new OperatorLRDiag(eig_ks.c, pX[0], this->nk, nocc[0], nvirt[0], /*add_on=*/true); + this->ops[3] = new OperatorLRDiag(eig_ks.c + this->nk * (nocc[0] + nvirt[0]), + pX[1], this->nk, nocc[1], nvirt[1], /*add_on=*/true); + + // 2. $H_{ai\sigma}[D^Z]=2\sum_{\sigma'}K_{ai\sigma}[D^Z_{\sigma'}]$ + auto newHxc = [&](const int sl, const int sr) + { + return new OperatorLRHxc(nspin, naos, nocc, nvirt, psi_ks, + *this->DM_trans, pot_hxc_gs, ucell, orb_cutoff, gd, kv, pX, pc, pmat, + { sl, sr }, T(2.0), ATYPE::CC_vo); + }; + this->ops[0]->add(newHxc(0, 0)); + this->ops[1] = newHxc(0, 1); + this->ops[2] = newHxc(1, 0); + this->ops[3]->add(newHxc(1, 1)); + +#ifdef __EXX + // exchange is spin-diagonal ($\delta_{\sigma\sigma'}$), so only blocks 0 and 3. + // Factor 2*alpha is unchanged from the closed-shell version: the EXX part of the + // kernel carries no singlet/triplet combination, only $H=2K$. + if (gs_is_hybrid(dft_functional)) + { + for (int is : {0, 1}) + { + this->psi_ks_spin_.push_back(LR_Util::get_psi_spin(psi_ks, is, this->nk)); + } + for (int is : {0, 1}) + { + this->ops[(is << 1) + is]->add(new OperatorLREXX(nspin, naos, nocc[is], nvirt[is], + ucell, this->psi_ks_spin_[is], *this->DM_trans, exx_lri, kv, pX[is], pc, pmat, + /*cal_force=*/true, 2.0 * exx_alpha, ATYPE_EXX::CC_vo)); + } + } +#endif + } + + protected: + void set_dm(const int is, const T* const X) override + { + const auto psi_ks_is = LR_Util::get_psi_spin(this->psi_ks_, is, this->nk); +#ifdef __MPI + this->dm_buf_ = cal_dm_trans_pblas(X, this->pX[is], psi_ks_is, this->pc_, + this->naos_, this->nocc[is], this->nvirt[is], this->pmat_); + for (auto& t : this->dm_buf_) { LR_Util::matsym(t.template data(), this->naos_, this->pmat_); } +#else + this->dm_buf_ = cal_dm_trans_blas(X, psi_ks_is, this->nocc[is], this->nvirt[is]); + for (auto& t : this->dm_buf_) { LR_Util::matsym(t.template data(), this->naos_); } +#endif + for (int ik = 0;ik < this->nk;++ik) + { + this->DM_trans->set_dmk_ptr(ik, this->dm_buf_[ik].template data()); + } + } + + private: + const int naos_ = 1; + const Parallel_2D& pc_; + const Parallel_Orbitals& pmat_; + const psi::Psi& psi_ks_; + std::vector> psi_ks_spin_; + std::unique_ptr> DM_trans; + /// the tensors `DM_trans` points into; kept alive for the whole `act` chain + std::vector dm_buf_; + }; +} + +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_HAMILT_ZEQ_L_H diff --git a/source/source_lcao/module_lr/hamilt_zeq_r.h b/source/source_lcao/module_lr/hamilt_zeq_r.h new file mode 100644 index 00000000000..e9a2c0d1cbe --- /dev/null +++ b/source/source_lcao/module_lr/hamilt_zeq_r.h @@ -0,0 +1,350 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_HAMILT_ZEQ_R_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_HAMILT_ZEQ_R_H +#include "source_hamilt/hamilt.h" +#include "source_estate/module_dm/density_matrix.h" +#include "source_lcao/module_lr/potentials/pot_grad_xc.h" +#include "source_lcao/module_lr/potentials/pot_hxc_lrtd.h" +#include "source_lcao/module_lr/operator_casida/operator_lr_hxc.h" +#include "hamilt_zequlr.h" +#include "source_lcao/module_lr/operator_casida/op_gxc_ulr.h" +#include "source_basis/module_ao/parallel_orbitals.h" +#ifdef __EXX +#include "source_lcao/module_lr/operator_casida/operator_lr_exx.h" +#endif +namespace LR +{ + template + class Z_vector_R : public HamiltLR + { + using ATYPE = typename OperatorLRHxc::MO_TO_AO_TYPE; +#ifdef __EXX + using ATYPE_EXX = typename OperatorLREXX::MO_TO_AO_TYPE; +#endif + public: + Z_vector_R(const std::string& xc_kernel, + const int& nspin, + const int& naos, + const std::vector& nocc, + const std::vector& nvirt, + const UnitCell& ucell, + const std::vector& orb_cutoff, + const Grid_Driver& gd, + const psi::Psi& psi_ks, + const ModuleBase::matrix& eig_ks, +#ifdef __EXX + std::weak_ptr> exx_lri, + const double& exx_alpha, +#endif + std::weak_ptr pot, + std::weak_ptr pot_hxc_gs, + const K_Vectors& kv, + const std::vector& pX, + const Parallel_2D& pc, + const Parallel_Orbitals& pmat, + const std::string& in_dir, + const std::string& out_dir, + const std::string& dft_functional, + const std::string& spin_type = "singlet") + : HamiltLR(xc_kernel, nspin, naos, nocc, nvirt, ucell, orb_cutoff, gd, psi_ks, eig_ks, +#ifdef __EXX + exx_lri, exx_alpha, +#endif + pot, kv, pX, pc, pmat, spin_type, in_dir, out_dir) + { + ModuleBase::TITLE("Z_vector_R", "Z_vector_R"); + + // Destroy the base chain while its borrowed density matrix is still alive. + delete this->ops; + this->ops = nullptr; + this->DM_trans = LR_Util::make_unique>(&pmat, 1, kv.kvec_d, this->nk); + LR_Util::initialize_DMR(*this->DM_trans, pmat, ucell, gd, orb_cutoff); + this->DM_diff = LR_Util::make_unique>(&pmat, 1, kv.kvec_d, this->nk); + LR_Util::initialize_DMR(*this->DM_diff, pmat, ucell, gd, orb_cutoff); + + // note: calculation_type cannot repeated, or it will be ignored in ops->add() + // 1. $2\sum_bX_{ib}K_{ab}[D^X]-2\sum_jX_{ja}K_{ij}[D^X]$ + // kernel: excited state + this->ops = new OperatorLRHxc(nspin, naos, nocc, nvirt, psi_ks, + *this->DM_trans, pot, ucell, orb_cutoff, gd, kv, pX, pc, pmat, + { 0 }, T(-2.0), ATYPE::CXC); +#ifdef __EXX + if (exx_kernel_list().count(xc_kernel)) + { + hamilt::Operator* op_hz_exx = new OperatorLREXX(nspin, naos, nocc[0], nvirt[0], ucell, psi_ks, + *this->DM_trans, exx_lri, kv, pX[0], pc, pmat, + /*cal_force=*/true, + -2.0 * exx_alpha, //alpha; H=2K when D is symmetrized + ATYPE_EXX::CXC, {}, hamilt::calculation_type::lr_dmtrans_exx); + this->ops->add(op_hz_exx); + } +#endif + // 2. $H_{ia}[T]$, equals to $2K_{ab}[T]$ when $T$ is symmetrized + // kernel: ground state + // Factor -4 (not -2), for the same reason as in `hamilt_zeq_left.h`: + // $K^S_\text{Hxc}=2$`pot_hxc_gs`, so $H^S=2K^S$ needs 4. + hamilt::Operator* op_ht = new OperatorLRHxc(nspin, naos, nocc, nvirt, psi_ks, + *this->DM_diff, pot_hxc_gs, ucell, orb_cutoff, gd, kv, pX, pc, pmat, + { 0 }, T(-4.0), ATYPE::CC_vo, hamilt::calculation_type::lr_dmdiff_hxc); + this->ops->add(op_ht); +#ifdef __EXX + if (gs_is_hybrid(dft_functional)) + { + hamilt::Operator* op_ht_exx = new OperatorLREXX(nspin, naos, nocc[0], nvirt[0], ucell, psi_ks, + *this->DM_diff, exx_lri, kv, pX[0], pc, pmat, + /*cal_force=*/true, + -2.0 * exx_alpha, //alpha; H=2K when D is symmetrized + ATYPE_EXX::CC_vo, {}, hamilt::calculation_type::lr_dmdiff_exx); + this->ops->add(op_ht_exx); + } +#endif + + // 3. $2\sum_{jb,kc} g^{xc}_{ia, jb, kc}X_{jb}X_{kc}$ + // NOT singlet-only for a local functional where the triplet kernel is + // $K^T_{xc}=f_{uu}-f_{ud}\ne0$ -- exactly what `PotHxcLR`'s `S2_triplet` branch + // evaluates -- so $\partial K^T$ carries a $g^{xc}$ term too, with the "-" spin + // combination. `PotGradXCLR` picks it via the `triplet` flag. + if (LR_Util::has_local_xc(xc_kernel)) + { + this->pot_grad = std::make_shared(pot.lock()->xc_kernel_components(), pot.lock()->get_rho_basis(), ucell, pot.lock()->nrxx, spin_type == "triplet"); + hamilt::Operator* op_gxc = new OperatorLRHxc(nspin, naos, nocc, nvirt, psi_ks, + *this->DM_trans, this->pot_grad, ucell, orb_cutoff, gd, kv, pX, pc, pmat, + { 0 }, T(-2.0), ATYPE::CC_vo, hamilt::calculation_type::lr_dmtrans_gxc); + assert(op_gxc != nullptr); + this->ops->add(op_gxc); + } + // // test: op_ht only + // delete this->ops; + // this->ops = new OperatorLRHxc(nspin, naos, nocc, nvirt, psi_ks, + // *this->DM_diff, pot_hxc_gs, ucell, orb_cutoff, gd, kv, pX, pc, pmat, + // { 0 }, T(-2.0), ATYPE::CC_vo); + + // $D^X$ is fed to the CXC operators TRANSPOSED. + // + // The RHS needs $K^{S/T}_{ba}[D^X]$ -- the same kernel the Casida equation was solved + // with, which `HamiltLR` builds from the un-symmetrized $D^X$ (`tdm_sym = false`). + // `CVCX_virt`/`CVCX_occ` give the kernel matrix with its two MO indices in + // the opposite order to what this term needs, and since + // $(K[D])^T = K[D^T]$, + // transposing the density matrix on the way in restores it. + this->cal_dm_trans = [&, this](const int& is, const T* X)->void + { + const auto psi_ks_is = LR_Util::get_psi_spin(psi_ks, is, this->nk); +#ifdef __MPI + std::vector dm_trans_2d = cal_dm_trans_pblas(X, this->pX[is], psi_ks_is, pc, naos, nocc[is], nvirt[is], pmat); + for (auto& t : dm_trans_2d) { LR_Util::mattrans(t.data(), naos, pmat); } +#else + std::vector dm_trans_2d = cal_dm_trans_blas(X, psi_ks_is, nocc[is], nvirt[is]); + for (auto& t : dm_trans_2d) + { + T* d = t.data(); + for (int u = 0;u < naos;++u) { for (int v = u + 1;v < naos;++v) { std::swap(d[u * naos + v], d[v * naos + u]); } } + } +#endif + for (int ik = 0;ik < this->nk;++ik) { this->DM_trans->set_dmk_ptr(ik, dm_trans_2d[ik].data()); } + }; + + this->cal_dm_diff = [&, this](const int& is, const T* const X)->void + { + const auto psi_ks_is = LR_Util::get_psi_spin(psi_ks, is, this->nk); +#ifdef __MPI + std::vector dm_diff_2d = cal_dm_diff_pblas(X, this->pX[is], psi_ks_is, pc, naos, nocc[is], nvirt[is], pmat); + for (auto& t : dm_diff_2d) LR_Util::matsym(t.data(), naos, pmat); +#else + std::vector dm_diff_2d = cal_dm_diff_blas(X, psi_ks_is, naos, nocc[is], nvirt[is]); + for (auto& t : dm_diff_2d) LR_Util::matsym(t.data(), naos); +#endif + for (int ik = 0;ik < this->nk;++ik) { this->DM_diff->set_dmk_ptr(ik, dm_diff_2d[ik].data()); } + // std::cout << "difference density matrix" << std::endl; + // for (int ik = 0;ik < this->nk;++ik) { LR_Util::print_value(dm_diff_2d[ik].data(), naos, naos); } + // std::cout << "test: set dm_diff to zero" << std::endl; + // for (int ik = 0;ik < this->nk;++ik) { dm_diff_2d[ik].zero(); } + }; + } + virtual void hPsi(const T* const psi, T* const hpsi, const int ld_psi, const int nband) const override + { + assert(ld_psi == this->nk * this->pX[0].get_local_size()); + for (int ib = 0;ib < nband;++ib) + { + const int offset = ib * ld_psi; + this->cal_dm_trans(0, psi + offset); // transition density matrix, only for test + this->cal_dm_diff(0, psi + offset); // difference density matrix + hamilt::Operator* node(this->ops); + while (node != nullptr) + { + node->act(/*nband=*/1, ld_psi, /*npol=*/1, psi + offset, hpsi + offset); + node = (hamilt::Operator*)(node->next_op); + } + } + } + + private: + std::unique_ptr> DM_diff; + std::function cal_dm_diff; + std::shared_ptr pot_grad; + }; + + /// @brief Open-shell (spin-unrestricted) counterpart of `Z_vector_R`. + /// + /// Right-hand side of the Z-vector equation (the operators produce $-R$): + /// $$R_{ai\sigma}=H_{ai\sigma}[T]+\sum_bX_{bi\sigma}H_{ba\sigma}[D^X] + /// -\sum_kX_{ak\sigma}H_{ik\sigma}[D^X] + /// +2\sum_{\kappa\lambda\sigma'}\sum_{\alpha\beta\sigma''}D^X D^X g^{xc}.$$ + /// + /// Factors relative to the closed-shell `Z_vector_R`: + /// - the $K[D^X]$ term keeps -2.0. There `pot` is `S2_singlet`, which already IS $K^S$, + /// so the factor is just $H=2K$; here `pot` is `S2_updown`, which is one + /// $K_{\sigma\sigma'}$ component, and the $\sum_{\sigma'}$ is done by the block loop. + /// - the $H[T]$ term goes -4.0 -> -2.0, because `pot_hxc_gs` is no longer the halved + /// `S2_gs` but `S2_updown` (same argument as in `Z_vector_UL`). + template + class Z_vector_UR : public ZeqULR + { + using ATYPE = typename OperatorLRHxc::MO_TO_AO_TYPE; +#ifdef __EXX + using ATYPE_EXX = typename OperatorLREXX::MO_TO_AO_TYPE; +#endif + public: + Z_vector_UR(const std::string& xc_kernel, + const int& nspin, + const int& naos, + const std::vector& nocc, + const std::vector& nvirt, + const UnitCell& ucell, + const std::vector& orb_cutoff, + const Grid_Driver& gd, + const psi::Psi& psi_ks, + const ModuleBase::matrix& eig_ks, +#ifdef __EXX + std::weak_ptr> exx_lri, + const double& exx_alpha, +#endif + std::weak_ptr pot, + std::weak_ptr pot_hxc_gs, + const K_Vectors& kv, + const std::vector& pX, + const Parallel_2D& pc, + const Parallel_Orbitals& pmat, + const std::string& ks_solver, + const std::string& dft_functional) + : ZeqULR(nocc, nvirt, pX, kv.get_nks() / nspin), + naos_(naos), pc_(pc), pmat_(pmat), psi_ks_(psi_ks) + { + ModuleBase::TITLE("Z_vector_UR", "Z_vector_UR"); + this->DM_trans = LR_Util::make_unique>(&pmat, 1, kv.kvec_d, this->nk); + LR_Util::initialize_DMR(*this->DM_trans, pmat, ucell, gd, orb_cutoff); + this->DM_diff = LR_Util::make_unique>(&pmat, 1, kv.kvec_d, this->nk); + LR_Util::initialize_DMR(*this->DM_diff, pmat, ucell, gd, orb_cutoff); + + // 1. $\sum_bX_{bi\sigma}H_{ba\sigma}[D^X]-\sum_kX_{ak\sigma}H_{ik\sigma}[D^X]$, + // excited-state kernel, $D^X$ fed TRANSPOSED (see `set_dm` below) + auto newCXC = [&](const int sl, const int sr) + { + return new OperatorLRHxc(nspin, naos, nocc, nvirt, psi_ks, + *this->DM_trans, pot, ucell, orb_cutoff, gd, kv, pX, pc, pmat, + { sl, sr }, T(-2.0), ATYPE::CXC); + }; + this->ops[0] = newCXC(0, 0); + this->ops[1] = newCXC(0, 1); + this->ops[2] = newCXC(1, 0); + this->ops[3] = newCXC(1, 1); + + // 2. $H_{ai\sigma}[T]$, ground-state kernel + auto newHT = [&](const int sl, const int sr) + { + return new OperatorLRHxc(nspin, naos, nocc, nvirt, psi_ks, + *this->DM_diff, pot_hxc_gs, ucell, orb_cutoff, gd, kv, pX, pc, pmat, + { sl, sr }, T(-2.0), ATYPE::CC_vo, hamilt::calculation_type::lr_dmdiff_hxc); + }; + this->ops[0]->add(newHT(0, 0)); + this->ops[1]->add(newHT(0, 1)); + this->ops[2]->add(newHT(1, 0)); + this->ops[3]->add(newHT(1, 1)); + + // 3. $2\sum_{\sigma'\sigma''}D^X_{\sigma'}D^X_{\sigma''}g^{xc}_{\dots,ai\tau}$. + // Quadratic in $D^X$, so it lives outside the 2x2 block structure (see `band_extra`). + // Factor -2.0, straight from the spin-orbital formula -- the closed-shell code uses the + // same number only because its `PotGradXCLR` carries the doubled S/T combination. + if (LR_Util::has_local_xc(xc_kernel)) + { + this->gxc_ = LR_Util::make_unique>(pot.lock()->xc_kernel_components(), + pot.lock()->get_rho_basis(), ucell, orb_cutoff, gd, kv, pmat, pc, psi_ks, + nocc, nvirt, naos, pX, pX, LR_Util::MO_TYPE::VO, T(-2.0), nspin, ks_solver); + } + +#ifdef __EXX + for (int is : {0, 1}) { this->psi_ks_spin_.push_back(LR_Util::get_psi_spin(psi_ks, is, this->nk)); } + if (exx_kernel_list().count(xc_kernel)) + { + for (int is : {0, 1}) + { + this->ops[(is << 1) + is]->add(new OperatorLREXX(nspin, naos, nocc[is], nvirt[is], + ucell, this->psi_ks_spin_[is], *this->DM_trans, exx_lri, kv, pX[is], pc, pmat, + /*cal_force=*/true, -2.0 * exx_alpha, ATYPE_EXX::CXC, {}, hamilt::calculation_type::lr_dmtrans_exx)); + } + } + if (gs_is_hybrid(dft_functional)) + { + for (int is : {0, 1}) + { + this->ops[(is << 1) + is]->add(new OperatorLREXX(nspin, naos, nocc[is], nvirt[is], + ucell, this->psi_ks_spin_[is], *this->DM_diff, exx_lri, kv, pX[is], pc, pmat, + /*cal_force=*/true, -2.0 * exx_alpha, ATYPE_EXX::CC_vo, {}, hamilt::calculation_type::lr_dmdiff_exx)); + } + } +#endif + } + + protected: + void band_extra(const T* const X, T* const out) const override + { + if (this->gxc_) { this->gxc_->act(X, out); } + } + + /// Rebuild BOTH density matrices from the `is` block of X: the CXC operators read + /// $D^X$ (transposed, un-symmetrized -- see the closed-shell `Z_vector_R` for why), + /// the CC_vo operators read the difference density matrix $T$ (symmetrized). + void set_dm(const int is, const T* const X) override + { + const auto psi_ks_is = LR_Util::get_psi_spin(this->psi_ks_, is, this->nk); +#ifdef __MPI + this->dmx_buf_ = cal_dm_trans_pblas(X, this->pX[is], psi_ks_is, this->pc_, + this->naos_, this->nocc[is], this->nvirt[is], this->pmat_); + for (auto& t : this->dmx_buf_) { LR_Util::mattrans(t.template data(), this->naos_, this->pmat_); } + this->dmd_buf_ = cal_dm_diff_pblas(X, this->pX[is], psi_ks_is, this->pc_, + this->naos_, this->nocc[is], this->nvirt[is], this->pmat_); + for (auto& t : this->dmd_buf_) { LR_Util::matsym(t.template data(), this->naos_, this->pmat_); } +#else + this->dmx_buf_ = cal_dm_trans_blas(X, psi_ks_is, this->nocc[is], this->nvirt[is]); + for (auto& t : this->dmx_buf_) + { + T* d = t.template data(); + for (int u = 0;u < this->naos_;++u) + { + for (int v = u + 1;v < this->naos_;++v) { std::swap(d[u * this->naos_ + v], d[v * this->naos_ + u]); } + } + } + this->dmd_buf_ = cal_dm_diff_blas(X, psi_ks_is, this->naos_, this->nocc[is], this->nvirt[is]); + for (auto& t : this->dmd_buf_) { LR_Util::matsym(t.template data(), this->naos_); } +#endif + for (int ik = 0;ik < this->nk;++ik) + { + this->DM_trans->set_dmk_ptr(ik, this->dmx_buf_[ik].template data()); + this->DM_diff->set_dmk_ptr(ik, this->dmd_buf_[ik].template data()); + } + } + + private: + const int naos_ = 1; + const Parallel_2D& pc_; + const Parallel_Orbitals& pmat_; + const psi::Psi& psi_ks_; + std::vector> psi_ks_spin_; + std::unique_ptr> DM_trans; + std::unique_ptr> DM_diff; + std::vector dmx_buf_; + std::vector dmd_buf_; + std::unique_ptr> gxc_; + }; +} + +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_HAMILT_ZEQ_R_H diff --git a/source/source_lcao/module_lr/hamilt_zequlr.h b/source/source_lcao/module_lr/hamilt_zequlr.h new file mode 100644 index 00000000000..717f7acf31e --- /dev/null +++ b/source/source_lcao/module_lr/hamilt_zequlr.h @@ -0,0 +1,168 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_HAMILT_ZEQULR_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_HAMILT_ZEQULR_H +#include "source_hamilt/hamilt.h" +#include "source_lcao/module_lr/operator_casida/operator_lr_hxc.h" +#include "source_basis/module_ao/parallel_orbitals.h" +#include "source_lcao/module_lr/utils/lr_util.h" +#include +#include +#include +#include +#include +#include + +namespace LR +{ + /// @brief Common skeleton of the open-shell (spin-unrestricted) Z-vector operators. + /// + /// Both sides of the Z-vector equation have the same block structure as the open-shell + /// Casida Hamiltonian `HamiltULR`: the vector is the concatenation [up-block | down-block] + /// and the operator is a 2x2 array of spin blocks + /// ops[(s_out << 1) + s_in], + /// because every kernel action carries a free outer spin and a summed inner spin, + /// $K_{pq\sigma}[D]=\sum_{\kappa\lambda\sigma'}K_{pq\sigma,\kappa\lambda\sigma'}D_{\kappa\lambda\sigma'}$. + /// The exchange part is $\propto\delta_{\sigma\sigma'}$, so EXX operators only ever go on + /// the diagonal blocks 0 and 3. + /// + /// Derived classes fill `ops` in their constructor and implement `set_dm`, which rebuilds + /// whatever density matrices the operators read from the `s_in` block of X. + template + class ZeqULR + { + public: + ZeqULR(const std::vector& nocc_in, + const std::vector& nvirt_in, + const std::vector& pX_in, + const int nk_in) + : nocc(nocc_in), nvirt(nvirt_in), pX(pX_in), nk(nk_in), + ldim(nk_in* (pX_in[0].get_local_size() + pX_in[1].get_local_size())), + gdim(nk_in* (nocc_in[0] * nvirt_in[0] + nocc_in[1] * nvirt_in[1])), + ops(4, nullptr) + { + } + virtual ~ZeqULR() { for (auto& op : this->ops) { delete op; } } + + void hPsi(const T* const psi_in, T* const hpsi, const int ld_psi, const int nband) + { + assert(ld_psi == this->ldim); + const std::vector ldim_is = { static_cast(nk * pX[0].get_local_size()), static_cast(nk * pX[1].get_local_size()) }; + for (int ib = 0;ib < nband;++ib) + { + const int offset_band = ib * ld_psi; + // All four spin blocks accumulate into this band's output, and the diagonal + // `OperatorLRDiag`s now add rather than assign (see `operator_lr_diag.h`), so the + // buffer has to start clean. + std::fill(hpsi + offset_band, hpsi + offset_band + ld_psi, T(0)); + // Terms that are not bilinear in a (out, in) spin pair -- currently only the + // $g^{xc}$ term of the right-hand side, which is quadratic in $D^X$ and needs + // both transition-density channels on the grid at once. + this->band_extra(psi_in + offset_band, hpsi + offset_band); + for (int is_in : {0, 1}) + { + const int offset_in = offset_band + is_in * ldim_is[0]; + this->set_dm(is_in, psi_in + offset_in); + typename OperatorLRHxc::TransitionDensityCache density_cache; + for (int is_out : {0, 1}) + { + const int offset_out = offset_band + is_out * ldim_is[0]; + hamilt::Operator* node(this->ops[(is_out << 1) + is_in]); + while (node != nullptr) + { + // `is_in` picks the density matrix (already done by `set_dm`), but the + // vector an operator consumes directly belongs to the OUT channel: + // $R_{ia\sigma}=-2\sum_{\sigma'}\sum_b X_{ib\sigma} + // K_{ab,\sigma\sigma'}[D^X_{\sigma'}]+\dots$ + // -- only $D^X$ carries the summed spin. `OperatorLRHxc`'s CXC branch + // says the same thing in code: it reads `psi_in` through `pX[sl]`, the + // OUT channel's distribution. + const T* input = psi_in + offset_out; + T* output = hpsi + offset_out; + act_with_shared_density(node, ldim_is[is_out], input, output, density_cache); + node = (hamilt::Operator*)(node->next_op); + } + } + } + } + } + + /// @brief The full (replicated) matrix, column by column. Only used by the LAPACK solver. + std::vector matrix() + { + ModuleBase::TITLE("ZeqULR", "matrix"); + const std::vector npairs = { nocc[0] * nvirt[0], nocc[1] * nvirt[1] }; + const std::vector ldim_is = { static_cast(nk * pX[0].get_local_size()), static_cast(nk * pX[1].get_local_size()) }; + const std::vector gdim_is = { nk * npairs[0], nk * npairs[1] }; + std::vector mat_full(static_cast(gdim) * gdim, T(0)); + for (int is_in : {0, 1}) + { + const auto& px = this->pX[is_in]; + const int loffset_in = is_in * ldim_is[0]; + const int goffset_in = is_in * gdim_is[0]; + for (int ik_in = 0;ik_in < nk;++ik_in) + { + for (int j = 0;j < nocc[is_in];++j) + { + for (int b = 0;b < nvirt[is_in];++b) + { + const int gcol = goffset_in + ik_in * npairs[is_in] + j * nvirt[is_in] + b; + std::vector X_col(this->ldim, T(0)); + const int lj = px.global2local_col(j); + const int lb = px.global2local_row(b); + if (px.in_this_processor(b, j)) + { + X_col[loffset_in + ik_in * px.get_local_size() + lj * px.get_row_size() + lb] = T(1); + } + this->set_dm(is_in, X_col.data() + loffset_in); + std::vector col(this->ldim, T(0)); + for (int is_out : {0, 1}) + { + const int loffset_out = is_out * ldim_is[0]; + const int goffset_out = is_out * gdim_is[0]; + const auto& pax = this->pX[is_out]; + hamilt::Operator* node(this->ops[(is_out << 1) + is_in]); + while (node != nullptr) + { + node->act(1, ldim_is[is_in], /*npol=*/1, + X_col.data() + loffset_in, col.data() + loffset_out); + node = (hamilt::Operator*)(node->next_op); + } +#ifdef __MPI + for (int ik_out = 0;ik_out < this->nk;++ik_out) + { + LR_Util::gather_2d_to_full(pax, + col.data() + loffset_out + ik_out * pax.get_local_size(), + mat_full.data() + static_cast(gcol) * gdim + + goffset_out + ik_out * npairs[is_out], + false, nvirt[is_out], nocc[is_out]); + } +#else + std::memcpy(mat_full.data() + static_cast(gcol) * gdim + goffset_out, + col.data() + loffset_out, gdim_is[is_out] * sizeof(T)); +#endif + } + } + } + } + } + return mat_full; + } + + const std::vector& nocc; + const std::vector& nvirt; + const std::vector& pX; + const int nk = 1; + const int ldim = 1; + const int gdim = 1; + + protected: + /// @brief rebuild the density matrices the operators read, from the `is` block of X + virtual void set_dm(const int is, const T* const X) = 0; + /// @brief optional per-band contribution that does not fit the 2x2 block structure. + /// NOTE `matrix()` deliberately does NOT call this: it exists for the right-hand side, + /// which is not a linear operator, while `matrix()` is only ever used for the LHS. + virtual void band_extra(const T* const X, T* const out) const {} + std::vector*> ops; + }; +} + +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_HAMILT_ZEQULR_H diff --git a/source/source_lcao/module_lr/hsolver_lrtd.hpp b/source/source_lcao/module_lr/hsolver_lrtd.hpp index 55880f0d294..40d95e2b126 100644 --- a/source/source_lcao/module_lr/hsolver_lrtd.hpp +++ b/source/source_lcao/module_lr/hsolver_lrtd.hpp @@ -38,9 +38,11 @@ namespace LR template inline void print_eigs(const std::vector& eigs, const std::string& label = "", const double factor = 1.0) { - std::cout << label << std::endl; + std::streamsize old = std::cout.precision(); + std::cout << label << std::setprecision(8) << std::endl; for (auto& e : eigs) { std::cout << e * factor << " "; } std::cout << std::endl; + std::cout.precision(old); } /// eigensolver for common Hamilt @@ -60,6 +62,7 @@ namespace LR const bool hermitian = true) { ModuleBase::TITLE("HSolverLR", "solve"); + ModuleBase::timer::start("HSolverLR", "solve"); const std::vector spin_types = { "singlet", "triplet" }; // note: if not TDA, the eigenvalues will be complex // then we will need a new constructor of DiagoDavid @@ -178,6 +181,7 @@ namespace LR // output iters std::cout << " Average iterative diagonalization steps: " << hsolver::DiagoIterAssist::avg_iter << "; current threshold: " << diag_ethr << std::endl; + ModuleBase::timer::end("HSolverLR", "solve"); } } } diff --git a/source/source_lcao/module_lr/lr_amp.h b/source/source_lcao/module_lr/lr_amp.h new file mode 100644 index 00000000000..6313180207f --- /dev/null +++ b/source/source_lcao/module_lr/lr_amp.h @@ -0,0 +1,149 @@ +#ifndef ABACUS_LR_GRADIENT_AMPLITUDES_H +#define ABACUS_LR_GRADIENT_AMPLITUDES_H +#include "utils/lr_util.h" +#include "source_base/parallel_reduce.h" +#include +#include +namespace LR +{ +// Open-shell eigenvectors store all up-spin k blocks before all down-spin k blocks. +inline int electron_hole_offset(const int state, const int nk, const int up_size, + const int down_size, const int ispin, const bool openshell) +{ + if (openshell) { return state * nk * (up_size + down_size) + ispin * nk * up_size; } + const int channel_size = ispin == 0 ? up_size : down_size; + return state * nk * channel_size; +} + +template +void save_mixed_root(const T* Xall, const int n, const std::vector& group, + const std::vector& mixing, std::vector& previous) +{ + if (group.size() != mixing.size()) + { + throw std::invalid_argument("JT mixing must contain one coefficient per root"); + } + previous.assign(n, T(0)); + for (size_t k = 0; k < group.size(); ++k) + { + for (int i = 0; i < n; ++i) { previous[i] += mixing[k] * Xall[group[k] * n + i]; } + } + double norm_squared = 0.0; + for (const T& value : previous) { norm_squared += std::real(LR_Util::get_conj(value) * value); } + Parallel_Reduce::reduce_all(norm_squared); + if (norm_squared <= 0.0) { throw std::runtime_error("JT reference has zero norm"); } + const double scale = 1.0 / std::sqrt(norm_squared); + for (T& value : previous) { value *= scale; } +} + +template +void follow_root(const T* Xall, const int n, const int nstates, const int initial_state, + int& target_state, std::vector& previous, std::ostream& ofs) +{ + + if (target_state < 0) { target_state = initial_state; } + + // Ranks with no local pairs still participate in the global overlap collective. + int reference_size = previous.size(); + Parallel_Reduce::reduce_all(reference_size); + if (reference_size > 0) + { + std::vector ov(nstates, T(0)); + for (int j = 0; j < nstates; ++j) + { + const T* const Xj = Xall + j * n; + T acc = T(0); + for (int i = 0; i < n; ++i) { acc += LR_Util::get_conj(previous[i]) * Xj[i]; } + ov[j] = acc; + } + // X is distributed over the same 2D grid as the particle-hole pairs, so the inner + // product is only complete after summing over that grid. Reduce the accumulators + // themselves (real and imaginary parts alike) and take the modulus afterwards -- + // reducing |partial| would be wrong. `reduce_all` is the guarded wrapper and is a no-op + // in a serial build, so no `#ifdef __MPI` is needed around it. + Parallel_Reduce::reduce_all(ov.data(), nstates); + int best = 0; + double best_ov = -1.0; + for (int j = 0; j < nstates; ++j) + { + const double a = std::abs(ov[j]); + if (a > best_ov) { best_ov = a; best = j; } + } + + if (best != target_state) + { + ofs << " EXCITED-STATE RELAX: followed root moved from index " + << target_state << " to " << best << " (overlap " << best_ov + << "); the states crossed and the index no longer names the same state." + << std::endl; + } + // A low best overlap means no current root resembles the one being followed -- the step + // was too large, or the state left the solved window. Say so: the relaxation continues + // but the surface it follows is no longer guaranteed continuous. + if (best_ov < 0.5) + { + ofs << " WARNING: largest amplitude overlap with the previous step is" + " only " << best_ov << ". The followed state may have left the window spanned by" + " lr_nstates; consider raising lr_nstates or reducing the ionic step." << std::endl; + } + target_state = best; + } + + if (n > 0) { previous.assign(Xall + target_state * n, Xall + (target_state + 1) * n); } + else { previous.clear(); } +} + +template +ct::Tensor pad_amplitudes(const std::vector& X, const bool openshell, + const std::vector& px, const std::vector& px_z, + const std::vector& nocc, const std::vector& nvirt, + const int nk, const int nloc, const int nloc_z, + const int ispin, const int istate_begin, const int nst) +{ + ct::Tensor Xz = LR_Util::newTensor({ nst, nloc_z }); + Xz.zero(); + // Closed shell: one channel, and `ispin` selects the spin COMBINATION (singlet/triplet) + // whose X is being widened -- `nocc`/`nvirt`/`paraX_` are the same for both. + // Open shell: one eigenvector holding both channels back to back, so both sub-blocks are + // widened and the second one is re-based, because widening the first moves where it starts. + const std::vector chan = openshell ? std::vector{ 0, 1 } + : std::vector{ ispin }; + const int ix = openshell ? 0 : ispin; // which entry of `X` + int soff_ch = 0; + int zoff_ch = 0; // channel offsets inside one state's block + for (const int is : chan) + { + const Parallel_2D& pxs = px[is]; + const Parallel_2D& pxz = px_z[is]; + for (int ist = 0;ist < nst;++ist) + { + const T* const src = X[ix].template data() + + (istate_begin + ist) * nloc + soff_ch; + T* const dst = Xz.data() + ist * nloc_z + zoff_ch; + for (int ik = 0;ik < nk;++ik) + { + const int soff = ik * pxs.get_local_size(); + const int zoff = ik * pxz.get_local_size(); + // Both windows are block-cyclic with nb2d = 1 on the same process grid and share + // the occupied dimension, so a global (virt, occ) element lives on the same + // process in both -- only its local index differs. The widening is therefore a + // purely local copy. + for (int o = 0;o < nocc[is];++o) + { + for (int v = 0;v < nvirt[is];++v) + { + if (!pxs.in_this_processor(v, o)) { continue; } + dst[zoff + pxz.global2local_col(o) * pxz.get_row_size() + pxz.global2local_row(v)] + = src[soff + pxs.global2local_col(o) * pxs.get_row_size() + pxs.global2local_row(v)]; + } + } + } + } + soff_ch += nk * pxs.get_local_size(); + zoff_ch += nk * pxz.get_local_size(); + } + return Xz; +} + +} +#endif diff --git a/source/source_lcao/module_lr/lr_density.hpp b/source/source_lcao/module_lr/lr_density.hpp new file mode 100644 index 00000000000..fb6cfb25d65 --- /dev/null +++ b/source/source_lcao/module_lr/lr_density.hpp @@ -0,0 +1,123 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_LR_DENSITY_HPP +#define ABACUS_SOURCE_LCAO_MODULE_LR_LR_DENSITY_HPP +#include "source_hamilt/module_gint/gint_interface.h" +#include "source_psi/psi.h" +#include "source_lcao/module_lr/dm_trans/dm_diff.h" +#include "source_lcao/module_lr/utils/lr_util_hcontainer.h" +#include "source_io/module_output/cube_io.h" +#include "lr_amp.h" +namespace LR +{ + template + class LR_Density + { + const UnitCell& ucell_; + const K_Vectors& kv_; + const Grid_Driver& gd_; + const psi::Psi& psi_ks_; + const std::vector& orb_cutoff_; + const Parallel_Grid& pgrid_; + const int nspin_; + const std::vector& nocc_; + const std::vector& nvirt_; + const int nao_; + const int nk_; + const std::vector& pX_; + const Parallel_2D& pc_; + const Parallel_Orbitals& pmat_; + const bool openshell_; + const std::vector spintype_; + /// cached aliases, read once here instead of at every call site below + const int out_chg_precision_ = PARAM.inp.out_chg[1]; + const std::string global_out_dir_ = PARAM.globalv.global_out_dir; + + inline void dm_to_density(module_dm::DensityMatrix& dm, double** density) + { + ModuleBase::TITLE("LR_Density", "dm_to_density"); + ModuleGint::cal_gint_rho(dm.get_dmr_vec(), 1, density, false); + } + inline void dm_to_density(module_dm::DensityMatrix, std::complex>& dm, double** density) + { + ModuleBase::TITLE("LR_Density", "dm_to_density"); + auto dm_to_density_real = [&](const char& part) -> void + { + module_dm::DensityMatrix, double> dm_real(&pmat_, 1, kv_.kvec_d, nk_); + LR_Util::initialize_DMR, double>(dm_real, pmat_, ucell_, gd_, orb_cutoff_); + LR_Util::get_DMR_real_imag_part(dm, dm_real, part); + ModuleGint::cal_gint_rho(dm_real.get_dmr_vec(), 1, density, false); // add-on + }; + dm_to_density_real('R'); + dm_to_density_real('I'); + } + + public: + LR_Density(const UnitCell& ucell, + const K_Vectors& kv, + const Grid_Driver& gd, + const psi::Psi& psi_ks, + const std::vector& orb_cutoff, + const Parallel_Grid& pgrid, + const int& nspin, + const std::vector& nocc, + const std::vector& nvirt, + const int& nao, + const std::vector& pX, + const Parallel_2D& pc, + const Parallel_Orbitals& pmat, + const bool openshell = false) : + ucell_(ucell), kv_(kv), gd_(gd), psi_ks_(psi_ks), + orb_cutoff_(orb_cutoff), pgrid_(pgrid), nspin_(nspin), nk_(kv.get_nks() / nspin), + nocc_(nocc), nvirt_(nvirt), nao_(nao), pX_(pX), pc_(pc), pmat_(pmat), + openshell_(openshell), spintype_(openshell ? std::vector({ "up", "down" }) : std::vector({ "singlet", "triplet" })) + { + } + + /// @brief calculate the electron density from the density matrix in 2d-block distribution + void cal_eh_density_single_state(const T* const X_istate, const int ispin, double** density) + { + ModuleBase::TITLE("LR_Density", "cal_eh_density_single_state"); + ModuleBase::GlobalFunc::ZEROS(density[0], this->pgrid_.get_nrxx()); + // 1. calculate the density matrix in AO basis + auto c_spin = LR_Util::get_psi_spin(psi_ks_, ispin,nk_); +#ifdef __MPI + const std::vector dm_diff_k = + cal_dm_diff_pblas(X_istate, pX_[ispin], c_spin, pc_, nao_, nocc_[ispin], nvirt_[ispin], pmat_); +#else + const std::vector dm_diff_k = + cal_dm_diff_blas(X_istate, c_spin, nao_, nocc_[ispin], nvirt_[ispin]); +#endif + // 2. calculate DM(R) + module_dm::DensityMatrix dm_diff= + LR_Util::build_dm_from_dmk(dm_diff_k, + this->pmat_, this->nk_, this->kv_.kvec_d, this->ucell_, this->gd_, this->orb_cutoff_); + // 3. calculate electron density from DM(R) + this->dm_to_density(dm_diff, density); + } + + void write_density_single_state(const double* const* const density, const std::string& filepath) + { + ModuleIO::write_vdata_palgrid(pgrid_, density[0], 0, 1, 0, filepath, 0.0, &ucell_, this->out_chg_precision_, 0, false, true); + } + + void output_eh_density_all_states(const T* const X, const int ispin, const int nstate) + { + ModuleBase::TITLE("LR_Density", "cal_eh_density_all_states"); + const int up_size = this->pX_[0].get_local_size(); + const int down_channel = openshell_ ? 1 : ispin; + const int down_size = this->pX_[down_channel].get_local_size(); + double** density; + LR_Util::_allocate_2order_nested_ptr(density, 1, pgrid_.get_nrxx()); + for (int istate = 0;istate < nstate;++istate) + { + const int offset = LR::electron_hole_offset(istate, nk_, up_size, down_size, ispin, openshell_); + const T* const amplitudes = X + offset; + this->cal_eh_density_single_state(amplitudes, ispin, density); + const std::string filepath = this->global_out_dir_ + "LR_e-h_density_" + spintype_[ispin] + "_" + std::to_string(istate + 1) + ".cube"; + this->write_density_single_state(density, filepath); + } + LR_Util::_deallocate_2order_nested_ptr(density, 1); + } + }; + +} +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_LR_DENSITY_HPP diff --git a/source/source_lcao/module_lr/lr_force.cpp b/source/source_lcao/module_lr/lr_force.cpp new file mode 100644 index 00000000000..dd75a016d36 --- /dev/null +++ b/source/source_lcao/module_lr/lr_force.cpp @@ -0,0 +1,380 @@ +#include "lr_force.h" +#include "cal_hs_grad.h" +#include "pulay_hc.h" +#include "source_lcao/pulay_fs.h" // only for gint terms +#include "source_hamilt/module_gint/gint_interface.h" +#include "source_lcao/module_lr/utils/lr_util.h" +// #include "source_lcao/module_lr/utils/lr_util_hcontainer.h" +namespace LR +{ + template + ModuleBase::matrix LR_Force::cal_force_overlap_edm(const module_dm::DensityMatrix& edm) + { + // const double* dS[3] = { dSloc_x, dSloc_y, dSloc_z }; + std::vector> dS = cal_hs_grad('S', this->ucell_, this->pv_, this->gd_, this->two_center_bundle_); + // test: output dS + // std::cout << "dS in 3 directions:\n"; + // for (int i = 0;i < 3;++i) { LR_Util::print_HR(dS.at(i), this->ucell_.nat, "dS" + std::to_string(i)); } + ModuleBase::matrix foverlap = PulayForceStress::cal_pulay_fs(edm, this->ucell_, dS, -1.); + if (this->test_force_) + { + ModuleIO::print_force(this->ofs_running_, this->ucell_, "OVERLAP FORCE (eV/Angstrom)", foverlap, false); + } + return foverlap; + } + + /// The LR density matrices ($D^X$, $T+D^Z$, EDM) carry one channel in the closed-shell + /// singlet/triplet algorithm and two independent channels in the open-shell one. + template + inline bool is_openshell_dm(const module_dm::DensityMatrix& dm) + { + return dm.get_dmr_vec().size() == 2; + } + + template + void LR_Force::dm_to_charge(const module_dm::DensityMatrix& dm, Charge& chr_out) + { + const int& nspin_dm = dm.get_dmr_vec().size(); + const int& nspin_global = this->nspin_; + chr_out.set_rhopw(const_cast(&this->rhopw_)); + chr_out.allocate(nspin_global, /*kin_den=*/false, /*meta_gga=*/false, /*test_charge=*/0); //chr still needs global nspin, because Forces (PW) depends on it + // So huge a Charge class... + // 1. Using a (private) `allocate_rho` to control whether to delete will definately cause memory leak here. No need for such judgement. + // 2. Charge-dependent interfaces need refactor: only rhopw_ and rho dependence are enough. + + ModuleGint::cal_gint_rho(dm.get_dmr_vec(), nspin_dm, chr_out.rho, false); + // if (nspin_dm == 1 && nspin_global == 2), chr_out.rho[1][irxx]=0 has been set in Charge::allocate() + } + + template + std::unique_ptr LR_Force::dm_to_hxc_potential(const module_dm::DensityMatrix& dm) + { + std::unique_ptr pot(new elecstate::Potential(&this->rhodpw_, &this->rhopw_, &this->ucell_, + &this->locpp_.vloc, const_cast(&this->sf_), + nullptr/*surchem*/, &this->etxc_, &this->vtxc_)); + if (this->vh_in_h_) { pot->pot_register({ "hartree", "xc" }); } + else { pot->pot_register({ "xc" }); } + Charge charge; + this->dm_to_charge(dm, charge); + pot->init_pot(&charge); // call update_from_charge inside + return pot; + } + + template + std::unique_ptr LR_Force::local_potential() + { + std::unique_ptr pot(new elecstate::Potential(&this->rhodpw_, &this->rhopw_, &this->ucell_, + &this->locpp_.vloc, const_cast(&this->sf_), + nullptr/*surchem*/, nullptr/*etxc*/, nullptr/*vtxc*/)); + pot->pot_register({ "local" }); + pot->init_pot(nullptr); + return pot; + } + + template + ModuleBase::matrix LR_Force::cal_force_hamilt_gs_dm_relaxed_diff(const module_dm::DensityMatrix& relax_diff_dm, + const module_dm::DensityMatrix& dm_gs, + const bool reproduce_gs, + const PotHxcLR* pot_hxc_gs) + { + const bool with_ewald = reproduce_gs; + // Two independent spin channels in `relax_diff_dm` <=> open-shell (spin-unrestricted) LR. + // The closed-shell singlet/triplet algorithm always builds a single-channel LR density + // matrix, even at nspin=2. + const bool openshell = is_openshell_dm(relax_diff_dm); + Charge chr_diff_relaxed; + this->dm_to_charge(relax_diff_dm, chr_diff_relaxed); + + // 1. local pp (Hellmann-Feynman)(fvl_dvl) + ewald + core correction (+ self-consistent charge) + ModuleBase::matrix f_pw = this->vl_in_h_ ? + ForcePWTerms()(this->ucell_, chr_diff_relaxed, this->rhopw_, this->locpp_, this->sf_, + this->nspin_, this->test_force_, this->ofs_running_, with_ewald) : + ModuleBase::matrix(this->ucell_.nat, 3); + + // 2. nonlocal pp (Hellmann-Feynman + Pulay) + ModuleBase::matrix fvnl = cal_force_nonlocal(this->ucell_, this->kvec_d_, this->gd_, this->two_center_bundle_, relax_diff_dm); + + // // 3. local pp (Pulay) + Hartree + xc (grid integration) + // ModuleBase::matrix fvl_dphi(this->ucell_.nat, 3); + // ModuleBase::matrix stress_tmp; // no use now, only for passing into interfaces + // PulayForceStress::cal_pulay_fs(relax_diff_dm.get_dmr_vec().size()/*nspin*/, fvl_dphi, stress_tmp, + // relax_diff_dm, this->ucell_, &pot_gs, true, false); + + // 3.1. local pp (Pulay) + ModuleBase::matrix fvl_dphi(this->ucell_.nat, 3); + ModuleBase::matrix stress_tmp; // no use now, only for passing into interfaces + std::unique_ptr pot_loc = this->local_potential(); + PulayForceStress::cal_pulay_fs(relax_diff_dm.get_dmr_vec().size()/*nspin*/, fvl_dphi, stress_tmp, + relax_diff_dm, this->ucell_, pot_loc.get(), true, false); + // the grid-based `cal_pulay_fs` only sums this rank's share of the real-space grid; + // core ABACUS always reduces right after it (see force_stress_lcao.cpp). + Parallel_Reduce::reduce_pool(fvl_dphi.c, fvl_dphi.nr * fvl_dphi.nc); + + // 3.2. Hartree + xc (Pulay) + // method 1 + // ModuleBase::matrix fgs_dphi(this->ucell_.nat, 3); + // PulayForceStress::cal_pulay_fs(relax_diff_dm.get_dmr_vec().size()/*nspin*/, fgs_dphi, stress_tmp, + // relax_diff_dm, this->ucell_, &pot_gs, true, false); + // ModuleBase::matrix fhxc_dphi = (fgs_dphi - fvl_dphi) * 0.5; // avoid double count of hxc Pulay term + // method 2 + ModuleBase::matrix fhxc_dphi(this->ucell_.nat, 3); + std::unique_ptr pot_hxc = this->dm_to_hxc_potential(dm_gs); + // `cal_pulay_fs` calculates 1*Pulay-term. + // For ground-state DFT, Pulay term = Hellmann-Feynman term, F = 1/2(Pulay + H-F) = Pulay, so directly call it once gives correct result. + PulayForceStress::cal_pulay_fs(relax_diff_dm.get_dmr_vec().size()/*nspin*/, fhxc_dphi, stress_tmp, + relax_diff_dm, this->ucell_, pot_hxc.get(), true, false); + Parallel_Reduce::reduce_pool(fhxc_dphi.c, fhxc_dphi.nr * fhxc_dphi.nc); // see `fvl_dphi` above + if (reproduce_gs) {fhxc_dphi *= 0.5;} // avoid double count + + // 3.3 Hartree + xc (Hellmann-Feynman) + ModuleBase::matrix fhxc_dvhxc(this->ucell_.nat, 3); + // The potential here must be the *linear response* of $V^\text{Hxc}$ to the difference + // density, i.e. $v_H[\rho^{T+Z}] + f_{xc}[\rho^\text{gs}]\,\rho^{T+Z}$ -- NOT + // $v_\text{Hxc}[\rho^{T+Z}]$. This term is + // $\sum_{\kappa\lambda}(T{+}D^Z)_{\kappa\lambda}\int\phi_\kappa\phi_\lambda\, + // f_{xc}\sum_{\alpha\beta}D^\text{gs}_{\alpha\beta}(\phi_\alpha\phi_\beta)^x$, + // the half of $\partial_x V^\text{Hxc}$ whose basis derivative falls on the *ground-state* + // pair. Hartree is linear in the density so feeding it $\rho^{T+Z}$ happens to be right; + // xc is not -- $\rho^{T+Z}$ is not even positive everywhere, while LDA has + // $v_{xc}\propto-\rho^{1/3}$. + // + // `pot_hxc_gs` supplies exactly this object (Hartree weight 1, xc = $(f_{uu}+f_{ud})/2$ at + // nspin=2), with no extra factor. Verified on H2/SZ TDRPA@LDA, where $K^T\equiv0$ makes the + // triplet gradient identical to $d(\varepsilon_a-\varepsilon_i)/dx$: analytic 28.5982 vs the + // KS-gap finite difference 28.598156. It used to be off by -3.5 eV/Ang. + // + // `reproduce_gs` is the exception: there `relax_diff_dm` *is* the ground-state density + // matrix and the term being checked is the true ground-state force, for which + // $v_\text{Hxc}[\rho^\text{gs}]$ is the correct potential. + if (reproduce_gs || pot_hxc_gs == nullptr) + { + std::unique_ptr pot_hxc_relaxed_diff = this->dm_to_hxc_potential(relax_diff_dm); + //`cal_pulay_fs` calculates only one spin channel because `relax_diff_dm` has only one. + PulayForceStress::cal_pulay_fs(1/*nspin*/, fhxc_dvhxc, stress_tmp, + dm_gs, this->ucell_, pot_hxc_relaxed_diff.get(), true, false); + fhxc_dvhxc *= gs_dm_channel_factor(this->nspin_); + } + else if (openshell) + { + // $v_\sigma=\sum_{\sigma'}f^{\sigma\sigma'}\rho^{T+Z}_{\sigma'}$ (+ the full Hartree), + // contracted with $D^\text{gs}_\sigma$. No factor 2 and no `gs_dm_channel_factor`: + // the two ground-state channels are summed explicitly by `cal_gint_fvl` below, and + // each carries occupation 1 rather than 2. + constexpr int nspin_dm = 2; + std::vector v_lin(nspin_dm, ModuleBase::matrix(1, this->rhopw_.nrxx)); + for (int sl = 0; sl < nspin_dm; ++sl) + { + for (int sr = 0; sr < nspin_dm; ++sr) + { + double* rho_in[1] = { const_cast(chr_diff_relaxed.rho[sr]) }; + pot_hxc_gs->cal_v_eff(rho_in, this->ucell_, v_lin[sl], { sl, sr }); + } + } + std::vector vr_eff(nspin_dm); + for (int is = 0; is < nspin_dm; ++is) { vr_eff[is] = v_lin[is].c; } + ModuleGint::cal_gint_fvl(nspin_dm, vr_eff, dm_gs.get_dmr_vec(), true, false, &fhxc_dvhxc, &stress_tmp); + } + else + { + ModuleBase::matrix v_lin(1, this->rhopw_.nrxx); // zero-initialized + double* rho_in[1] = { const_cast(chr_diff_relaxed.rho[0]) }; + pot_hxc_gs->cal_v_eff(rho_in, this->ucell_, v_lin); + std::vector vr_eff = { v_lin.c }; + ModuleGint::cal_gint_fvl(1, vr_eff, dm_gs.get_dmr_vec(), true, false, &fhxc_dvhxc, &stress_tmp); + fhxc_dvhxc *= 2; // for the two channels of the ground-state dm. + fhxc_dvhxc *= gs_dm_channel_factor(this->nspin_); + } + // all three branches above compute `fhxc_dvhxc` via `cal_gint_fvl`/the grid-based + // `cal_pulay_fs`, neither of which reduces internally (see `fvl_dphi` above). + Parallel_Reduce::reduce_pool(fhxc_dvhxc.c, fhxc_dvhxc.nr * fhxc_dvhxc.nc); + + // 4. kinetic (Pulay) + std::vector> dT = cal_hs_grad('T', this->ucell_, this->pv_, this->gd_, this->two_center_bundle_); + ModuleBase::matrix ft_dphi = PulayForceStress::cal_pulay_fs(relax_diff_dm, this->ucell_, dT); + + if (this->test_force_) + { + ModuleIO::print_force(this->ofs_running_, this->ucell_, "PW FORCE (eV/Angstrom)", f_pw, false); + ModuleIO::print_force(this->ofs_running_, this->ucell_, "NONLOCAL FORCE (eV/Angstrom)", fvnl, false); + ModuleIO::print_force(this->ofs_running_, this->ucell_, "KINETIC FORCE (eV/Angstrom)", ft_dphi, false); + ModuleIO::print_force(this->ofs_running_, this->ucell_, "LOCAL-PP Pulay FORCE (eV/Angstrom)", fvl_dphi, false); + ModuleIO::print_force(this->ofs_running_, this->ucell_, "HARTREE+XC Pulay FORCE (eV/Angstrom)", fhxc_dphi, false); + ModuleIO::print_force(this->ofs_running_, this->ucell_, "HARTREE+XC Hellmann-Feynman FORCE (eV/Angstrom)", fhxc_dvhxc, false); + } + + // from the formula, we do not need the non-ortho term (overlap*edm) here. + return f_pw + fvnl + ft_dphi + fvl_dphi + fhxc_dphi + fhxc_dvhxc; + } + template + ModuleBase::matrix LR_Force::cal_force_hxc_dmtrans(const module_dm::DensityMatrix& dm_trans, const PotHxcLR& pot_hxc) + { + // `dm_trans` (D^X) must be SYMMETRIZED before entering here: `cal_pulay_fs` builds v from + // rho[D^X] (which only sees the symmetric part) but contracts with D^X as passed, so an + // un-symmetrized D^X makes the two slots of the bilinear form Tr[D^X d(K_H)[D^X]] disagree. + // + // `cal_pulay_fs` returns 2 * sum_{mn} D_{mn} \int (d phi_m) v phi_n, where the factor 2 is + // `cal_gint_fvl`'s internal m<->n doubling, i.e. it is exactly the *bra-pair* derivative (Pulay). + // The *ket-pair* derivative (Hellmann-Feynman) is equal to it (both slots hold the same D^X), + // so the total needs one more factor 2 (Pulay -> Pulay + Hellmann-Feynman). + // Verified on H2/SZ against the analytic 4-center derivative: 4*sum D^sym P = 16.8653548 + // vs 2*d(ai|ia)/dz = 16.865355 eV/Ang (7 digits). + const double pulay_to_total_sym = 2.0; + // Open shell: the kernel couples the channels, so the density has to be built per channel + // and the potential accumulated over the summed spin before contracting with $D^X_\sigma$. + // `pulay_to_total_sym` is a Pulay -> Pulay+Hellmann-Feynman factor and is spin-independent. + if (is_openshell_dm(dm_trans)) + { + return PulayForceStress::cal_pulay_fs_openshell(dm_trans, this->ucell_, &pot_hxc) * pulay_to_total_sym; + } + return PulayForceStress::cal_pulay_fs(dm_trans, this->ucell_, &pot_hxc) * pulay_to_total_sym; + } + + template + ModuleBase::matrix LR_Force::cal_force_gxc_dmtrans(const module_dm::DensityMatrix& dm_trans, + const module_dm::DensityMatrix& dm_gs, const PotGradXCLR& pot_grad) + { + // The third source of position dependence in + // $f^{xc}_{\kappa\lambda,\alpha\beta} + // =\int\phi_\kappa\phi_\lambda\,f_{xc}[\rho^\text{gs}(r)]\,\phi_\alpha\phi_\beta$. + // `cal_force_hxc_dmtrans` above differentiates the two basis pairs of $D^X$, which for the + // Hartree kernel $(\kappa\lambda|\alpha\beta)$ is everything. The xc kernel additionally + // depends on the nuclear positions through $\rho^\text{gs}$ itself, and that derivative is + // $\int\rho^X g^{xc}\rho^X\,\partial_x\rho^\text{gs}|_\text{basis} + // =\int v^{(2)}[\rho^X,\rho^X]\,\partial_x\rho^\text{gs}|_\text{basis}$, + // i.e. the Pulay derivative of the *ground-state* density against the $g^{xc}$ potential. + // + // No spin factor: this is the sibling of `cal_force_hxc_dmtrans` above, which likewise has + // none, because `PotGradXCLR` (like `pot[ispin]` there) already carries the S2_singlet / + // S2_triplet spin combination. The superficially similar Hellmann-Feynman half of + // `cal_force_hamilt_gs_dm_relaxed_diff` DOES need a factor 2, but only because it is built + // from `pot_hxc_gs`, which is normalized as S2_gs = S2_singlet/2. + // Confirmed numerically on H2/SZ TDLDA (see `cal_multiplier_w_from_z.h`). + Charge chr_x; + this->dm_to_charge(dm_trans, chr_x); + ModuleBase::matrix v2(1, this->rhopw_.nrxx); // zero-initialized + double* rho_in[1] = { const_cast(chr_x.rho[0]) }; + pot_grad.cal_v_eff(rho_in, this->ucell_, v2); + + ModuleBase::matrix f(this->ucell_.nat, 3); + ModuleBase::matrix stress_tmp; + std::vector vr_eff = { v2.c }; + ModuleGint::cal_gint_fvl(1, vr_eff, dm_gs.get_dmr_vec(), true, false, &f, &stress_tmp); + Parallel_Reduce::reduce_pool(f.c, f.nr * f.nc); // see `fvl_dphi` in cal_force_hamilt_gs_dm_relaxed_diff + f *= gs_dm_channel_factor(this->nspin_); + return f; + } + + template + ModuleBase::matrix LR_Force::cal_force_gxc_dmtrans_openshell( + const module_dm::DensityMatrix& dm_trans, + const module_dm::DensityMatrix& dm_gs, const PotGradXCLR& pot_grad) + { + // Same term as `cal_force_gxc_dmtrans`, spin-resolved: the free index $\tau$ (spin channel) + // of $v^{(2)}_\tau$ is contracted with $D^\text{gs}_\tau$, and `cal_gint_fvl` does the + // $\sum_\tau$. No `gs_dm_channel_factor` here -- both ground-state channels are summed + // explicitly and each carries occupation 1. + constexpr int nspin_dm = 2; + assert(dm_trans.get_dmr_vec().size() == nspin_dm); + Charge chr_x; + this->dm_to_charge(dm_trans, chr_x); + const double* rho_in[nspin_dm] = { chr_x.rho[0], chr_x.rho[1] }; + + std::vector v2(nspin_dm, ModuleBase::matrix(1, this->rhopw_.nrxx)); + for (int tau = 0; tau < nspin_dm; ++tau) { pot_grad.cal_v_eff_openshell(rho_in, this->ucell_, v2[tau], tau); } + + ModuleBase::matrix f(this->ucell_.nat, 3); + ModuleBase::matrix stress_tmp; + std::vector vr_eff(nspin_dm); + for (int is = 0; is < nspin_dm; ++is) { vr_eff[is] = v2[is].c; } + ModuleGint::cal_gint_fvl(nspin_dm, vr_eff, dm_gs.get_dmr_vec(), true, false, &f, &stress_tmp); + Parallel_Reduce::reduce_pool(f.c, f.nr * f.nc); // see `fvl_dphi` in cal_force_hamilt_gs_dm_relaxed_diff + return f; + } + +#ifdef __EXX + template + ModuleBase::matrix LR_Force::cal_force_exx_dm_trans( + const std::map>>& dm_trans, + const double& alpha, + const std::string& spin_suffix) + { + ModuleBase::matrix f_exx_dmtrans(this->ucell_.nat, 3); + auto& exx_lri_kernel = this->exx_lri_.lock()->get(); + exx_lri_kernel.set_Ds(dm_trans, this->exx_lri_.lock()->get_info().dm_threshold, spin_suffix); + exx_lri_kernel.cal_Hs({ "", "", spin_suffix }); + exx_lri_kernel.cal_force({ "", "", spin_suffix, "", "" });// using dm_trans,Pulay term only + for (std::size_t idim = 0; idim < 3; ++idim) + for (const auto& force_item : exx_lri_kernel.force[idim]) + f_exx_dmtrans(force_item.first, idim) = std::real(force_item.second); + const double fac = -2.0 * alpha; //-2 is the same as post_process_Hexx, Hartree to Ry (which didn't act on Hs) + const double pulay_to_total_sym = 2.0; // Pulay -> Pulay + Hellmann-Feynman, only when Ds_left and Ds_right are equal + return f_exx_dmtrans * fac * pulay_to_total_sym; // dm_trans (DX) already contain the spin channel (sqrt(2) times of up/down channel DX) + // return f_exx_dmtrans * fac * 2; // 2 is the same in post_process_Eexx at nspin=1 ( up->up + down->down, 2 spin-conserving transitions) + } + + template + ModuleBase::matrix LR_Force::cal_force_exx_gs_dm_relaxed_diff( + const std::map>>& dm_gs, + const std::map>>& relaxed_diff_dm, + const double& alpha, + const std::string& spin_suffix) + { + ModuleBase::matrix f_exx_gs_diff(this->ucell_.nat, 3); + auto& exx_lri_kernel = this->exx_lri_.lock()->get(); + ExxForceTwoDM lr_exx_kernel(std::move(exx_lri_kernel)); + + auto add_force_from_kernel = [&]() { + for (std::size_t idim = 0; idim < 3; ++idim) + for (const auto& force_item : lr_exx_kernel.force[idim]) + f_exx_gs_diff(force_item.first, idim) += std::real(force_item.second); + }; + + auto transpose_dm = [](const std::map>>& dm) + -> std::map>> + { + std::map>> dm_transpose; + for (const auto& pair0 : dm) + for (const auto& pair1 : pair0.second) + { + const int& iat0 = pair0.first; + const int& iat1 = pair1.first.first; + const auto& R = pair1.first.second; + dm_transpose[iat1][{iat0, { -R[0], -R[1], -R[2] }}] = pair1.second.transpose(); + } + return dm_transpose; + }; + + // `cal_force` calculates Pulay term. + // If D_IJ = D_KL(H - F = Pulay), it caluclates 0.5 * d(ik | jl). + // Multiply spin factor (outside) on it gives the final result. + // The spin factor is not hard-coded in this function. + + // 1. Pulay term + lr_exx_kernel.set_Ds(dm_gs, this->exx_lri_.lock()->get_info().dm_threshold, spin_suffix); + lr_exx_kernel.cal_Hs({ "", "", spin_suffix }); // using dm_gs as D_KL + lr_exx_kernel.cal_force(relaxed_diff_dm, { "", "", spin_suffix , "", "" }); // using relaxed_diff_dm as D_IJ + add_force_from_kernel(); + + // 2. Hellmann-Feynman term + const auto& dm_gs_transpose = transpose_dm(dm_gs); + const auto& relaxed_diff_dm_transpose = transpose_dm(relaxed_diff_dm); + lr_exx_kernel.set_Ds(relaxed_diff_dm_transpose, this->exx_lri_.lock()->get_info().dm_threshold, spin_suffix); + lr_exx_kernel.cal_Hs({ "", "", spin_suffix }); // using relaxed_diff_dm as D_KL + lr_exx_kernel.cal_force(dm_gs_transpose, { "", "", spin_suffix , "", "" }); // using dm_gs as D_IJ + add_force_from_kernel(); + + // move back + exx_lri_kernel = std::move(lr_exx_kernel); + + // -2 * 0.5 * alpha + // -2 is the same as post_process_Hexx (a.u. to Ry, which didn't act on Hs) + // 0.5 is the 2-electron integral prefactor,used in ground-state energy/force where two density matrix are identical + // But the LR-grad Lagrangian/force 2-e term here Tr[(T+Z)H[D]] does not have 1/2 factor. + const double fac = -2 * alpha; + return f_exx_gs_diff * fac; + } +#endif +} + +template class LR::LR_Force; +template class LR::LR_Force>; \ No newline at end of file diff --git a/source/source_lcao/module_lr/lr_force.h b/source/source_lcao/module_lr/lr_force.h new file mode 100644 index 00000000000..de614399b41 --- /dev/null +++ b/source/source_lcao/module_lr/lr_force.h @@ -0,0 +1,126 @@ +#ifndef ABACUS_LR_FORCE_H +#define ABACUS_LR_FORCE_H +#include +#include "force_funcs.h" +#include "source_lcao/module_lr/potentials/pot_hxc_lrtd.h" +#include "source_lcao/module_lr/potentials/pot_grad_xc.h" +// free functions, usefull for both ground and excited state +#ifdef __EXX +#include "source_lcao/module_ri/exx_lri.h" +#include "exx_force_dm.h" +using TAC = std::pair>; +#endif +namespace LR +{ + /// `dm_gs` carries the ground-state occupations, at nspin=1 they are 2 for fully occupied + /// bands, so any term that contracts against ONE channel of `dm_gs` needs this factor. + /// Closed-shell bookkeeping only: the open-shell path sums the two channels explicitly and + /// must not apply it. + inline double gs_dm_channel_factor(const int& nspin) { return (nspin == 1) ? 0.5 : 1.0; } + + template + class LR_Force + { + public: + LR_Force(const UnitCell& ucell, + const std::vector>& kvec_d, + const Parallel_Orbitals& pv, + const ModulePW::PW_Basis& rhodpw, + const ModulePW::PW_Basis& rhopw, + const pseudopot_cell_vl& locpp, + const Structure_Factor& sf, + const Grid_Driver& gd, + const TwoCenterBundle& two_center_bundle +#ifdef __EXX + , std::weak_ptr> exx_lri_in, + const double& alpha +#endif + )///< for 2-center integrals + : ucell_(ucell), kvec_d_(kvec_d), pv_(pv), + rhodpw_(rhodpw), rhopw_(rhopw), sf_(sf), locpp_(locpp), gd_(gd), + two_center_bundle_(two_center_bundle) +#ifdef __EXX + , exx_lri_(exx_lri_in), alpha_(alpha) +#endif + { + } + + /// 1. $Tr[H_{GS}^x * (T+D^Z)]$, where GS=groud state and $(T+D^Z)$ is the relaxed difference density matrix + ModuleBase::matrix cal_force_hamilt_gs_dm_relaxed_diff(const module_dm::DensityMatrix& relaxed_diff_dm, + const module_dm::DensityMatrix& dm_gs, const bool reproduce_gs = false, + const PotHxcLR* pot_hxc_gs = nullptr); + + /// 2. $Tr[S^x * (EDM)] + ModuleBase::matrix cal_force_overlap_edm(const module_dm::DensityMatrix& edm); + + /// 3. $\sum_{mnkl}(mn|f_{Hxc}|kl)^x *D^X *D^X$ + ModuleBase::matrix cal_force_hxc_dmtrans(const module_dm::DensityMatrix& dm_trans, const PotHxcLR& pot_hxc); + + /// 3b. the $g^{xc}$ half of $\partial_x K^{S/T}[D^X]D^X$: + /// $\int v^{(2)}[\rho^X,\rho^X](r)\,\partial_x\rho^\text{gs}(r)|_\text{basis}$ + ModuleBase::matrix cal_force_gxc_dmtrans(const module_dm::DensityMatrix& dm_trans, + const module_dm::DensityMatrix& dm_gs, const PotGradXCLR& pot_grad); + + /// 3b'. open-shell version: $\sum_\tau\int v^{(2)}_\tau[\rho^X,\rho^X]\, + /// \partial_x\rho^\text{gs}_\tau|_\text{basis}$. Both transition-density channels + /// enter each $v^{(2)}_\tau$, so this cannot be a per-channel loop over the above. + ModuleBase::matrix cal_force_gxc_dmtrans_openshell(const module_dm::DensityMatrix& dm_trans, + const module_dm::DensityMatrix& dm_gs, const PotGradXCLR& pot_grad); + +#ifdef __EXX + // auto* lrexx_ptr = dynamic_cast, 3, TK>*>(&exx_lri_in.get()); + /// 4. $\alpha \sum_{mnkl}(mk|nl)^x *D^X *D^X$ + ModuleBase::matrix cal_force_exx_dm_trans( + const std::map>>& dm_trans, + const double& alpha, + const std::string& spin_suffix = ""); + ModuleBase::matrix cal_force_exx_gs_dm_relaxed_diff( + const std::map>>& dm_gs, + const std::map>>& relaxed_diff_dm, + const double& alpha, + const std::string& spin_suffix = ""); +#endif + + // test functions + /// reproduce the force of the ground state + ModuleBase::matrix reproduce_force_gs(const K_Vectors& kv, + const module_dm::DensityMatrix& dm_gs, + const module_dm::DensityMatrix& edm_gs); + + /// repreduce the ground state local term + ModuleBase::matrix reproduce_force_gs_loc(const module_dm::DensityMatrix& dm_gs, + const elecstate::Potential& pot_gs); + + protected: + const UnitCell& ucell_; + const std::vector>& kvec_d_; + const Parallel_Orbitals& pv_; + const ModulePW::PW_Basis& rhodpw_; + const ModulePW::PW_Basis& rhopw_; + const pseudopot_cell_vl& locpp_; + const Structure_Factor& sf_; + const Grid_Driver& gd_; + const TwoCenterBundle& two_center_bundle_; +#ifdef __EXX + std::weak_ptr> exx_lri_; + const double alpha_; +#endif + /// cached aliases, read once here instead of at every log/flag check below + std::ofstream& ofs_running_ = GlobalV::ofs_running; + const int nspin_ = PARAM.inp.nspin; + const bool test_force_ = PARAM.inp.test_force; + const bool vl_in_h_ = PARAM.inp.vl_in_h; + const bool vh_in_h_ = PARAM.inp.vh_in_h; + const std::string dft_functional_ = PARAM.inp.dft_functional; + + // PotXC borrows these buffers; they outlive every local potential created here. + double etxc_ = 0.0; + double vtxc_ = 0.0; + + void dm_to_charge(const module_dm::DensityMatrix& dm, Charge& chr_out); + std::unique_ptr dm_to_hxc_potential(const module_dm::DensityMatrix& dm); + std::unique_ptr local_potential(); + }; +} + +#endif diff --git a/source/source_lcao/module_lr/lr_force_aux.cpp b/source/source_lcao/module_lr/lr_force_aux.cpp new file mode 100644 index 00000000000..c1cec07acff --- /dev/null +++ b/source/source_lcao/module_lr/lr_force_aux.cpp @@ -0,0 +1,59 @@ +#include "lr_force.h" +#include "source_lcao/pulay_fs.h" +#include "source_lcao/module_lr/utils/lr_util_hcontainer.h" +#include "source_io/module_output/output_log.h" +#ifdef __EXX +#include "operator_casida/operator_lr_exx.h" // gs_is_hybrid +#endif +namespace LR +{ + template + ModuleBase::matrix LR_Force::reproduce_force_gs(const K_Vectors& kv, + const module_dm::DensityMatrix& dm_gs, + const module_dm::DensityMatrix& edm_gs) + { + // local + Hartree + xc term, including Hellmann-Feynman and Pulay + ModuleBase::matrix f_gs_hf_pulay = cal_force_hamilt_gs_dm_relaxed_diff(dm_gs, dm_gs, true); // pw(vl_dvl+ewald)+vnl+t_dphi+vl_dphi + // edm term + ModuleBase::matrix f_nonortho = cal_force_overlap_edm(edm_gs); // overlap +#ifdef __EXX + if (gs_is_hybrid(this->dft_functional_)) + { + const auto& Ds_gs = LR_Util::get_exx_Ds_gs(dm_gs, ucell_, kv, pv_); + const auto& Ds_gs_2 = LR_Util::get_exx_Ds_gs(dm_gs, ucell_, kv, pv_); + // at nspin=1 there is only one channel, and `get_exx_Ds_gs` already returns 0.5*D + // (= D_up), so both halves below reuse channel 0. + const int is_2nd = (Ds_gs.size() > 1) ? 1 : 0; + ModuleBase::matrix f_gs_exx(ucell_.nat, 3); + // test the two function using the two spin channels respectively + // 0.5 is from dE = 0.5 dTr[D(HD)]. No 0.5 in excited-state calculateion of dTr[(T+Z)(HD)] + f_gs_exx += cal_force_exx_gs_dm_relaxed_diff(Ds_gs.at(0), Ds_gs_2.at(0), alpha_, std::to_string(0)) * 0.5; // test passed, = 0.5 groud-state EXX force + f_gs_exx += cal_force_exx_dm_trans(Ds_gs.at(is_2nd), alpha_, std::to_string(is_2nd)) * 0.5; + if (this->test_force_) + ModuleIO::print_force(this->ofs_running_, ucell_, "EXX GS FORCE reproduce (eV/Angstrom)", f_gs_exx, false); + f_gs_hf_pulay += f_gs_exx; + } +#endif + return f_gs_hf_pulay + f_nonortho; + } + + template + ModuleBase::matrix LR_Force::reproduce_force_gs_loc( + const module_dm::DensityMatrix& dm_gs, + const elecstate::Potential& pot_gs) + { + Charge chr_gs; + this->dm_to_charge(dm_gs, chr_gs); + // local pp (Pulay) + Hartree + xc (grid integration) + ModuleBase::matrix fvl_dphi(this->ucell_.nat, 3); + ModuleBase::matrix stress_tmp; // no use now, only for passing into interfaces + PulayForceStress::cal_pulay_fs(dm_gs.get_dmr_vec().size()/*nspin*/, fvl_dphi, stress_tmp, + dm_gs, this->ucell_, &pot_gs, true, false); + Parallel_Reduce::reduce_pool(fvl_dphi.c, fvl_dphi.nr * fvl_dphi.nc); // see lr_force.cpp's `fvl_dphi` + return fvl_dphi; + } + +} + +template class LR::LR_Force; +template class LR::LR_Force>; diff --git a/source/source_lcao/module_lr/lr_grad_cs.cpp b/source/source_lcao/module_lr/lr_grad_cs.cpp new file mode 100644 index 00000000000..70142778b2b --- /dev/null +++ b/source/source_lcao/module_lr/lr_grad_cs.cpp @@ -0,0 +1,245 @@ +#include "gradient_inputs.h" +#include "gradient_output.h" +#include "gradient_checks.h" +#include "cal_edm.h" +#include "dm_trans/dm_diff.h" +#include "grad_degen.h" +#include "source_base/timer.h" +#include "source_io/module_output/output_log.h" +namespace LR +{ +template +std::vector evaluate_closed_shell_force( + const GradientInputs& inputs, LR_Force& lr_force, + const module_dm::DensityMatrix& dm_gs, + const ct::Tensor& Xz, const ct::Tensor& Z, + const std::vector& omega, const int label_begin, const int ispin) +{ + ModuleBase::timer::start("LR", "evaluate_closed_shell_force"); + // for each block, calculate dm_trans, dm_relaxed_diff, edm and force + const int nst = static_cast(omega.size()); + assert(static_cast(Xz.shape().dim_size(0)) == nst); + // `ist_begin`/`ist_end` label the blocks in the output only; the gradient itself never looks + // up a state, it only uses `omega[i]`. That is what lets a caller pass excitation vectors + // that are not the stored eigenvectors (see `cal_grad_matrix_degenerate`). + const int ist_begin = label_begin; + const int ist_end = label_begin + nst; + const std::vector& nvirt_g = inputs.nvirt; + const std::vector& paraX_g = inputs.px; + const int nloc_g = inputs.nloc; + + + ModuleBase::TITLE("ESolver_LR", "cal_force"); + // Spin channel 0, NOT `ispin`. `ispin` indexes `spin_types` = {singlet, triplet}. + // Closed shell always use spin-up channel of psi_ks, i.e. psi_ks(0). + const auto& c = LR_Util::get_psi_spin(inputs.psi_ks, 0, inputs.nk); // wavefunction coefficients of ground state + + // calculate the force (the partial gradient of Lagrangian) + inputs.ofs << "Start to calculate excited-state force of " << inputs.spin_types[ispin] << std::endl; + // ground state dm for currrent spin (only for test the correctness of the force) + // module_dm::DensityMatrix dm_gs(inputs.pmat, 1, inputs.kv.kvec_d, inputs.nk); + + std::vector forces(ist_end - ist_begin); + for (int istate = ist_begin;istate < ist_end;++istate) + { + const int offset = (istate - ist_begin) * nloc_g; // block of X (widened into the Z window) + const int zoffset = offset; // block of Z + // The imag part will be cancelled in the force calculation, so we use double DM(R) to calculate force. + // But complex transition DM(R) is still used in energy density matrix calculation. +#ifdef __MPI + const auto& dm_trans_k = cal_dm_trans_pblas(Xz.data() + offset, paraX_g[ispin], c, inputs.pc, inputs.nbasis, inputs.nocc[ispin], nvirt_g[ispin], inputs.pmat); +#else + const auto& dm_trans_k = cal_dm_trans_blas(Xz.data() + offset, c, inputs.nocc[ispin], nvirt_g[ispin]); +#endif + // D(X) complex, for the EXX (LibRI) force. Built FIRST and left UN-symmetrized: + // the exchange kernel (mu kappa | nu lambda) puts the two indices of one D^X into + // different electron coordinates, so Tr[D^X D^X K_exx] = (aa|ii) requires the full + // non-symmetric D^X. Symmetrizing would give 1/2[(aa|ii)+(ai|ia)], which is wrong. + // (`cal_force_exx_dm_trans` feeds the same tensor to both slots, so it is consistent.) + auto dm_trans = // D(X) complex + LR_Util::build_dm_from_dmk(dm_trans_k, + inputs.pmat, inputs.nk, inputs.kv.kvec_d, inputs.ucell, inputs.gd, inputs.orb_cutoff); + LR_Util::transpose_DMR(dm_trans, inputs.pmat); + // D(X) real, for the grid Hxc force. The Coulomb kernel (mu nu | kappa lambda) is + // symmetric within each index pair, so it only ever sees the symmetric part of D^X. + // In `PulayForceStress::cal_pulay_fs`, `cal_gint_rho` (which builds v) symmetrizes + // implicitly, while `cal_gint_fvl`'s internal factor 2 assumes D_{mu nu} = D_{nu mu}. + // Passing an un-symmetrized D^X makes the two slots of the bilinear form disagree. + // NOTE: `build_dm_from_dmk` symmetrizes `dm_trans_k` IN PLACE, hence the ordering. + auto dm_trans_real = // D(X), double (FIXME: not enough for periodic system!) + LR_Util::build_dm_from_dmk(dm_trans_k, + inputs.pmat, inputs.nk, inputs.kv.kvec_d, inputs.ucell, inputs.gd, inputs.orb_cutoff, + /*symmetrize=*/true); + LR_Util::transpose_DMR(dm_trans_real, inputs.pmat); + // LR_Util::print_DMR(dm_trans, "dm_trans of istate " + std::to_string(istate)); + // difference density matrix +#ifdef __MPI + std::vector dm_diff_k = cal_dm_diff_pblas(Xz.data() + offset, paraX_g[ispin], c, inputs.pc, inputs.nbasis, inputs.nocc[ispin], nvirt_g[ispin], inputs.pmat); +#else + std::vector dm_diff_k = cal_dm_diff_blas(Xz.data() + offset, c, inputs.nbasis, inputs.nocc[ispin], nvirt_g[ispin]); +#endif + // std::cout << "dm_diff_k T(k) before symmetrization, istate " + std::to_string(istate) << std::endl; + // LR_Util::print_value(dm_diff_k[0].data(), inputs.pmat.get_col_size(), inputs.pmat.get_row_size()); + // for (auto& d : dm_diff_k) { LR_Util::matsym(d.data(), inputs.nbasis, inputs.pmat); } // symmetrize + // std::cout << "dm_diff_k T(k) after symmetrization, istate " + std::to_string(istate) << std::endl; + // LR_Util::print_value(dm_diff_k[0].data(), inputs.pmat.get_col_size(), inputs.pmat.get_row_size()); + +#ifdef __MPI + const std::vector& dm_relaxed_k = cal_dm_trans_pblas(Z.template data() + zoffset, paraX_g[ispin], c, inputs.pc, inputs.nbasis, inputs.nocc[ispin], nvirt_g[ispin], inputs.pmat); +#else + const std::vector& dm_relaxed_k = cal_dm_trans_blas(Z.template data() + zoffset, c, inputs.nocc[ispin], nvirt_g[ispin]); +#endif + // std::cout << "dm_relaxed_k Z(k) before symmetrization, istate " + std::to_string(istate) << std::endl; + // LR_Util::print_value(dm_relaxed_k[0].data(), inputs.pmat.get_col_size(), inputs.pmat.get_row_size()); +#ifdef __MPI + for (auto& d : dm_relaxed_k) { LR_Util::matsym(d.data(), inputs.nbasis, inputs.pmat); } // symmetrize +#else + for (auto& d : dm_relaxed_k) { LR_Util::matsym(d.data(), inputs.nbasis); } // symmetrize +#endif + // std::cout << "dm_relaxed_k Z(k) after symmetrization, istate " + std::to_string(istate) << std::endl; + // LR_Util::print_value(dm_relaxed_k[0].data(), inputs.pmat.get_col_size(), inputs.pmat.get_row_size()); + // relaxed difference density matrix + const std::vector& relaxed_diff_dm_k = dm_diff_k + dm_relaxed_k; + const module_dm::DensityMatrix& diff_dm = + LR_Util::build_dm_from_dmk(dm_diff_k, + inputs.pmat, inputs.nk, inputs.kv.kvec_d, inputs.ucell, inputs.gd, inputs.orb_cutoff); + const module_dm::DensityMatrix& relaxed_diff_dm = + LR_Util::build_dm_from_dmk(relaxed_diff_dm_k, + inputs.pmat, inputs.nk, inputs.kv.kvec_d, inputs.ucell, inputs.gd, inputs.orb_cutoff); + // LR_Util::print_DMR(relaxed_diff_dm, "relaxed_diff_dm T+Z (Z symmetrized) of istate " + std::to_string(istate)); + + // module_dm::DensityMatrix relaxed_diff_dm = // T+D(Z), (R) can be complex + // LR_Util::build_dm_from_dmk( + // // LR_Util::operator+( + // cal_dm_diff_pb las(Xz.data() + offset, paraX_g[ispin], c, inputs.pc, inputs.nbasis, inputs.nocc[ispin], nvirt_g[ispin], inputs.pmat) + // + cal_dm_trans_pblas(Z.template data() + offset, paraX_g[ispin], c, inputs.pc, inputs.nbasis, inputs.nocc[ispin], nvirt_g[ispin], inputs.pmat) + // ,// ), + // inputs.pmat, inputs.nk, inputs.kv.kvec_d, inputs.ucell, inputs.gd, inputs.orb_cutoff); + // LR_Util::print_DMR(relaxed_diff_dm, "relaxed_diff_dm of istate " + std::to_string(istate)); + module_dm::DensityMatrix relaxed_diff_dm_real(&inputs.pmat, 1, inputs.kv.kvec_d, inputs.nk); + LR_Util::initialize_DMR(relaxed_diff_dm_real, inputs.pmat, inputs.ucell, inputs.gd, inputs.orb_cutoff); + LR_Util::get_DMR_real_imag_part(relaxed_diff_dm, relaxed_diff_dm_real, 'R'); + + // get edm of type DensityMatrix + // weak_ptr here is to avoid "could not match 'weak_ptr' against 'shared_ptr'" + // but why there're no bug in the previous code (HamiltLR and HamiltULF)? + // seems because those two are classes having constructors + // but `cal_edm_from_XZ_istate` here is a functions + std::weak_ptr pot_weak = inputs.pot[ispin]; + const T* const x_istate = Xz.data() + offset; + const T* const z_istate = Z.template data() + zoffset; + const double omega_istate = omega[istate - ist_begin]; + const std::vector& edm_k = cal_edm_from_XZ_istate( + inputs, x_istate, z_istate, omega_istate, inputs.eig_ks.c, + dm_trans, c, pot_weak, inputs.spin_types[ispin]); + if (inputs.test_force && inputs.nocc[0] == 1 && nvirt_g[0] == 1) + { +#ifdef __MPI + const std::vector& dm_diff = cal_dm_diff_pblas(Xz.data() + offset, paraX_g[0], c, inputs.pc, inputs.nbasis, inputs.nocc[0], nvirt_g[0], inputs.pmat); +#else + const std::vector& dm_diff = cal_dm_diff_blas(Xz.data() + offset, c, inputs.nbasis, inputs.nocc[0], nvirt_g[0]); +#endif + // test_dm_diff_H2(relaxed_diff_dm.get_dmk_ptr(0), c, inputs.nbasis); + test_dm_diff_H2(dm_diff[0].data(), c, inputs.nbasis); + test_edm_H2(edm_k[0].data(), inputs.eig_ks.c, c, inputs.nbasis); + } + module_dm::DensityMatrix edm_real = LR_Util::build_dm_from_dmk(edm_k, + inputs.pmat, inputs.nk, inputs.kv.kvec_d, inputs.ucell, inputs.gd, inputs.orb_cutoff, /*symmetrize=*/true); + // print edm_real (R) + if (inputs.test_force) + { + LR_Util::save_DMR(edm_real, "data-EDMR-sparse" + std::string(inputs.excited_relax ? "_state" + std::to_string(istate) : ""), inputs.pmat, inputs.out_dir, inputs.nbasis, inputs.my_rank); + // LR_Util::print_DMR(edm_real, "edm_real (R) of istate " + std::to_string(istate)); + } + + ModuleBase::matrix force_hxc_dmtrans = lr_force.cal_force_hxc_dmtrans(dm_trans_real, *inputs.pot[ispin]); + if (inputs.test_force) + ModuleIO::print_force(inputs.ofs, inputs.ucell, "HXC DMTRANS FORCE (eV/Angstrom)", force_hxc_dmtrans, false); + + + // the $g^{xc}$ half of $\partial_x K[D^X]D^X$, i.e. the derivative of the xc kernel through + // the ground-state density (see `cal_force_gxc_dmtrans`). Only for local kernels. + if (LR_Util::has_local_xc(inputs.xc_kernel)) + { + // The density dependence of K[D^X]D^X belongs to the LR functional. + const std::shared_ptr& pot_lr = inputs.pot[ispin]; + const bool triplet = inputs.spin_types[ispin] == "triplet"; + PotGradXCLR pot_grad(pot_lr->xc_kernel_components(), pot_lr->get_rho_basis(), + inputs.ucell, pot_lr->nrxx, triplet); + ModuleBase::matrix force_gxc_dmtrans = lr_force.cal_force_gxc_dmtrans(dm_trans_real, dm_gs, pot_grad); + if (inputs.test_force) + ModuleIO::print_force(inputs.ofs, inputs.ucell, "GXC DMTRANS FORCE (eV/Angstrom)", force_gxc_dmtrans, false); + force_hxc_dmtrans += force_gxc_dmtrans; + } + ModuleBase::matrix force_hamiltgs_relaxed_diff = lr_force.cal_force_hamilt_gs_dm_relaxed_diff(relaxed_diff_dm_real, dm_gs, false, inputs.pot_hxc_gs.get()); + if (inputs.test_force) + ModuleIO::print_force(inputs.ofs, inputs.ucell, "H_GS-(T+Z) FORCE (without EXX) (eV/Angstrom)", force_hamiltgs_relaxed_diff, false); + + ModuleBase::matrix force_overlap_edm = lr_force.cal_force_overlap_edm(edm_real); // "-" sign has been included in the force factor + if (inputs.test_force) + ModuleIO::print_force(inputs.ofs, inputs.ucell, "OVERLAP-EDM FORCE (eV/Angstrom)", force_overlap_edm, false); + + if (inputs.test_force) + { + // test H[T] force (Z=0), non-EXX part + module_dm::DensityMatrix diff_dm_real(&inputs.pmat, 1, inputs.kv.kvec_d, inputs.nk); + LR_Util::initialize_DMR(diff_dm_real, inputs.pmat, inputs.ucell, inputs.gd, inputs.orb_cutoff); + LR_Util::get_DMR_real_imag_part(diff_dm, diff_dm_real, 'R'); + + inputs.ofs << "========== [TEST H_GS-(T) force (Z=0), non-EXX part] ===========" << std::endl; + ModuleBase::matrix force_hamiltgs_diff = lr_force.cal_force_hamilt_gs_dm_relaxed_diff(diff_dm_real, dm_gs); + ModuleIO::print_force(inputs.ofs, inputs.ucell, "H_GS-T FORCE (without EXX) (eV/Angstrom)", force_hamiltgs_diff, false); + inputs.ofs << "========== [\\TEST H_GS-(T) force (Z=0), non-EXX part] ===========" << std::endl; + } + + +#ifdef __EXX + const double& alpha = inputs.hybrid_alpha; + + if (LR::exx_kernel_list().count(inputs.xc_kernel)) + { + const auto& Ds_trans = LR_Util::get_exx_Ds_spin1(dm_trans, inputs.ucell, inputs.kv, inputs.pmat); + ModuleBase::matrix force_exx_dmtrans = lr_force.cal_force_exx_dm_trans(Ds_trans, alpha * 4.0); // cancel the two 0.5s in Ds + if (inputs.test_force) + ModuleIO::print_force(inputs.ofs, inputs.ucell, "EXX DMTRANS FORCE (eV/Angstrom)", force_exx_dmtrans, false); + force_hxc_dmtrans += force_exx_dmtrans; + + } + + if (LR::gs_is_hybrid(inputs.dft_functional)) + { + const auto& Ds_gs = LR_Util::get_exx_Ds_spin1(dm_gs, inputs.ucell, inputs.kv, inputs.pmat); // returns 0.5*D[0] + const auto& Ds_relaxed_diff = LR_Util::get_exx_Ds_spin1(relaxed_diff_dm, inputs.ucell, inputs.kv, inputs.pmat); // returns 0.5*D[0] + // LR_Util::print_CV(Ds_relaxed_diff, "Ds_relaxed_diff for EXX force"); + // `get_exx_Ds_spin1` feeds `split_m2D_ktoR(..., nspin=1)`, which reads only channel 0 + // with a 0.5 prefactor. For `dm_gs` that channel is $D^\text{gs}_\uparrow$ at nspin=2 + // but the spin-summed $D^\text{gs}$ at nspin=1, i.e. twice as large. + ModuleBase::matrix force_exx_gs_relaxed_diff = lr_force.cal_force_exx_gs_dm_relaxed_diff(Ds_gs, Ds_relaxed_diff, alpha * 4.0) * gs_dm_channel_factor(inputs.nspin); // cancel the two 0.5s in Ds + if (inputs.test_force) + ModuleIO::print_force(inputs.ofs, inputs.ucell, "EXX GS-(T+Z) FORCE (eV/Angstrom)", force_exx_gs_relaxed_diff, false); + force_hamiltgs_relaxed_diff += force_exx_gs_relaxed_diff; + + if (inputs.test_force) + { + // test H[T] force (Z=0), EXX part + const auto& Ds_diff = LR_Util::get_exx_Ds_spin1(diff_dm, inputs.ucell, inputs.kv, inputs.pmat); // returns 0.5*D[0] + inputs.ofs << "========== [TEST H_GS-(T) force (Z=0), EXX part] ===========" << std::endl; + ModuleBase::matrix force_exx_gs_diff = lr_force.cal_force_exx_gs_dm_relaxed_diff(Ds_gs, Ds_diff, alpha * 4.0) * gs_dm_channel_factor(inputs.nspin); // cancel the two 0.5s in Ds + ModuleIO::print_force(inputs.ofs, inputs.ucell, "H_GS-T EXX FORCE (Z=0) (eV/Angstrom)", force_exx_gs_diff, false); + inputs.ofs << "========== [\\TEST H_GS-(T) force (Z=0), EXX part] ===========" << std::endl; + } + } +#endif + forces[istate - ist_begin] = force_hxc_dmtrans + force_hamiltgs_relaxed_diff + force_overlap_edm; + } + // total force + print_lr_force(forces, std::cout, ist_begin); + print_lr_force(forces, inputs.ofs, ist_begin); + ModuleBase::timer::end("LR", "evaluate_closed_shell_force"); + return forces; +} + + +template std::vector evaluate_closed_shell_force(const GradientInputs&, LR_Force&, const module_dm::DensityMatrix&, const ct::Tensor&, const ct::Tensor&, const std::vector&, int, int); +template std::vector evaluate_closed_shell_force>(const GradientInputs>&, LR_Force>&, const module_dm::DensityMatrix, double>&, const ct::Tensor&, const ct::Tensor&, const std::vector&, int, int); +} diff --git a/source/source_lcao/module_lr/lr_grad_os.cpp b/source/source_lcao/module_lr/lr_grad_os.cpp new file mode 100644 index 00000000000..2fc847e53f0 --- /dev/null +++ b/source/source_lcao/module_lr/lr_grad_os.cpp @@ -0,0 +1,176 @@ +#include "gradient_inputs.h" +#include "gradient_output.h" +#include "gradient_checks.h" +#include "cal_edm.h" +#include "dm_trans/dm_diff.h" +#include "grad_degen.h" +#include "source_base/timer.h" +#include "source_io/module_output/output_log.h" +namespace LR +{ +template +std::vector evaluate_open_shell_force( + const GradientInputs& inputs, LR_Force& lr_force, + const module_dm::DensityMatrix& dm_gs, + const ct::Tensor& Xz, const ct::Tensor& Z, + const std::vector& omega, const int label_begin) +{ + ModuleBase::timer::start("LR", "evaluate_open_shell_force"); + const int nst = static_cast(omega.size()); + assert(static_cast(Xz.shape().dim_size(0)) == nst); + const int ist_begin_ = label_begin; + const int ist_end_ = label_begin + nst; + const std::vector& nvirt_g = inputs.nvirt; + const std::vector& paraX_g = inputs.px; + const int nloc_g = inputs.nloc; + + + const std::vector ld_x = { static_cast(inputs.nk * paraX_g[0].get_local_size()), + static_cast(inputs.nk * paraX_g[1].get_local_size()) }; + const std::vector off_x = { 0, ld_x[0] }; + std::vector> c_spin; + for (int is : {0, 1}) { c_spin.push_back(LR_Util::get_psi_spin(inputs.psi_ks, is, inputs.nk)); } + + inputs.ofs << "Start to calculate excited-state force of updown (open shell)" << std::endl; + + const int ist_begin = ist_begin_; + const int ist_end = ist_end_; + std::vector forces(ist_end - ist_begin); + for (int istate = ist_begin;istate < ist_end;++istate) + { + const int offset = (istate - ist_begin) * nloc_g; // X widened into the Z window + const T* const X_istate = Xz.data() + offset; + const T* const Z_istate = Z.template data() + offset; + + // 1. the k-space blocks of each spin channel + std::vector> dmx_k(2), dmdiff_k(2), relaxed_k(2); + for (int is : {0, 1}) + { +#ifdef __MPI + dmx_k[is] = cal_dm_trans_pblas(X_istate + off_x[is], paraX_g[is], c_spin[is], inputs.pc, + inputs.nbasis, inputs.nocc[is], nvirt_g[is], inputs.pmat); + dmdiff_k[is] = cal_dm_diff_pblas(X_istate + off_x[is], paraX_g[is], c_spin[is], inputs.pc, + inputs.nbasis, inputs.nocc[is], nvirt_g[is], inputs.pmat); + std::vector dmz_k = cal_dm_trans_pblas(Z_istate + off_x[is], paraX_g[is], c_spin[is], + inputs.pc, inputs.nbasis, inputs.nocc[is], nvirt_g[is], inputs.pmat); + for (auto& d : dmz_k) { LR_Util::matsym(d.template data(), inputs.nbasis, inputs.pmat); } +#else + dmx_k[is] = cal_dm_trans_blas(X_istate + off_x[is], c_spin[is], inputs.nocc[is], nvirt_g[is]); + dmdiff_k[is] = cal_dm_diff_blas(X_istate + off_x[is], c_spin[is], inputs.nbasis, inputs.nocc[is], nvirt_g[is]); + std::vector dmz_k = cal_dm_trans_blas(Z_istate + off_x[is], c_spin[is], inputs.nocc[is], nvirt_g[is]); + for (auto& d : dmz_k) { LR_Util::matsym(d.template data(), inputs.nbasis); } +#endif + relaxed_k[is] = dmdiff_k[is] + dmz_k; + } + + // 2. $D^X$. Complex and UN-symmetrized first (the EXX kernel needs the full + // non-symmetric $D^X$), then the real symmetrized copy for the grid Hxc force -- + // `build_dm_from_dmk_spin` symmetrizes IN PLACE, hence the ordering. + auto dm_trans = LR_Util::build_dm_from_dmk_spin(dmx_k, + inputs.pmat, inputs.nk, inputs.kv.kvec_d, inputs.ucell, inputs.gd, inputs.orb_cutoff); + LR_Util::transpose_DMR(dm_trans, inputs.pmat); + auto dm_trans_real = LR_Util::build_dm_from_dmk_spin(dmx_k, + inputs.pmat, inputs.nk, inputs.kv.kvec_d, inputs.ucell, inputs.gd, inputs.orb_cutoff, + /*symmetrize=*/true); + LR_Util::transpose_DMR(dm_trans_real, inputs.pmat); + + // 3. the relaxed difference density matrix $T+D^Z$ + const module_dm::DensityMatrix& relaxed_diff_dm = + LR_Util::build_dm_from_dmk_spin(relaxed_k, + inputs.pmat, inputs.nk, inputs.kv.kvec_d, inputs.ucell, inputs.gd, inputs.orb_cutoff); + module_dm::DensityMatrix relaxed_diff_dm_real(&inputs.pmat, 2, inputs.kv.kvec_d, inputs.nk); + LR_Util::initialize_DMR(relaxed_diff_dm_real, inputs.pmat, inputs.ucell, inputs.gd, inputs.orb_cutoff); + LR_Util::get_DMR_real_imag_part(relaxed_diff_dm, relaxed_diff_dm_real, 'R'); + + // 4. the energy-weighted density matrix + std::weak_ptr pot_weak = inputs.pot[0]; + const double omega_istate = omega[istate - ist_begin]; + const std::vector>& edm_k = cal_edm_from_XZ_istate_openshell( + inputs, X_istate, Z_istate, omega_istate, inputs.eig_ks.c, dm_trans, pot_weak); + module_dm::DensityMatrix edm_real = LR_Util::build_dm_from_dmk_spin(edm_k, + inputs.pmat, inputs.nk, inputs.kv.kvec_d, inputs.ucell, inputs.gd, inputs.orb_cutoff, + /*symmetrize=*/true); + + // 5. the force terms + ModuleBase::matrix force_hxc_dmtrans = lr_force.cal_force_hxc_dmtrans(dm_trans_real, *inputs.pot[0]); + if (inputs.test_force) + ModuleIO::print_force(inputs.ofs, inputs.ucell, "HXC DMTRANS FORCE (eV/Angstrom)", force_hxc_dmtrans, false); + + + // the $g^{xc}$ half of $\partial_x K[D^X]D^X$, i.e. the derivative of the xc kernel + // through the ground-state density. Only for local kernels. + if (LR_Util::has_local_xc(inputs.xc_kernel)) + { + // The density dependence of K[D^X]D^X belongs to the LR functional. + const std::shared_ptr& pot_lr = inputs.pot[0]; + PotGradXCLR pot_grad(pot_lr->xc_kernel_components(), pot_lr->get_rho_basis(), + inputs.ucell, pot_lr->nrxx, /*triplet=*/false); + ModuleBase::matrix force_gxc_dmtrans = + lr_force.cal_force_gxc_dmtrans_openshell(dm_trans_real, dm_gs, pot_grad); + if (inputs.test_force) + ModuleIO::print_force(inputs.ofs, inputs.ucell, "GXC DMTRANS FORCE (eV/Angstrom)", force_gxc_dmtrans, false); + force_hxc_dmtrans += force_gxc_dmtrans; + } + + ModuleBase::matrix force_hamiltgs_relaxed_diff = lr_force.cal_force_hamilt_gs_dm_relaxed_diff( + relaxed_diff_dm_real, dm_gs, false, inputs.pot_hxc_gs.get()); + if (inputs.test_force) + ModuleIO::print_force(inputs.ofs, inputs.ucell, "H_GS-(T+Z) FORCE (without EXX) (eV/Angstrom)", force_hamiltgs_relaxed_diff, false); + + ModuleBase::matrix force_overlap_edm = lr_force.cal_force_overlap_edm(edm_real); + if (inputs.test_force) + ModuleIO::print_force(inputs.ofs, inputs.ucell, "OVERLAP-EDM FORCE (eV/Angstrom)", force_overlap_edm, false); + +#ifdef __EXX + const double& alpha = inputs.hybrid_alpha; + // Exchange is spin-diagonal, so each channel is done independently and summed. + // `get_exx_Ds_gs` returns the channels unscaled (SPIN_multiple = 1 at nspin=2), unlike + // the closed-shell `get_exx_Ds_spin1` which returns 0.5*D. Counting the closed-shell + // `alpha*4.0` back down for both terms: + // $D^XD^X$: each slot goes 0.5*D^X_tot -> D^X_is, i.e. x2 each, and the explicit + // sum over is adds another x2 -- but $D^X_\text{tot}=\sqrt2 D^X_\sigma$ eats one, + // so 4/(2*2) * ... = `alpha`. + // $D^\text{gs}(T{+}D^Z)$: the left slot goes 0.5*D_up -> D_is (x2); the right slot + // goes 0.5*(T+Z)_tot = (T+Z)_up -> (T+Z)_is (x1, no sqrt2 here); the explicit sum + // over is adds x2. So 4/(2*1*2) = `alpha` as well -- NOT `2*alpha`: the earlier + // comment forgot that the closed-shell right slot is already the spin SUM, which is + // exactly what the `for (is)` loop below now supplies. + if (LR::exx_kernel_list().count(inputs.xc_kernel)) + { + const auto& Ds_trans = LR_Util::get_exx_Ds_gs(dm_trans, inputs.ucell, inputs.kv, inputs.pmat); + ModuleBase::matrix force_exx_dmtrans(inputs.ucell.nat, 3); + for (int is : {0, 1}) + { + force_exx_dmtrans += lr_force.cal_force_exx_dm_trans(Ds_trans[is], alpha, std::to_string(is)); + } + if (inputs.test_force) + ModuleIO::print_force(inputs.ofs, inputs.ucell, "EXX DMTRANS FORCE (eV/Angstrom)", force_exx_dmtrans, false); + force_hxc_dmtrans += force_exx_dmtrans; + } + if (LR::gs_is_hybrid(inputs.dft_functional)) + { + const auto& Ds_gs = LR_Util::get_exx_Ds_gs(dm_gs, inputs.ucell, inputs.kv, inputs.pmat); + const auto& Ds_relaxed_diff = LR_Util::get_exx_Ds_gs(relaxed_diff_dm, inputs.ucell, inputs.kv, inputs.pmat); + ModuleBase::matrix force_exx_gs_relaxed_diff(inputs.ucell.nat, 3); + for (int is : {0, 1}) + { + force_exx_gs_relaxed_diff += lr_force.cal_force_exx_gs_dm_relaxed_diff( + Ds_gs[is], Ds_relaxed_diff[is], alpha, std::to_string(is)); + } + if (inputs.test_force) + ModuleIO::print_force(inputs.ofs, inputs.ucell, "EXX GS-(T+Z) FORCE (eV/Angstrom)", force_exx_gs_relaxed_diff, false); + force_hamiltgs_relaxed_diff += force_exx_gs_relaxed_diff; + } +#endif + forces[istate - ist_begin] = force_hxc_dmtrans + force_hamiltgs_relaxed_diff + force_overlap_edm; + } + print_lr_force(forces, std::cout, ist_begin); + print_lr_force(forces, inputs.ofs, ist_begin); + ModuleBase::timer::end("LR", "evaluate_open_shell_force"); + return forces; +} + + +template std::vector evaluate_open_shell_force(const GradientInputs&, LR_Force&, const module_dm::DensityMatrix&, const ct::Tensor&, const ct::Tensor&, const std::vector&, int); +template std::vector evaluate_open_shell_force>(const GradientInputs>&, LR_Force>&, const module_dm::DensityMatrix, double>&, const ct::Tensor&, const ct::Tensor&, const std::vector&, int); +} diff --git a/source/source_lcao/module_lr/lr_spectrum.cpp b/source/source_lcao/module_lr/lr_spectrum.cpp index 7f38537eab1..2cc82c05817 100644 --- a/source/source_lcao/module_lr/lr_spectrum.cpp +++ b/source/source_lcao/module_lr/lr_spectrum.cpp @@ -100,13 +100,13 @@ ModuleBase::Vector3> LR::LR_Spectrum>: LR_Util::initialize_DMR(DM_trans_real_imag, this->pmat, this->ucell, this->gd_, this->orb_cutoff_); // real part - LR_Util::get_DMR_real_imag_part(DM_trans, DM_trans_real_imag, ucell.nat, is, 'R'); + LR_Util::get_DMR_real_imag_part(DM_trans, DM_trans_real_imag, is, 'R'); ModuleBase::GlobalFunc::ZEROS(rho_trans_real[0], this->rho_basis.nrxx); ModuleGint::cal_gint_rho(DM_trans_real_imag.get_dmr_vec(), 1, rho_trans_real, false); // LR_Util::print_grid_nonzero(rho_trans_real[0], this->rho_basis.nrxx, 10, "rho_trans"); // imag part - LR_Util::get_DMR_real_imag_part(DM_trans, DM_trans_real_imag, ucell.nat, is, 'I'); + LR_Util::get_DMR_real_imag_part(DM_trans, DM_trans_real_imag, is, 'I'); ModuleBase::GlobalFunc::ZEROS(rho_trans_imag[0], this->rho_basis.nrxx); ModuleGint::cal_gint_rho(DM_trans_real_imag.get_dmr_vec(), 1, rho_trans_imag, false); // LR_Util::print_grid_nonzero(rho_trans_imag[0], this->rho_basis.nrxx, 10, "rho_trans"); diff --git a/source/source_lcao/module_lr/lr_spectrum_velocity.cpp b/source/source_lcao/module_lr/lr_spectrum_velocity.cpp index b94b8d3bd97..96a0a4612e4 100644 --- a/source/source_lcao/module_lr/lr_spectrum_velocity.cpp +++ b/source/source_lcao/module_lr/lr_spectrum_velocity.cpp @@ -92,7 +92,7 @@ namespace LR { for (int is = 0;is < this->nspin_x; ++is) { - trans_dipole[i] += LR_Util::dot_R_matrix(*vR.get_current_term_pointer(i), *DM_trans.get_dmr_ptr(is + 1), ucell.nat) * fac; + trans_dipole[i] += LR_Util::dot_R_matrix(*vR.get_current_term_pointer(i), *DM_trans.get_dmr_ptr(is + 1)) * fac; } // end for spin_x, only matter in open-shell system trans_dipole[i] *= static_cast(this->nk); // nk is divided inside DM_trans, now recover it if (this->nspin_x == 1) { trans_dipole[i] *= sqrt(2.0); } // *2 for 2 spins, /sqrt(2) for the halfed dimension of X in the normalizaiton @@ -261,7 +261,9 @@ namespace LR std::vector eig_ks_diff(this->ldim); for (int is = 0;is < this->nspin_x;++is) { - cal_eig_ks_diff(eig_ks_diff.data() + is * nk * pX[0].get_local_size(), eig_ks, pX[is], nk, nocc[is], nvirt[is]); + const int nbands = nocc[is] + nvirt[is]; + const double* eig_spin = eig_ks + is * nk * nbands; + cal_eig_ks_diff(eig_ks_diff.data() + is * nk * pX[0].get_local_size(), eig_spin, pX[is], nk, nocc[is], nvirt[is]); } // X/(ec-ev) diff --git a/source/source_lcao/module_lr/operator_casida/op_gxc_ulr.h b/source/source_lcao/module_lr/operator_casida/op_gxc_ulr.h new file mode 100644 index 00000000000..a4b77c8ff18 --- /dev/null +++ b/source/source_lcao/module_lr/operator_casida/op_gxc_ulr.h @@ -0,0 +1,146 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_OPERATOR_CASIDA_OPERATOR_GXC_ULR_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_OPERATOR_CASIDA_OPERATOR_GXC_ULR_H +#include "source_lcao/module_lr/potentials/pot_grad_xc.h" +#include "source_cell/klist.h" +#include "source_estate/module_dm/density_matrix.h" +#include "source_lcao/module_lr/dm_trans/dm_trans.h" +#include "source_lcao/module_lr/utils/lr_util.h" +#include "source_lcao/module_lr/utils/lr_util_hcontainer.h" +#include "source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo.h" +#include "source_hamilt/module_gint/gint_interface.h" +#include "source_hamilt/module_hcontainer/hcontainer_funcs.h" + +namespace LR +{ + /// @brief The open-shell $g^{xc}$ term, i.e. + /// $\sum_{\kappa\lambda\sigma'}\sum_{\alpha\beta\sigma''}D^X_{\sigma'}D^X_{\sigma''} + /// g^{xc}_{\kappa\lambda\sigma',\alpha\beta\sigma'',pq\tau}$ + /// projected onto the MO block selected by `mo_type` (VO for the Z-vector right-hand side, + /// OO for $W^c$). + /// + /// This cannot be an `OperatorLRHxc` block like the other kernel terms: it is quadratic in + /// $D^X$ rather than bilinear in a (out, in) spin pair, so BOTH transition-density channels + /// have to be on the grid at the same time while the free spin $\tau$ is the output index. + /// Hence the density is built once for the whole X vector and the potential is evaluated + /// twice, once per $\tau$. + /// + /// Only the real (gamma-only) path is implemented, matching the rest of the LR gradient -- + /// `solve_Z_CG` has no complex version either. + template + class OperatorGxcULR + { + public: + OperatorGxcULR(const KernelXC& kxc, + const ModulePW::PW_Basis& rho_basis, + const UnitCell& ucell, + const std::vector& orb_cutoff, + const Grid_Driver& gd, + const K_Vectors& kv, + const Parallel_Orbitals& pmat, + const Parallel_2D& pc, + const psi::Psi& psi_ks, + const std::vector& nocc, + const std::vector& nvirt, + const int naos, + const std::vector& pX, ///< layout of the X blocks (always VO) + const std::vector& pout, ///< layout of the output blocks (VO or OO) + const LR_Util::MO_TYPE mo_type, + const T factor, + const int& nspin, + const std::string& ks_solver) + : pot_grad_(kxc, rho_basis, ucell, rho_basis.nrxx, /*triplet=*/false), + ucell_(ucell), gd_(gd), kv_(kv), pmat_(pmat), pc_(pc), psi_ks_(psi_ks), + nocc_(nocc), nvirt_(nvirt), naos_(naos), pX_(pX), pout_(pout), + orb_cutoff_(orb_cutoff), mo_type_(mo_type), factor_(factor), + nk_(kv.get_nks() / nspin), nrxx_(rho_basis.nrxx), ks_solver_(ks_solver) + { + for (int is : {0, 1}) { this->psi_spin_.push_back(LR_Util::get_psi_spin(psi_ks, is, this->nk_)); } + this->hR_ = LR_Util::make_unique>(&pmat); + LR_Util::initialize_HR(*this->hR_, ucell, gd, orb_cutoff); + } + + /// @brief `out += factor *

` for both spin channels + void act(const T* const X_full, T* const out_full) const + { + ModuleBase::TITLE("OperatorGxcULR", "act"); + ModuleBase::timer::start("OperatorGxcULR", "act"); + const std::vector off_x = { 0, static_cast(nk_ * pX_[0].get_local_size()) }; + const std::vector off_out = { 0, static_cast(nk_ * pout_[0].get_local_size()) }; + + // 1. the two transition-density channels, on the grid together + std::vector> dmk(2); + for (int is : {0, 1}) + { +#ifdef __MPI + dmk[is] = cal_dm_trans_pblas(X_full + off_x[is], pX_[is], psi_spin_[is], pc_, + naos_, nocc_[is], nvirt_[is], pmat_); + for (auto& t : dmk[is]) { LR_Util::matsym(t.template data(), naos_, pmat_); } +#else + dmk[is] = cal_dm_trans_blas(X_full + off_x[is], psi_spin_[is], nocc_[is], nvirt_[is]); + for (auto& t : dmk[is]) { LR_Util::matsym(t.template data(), naos_); } +#endif + } + module_dm::DensityMatrix dm = LR_Util::build_dm_from_dmk_spin(dmk, + pmat_, nk_, kv_.kvec_d, ucell_, gd_, orb_cutoff_); + + double** rho1 = nullptr; + LR_Util::_allocate_2order_nested_ptr(rho1, 2, nrxx_); + for (int is : {0, 1}) { ModuleBase::GlobalFunc::ZEROS(rho1[is], nrxx_); } + ModuleGint::cal_gint_rho(dm.get_dmr_vec(), 2, rho1, false); + + // 2. one potential per output spin, then AO -> MO. + // The V(R) container is real, so for complex T only the gamma-only case is right -- + // the multi-k real/imag split that `OperatorLRHxc` does is not reproduced here. That + // is not a new restriction: `solve_Z_CG` has no complex version either. + for (int tau : {0, 1}) + { + ModuleBase::matrix v2(1, nrxx_); // zero-initialized + pot_grad_.cal_v_eff_openshell(rho1, ucell_, v2, tau); + + this->hR_->set_zero(); + ModuleGint::cal_gint_vl(v2.c, this->hR_.get()); + std::vector v_2d(nk_, LR_Util::newTensor({ pmat_.get_col_size(), pmat_.get_row_size() })); + for (auto& v : v_2d) { v.zero(); } + const int nrow = ModuleBase::GlobalFunc::IS_COLUMN_MAJOR_KS_SOLVER(this->ks_solver_) + ? pmat_.get_row_size() : pmat_.get_col_size(); + for (int ik = 0;ik < nk_;++ik) + { + folding_HR(*this->hR_, v_2d[ik].template data(), kv_.kvec_d[ik], nrow, 1); + } +#ifdef __MPI + ao_to_mo_pblas(v_2d, pmat_, psi_spin_[tau], pc_, naos_, nocc_[tau], nvirt_[tau], + pout_[tau], out_full + off_out[tau], /*add_on=*/true, mo_type_, factor_); +#else + ao_to_mo_blas(v_2d, psi_spin_[tau], nocc_[tau], nvirt_[tau], + out_full + off_out[tau], /*add_on=*/true, mo_type_, factor_); +#endif + } + LR_Util::_deallocate_2order_nested_ptr(rho1, 2); + ModuleBase::timer::end("OperatorGxcULR", "act"); + } + + private: + PotGradXCLR pot_grad_; + const UnitCell& ucell_; + const Grid_Driver& gd_; + const K_Vectors& kv_; + const Parallel_Orbitals& pmat_; + const Parallel_2D& pc_; + const psi::Psi& psi_ks_; + std::vector> psi_spin_; + const std::vector& nocc_; + const std::vector& nvirt_; + const int naos_ = 1; + const std::vector& pX_; + const std::vector& pout_; + std::vector orb_cutoff_; + const LR_Util::MO_TYPE mo_type_ = LR_Util::VO; + const T factor_ = T(1); + const int nk_ = 1; + const int nrxx_ = 1; + const std::string ks_solver_; + std::unique_ptr> hR_; + }; +} + +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_OPERATOR_CASIDA_OPERATOR_GXC_ULR_H diff --git a/source/source_lcao/module_lr/operator_casida/operator_lr_diag.h b/source/source_lcao/module_lr/operator_casida/operator_lr_diag.h index 041a43c2063..4e895830911 100644 --- a/source/source_lcao/module_lr/operator_casida/operator_lr_diag.h +++ b/source/source_lcao/module_lr/operator_casida/operator_lr_diag.h @@ -12,8 +12,9 @@ namespace LR class OperatorLRDiag : public hamilt::Operator { public: - OperatorLRDiag(const double* eig_ks, const Parallel_2D& pX_in, const int& nk_in, const int& nocc_in, const int& nvirt_in) - : pX(pX_in), nk(nk_in), nocc(nocc_in), nvirt(nvirt_in) + OperatorLRDiag(const double* eig_ks, const Parallel_2D& pX_in, const int& nk_in, const int& nocc_in, const int& nvirt_in, + const bool add_on = false) + : pX(pX_in), nk(nk_in), nocc(nocc_in), nvirt(nvirt_in), add_on_(add_on) { // calculate the difference of eigenvalues ModuleBase::TITLE("OperatorLRDiag", "OperatorLRDiag"); const int nbands = nocc + nvirt; @@ -35,8 +36,12 @@ namespace LR }; void init(const int ik_in) override {}; - /// caution: put this operator at the head of the operator list, - /// because vector_mul_vector_op directly assign to (rather than add on) psi_out. + /// By default this ASSIGNS to `hpsi`, so it has to be the head of its operator list. + /// That is fine for a single-chain Hamiltonian, but an open-shell 2x2 spin-block + /// operator writes into the same output buffer from four chains: the diagonal block of + /// the DOWN channel runs last and its assignment wipes the up->down contribution that + /// the off-diagonal chain wrote earlier. Those callers pass `add_on = true` and zero the + /// output themselves. virtual void act(const int nbands, const int nbasis, const int npol, @@ -51,15 +56,16 @@ namespace LR hpsi, psi_in, this->eig_ks_diff.c, - false); + this->add_on_); ModuleBase::timer::end("OperatorLRDiag", "act"); } private: const Parallel_2D& pX; ModuleBase::matrix eig_ks_diff; - const int& nk; - const int& nocc; - const int& nvirt; + const int nk = 1; + const int nocc = 1; + const int nvirt = 1; + const bool add_on_ = false; Device* ctx = {}; }; } diff --git a/source/source_lcao/module_lr/operator_casida/operator_lr_exx.cpp b/source/source_lcao/module_lr/operator_casida/operator_lr_exx.cpp index 7ffc500f946..cce6e52949d 100644 --- a/source/source_lcao/module_lr/operator_casida/operator_lr_exx.cpp +++ b/source/source_lcao/module_lr/operator_casida/operator_lr_exx.cpp @@ -1,161 +1,204 @@ #ifdef __EXX #include "operator_lr_exx.h" +#include +#include +#include +#include #include "source_lcao/module_lr/dm_trans/dm_trans.h" #include "source_lcao/module_lr/utils/lr_util.h" -#include "source_lcao/module_lr/utils/lr_util_print.h" -#include "source_lcao/module_lr/ri_benchmark/ri_benchmark.h" +#include "source_lcao/module_lr/exx_proj.h" +#include "source_base/parallel_reduce.h" namespace LR { - template - void OperatorLREXX::allocate_Ds_onebase() + namespace { - ModuleBase::TITLE("OperatorLREXX", "allocate_Ds_onebase"); - for (int iat1 = 0;iat1 < ucell.nat;++iat1) { - const int it1 = ucell.iat2it[iat1]; - for (int iat2 = 0;iat2 < ucell.nat;++iat2) { - const int it2 = ucell.iat2it[iat2]; - for (auto cell : this->BvK_cells) { - this->Ds_onebase[iat1][std::make_pair(iat2, cell)] = - RI::Tensor({ static_cast(ucell.atoms[it1].nw), static_cast(ucell.atoms[it2].nw) }); - } - } + template + T projection_phase(const ModuleBase::Vector3& k, + const std::array& cell, + const ModuleBase::Matrix3& lattice); + + template<> + double projection_phase(const ModuleBase::Vector3&, + const std::array&, + const ModuleBase::Matrix3&) + { + return 1.0; } - } - template<> - void OperatorLREXX::cal_DM_onebase(const int io, const int iv, const int ik) const - { - ModuleBase::TITLE("OperatorLREXX", "cal_DM_onebase"); - // NOTICE: DM_onebase will be passed into `cal_energy` interface and conjugated by "zdotc". - // So the formula should be the same as RHS. instead of LHS of the A-matrix, - // i.e. c1v · conj(c2o) · e^{-ik(R2-R1)} - assert(ik == 0); - for (auto cell : this->BvK_cells) + template<> + std::complex projection_phase>( + const ModuleBase::Vector3& k, + const std::array& cell, + const ModuleBase::Matrix3& lattice) { - for (int it1 = 0;it1 < ucell.ntype;++it1) - for (int ia1 = 0; ia1 < ucell.atoms[it1].na; ++ia1) - for (int it2 = 0;it2 < ucell.ntype;++it2) - for (int ia2 = 0;ia2 < ucell.atoms[it2].na;++ia2) - { - int iat1 = ucell.itia2iat(it1, ia1); - int iat2 = ucell.itia2iat(it2, ia2); - auto& D2d = this->Ds_onebase[iat1][std::make_pair(iat2, cell)]; - const int nw1 = ucell.atoms[it1].nw; - const int nw2 = ucell.atoms[it2].nw; - for (int iw1 = 0;iw1 < nw1;++iw1) - for (int iw2 = 0;iw2 < nw2;++iw2) - { - const int iwt1 = ucell.itiaiw2iwt(it1, ia1, iw1); - const int iwt2 = ucell.itiaiw2iwt(it2, ia2, iw2); - if (this->pmat.in_this_processor(iwt1, iwt2)) - D2d(iw1, iw2) = this->psi_ks_full(ik, io, iwt1) * this->psi_ks_full(ik, nocc + iv, iwt2); - } - } + const auto translation = RI_Util::array3_to_Vector3(cell) * lattice; + const double angle = ModuleBase::TWO_PI * (k * translation); + // The legacy probe contraction conjugated the negative Bloch phase. + return std::exp(ModuleBase::IMAG_UNIT * angle); } } - template<> - void OperatorLREXX>::cal_DM_onebase(const int io, const int iv, const int ik) const + template + void OperatorLREXX::project_k(const T* psi_in, T* hpsi) const { - ModuleBase::TITLE("OperatorLREXX", "cal_DM_onebase"); - // NOTICE: DM_onebase will be passed into `cal_energy` interface and conjugated by "zdotc". - // So the formula should be the same as RHS. instead of LHS of the A-matrix, - // i.e. c1v · conj(c2o) · e^{-ik(R2-R1)} - for (auto cell : this->BvK_cells) + ModuleBase::timer::start("OperatorLREXX", "project_k"); + const bool occupied = dm_pq_ == MO_TO_AO_TYPE::CC_oo; + const int nright = occupied ? nocc : nvirt; + // Keep the historical aims probe indices, including its complex CC_oo + // convention, without retaining per-pair density construction. + const bool complex_occupied = occupied && std::is_same>::value; + const bool benchmark_probe = !aims_nbasis.empty() + && (dm_pq_ == MO_TO_AO_TYPE::CC_vo || complex_occupied); + if (benchmark_probe) + { + if (aims_nbasis.size() != static_cast(ucell.ntype)) + { + throw std::invalid_argument("EXX benchmark probe requires one basis index per atom type"); + } + for (const int index : aims_nbasis) + { + if (index < 0 || index >= naos) + { + throw std::out_of_range("EXX benchmark probe basis index is outside the AO matrix"); + } + } + } + std::vector h(static_cast(naos) * naos); + std::vector result(static_cast(nocc) * nright); + std::vector scratch(static_cast(naos) * nocc); + const double factor = 2.0 * alpha; + const auto lri = this->exx_lri.lock(); + psi::Psi coxt_full; + psi::Psi cvx_full; + if (dm_pq_ == MO_TO_AO_TYPE::CXC || dm_pq_ == MO_TO_AO_TYPE::CXC_o) { - std::complex frac = RI::Global_Func::convert>(std::exp( - -ModuleBase::TWO_PI * ModuleBase::IMAG_UNIT * (this->kv.kvec_c.at(ik) * (RI_Util::array3_to_Vector3(cell) * ucell.latvec)))); - for (int it1 = 0;it1 < ucell.ntype;++it1) - for (int ia1 = 0; ia1 < ucell.atoms[it1].na; ++ia1) - for (int it2 = 0;it2 < ucell.ntype;++it2) - for (int ia2 = 0;ia2 < ucell.atoms[it2].na;++ia2) + coxt_full.resize(this->nk, nvirt, this->naos); + cvx_full.resize(this->nk, nocc, this->naos); + this->cal_coxt_cvx(psi_in, coxt_full, cvx_full); + } + for (int ik = 0; ik < nk; ++ik) + { + std::fill(h.begin(), h.end(), T(0)); + std::fill(result.begin(), result.end(), T(0)); + // Hexxs contains communicated atom blocks. Keep uniquely owned AO + // elements used by the probe contraction; reduce the smaller MO result below. + for (const auto& atom : lri->Hexxs[0]) + { + const int it1 = ucell.iat2it[atom.first]; + const int ia1 = ucell.iat2ia[atom.first]; + for (const auto& pair : atom.second) + { + const int it2 = ucell.iat2it[pair.first.first]; + const int ia2 = ucell.iat2ia[pair.first.first]; + const T phase = projection_phase(kv.kvec_c[ik], pair.first.second, ucell.latvec); + const auto& block = pair.second; + for (int mu = 0; mu < ucell.atoms[it1].nw; ++mu) + { + const int native_row = ucell.itiaiw2iwt(it1, ia1, mu); + const int row = benchmark_probe ? aims_nbasis[it1] : native_row; + for (int nu = 0; nu < ucell.atoms[it2].nw; ++nu) { - int iat1 = ucell.itia2iat(it1, ia1); - int iat2 = ucell.itia2iat(it2, ia2); - auto& D2d = this->Ds_onebase[iat1][std::make_pair(iat2, cell)]; - const int nw1 = ucell.atoms[it1].nw; - const int nw2 = ucell.atoms[it2].nw; - for (int iw1 = 0;iw1 < nw1;++iw1) - for (int iw2 = 0;iw2 < nw2;++iw2) - { - const int iwt1 = ucell.itiaiw2iwt(it1, ia1, iw1); - const int iwt2 = ucell.itiaiw2iwt(it2, ia2, iw2); - if (this->pmat.in_this_processor(iwt1, iwt2)) - D2d(iw1, iw2) = frac * std::conj(this->psi_ks_full(ik, io, iwt2)) * this->psi_ks_full(ik, nocc + iv, iwt1); - } + const int native_col = ucell.itiaiw2iwt(it2, ia2, nu); + const int col = benchmark_probe ? aims_nbasis[it2] : native_col; + if (pmat.in_this_processor(row, col)) + { + h[static_cast(col) * naos + row] += phase * block(mu, nu); + } } + } + } + } + const T* co = &psi_ks_full(ik, 0, 0); + const T* cv = &psi_ks_full(ik, nocc, 0); + if (dm_pq_ == MO_TO_AO_TYPE::CC_vo || occupied) + { + const T* right = occupied ? co : cv; + project_exx(h.data(), co, right, naos, nocc, nright, factor, + scratch.data(), result.data()); + } + else + { + const T* coxt = &coxt_full(ik, 0, 0); + if (dm_pq_ == MO_TO_AO_TYPE::CXC) + { + const T* cvx = &cvx_full(ik, 0, 0); + project_exx(h.data(), cvx, cv, naos, nocc, nvirt, factor, + scratch.data(), result.data()); + } + const double occ_factor = dm_pq_ == MO_TO_AO_TYPE::CXC ? -factor : factor; + project_exx(h.data(), co, coxt, naos, nocc, nvirt, occ_factor, + scratch.data(), result.data()); + } + const int result_size = static_cast(result.size()); + Parallel_Reduce::reduce_all(result.data(), result_size); + const bool transpose = occupied && std::is_same>::value; + const int start = ik * pX.get_local_size(); + for (int io = 0; io < nocc; ++io) + { + for (int iv = 0; iv < nright; ++iv) + { + if (pX.in_this_processor(iv, io)) + { + const int local = pX.global2local_col(io) * pX.get_row_size() + + pX.global2local_row(iv); + // The legacy complex CC_oo probe orders (io,iv), unlike + // CC_vo/CXC_o; preserve it even for nonsymmetric H(k). + const int index = transpose ? iv * nocc + io : io * nright + iv; + hpsi[start + local] += result[index]; + } + } + } } + ModuleBase::timer::end("OperatorLREXX", "project_k"); } template - void OperatorLREXX::act(const int nbands, - const int nbasis, - const int npol, - const T* psi_in, - T* hpsi, - const int ngk_ik, - const bool is_first_node)const + void OperatorLREXX::cal_Hs() const { - ModuleBase::TITLE("OperatorLREXX", "act"); - ModuleBase::timer::start("OperatorLREXX", "act"); - + ModuleBase::TITLE("OperatorLREXX", "cal_Hs"); + ModuleBase::timer::start("OperatorLREXX", "cal_Hs"); // convert parallel info to LibRI interfaces - std::vector, std::set>> judge = RI_2D_Comm::get_2D_judge(ucell,this->pmat); + std::vector, std::set>> judge = RI_2D_Comm::get_2D_judge(ucell,this->pmat); // suppose Cs,Vs, have already been calculated in the ion-step of ground state // and DM_trans has been calculated in hPsi() outside. // 1. set_Ds (once) // convert to vector for the interface of RI_2D_Comm::split_m2D_ktoR (interface will be unified to ct::Tensor) - std::vector> DMk_trans_vector = this->DM_trans->get_dmk_vec(); + std::vector> DMk_trans_vector = this->DM_trans.get_dmk_vec(); // assert(DMk_trans_vector.size() == nk); std::vector*> DMk_trans_pointer(nk); for (int ik = 0;ik < nk;++ik) { DMk_trans_pointer[ik] = &DMk_trans_vector[ik]; } // if multi-k, DM_trans(TR=double) -> Ds_trans(TR=T=complex) - std::vector>>> Ds_trans = - // aims_nbasis.empty() ? // ucell.nw is updated, abandoned 25-05-23 + auto Ds_trans = RI_2D_Comm::split_m2D_ktoR(ucell,this->kv, DMk_trans_pointer, this->pmat, 1); - //: RI_Benchmark::split_Ds(DMk_trans_vector, aims_nbasis, ucell); //0.5 will be multiplied - // LR_Util::print_CV(Ds_trans[0], "Ds_trans in OperatorLREXX", 1e-10); // 2. cal_Hs - ModuleBase::timer::start("OperatorLREXX", "cal_Hs"); auto lri = this->exx_lri.lock(); lri->exx_lri.set_Ds(std::move(Ds_trans[0]), lri->info.dm_threshold); lri->exx_lri.cal_Hs(); lri->Hexxs[0] = RI::Communicate_Tensors_Map_Judge::comm_map2_first( lri->mpi_comm, std::move(lri->exx_lri.Hs), std::get<0>(judge[0]), std::get<1>(judge[0])); lri->post_process_Hexx(lri->Hexxs[0]); - // LR_Util::print_CV(lri->Hexxs[0], "Hexxs in OperatorLREXX", 1e-10); ModuleBase::timer::end("OperatorLREXX", "cal_Hs"); + } - // 3. set [AX]_iak = DM_onbase * Hexxs for each occ-virt pair and each k-point - // caution: parrallel - ModuleBase::timer::start("OperatorLREXX", "cal_energy"); - for (int ik = 0;ik < nk;++ik) - { - for (int io = 0;io < this->nocc;++io) - { - for (int iv = 0;iv < this->nvirt;++iv) - { - const int xstart_bk = ik * pX.get_local_size(); - this->cal_DM_onebase(io, iv, ik); //set Ds_onebase for all e-h pairs (not only on this processor) - // LR_Util::print_CV(Ds_onebase, "Ds_onebase of occ " + std::to_string(io) + ", virtual " + std::to_string(iv) + " in OperatorLREXX", 1e-10); - const T& ene = 2 * alpha * //minus for exchange(but here plus, since `post_process_Hexx` has taken minus), 2 for Hartree to Ry - lri->exx_lri.post_2D.cal_energy(this->Ds_onebase, lri->Hexxs[0]); - if (this->pX.in_this_processor(iv, io)) - { - hpsi[xstart_bk + this->pX.global2local_col(io) * this->pX.get_row_size() + this->pX.global2local_row(iv)] += ene; - } - //for debug - GlobalV::ofs_running << "Direct term: ik="< + void OperatorLREXX::act(const int nbands, + const int nbasis, + const int npol, + const T* psi_in, + T* hpsi, + const int ngk_ik, + const bool is_first_node) const + { + ModuleBase::TITLE("OperatorLREXX", "act"); + ModuleBase::timer::start("OperatorLREXX", "act"); + this->cal_Hs(); + this->project_k(psi_in, hpsi); ModuleBase::timer::end("OperatorLREXX", "act"); } template class OperatorLREXX; template class OperatorLREXX>; } -#endif \ No newline at end of file +#endif diff --git a/source/source_lcao/module_lr/operator_casida/operator_lr_exx.h b/source/source_lcao/module_lr/operator_casida/operator_lr_exx.h index d426cdf6d39..2998f92bd48 100644 --- a/source/source_lcao/module_lr/operator_casida/operator_lr_exx.h +++ b/source/source_lcao/module_lr/operator_casida/operator_lr_exx.h @@ -6,40 +6,54 @@ #include "source_estate/module_dm/density_matrix.h" #include "source_lcao/module_ri/exx_lri.h" #include "source_lcao/module_lr/utils/lr_util.h" +#include "source_lcao/module_lr/dm_trans/dm_diff.h" namespace LR { /// @brief Exx part of A operator + /// The single source of truth is `LR_Util::hybrid_xc_list` -- see there for what the list + /// does and does not decide. + inline const std::set& exx_kernel_list() { return LR_Util::hybrid_xc_list(); }; + + /// @brief does the GROUND STATE carry exact exchange, i.e. is `dft_functional` a hybrid? + inline bool gs_is_hybrid(const std::string& dft_functional) { return exx_kernel_list().count(LR_Util::tolower(dft_functional)) > 0; }; + template class OperatorLREXX : public hamilt::Operator { - using TA = int; - static const size_t Ndim = 3; - using TC = std::array; - using TAC = std::pair; - public: + /// @brief type of molecular orbital to atomic orbital transformation: + /// CC_vo: MO = C_v^* AO C_o; + /// CC_oo: MO = C_o^* AO C_o; + /// CXC: MO = C_v^* AO X C_v- C_o^* X^* AO C_o + /// CXC_o: MO = C_o^* X^* AO C_o + enum class MO_TO_AO_TYPE { CC_vo, CC_oo, CXC, CXC_o }; + OperatorLREXX(const int& nspin, const int& naos, const int& nocc, const int& nvirt, const UnitCell& ucell_in, const psi::Psi& psi_ks_in, - std::unique_ptr>& DM_trans_in, + const module_dm::DensityMatrix& DM_trans_in, // HContainer* hR_in, std::weak_ptr> exx_lri_in, const K_Vectors& kv_in, const Parallel_2D& pX_in, const Parallel_2D& pc_in, const Parallel_Orbitals& pmat_in, - const double& alpha = 1.0) + const bool cal_force, + const double& alpha = 1.0, + const MO_TO_AO_TYPE dm_pq_in = MO_TO_AO_TYPE::CC_vo, + const std::vector& aims_nbasis = {}, + const hamilt::calculation_type cal_type_in = hamilt::calculation_type::lr_dmtrans_exx) : nspin(nspin), naos(naos), nocc(nocc), nvirt(nvirt), nk(kv_in.get_nks() / nspin), psi_ks(psi_ks_in), DM_trans(DM_trans_in), exx_lri(exx_lri_in), kv(kv_in), - pX(pX_in), pc(pc_in), pmat(pmat_in), ucell(ucell_in), alpha(alpha) + pX(pX_in), pc(pc_in), pmat(pmat_in), ucell(ucell_in), alpha(alpha), dm_pq_(dm_pq_in), + aims_nbasis(aims_nbasis) { ModuleBase::TITLE("OperatorLREXX", "OperatorLREXX"); - std::cout<<"Initializing OperatorLREXX"<cal_type = hamilt::calculation_type::lcao_exx; + this->cal_type = cal_type_in; this->is_first_node = false; // reduce psi_ks for later use @@ -49,13 +63,10 @@ namespace LR { LR_Util::gather_2d_to_full(this->pc, &this->psi_ks(ik, 0, 0), &this->psi_ks_full(ik, 0, 0), false, this->naos, nocc + nvirt); } - - // get cells in BvK supercell - const TC period = RI_Util::get_Born_vonKarmen_period(kv_in); - this->BvK_cells = RI_Util::get_Born_von_Karmen_cells(period); - - this->allocate_Ds_onebase(); - this->exx_lri.lock()->Hexxs.resize(1); + if (!this->exx_lri.expired()) + { + this->exx_lri.lock()->Hexxs.resize(1); + } }; void init(const int ik_in) override {}; @@ -76,6 +87,7 @@ namespace LR const int nvirt = 1; const int nk = 1; ///< number of k-points const double alpha = 1.0; //(allow non-ref constant) + MO_TO_AO_TYPE dm_pq_ = MO_TO_AO_TYPE::CC_vo; const bool cal_dm_trans = false; const bool tdm_sym = false; ///< whether transition density matrix is symmetric const K_Vectors& kv; @@ -84,23 +96,11 @@ namespace LR psi::Psi psi_ks_full; /// transition density matrix - std::unique_ptr>& DM_trans; - - /// density matrix of a certain (i, a, k), with full naos*naos size for each key - /// D^{iak}_{\mu\nu}(k): 1/N_k * c_{ak,\mu} c^*_{ik,\nu} - /// D^{iak}_{\mu\nu}(R): D^{iak}_{\mu\nu}(k)e^{-ikR} - // module_dm::DensityMatrix* DM_onebase; - mutable std::map>> Ds_onebase; - - // cells in the Born von Karmen supercell (direct) - std::vector> BvK_cells; - - /// transition hamiltonian in AO representation - // hamilt::HContainer* hR = nullptr; + const module_dm::DensityMatrix& DM_trans; /// C, V tensors of RI, and LibRI interfaces /// gamma_only: T=double, Tpara of exx (equal to Tpara of Ds(R) ) is also double - ///.multi-k: T=complex, Tpara of exx here must be complex, because Ds_onebase is complex + /// multi-k: both the transition density and the exchange response are complex /// so TR in DensityMatrix and Tdata in Exx_LRI are all equal to T std::weak_ptr> exx_lri; @@ -111,10 +111,45 @@ namespace LR const Parallel_2D& pX; const Parallel_Orbitals& pmat; - // allocate Ds_onebase - void allocate_Ds_onebase(); + /// number of basis functions per type in the aims benchmark (empty for the normal case) + const std::vector aims_nbasis; + + /// only for gradient calculation + + + /// Build and communicate the exchange response once per input density. + void cal_Hs() const; + + /// Gamma and complex Bloch matrix projection, including benchmark indices. + void project_k(const T* psi_in, T* hpsi) const; - void cal_DM_onebase(const int io, const int iv, const int ik) const; + void cal_coxt_cvx(const T* x_istate, psi::Psi& coxt_full, psi::Psi& cvx_full) const // C_o X^T, C_v X (only for gradients) + { + ModuleBase::TITLE("OperatorLREXX", "cal_coxt_cvx"); + const auto& c = this->psi_ks; + // allocate local cvx + Parallel_2D pcx; + LR_Util::setup_2d_division(pcx, pX.get_block_size(), naos, nocc, pX.blacs_ctxt); + ct::Tensor cvx(ct::DataTypeToEnum::value, DEV::CpuDevice, { pcx.get_col_size(), pcx.get_row_size() }); + + // allocate local coxt + Parallel_2D pcxt; + LR_Util::setup_2d_division(pcxt, pX.get_block_size(), naos, nvirt, pX.blacs_ctxt); + ct::Tensor coxt(ct::DataTypeToEnum::value, DEV::CpuDevice, { pcxt.get_col_size(), pcxt.get_row_size() }); + + // calculate global coxt_full, cvx_full + cvx_full.zero_out(); + coxt_full.zero_out(); + for (int ik = 0;ik < nk;++ik) + { + c.fix_k(ik); + const int start = ik * pX.get_local_size(); + CvX(c.get_pointer(), pc, x_istate + start, pX, naos, nocc, nvirt, cvx.data(), pcx); + LR_Util::gather_2d_to_full(pcx, cvx.data(), &cvx_full(ik, 0, 0), false, naos, nocc); + CoXT(c.get_pointer(), pc, x_istate + start, pX, naos, nocc, nvirt, coxt.data(), pcxt); + LR_Util::gather_2d_to_full(pcxt, coxt.data(), &coxt_full(ik, 0, 0), false, naos, nvirt); + } + } }; } diff --git a/source/source_lcao/module_lr/operator_casida/operator_lr_hxc.cpp b/source/source_lcao/module_lr/operator_casida/operator_lr_hxc.cpp index 64c38a184bb..da5f82111bb 100644 --- a/source/source_lcao/module_lr/operator_casida/operator_lr_hxc.cpp +++ b/source/source_lcao/module_lr/operator_casida/operator_lr_hxc.cpp @@ -1,4 +1,5 @@ #include "operator_lr_hxc.h" +#include #include #include "source_io/module_parameter/parameter.h" #include "source_base/timer.h" @@ -9,6 +10,7 @@ #include "source_hamilt/module_hcontainer/hcontainer_funcs.h" #include "source_lcao/module_lr/ao_to_mo_transformer/ao_to_mo.h" #include "source_hamilt/module_gint/gint_interface.h" +#include "source_lcao/module_lr/ao_to_mo_transformer/cvcx.h" inline double conj(double a) { return a; } inline std::complex conj(std::complex a) { return std::conj(a); } @@ -18,18 +20,28 @@ namespace LR template void OperatorLRHxc::act(const int nbands, const int nbasis, const int npol, const T* psi_in, T* hpsi, const int ngk_ik, const bool is_first_node)const { - ModuleBase::TITLE("OperatorLRHxc", "act"); - ModuleBase::timer::start("OperatorLRHxc", "act"); + TransitionDensityCache density_cache; + this->act_with_shared_density(psi_in, hpsi, density_cache); + } + + template + void OperatorLRHxc::act_with_shared_density( + const T* psi_in, T* hpsi, TransitionDensityCache& density_cache) const + { + ModuleBase::TITLE("OperatorLRHxc", "act_with_shared_density"); + ModuleBase::timer::start("OperatorLRHxc", "act_with_shared_density"); const int& sl = ispin_ks[0]; const auto psil_ks = LR_Util::get_psi_spin(psi_ks, sl, nk); - this->DM_trans->cal_dmr(-1); //DM_trans->get_dmr_vec() is 2d-block parallized - // LR_Util::print_DMR(*DM_trans, ucell.nat, "DMR"); - - // ========================= begin grid calculation========================= - this->grid_calculation(nbands); //DM(R) to H(R) - // ========================= end grid calculation ========================= + if (density_cache.count(&this->DM_trans) == 0) + { + this->DM_trans.cal_dmr(-1); // DM_trans.get_dmr_vec() is 2D-block parallelized. + // cal_dmr preserves AO indices while converting DMK's storage layout to DMR. + } + // ========================= begin grid calculation ========================= + this->grid_calculation(density_cache); // DM(R) -> rho(r) -> V_Hxc(r) -> H(R) + // ========================= end grid calculation =========================== // V(R)->V(k) std::vector v_hxc_2d(nk, LR_Util::newTensor({ pmat.get_col_size(), pmat.get_row_size() })); @@ -41,37 +53,79 @@ namespace LR // for (int ik = 0;ik < nk;++ik) // LR_Util::print_tensor(v_hxc_2d[ik], "4.V(k)[ik=" + std::to_string(ik) + "]", &this->pmat); - // 5. [AX]^{Hxc}_{ai}=\sum_{\mu,\nu}c^*_{a,\mu,}V^{Hxc}_{\mu,\nu}c_{\nu,i} + // 5. AO to MO transformation + switch (this->dm_pq_) + { + case MO_TO_AO_TYPE::CC_vo: //[AX]^{Hxc}_{ai}=\sum_{\mu,\nu}c^*_{\mu,a}V^{Hxc}_{\mu,\nu}c_{\nu,i} +#ifdef __MPI + ao_to_mo_pblas(v_hxc_2d, this->pmat, psil_ks, this->pc, this->naos, this->nocc[sl], this->nvirt[sl], this->pX[sl], hpsi, /*add_on=*/true, LR_Util::MO_TYPE::VO, this->factor_); +#else + ao_to_mo_blas(v_hxc_2d, psil_ks, this->nocc[sl], this->nvirt[sl], hpsi, /*add_on=*/true, LR_Util::MO_TYPE::VO, this->factor_); +#endif + break; + case MO_TO_AO_TYPE::CC_oo: //[AX]^{Hxc}_{ij}=\sum_{\mu,\nu}c^*_{\mu,i}V^{Hxc}_{\mu,\nu}c_{\nu,j} #ifdef __MPI - ao_to_mo_pblas(v_hxc_2d, this->pmat, psil_ks, this->pc, naos, nocc[sl], nvirt[sl], this->pX[sl], hpsi); + ao_to_mo_pblas(v_hxc_2d, this->pmat, psil_ks, this->pc, this->naos, this->nocc[sl], this->nvirt[sl], this->pX[sl], hpsi, /*add_on=*/true, LR_Util::MO_TYPE::OO, this->factor_); #else - ao_to_mo_blas(v_hxc_2d, psil_ks, nocc[sl], nvirt[sl], hpsi); + ao_to_mo_blas(v_hxc_2d, psil_ks, this->nocc[sl], this->nvirt[sl], hpsi, /*add_on=*/true, LR_Util::MO_TYPE::OO, this->factor_); #endif + break; + case MO_TO_AO_TYPE::CXC: + { +#ifdef __MPI + CVCX_virt_pblas(v_hxc_2d, this->pmat, psil_ks, this->pc, psi_in, this->pX[sl], + this->naos, this->nocc[sl], this->nvirt[sl], hpsi, /*add_on=*/true, this->factor_); + CVCX_occ_pblas(v_hxc_2d, this->pmat, psil_ks, this->pc, psi_in, this->pX[sl], + this->naos, this->nocc[sl], this->nvirt[sl], hpsi, /*add_on=*/true, -this->factor_); +#else + CVCX_virt_blas(v_hxc_2d, psil_ks, psi_in, this->naos, this->nocc[sl], this->nvirt[sl], hpsi, /*add_on=*/true, this->factor_); + CVCX_occ_blas(v_hxc_2d, psil_ks, psi_in, this->naos, this->nocc[sl], this->nvirt[sl], hpsi, /*add_on=*/true, -this->factor_); +#endif + break; + } + case MO_TO_AO_TYPE::CXC_o: +#ifdef __MPI + CVCX_occ_pblas(v_hxc_2d, this->pmat, psil_ks, this->pc, psi_in, this->pX[sl], + this->naos, this->nocc[sl], this->nvirt[sl], hpsi, /*add_on=*/true, this->factor_); +#else + CVCX_occ_blas(v_hxc_2d, psil_ks, psi_in, this->naos, this->nocc[sl], this->nvirt[sl], hpsi, /*add_on=*/true, this->factor_); +#endif + break; + default: + throw std::runtime_error("Unknown DM_TYPE"); + break; + } // for debug //std::cout << "After Hxc, hpsi: [nvirt= " << nvirt[sl] << " nocc= " << nocc[sl] << " nk= " << nk << " ]" << std::endl; //LR_Util::print_value(hpsi, nk, nocc[sl], nvirt[sl]); - ModuleBase::timer::end("OperatorLRHxc", "act"); + ModuleBase::timer::end("OperatorLRHxc", "act_with_shared_density"); } template<> - void OperatorLRHxc::grid_calculation(const int& nbands) const + void OperatorLRHxc::grid_calculation(TransitionDensityCache& density_cache) const { ModuleBase::TITLE("OperatorLRHxc", "grid_calculation(real)"); ModuleBase::timer::start("OperatorLRHxc", "grid_calculation"); // 2. transition electron density // \f[ \tilde{\rho}(r)=\sum_{\mu_j, \mu_b}\tilde{\rho}_{\mu_j,\mu_b}\phi_{\mu_b}(r)\phi_{\mu_j}(r) \f] - double** rho_trans = nullptr; - const int& nrxx = this->pot.lock()->nrxx; - LR_Util::_allocate_2order_nested_ptr(rho_trans, 1, nrxx); // currently gint_kernel_rho uses PARAM.inp.nspin, it needs refactor - ModuleBase::GlobalFunc::ZEROS(rho_trans[0], nrxx); - ModuleGint::cal_gint_rho(this->DM_trans->get_dmr_vec(), 1, rho_trans, false); - // 3. v_hxc = f_hxc * rho_trans - ModuleBase::matrix vr_hxc(1, nrxx); //grid - this->pot.lock()->cal_v_eff(rho_trans, ucell, vr_hxc, ispin_ks); - LR_Util::_deallocate_2order_nested_ptr(rho_trans, 1); + // For a fixed input spin, rho is shared by both output-spin kernels. + const int nrxx = this->pot.lock()->nrxx; + auto& density = density_cache[&this->DM_trans]; + if (density.empty()) + { + const std::vector zeros(nrxx, 0.0); + density.assign(1, zeros); // nspin=1 for transition density + double* rho = density[0].data(); + const auto& dmr = this->DM_trans.get_dmr_vec(); + ModuleGint::cal_gint_rho(dmr, 1, &rho, false); + } + // 3. v_hxc = f_hxc * rho_trans, evaluated separately for each output spin + double* rho = density[0].data(); + ModuleBase::matrix vr_hxc(1, nrxx); // grid + this->pot.lock()->cal_v_eff(&rho, ucell, vr_hxc, ispin_ks); // 4. V^{Hxc}_{\mu,\nu}=\int{dr} \phi_\mu(r) v_{Hxc}(r) \phi_\nu(r) this->hR->set_zero(); // clear hR for each bands @@ -80,47 +134,49 @@ namespace LR } template<> - void OperatorLRHxc, base_device::DEVICE_CPU>::grid_calculation(const int& nbands) const + void OperatorLRHxc, base_device::DEVICE_CPU>::grid_calculation(TransitionDensityCache& density_cache) const { ModuleBase::TITLE("OperatorLRHxc", "grid_calculation(complex)"); ModuleBase::timer::start("OperatorLRHxc", "grid_calculation"); - module_dm::DensityMatrix, double> DM_trans_real_imag(&pmat, 1, kv.kvec_d, kv.get_nks() / nspin); - DM_trans_real_imag.init_dmr(*this->hR); - hamilt::HContainer HR_real_imag(ucell, &this->pmat); - LR_Util::initialize_HR, double>(HR_real_imag, ucell, gd, orb_cutoff_); - - auto dmR_to_hR = [&, this](const char& type) -> void + // 2. transition electron density, evaluated for each real/imaginary part + // \f[ \tilde{\rho}(r)=\sum_{\mu_j, \mu_b}\tilde{\rho}_{\mu_j,\mu_b}\phi_{\mu_b}(r)\phi_{\mu_j}(r) \f] + const int nrxx = this->pot.lock()->nrxx; + const int nparts = nk > 1 ? 2 : 1; // real; also imaginary for multi-k + auto& density = density_cache[&this->DM_trans]; + if (density.empty()) + { + module_dm::DensityMatrix, double> dm_real_imag(&pmat, 1, kv.kvec_d, nk); + dm_real_imag.init_dmr(*this->hR); + const std::vector zeros(nrxx, 0.0); + density.assign(nparts, zeros); + for (int ipart = 0; ipart < nparts; ++ipart) { - LR_Util::get_DMR_real_imag_part(*this->DM_trans, DM_trans_real_imag, ucell.nat, type); - // if (this->first_print)LR_Util::print_DMR(DM_trans_real_imag, ucell.nat, "DMR(2d, real)"); - - - // 2. transition electron density - double** rho_trans = nullptr; - const int& nrxx = this->pot.lock()->nrxx; - - LR_Util::_allocate_2order_nested_ptr(rho_trans, 1, nrxx); // nspin=1 for transition density - ModuleBase::GlobalFunc::ZEROS(rho_trans[0], nrxx); - ModuleGint::cal_gint_rho(DM_trans_real_imag.get_dmr_vec(), 1, rho_trans, false); - // print_grid_nonzero(rho_trans[0], nrxx, 10, "rho_trans"); - - // 3. v_hxc = f_hxc * rho_trans - ModuleBase::matrix vr_hxc(1, nrxx); //grid - this->pot.lock()->cal_v_eff(rho_trans, ucell, vr_hxc, ispin_ks); - // print_grid_nonzero(vr_hxc.c, this->poticab->nrxx, 10, "vr_hxc"); - - LR_Util::_deallocate_2order_nested_ptr(rho_trans, 1); - - // 4. V^{Hxc}_{\mu,\nu}=\int{dr} \phi_\mu(r) v_{Hxc}(r) \phi_\nu(r) - HR_real_imag.set_zero(); - ModuleGint::cal_gint_vl(vr_hxc.c, &HR_real_imag); - // LR_Util::print_HR(HR_real_imag, this->ucell.nat, "VR(real, 2d)"); - LR_Util::set_HR_real_imag_part(HR_real_imag, *this->hR, ucell.nat, type); - }; - this->hR->set_zero(); - dmR_to_hR('R'); //real - if (kv.get_nks() / this->nspin > 1) { dmR_to_hR('I'); } //imag for multi-k + const char part = ipart == 0 ? 'R' : 'I'; + LR_Util::get_DMR_real_imag_part(this->DM_trans, dm_real_imag, part); + // nspin=1 for each real/imaginary transition-density component + double* rho = density[ipart].data(); + const auto& dmr = dm_real_imag.get_dmr_vec(); + ModuleGint::cal_gint_rho(dmr, 1, &rho, false); + } + } + + hamilt::HContainer hr_real_imag(ucell, &this->pmat); + LR_Util::initialize_HR, double>(hr_real_imag, ucell, gd, orb_cutoff_); + this->hR->set_zero(); // clear hR for each band + for (int ipart = 0; ipart < nparts; ++ipart) + { + const char part = ipart == 0 ? 'R' : 'I'; + // 3. v_hxc = f_hxc * rho_trans, evaluated separately for each output spin + double* rho = density[ipart].data(); + ModuleBase::matrix vr_hxc(1, nrxx); // grid + this->pot.lock()->cal_v_eff(&rho, ucell, vr_hxc, ispin_ks); + + // 4. V^{Hxc}_{\mu,\nu}=\int{dr} \phi_\mu(r) v_{Hxc}(r) \phi_\nu(r) + hr_real_imag.set_zero(); + ModuleGint::cal_gint_vl(vr_hxc.c, &hr_real_imag); + LR_Util::set_HR_real_imag_part(hr_real_imag, *this->hR, part); + } ModuleBase::timer::end("OperatorLRHxc", "grid_calculation"); } diff --git a/source/source_lcao/module_lr/operator_casida/operator_lr_hxc.h b/source/source_lcao/module_lr/operator_casida/operator_lr_hxc.h index 0abb437ef50..0cd1af890e6 100644 --- a/source/source_lcao/module_lr/operator_casida/operator_lr_hxc.h +++ b/source/source_lcao/module_lr/operator_casida/operator_lr_hxc.h @@ -1,9 +1,11 @@ #ifndef ABACUS_SOURCE_LCAO_MODULE_LR_OPERATOR_CASIDA_OPERATOR_LR_HXC_H #define ABACUS_SOURCE_LCAO_MODULE_LR_OPERATOR_CASIDA_OPERATOR_LR_HXC_H +#include #include "source_cell/klist.h" #include "source_hamilt/operator.h" #include "source_estate/module_dm/density_matrix.h" +#include "source_lcao/module_lr/potentials/pot_lr_base.h" #include "source_lcao/module_lr/potentials/pot_hxc_lrtd.h" #include "source_lcao/module_lr/utils/lr_util.h" #include "source_lcao/module_lr/utils/lr_util_hcontainer.h" @@ -14,29 +16,39 @@ namespace LR class OperatorLRHxc : public hamilt::Operator { public: + /// @brief type of molecular orbital to atomic orbital transformation: + /// CC_vo: MO = C_v^* AO C_o; + /// CC_oo: MO = C_o^* AO C_o; + /// CXC: MO = C_v^* AO X C_v- C_o^* X^* AO C_o + /// CXC_o: MO = C_o^* X^* AO C_o + enum class MO_TO_AO_TYPE { CC_vo, CC_oo, CXC, CXC_o }; + //when nspin=2, nks is 2 times of real number of k-points. else (nspin=1 or 4), nks is the real number of k-points - OperatorLRHxc(const int& nspin, - const int& naos, - const std::vector& nocc, - const std::vector& nvirt, - const psi::Psi& psi_ks_in, - std::unique_ptr>& DM_trans_in, - std::weak_ptr pot_in, - const UnitCell& ucell_in, - const std::vector& orb_cutoff, - const Grid_Driver& gd_in, - const K_Vectors& kv_in, - const std::vector& pX_in, - const Parallel_2D& pc_in, - const Parallel_Orbitals& pmat_in, - const std::vector& ispin_ks = {0}) - : nspin(nspin), naos(naos), nocc(nocc), nvirt(nvirt), nk(kv_in.get_nks() / nspin), psi_ks(psi_ks_in), + OperatorLRHxc(const int& nspin, + const int& naos, + const std::vector& nocc, + const std::vector& nvirt, + const psi::Psi& psi_ks_in, + module_dm::DensityMatrix& DM_trans_in, + std::weak_ptr pot_in, + const UnitCell& ucell_in, + const std::vector& orb_cutoff, + const Grid_Driver& gd_in, + const K_Vectors& kv_in, + const std::vector& pX_in, + const Parallel_2D& pc_in, + const Parallel_Orbitals& pmat_in, + const std::vector& ispin_ks = { 0 }, + const T factor_in = (T)1.0, + const MO_TO_AO_TYPE dm_pq_in = MO_TO_AO_TYPE::CC_vo, + const hamilt::calculation_type cal_type_in = hamilt::calculation_type::lr_dmtrans_hxc) + : nspin(nspin), naos(naos), nocc(nocc), nvirt(nvirt), nk(kv_in.get_nks() / nspin), psi_ks(psi_ks_in), DM_trans(DM_trans_in), pot(pot_in), ucell(ucell_in), orb_cutoff_(orb_cutoff), gd(gd_in), - kv(kv_in), pX(pX_in), pc(pc_in), pmat(pmat_in), ispin_ks(ispin_ks) - { + kv(kv_in), pX(pX_in), pc(pc_in), pmat(pmat_in), ispin_ks(ispin_ks), + factor_(factor_in), dm_pq_(dm_pq_in) + { ModuleBase::TITLE("OperatorLRHxc", "OperatorLRHxc"); - std::cout<<"Initializing OperatorLRHxc"<cal_type = hamilt::calculation_type::lcao_gint; + this->cal_type = cal_type_in; this->is_first_node = true; this->hR = std::unique_ptr>(new hamilt::HContainer(&pmat_in)); LR_Util::initialize_HR(*this->hR, ucell_in, gd_in, orb_cutoff); @@ -54,8 +66,14 @@ namespace LR const int ngk_ik = 0, const bool is_first_node = false) const override; + // Valid only while the input density matrices are unchanged. Callers create + // one cache per input spin and vector, and share it between output spins. + using TransitionDensityCache = std::map*, + std::vector>>; + void act_with_shared_density(const T* psi_in, T* hpsi, TransitionDensityCache& density_cache) const; + private: - void grid_calculation(const int& nbands)const; + void grid_calculation(TransitionDensityCache& density_cache) const; //global sizes const int& nspin; @@ -70,25 +88,48 @@ namespace LR const psi::Psi& psi_ks = nullptr; /// transition density matrix - std::unique_ptr>& DM_trans; + module_dm::DensityMatrix& DM_trans; /// transition hamiltonian in AO representation std::unique_ptr> hR = nullptr; /// parallel info const Parallel_2D& pc; - const std::vector& pX; + const std::vector& pX; // output vector, OV/OO/VV const Parallel_Orbitals& pmat; - std::weak_ptr pot; + std::weak_ptr pot; const UnitCell& ucell; std::vector orb_cutoff_; const Grid_Driver& gd; + MO_TO_AO_TYPE dm_pq_ = MO_TO_AO_TYPE::CC_vo; + const T factor_ = (T)1.0; + /// test mutable bool first_print = true; }; + + /// Apply one node, reusing only Hxc grid densities. Other operators retain + /// their original action and ordering (including diagonal and EXX terms). + template + void act_with_shared_density(hamilt::Operator* node, + const int nbasis, + const T* psi_in, + T* hpsi, + typename OperatorLRHxc::TransitionDensityCache& density_cache) + { + auto* hxc = dynamic_cast*>(node); + if (hxc != nullptr) + { + hxc->act_with_shared_density(psi_in, hpsi, density_cache); + } + else + { + node->act(1, nbasis, 1, psi_in, hpsi); + } + } } #endif // ABACUS_SOURCE_LCAO_MODULE_LR_OPERATOR_CASIDA_OPERATOR_LR_HXC_H diff --git a/source/source_lcao/module_lr/potentials/pot_grad_xc.cpp b/source/source_lcao/module_lr/potentials/pot_grad_xc.cpp new file mode 100644 index 00000000000..d3bd46fbb0d --- /dev/null +++ b/source/source_lcao/module_lr/potentials/pot_grad_xc.cpp @@ -0,0 +1,265 @@ +#include "pot_grad_xc.h" +#include +#include +#include "source_io/module_parameter/parameter.h" +#include "source_lcao/module_lr/potentials/xc_kernel.h" +#include "source_base/timer.h" +#include "source_hamilt/module_xc/xc_functional.h" +#include "source_lcao/module_lr/utils/lr_util.h" +#include "source_lcao/module_lr/utils/lr_util_xc.hpp" +#include +namespace LR +{ + using Vec3 = ModuleBase::Vector3; + PotGradXCLR::Scratch& PotGradXCLR::scratch() + { + static Scratch sc; // see the comment on `Scratch` in the header + return sc; + } + + void PotGradXCLR::Scratch::alloc(const int nrxx, const bool two_channel, const bool gga) + { + // resize() on an already-large vector is a no-op, so only the first call allocates. + if (static_cast(this->vtmp.size()) < nrxx) { this->vtmp.resize(nrxx); } + if (gga) + { + if (static_cast(this->gdot.size()) < nrxx) { this->gdot.resize(nrxx); } + if (static_cast(this->div.size()) < nrxx) { this->div.resize(nrxx); } + const int nch = two_channel ? 2 : 1; + for (int is = 0; is < nch; ++is) + { + if (static_cast(this->drho1[is].size()) < nrxx) { this->drho1[is].resize(nrxx); } + } + } + } + + // constructor for exchange-correlation kernel + PotGradXCLR::PotGradXCLR(const KernelXC& xc_kernel, const ModulePW::PW_Basis& rho_basis, const UnitCell& ucell, + const int& nrxx, const bool triplet) + :xc_kernel_components_(xc_kernel), triplet_(triplet), + PotLRBase(rho_basis, LR_Util::kernel_nspin(), nrxx, ucell.tpiba) + {} + + /// $v^{(2)}(r)=\iint dr'dr''\,g^{xc}(r,r',r'')\rho^1(r')\rho^1(r'')$, i.e. the third functional + /// derivative of $E_{xc}$ contracted twice with the transition density $\rho^1$ (no factor 1/2). + /// + /// All the coefficients and spin sums live in `KernelXC::GxcCoef`, so this is a plain + /// transcription of the boxed formula and is identical for nspin=1, singlet and triplet. + /// + /// Worth stating once: for a GGA, $g^{xc}$ is NOT the whole of $v^{(2)}$. Because + /// $\sigma=\nabla\rho\cdot\nabla\rho$ is quadratic in the density it has a non-vanishing + /// *second* derivative along $\rho^1$, which drags the second-order kernels $f^{\rho\sigma}$ + /// and $f^{\sigma\sigma}$ into the answer (the $a_q$, $\boldsymbol{e}_q$, $c_s$, $c_t$ terms). + void PotGradXCLR::cal_v_eff(double** rho, const UnitCell& ucell, ModuleBase::matrix& v_eff, const std::vector& ispin_op) const + { + ModuleBase::TITLE("PotGradXCLR", "cal_v_eff"); + ModuleBase::timer::start("PotGradXCLR", "cal_v_eff"); + const auto& kxc = this->xc_kernel_components_; + + if (kxc.openshell) + { + throw std::domain_error("open shell (S2_updown) unfinished in " + + std::string(__FILE__) + " line " + std::to_string(__LINE__)); + } + const auto& g = kxc.gxc(this->triplet_); + + // Branch on THIS kernel's own GGA-ness, not the ground state's: in a cross-functional + // run (e.g. TDLDA@PBE) `XC_Functional::get_func_type()` reflects `dft_functional` (PBE, + // GGA) while `kxc` was built for `xc_kernel` (lda) and never filled `drho_gs_`. + if (!kxc.is_gga()) // LDA: only the $g^{\rho\rho\rho}$ term survives + { + const double* const a_s2 = g.a_s2.data(); + const double* const r1 = rho[0]; + double* const v = v_eff.c; +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ir = 0;ir < nrxx_;++ir) + { + v[ir] += ModuleBase::e2 * a_s2[ir] * r1[ir] * r1[ir]; + } + } + else // GGA or HYB_GGA + { + scratch().alloc(nrxx_, /*two_channel=*/false, /*gga=*/true); + Vec3* const drho1 = scratch().drho1[0].data(); // transition density gradient + LR_Util::grad(rho[0], drho1, this->rho_basis_, this->tpiba_); + + double* const v_tmp = scratch().vtmp.data(); + Vec3* const gdot_terms = scratch().gdot.data(); + const Vec3* const dgs = kxc.drho_gs.at(0).data(); + const double* const r1 = rho[0]; + const double* const e_s2 = g.e_s2.data(); const double* const e_st = g.e_st.data(); + const double* const e_t2 = g.e_t2.data(); const double* const e_q = g.e_q.data(); + const double* const c_s = g.c_s.data(); const double* const c_t = g.c_t.data(); + const double* const a_s2 = g.a_s2.data(); const double* const a_st = g.a_st.data(); + const double* const a_t2 = g.a_t2.data(); const double* const a_q = g.a_q.data(); + + // 1. the vector under the divergence, accumulated negated so that `grad_dot` yields + // $-\nabla\cdot\boldsymbol{E}$. The four $e$ coefficients share the same + // $\nabla\rho^{gs}$ direction (see `KernelXC::GxcCoef`), so it is pulled out of + // their sum -- exact, and it keeps this bandwidth-bound loop reading 4 doubles per + // point instead of 12. +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ir = 0;ir < nrxx_;++ir) + { + const Vec3& drho = dgs[ir]; // $\nabla\rho$ + const double s = r1[ir]; // $\rho^1$ + const double t = drho * drho1[ir]; // $\nabla\rho\cdot\nabla\rho^1$ + const double q = drho1[ir] * drho1[ir]; // $\nabla\rho^1\cdot\nabla\rho^1$ + const double e = e_s2[ir] * (s * s) + e_st[ir] * (s * t) + + e_t2[ir] * (t * t) + e_q[ir] * q; + gdot_terms[ir] = -(drho * e + drho1[ir] * (c_s[ir] * s + c_t[ir] * t)); + } + XC_Functional::grad_dot(gdot_terms, v_tmp, &this->rho_basis_, this->tpiba_); + + // 2. the local terms $A$ +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ir = 0;ir < nrxx_;++ir) + { + const double s = r1[ir]; + const double t = dgs[ir] * drho1[ir]; + const double q = drho1[ir] * drho1[ir]; + v_tmp[ir] += a_s2[ir] * (s * s) + a_st[ir] * (s * t) + + a_t2[ir] * (t * t) + a_q[ir] * q; + } + BlasConnector::axpy(nrxx_, ModuleBase::e2, v_tmp, 1, v_eff.c, 1); + } + + ModuleBase::timer::end("PotGradXCLR", "cal_v_eff"); + } + + + /// $v^{(2)}_\tau=A_\tau-\nabla\cdot\boldsymbol{E}_\tau$ for the open-shell case, contracted + /// straight out of the raw libxc arrays (there is no useful pre-contraction: the free spin + /// $\tau$ stays open, so a `GxcCoef`-style cache would cost ~117 doubles per grid point + /// against the 35 the third-order arrays already occupy). + /// + /// With $s_\sigma=\rho^1_\sigma$, $t_{ab}=\nabla\rho_a\cdot\nabla\rho^1_b$ (NOT symmetric) + /// and $q_{ab}=\nabla\rho^1_a\cdot\nabla\rho^1_b$, the $\lambda$-derivatives of libxc's + /// three sigma variables are + /// $S=(2t_{uu},\;t_{ud}+t_{du},\;2t_{dd})$, $Q=(2q_{uu},\;2q_{ud},\;2q_{dd})$, + /// and with $D=\sum_\sigma s_\sigma\partial_{\rho_\sigma}+\sum_a S_a\partial_{\sigma_a}$, + /// $A_\tau = D^2 e^{\rho_\tau} + \sum_a Q_a e^{\rho_\tau\sigma_a}$, + /// $u'_a = D\,e^{\sigma_a}$, $u''_a = D^2 e^{\sigma_a}+\sum_b Q_b e^{\sigma_a\sigma_b}$, + /// $\boldsymbol{E}_\tau=\sum_a\theta^\tau_a + /// \big(u''_a\,\nabla\rho_{c(\tau,a)} + 2u'_a\,\nabla\rho^1_{c(\tau,a)}\big)$. + /// The channel selector $c(\tau,a)$ is what produces the closed-shell $\tilde\theta$: for the + /// triplet $\nabla\rho^1_d=-\nabla\rho^1_u$, which is invisible in any singlet-only test. + void PotGradXCLR::cal_v_eff_openshell(const double* const* const rho1, const UnitCell& ucell, + ModuleBase::matrix& v_eff, const int tau) const + { + ModuleBase::TITLE("PotGradXCLR", "cal_v_eff_openshell"); + ModuleBase::timer::start("PotGradXCLR", "cal_v_eff_openshell"); + using namespace LR::libxc_idx; + const auto& kxc = this->xc_kernel_components_; + assert(tau == 0 || tau == 1); + const std::vector& v2rs = kxc.v2rhosigma; + const std::vector& v2s2 = kxc.v2sigma2; + const std::vector& v3r3 = kxc.v3rho3; + const std::vector& v3r2s = kxc.v3rho2sigma; + const std::vector& v3rs2 = kxc.v3rhosigma2; + const std::vector& v3s3 = kxc.v3sigma3; + + // Branch on THIS kernel's own GGA-ness, not the ground state's (see `cal_v_eff`). + if (!kxc.is_gga()) // LDA: only $g^{\rho\rho\rho}$ survives + { + const double* const r1u = rho1[0]; const double* const r1d = rho1[1]; + const double* const g3 = v3r3.data(); + double* const v = v_eff.c; +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ir = 0;ir < nrxx_;++ir) + { + const double s[2] = { r1u[ir], r1d[ir] }; + double a = 0.; + for (int s0 = 0;s0 < 2;++s0) { + for (int s1 = 0;s1 < 2;++s1) { a += s[s0] * s[s1] * g3[ir * 4 + r3(tau, s0, s1)]; } } + v[ir] += ModuleBase::e2 * a; + } + ModuleBase::timer::end("PotGradXCLR", "cal_v_eff_openshell"); + return; + } + + // GGA / HYB_GGA + scratch().alloc(nrxx_, /*two_channel=*/true, /*gga=*/true); + Vec3* const drho1[2] = { scratch().drho1[0].data(), scratch().drho1[1].data() }; + for (int is : {0, 1}) { LR_Util::grad(rho1[is], drho1[is], this->rho_basis_, this->tpiba_); } + + double* const v_tmp = scratch().vtmp.data(); + Vec3* const gdot_terms = scratch().gdot.data(); +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ir = 0;ir < nrxx_;++ir) + { + const int o4 = ir * 4; + const int o6 = ir * 6; + const int o9 = ir * 9; + const int o10 = ir * 10; + const int o12 = ir * 12; + const ModuleBase::Vector3 drho[2] = { kxc.drho_gs[0][ir], kxc.drho_gs[1][ir] }; + const ModuleBase::Vector3 dr1[2] = { drho1[0][ir], drho1[1][ir] }; + const double s[2] = { rho1[0][ir], rho1[1][ir] }; + + // $t_{ab}=\nabla\rho_a\cdot\nabla\rho^1_b$, then $S_a=\mathrm{d}\sigma_a/\mathrm{d}\lambda$ + const double t00 = drho[0] * dr1[0]; + const double t01 = drho[0] * dr1[1]; + const double t10 = drho[1] * dr1[0]; + const double t11 = drho[1] * dr1[1]; + const double S[3] = { 2. * t00, t01 + t10, 2. * t11 }; + // $Q_a=\mathrm{d}^2\sigma_a/\mathrm{d}\lambda^2$ + const double Q[3] = { 2. * (dr1[0] * dr1[0]), 2. * (dr1[0] * dr1[1]), 2. * (dr1[1] * dr1[1]) }; + + // ---- the local part $A_\tau$ ---- + double A = 0.; + for (int s0 = 0;s0 < 2;++s0) { + for (int s1 = 0;s1 < 2;++s1) { A += s[s0] * s[s1] * v3r3[o4 + r3(tau, s0, s1)]; } } + for (int s0 = 0;s0 < 2;++s0) { + for (int a = 0;a < 3;++a) { A += 2. * s[s0] * S[a] * v3r2s[o9 + r2s(tau, s0, a)]; } } + for (int a = 0;a < 3;++a) { + for (int b = 0;b < 3;++b) { A += S[a] * S[b] * v3rs2[o12 + rs2(tau, a, b)]; } } + for (int a = 0;a < 3;++a) { A += Q[a] * v2rs[o6 + rs(tau, a)]; } + v_tmp[ir] = A; + + // ---- the divergence part $\boldsymbol{E}_\tau$ ---- + // $u'_a$ and $u''_a$ do not depend on $\tau$; only the $\theta$ weights and the + // channel selector below do. + ModuleBase::Vector3 E(0., 0., 0.); + for (int a = 0;a < 3;++a) + { + const double th = theta[tau][a]; + if (th == 0.) { continue; } + const int c = chan[tau][a]; + double up = 0.; + double upp = 0.; + for (int s0 = 0;s0 < 2;++s0) { up += s[s0] * v2rs[o6 + rs(s0, a)]; } + for (int b = 0;b < 3;++b) { up += S[b] * v2s2[o6 + p2[a][b]]; } + + for (int s0 = 0;s0 < 2;++s0) { + for (int s1 = 0;s1 < 2;++s1) { upp += s[s0] * s[s1] * v3r2s[o9 + r2s(s0, s1, a)]; } } + for (int s0 = 0;s0 < 2;++s0) { + for (int b = 0;b < 3;++b) { upp += 2. * s[s0] * S[b] * v3rs2[o12 + rs2(s0, a, b)]; } } + for (int b = 0;b < 3;++b) { + for (int cc = 0;cc < 3;++cc) { upp += S[b] * S[cc] * v3s3[o10 + p3[a][b][cc]]; } } + for (int b = 0;b < 3;++b) { upp += Q[b] * v2s2[o6 + p2[a][b]]; } + + E += th * (drho[c] * upp + dr1[c] * (2. * up)); + } + gdot_terms[ir] = -E; // `grad_dot` then yields $-\nabla\cdot\boldsymbol{E}_\tau$ + } + double* const div = scratch().div.data(); + XC_Functional::grad_dot(gdot_terms, div, &this->rho_basis_, this->tpiba_); + double* const vout = v_eff.c; +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ir = 0;ir < nrxx_;++ir) { vout[ir] += ModuleBase::e2 * (v_tmp[ir] + div[ir]); } + ModuleBase::timer::end("PotGradXCLR", "cal_v_eff_openshell"); + } +} diff --git a/source/source_lcao/module_lr/potentials/pot_grad_xc.h b/source/source_lcao/module_lr/potentials/pot_grad_xc.h new file mode 100644 index 00000000000..d71d86123c0 --- /dev/null +++ b/source/source_lcao/module_lr/potentials/pot_grad_xc.h @@ -0,0 +1,47 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_POTENTIALS_POT_GRAD_XC_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_POTENTIALS_POT_GRAD_XC_H +#include "source_lcao/module_lr/potentials/xc_kernel.h" +#include "source_lcao/module_lr/potentials/pot_lr_base.h" + +namespace LR +{ + /// the "potential" contributing to RHS of Z-vector equation + /// from the derivative of xc kernel + class PotGradXCLR : public PotLRBase + { + public: + /// constructor for exchange-correlation kernel + /// `triplet` selects which spin combination of $g^{xc}$ to use. It matters: + /// the triplet kernel $K^T_{xc}=f_{uu}-f_{ud}$ is NOT zero for a local functional. + PotGradXCLR(const KernelXC& xc_kernel_in, const ModulePW::PW_Basis& rho_basis, const UnitCell& ucell, + const int& nrxx, const bool triplet = false); + ~PotGradXCLR() {} + virtual void cal_v_eff(double** rho, const UnitCell& ucell, ModuleBase::matrix& v_eff, const std::vector& ispin_op = { 0,0 }) const override; + /// @brief Open-shell (spin-unrestricted) $v^{(2)}_\tau$. + /// Unlike the closed-shell case this is NOT a per-(sl,sr) block object: the free spin + /// $\tau$ is the output index, but BOTH channels of the transition density enter the + /// same expression, so `rho1` must carry `rho1[0]` and `rho1[1]` together. + /// $v^{(2)}_\tau = A_\tau - \nabla\cdot\boldsymbol{E}_\tau$ + void cal_v_eff_openshell(const double* const* const rho1, const UnitCell& ucell, + ModuleBase::matrix& v_eff, const int tau) const; + /// kernel components from PotHxcLR + const KernelXC& xc_kernel_components_; + const bool triplet_ = false; + + private: + /// Scratch, shared by every `PotGradXCLR` and grown on demand. These used to be + /// allocated (and value-initialized) on every call. Safe to share because + /// `cal_v_eff` is only ever entered from a single thread (all the OpenMP is inside). + struct Scratch + { + std::vector> drho1[2]; ///< $\nabla\rho^1_\sigma$ + std::vector> gdot; ///< integrand of the divergence + std::vector vtmp; ///< local part $A$ + std::vector div; ///< $-\nabla\cdot\boldsymbol{E}$ + void alloc(const int nrxx, const bool two_channel, const bool gga); + }; + static Scratch& scratch(); + }; + +} +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_POTENTIALS_POT_GRAD_XC_H diff --git a/source/source_lcao/module_lr/potentials/pot_hxc_lrtd.cpp b/source/source_lcao/module_lr/potentials/pot_hxc_lrtd.cpp index 9fa97e18899..01039bfb65c 100644 --- a/source/source_lcao/module_lr/potentials/pot_hxc_lrtd.cpp +++ b/source/source_lcao/module_lr/potentials/pot_hxc_lrtd.cpp @@ -4,49 +4,170 @@ #include "source_base/timer.h" #include "source_hamilt/module_xc/xc_functional.h" #include +#include +#include +#include #include "source_lcao/module_lr/utils/lr_util.h" #include "source_lcao/module_lr/utils/lr_util_xc.hpp" #define FXC_PARA_TYPE const double* const rho, ModuleBase::matrix& v_eff, const std::vector& ispin_op = { 0,0 } namespace LR { + using Vec3 = ModuleBase::Vector3; + std::shared_ptr PotHxcLR::make_kernel(const std::string& xc_kernel, + const ModulePW::PW_Basis& rho_basis, const UnitCell& ucell, const Charge& chg_gs, + const Parallel_Grid& pgrid, const bool openshell, const int gxc_spin, + const std::vector& lr_init_xc_kernel) + { //calls XC_Functional::set_func_type and libxc + return std::make_shared(rho_basis, ucell, chg_gs, pgrid, LR_Util::kernel_nspin(), + xc_kernel, lr_init_xc_kernel, openshell, gxc_spin); + } + + PotHxcLR::PotHxcLR(const std::string& xc_kernel, const ModulePW::PW_Basis& rho_basis, + const UnitCell& ucell, const Charge& chg_gs, const Parallel_Grid& pgrid, + const SpinType& st, const std::vector& lr_init_xc_kernel) + : PotHxcLR(xc_kernel, rho_basis, ucell, chg_gs, pgrid, st, + lr_init_xc_kernel, KernelXC::GxcSpin::NoGxc) + { + } + // constructor for exchange-correlation kernel PotHxcLR::PotHxcLR(const std::string& xc_kernel, const ModulePW::PW_Basis& rho_basis, const UnitCell& ucell, const Charge& chg_gs/*ground state*/, const Parallel_Grid& pgrid, - const SpinType& st, const std::vector& lr_init_xc_kernel) - :xc_kernel_(xc_kernel), tpiba_(ucell.tpiba), spin_type_(st), rho_basis_(rho_basis), nrxx_(chg_gs.nrxx), - nspin_(PARAM.inp.nspin == 1 || (PARAM.inp.nspin == 4 && !PARAM.globalv.domag && !PARAM.globalv.domag_z) ? 1 : 2), + const SpinType& st, const std::vector& lr_init_xc_kernel, const int gxc_spin) + :PotLRBase(rho_basis, LR_Util::kernel_nspin(), chg_gs.nrxx, ucell.tpiba), + xc_kernel_(xc_kernel), spin_type_(st), + pot_hartree_(LR_Util::make_unique(&rho_basis)), + xc_kernel_components_(make_kernel(xc_kernel, rho_basis, ucell, chg_gs, pgrid, (st == SpinType::S2_updown), gxc_spin, lr_init_xc_kernel)), + xc_type_(XCType(XC_Functional::get_func_type())) + { + if (LR_Util::has_local_xc(xc_kernel)) { this->set_integral_func(this->spin_type_, this->xc_type_); } + const bool gga = (this->xc_type_ == XCType::GGA || this->xc_type_ == XCType::HYB_GGA); + this->build_spin_combos(gga); + } + + PotHxcLR::PotHxcLR(std::shared_ptr kernel, const std::string& xc_kernel, + const ModulePW::PW_Basis& rho_basis, const UnitCell& ucell, const int nrxx, const SpinType& st) + :PotLRBase(rho_basis, LR_Util::kernel_nspin(), nrxx, ucell.tpiba), + xc_kernel_(xc_kernel), spin_type_(st), pot_hartree_(LR_Util::make_unique(&rho_basis)), - xc_kernel_components_(rho_basis, ucell, chg_gs, pgrid, nspin_, xc_kernel, lr_init_xc_kernel, (st == SpinType::S2_updown)), //call XC_Functional::set_func_type and libxc + xc_kernel_components_(std::move(kernel)), xc_type_(XCType(XC_Functional::get_func_type())) { - if (std::set({ "lda", "pwlda", "pbe", "hse" }).count(xc_kernel)) { this->set_integral_func(this->spin_type_, this->xc_type_); } + assert(this->xc_kernel_components_ != nullptr); + if (LR_Util::has_local_xc(xc_kernel)) { this->set_integral_func(this->spin_type_, this->xc_type_); } + const bool gga = (this->xc_type_ == XCType::GGA || this->xc_type_ == XCType::HYB_GGA); + this->build_spin_combos(gga); + } + + + PotHxcLR::Scratch& PotHxcLR::scratch() + { + static Scratch s; // see the comment on `Scratch` in the header + return s; + } + + void PotHxcLR::Scratch::alloc(const int nrxx, const int npw, const int nmaxgr, const bool gga) + { + // resize() on an already-large vector is a no-op, so only the first call allocates. + if (static_cast(this->vxc.size()) < nrxx) { this->vxc.resize(nrxx); } + if (static_cast(this->rhog.size()) < npw) { this->rhog.resize(npw); } + if (static_cast(this->vg.size()) < nmaxgr) { this->vg.resize(nmaxgr); } + if (gga) + { + if (static_cast(this->drho.size()) < nrxx) { this->drho.resize(nrxx); } + if (static_cast(this->gdot.size()) < nrxx) { this->gdot.resize(nrxx); } + } + } + + void PotHxcLR::build_spin_combos(const bool gga) + { + if (this->nspin != 2) { return; } // nspin=1 kernels have a single component already + double sign = 0.; + switch (this->spin_type_) + { + case SpinType::S2_singlet: case SpinType::S2_gs: sign = 1.; break; + case SpinType::S2_triplet: sign = -1.; break; + default: return; // S2_updown indexes single components, nothing to pre-contract + } + auto& fxc = *this->xc_kernel_components_; + if (this->v2rho2_comb_.empty() && !fxc.v2rho2.empty()) + { + this->v2rho2_comb_.resize(nrxx); + const double* const f = fxc.v2rho2.data(); + double* const c = this->v2rho2_comb_.data(); +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ir = 0; ir < nrxx; ++ir) { c[ir] = f[3 * ir] + sign * f[3 * ir + 1]; } + } + if (gga && this->vsigma_comb_.empty() && !fxc.vsigma.empty()) + { + this->vsigma_comb_.resize(nrxx); + const double* const f = fxc.vsigma.data(); + double* const c = this->vsigma_comb_.data(); +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ir = 0; ir < nrxx; ++ir) { c[ir] = 2. * f[3 * ir] + sign * f[3 * ir + 1]; } + } + } + + void PotHxcLR::add_v_hartree(const UnitCell& ucell, ModuleBase::matrix& v_eff, const double factor) const + { + const int npw = this->rho_basis_.npw; + const std::complex* const rhog = scratch().rhog.data(); + std::complex* const vg = scratch().vg.data(); + const int ig0 = this->rho_basis_.ig_gge0; + const double pre = ModuleBase::e2 * ModuleBase::FOUR_PI / ucell.tpiba2; +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ig = 0; ig < npw; ++ig) + { + vg[ig] = (ig == ig0) ? std::complex(0., 0.) // V(G=0) = 0 + : (pre / this->rho_basis_.gg[ig]) * rhog[ig]; + } + this->rho_basis_.recip2real(vg, vg); + double* const v = v_eff.c; +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ir = 0; ir < nrxx; ++ir) { v[ir] += factor * vg[ir].real(); } } - void PotHxcLR::cal_v_eff(double** rho, const UnitCell& ucell, ModuleBase::matrix& v_eff, const std::vector& ispin_op) + void PotHxcLR::cal_v_eff(double** rho, const UnitCell& ucell, ModuleBase::matrix& v_eff, const std::vector& ispin_op) const { ModuleBase::TITLE("PotHxcLR", "cal_v_eff"); ModuleBase::timer::start("PotHxcLR", "cal_v_eff"); - auto& fxc = this->xc_kernel_components_; - // Hartree + // Hartree prefactor (0 = this spin combination has no Hartree term at all) + double hartree_factor = 0.; switch (this->spin_type_) { - case SpinType::S1: case SpinType::S2_updown: - v_eff += elecstate::H_Hartree_pw::v_hartree(ucell, const_cast(&this->rho_basis_), 1, rho); - break; - case SpinType::S2_singlet: - v_eff += 2 * elecstate::H_Hartree_pw::v_hartree(ucell, const_cast(&this->rho_basis_), 1, rho); - break; - default: - break; + case SpinType::S1_gs: case SpinType::S2_updown: case SpinType::S2_gs: hartree_factor = 1.; break; + case SpinType::S1: case SpinType::S2_singlet: hartree_factor = 2.; break; + default: break; } + const bool local_xc = LR_Util::has_local_xc(this->xc_kernel_); + const bool gga = local_xc && (this->xc_type_ == XCType::GGA || this->xc_type_ == XCType::HYB_GGA); + + // $\rho^X(G)$ is needed by the Hartree term and by the GGA gradient, and used to be + // transformed twice (once inside `H_Hartree_pw::v_hartree`, once inside `LR_Util::grad`). + if (hartree_factor != 0. || gga) + { + scratch().alloc(nrxx, this->rho_basis_.npw, this->rho_basis_.nmaxgr, gga); + this->rho_basis_.real2recip(rho[0], scratch().rhog.data()); + } + if (hartree_factor != 0.) { this->add_v_hartree(ucell, v_eff, hartree_factor); } + // XC - if (this->xc_kernel_ == "rpa" || this->xc_kernel_ == "hf") { + if (!local_xc) + { ModuleBase::timer::end("PotHxcLR", "cal_v_eff"); - return; - } // no xc + return; + } #ifdef __LIBXC - this->kernel_to_potential_[spin_type_](rho[0], v_eff, ispin_op); + this->kernel_to_potential_.at(spin_type_)(rho[0], v_eff, ispin_op); #else throw std::domain_error("GlobalV::XC_Functional::get_func_type() =" + std::to_string(XC_Functional::get_func_type()) + " unfinished in " + std::string(__FILE__) + " line " + std::to_string(__LINE__)); @@ -54,244 +175,243 @@ namespace LR ModuleBase::timer::end("PotHxcLR", "cal_v_eff"); } + // In every integrand below the kernel arrays are read through raw pointers hoisted out of + // the loop, and the loops carry an OpenMP directive. Both matter: these are pure + // element-wise passes over nrxx that used to run single-threaded, with `.at()` (a bounds + // check per access, which also blocks vectorization) on every kernel component. void PotHxcLR::set_integral_func(const SpinType& s, const XCType& xc) { auto& funcs = this->kernel_to_potential_; - auto& fxc = this->xc_kernel_components_; - if (xc == XCType::LDA) { switch (s) - { - case SpinType::S1: - funcs[s] = [this, &fxc](FXC_PARA_TYPE)->void - { - for (int ir = 0;ir < nrxx;++ir) { v_eff(0, ir) += ModuleBase::e2 * fxc.v2rho2.at(ir) * rho[ir]; } - }; - break; - case SpinType::S2_singlet: - funcs[s] = [this, &fxc](FXC_PARA_TYPE)->void - { - for (int ir = 0;ir < nrxx;++ir) - { - const int irs0 = 3 * ir; - const int irs1 = irs0 + 1; - v_eff(0, ir) += ModuleBase::e2 * (fxc.v2rho2.at(irs0) + fxc.v2rho2.at(irs1)) * rho[ir]; - } - }; - break; - case SpinType::S2_triplet: - funcs[s] = [this, &fxc](FXC_PARA_TYPE)->void - { - for (int ir = 0;ir < nrxx;++ir) - { - const int irs0 = 3 * ir; - const int irs1 = irs0 + 1; - v_eff(0, ir) += ModuleBase::e2 * (fxc.v2rho2.at(irs0) - fxc.v2rho2.at(irs1)) * rho[ir]; - } - }; - break; - case SpinType::S2_updown: - funcs[s] = [this, &fxc](FXC_PARA_TYPE)->void - { - assert(ispin_op.size() >= 2); - const int is = ispin_op[0] + ispin_op[1]; - for (int ir = 0;ir < nrxx;++ir) { v_eff(0, ir) += ModuleBase::e2 * fxc.v2rho2.at(3 * ir + is) * rho[ir]; } - }; - break; - default: - throw std::domain_error("SpinType =" + std::to_string(static_cast(s)) - + " unfinished in " + std::string(__FILE__) + " line " + std::to_string(__LINE__)); - break; - } - } else if (xc == XCType::GGA || xc == XCType::HYB_GGA) { switch (s) - { - case SpinType::S1: - funcs[s] = [this, &fxc](FXC_PARA_TYPE)->void - { - // test: output drho - // double thr = 1e-1; - // auto out_thr = [this, &thr](const double* v) { - // for (int ir = 0;ir < nrxx;++ir) if (std::abs(v[ir]) > thr) std::cout << v[ir] << " "; - // std::cout << std::endl;}; - // auto out_thr3 = [this, &thr](const std::vector>& v) { - // for (int ir = 0;ir < nrxx;++ir) if (std::abs(v.at(ir).x) > thr) std::cout << v.at(ir).x << " "; - // std::cout << std::endl; - // for (int ir = 0;ir < nrxx;++ir) if (std::abs(v.at(ir).y) > thr) std::cout << v.at(ir).y << " "; - // std::cout << std::endl; - // for (int ir = 0;ir < nrxx;++ir) if (std::abs(v.at(ir).z) > thr) std::cout << v.at(ir).z << " "; - // std::cout << std::endl;}; - - std::vector> drho(nrxx); // transition density gradient - LR_Util::grad(rho, drho.data(), this->rho_basis_, this->tpiba_); - - std::vector vxc_tmp(nrxx, 0.0); - - //1. $\partial E/\partial\rho = 2f^{\rho\sigma}*\nabla\rho*\rho_1+4f^{\sigma\sigma}\nabla\rho(\nabla\rho\cdot\nabla\rho_1)+2v^\sigma\nabla\rho_1$ - std::vector> e_drho(nrxx); - for (int ir = 0;ir < nrxx;++ir) + auto& fxc = *this->xc_kernel_components_; + if (xc == XCType::LDA) { + switch (s) + { + case SpinType::S1: + case SpinType::S1_gs: + { + // S1_gs is exactly half of S1 (see the SpinType doc in the header). + const double prefac = (s == SpinType::S1_gs) ? 1.0 : 2.0; + funcs[s] = [this, &fxc, prefac](FXC_PARA_TYPE)->void { - e_drho[ir] = -(fxc.v2rhosigma_2drho.at(ir) * rho[ir] - + fxc.v2sigma2_4drho.at(ir) * (fxc.drho_gs.at(0).at(ir) * drho.at(ir)) - + drho.at(ir) * fxc.vsigma.at(ir) * 2.); - } - XC_Functional::grad_dot(e_drho.data(), vxc_tmp.data(), &this->rho_basis_, this->tpiba_); - - // 2. $f^{\rho\rho}\rho_1+2f^{\rho\sigma}\nabla\rho\cdot\nabla\rho_1$ - for (int ir = 0;ir < nrxx;++ir) - { - vxc_tmp[ir] += (fxc.v2rho2.at(ir) * rho[ir] - + fxc.v2rhosigma_2drho.at(ir) * drho.at(ir)); - } - BlasConnector::axpy(nrxx, ModuleBase::e2, vxc_tmp.data(), 1, v_eff.c, 1); - }; - break; - case SpinType::S2_singlet: - funcs[s] = [this, &fxc](FXC_PARA_TYPE)-> void - { - std::vector> drho(nrxx); // transition density gradient - LR_Util::grad(rho, drho.data(), this->rho_basis_, this->tpiba_); - - std::vector vxc_tmp(nrxx, 0.0); - - // 1. the terms in grad_dot int f_uu - std::vector> gdot_terms(nrxx); - for (int ir = 0;ir < nrxx;++ir) - { - // gdot terms in f_uu - gdot_terms[ir] = -(rho[ir] * fxc.v2rhosigma_drho_singlet.at(ir) - + (fxc.drho_gs.at(0).at(ir) * drho.at(ir)) * fxc.v2sigma2_drho_singlet.at(ir) - + drho.at(ir) * (fxc.vsigma.at(ir * 3) * 2. + fxc.vsigma.at(ir * 3 + 1))); - } - XC_Functional::grad_dot(gdot_terms.data(), vxc_tmp.data(), &this->rho_basis_, this->tpiba_); - - // 2. terms not in grad_dot - for (int ir = 0;ir < nrxx;++ir) + const double* const f = fxc.v2rho2.data(); + double* const v = v_eff.c; + const double fac = ModuleBase::e2 * prefac; +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ir = 0;ir < nrxx;++ir) { v[ir] += fac * f[ir] * rho[ir]; } + }; + break; + } + case SpinType::S2_singlet: + case SpinType::S2_gs: + case SpinType::S2_triplet: + { + // S2_gs is exactly half of S2_singlet (see the SpinType doc in the header). + // The singlet/triplet difference is entirely in `v2rho2_comb_`'s sign. + const double prefac = (s == SpinType::S2_gs) ? 0.5 : 1.0; + funcs[s] = [this, prefac](FXC_PARA_TYPE)->void { - vxc_tmp[ir] += rho[ir] * (fxc.v2rho2.at(ir * 3) + fxc.v2rho2.at(ir * 3 + 1)) - + drho.at(ir) * fxc.v2rhosigma_drho_singlet.at(ir); - } - BlasConnector::axpy(nrxx, ModuleBase::e2, vxc_tmp.data(), 1, v_eff.c, 1); - }; - break; - case SpinType::S2_triplet: - funcs[s] = [this, &fxc](FXC_PARA_TYPE)->void - { - std::vector> drho(nrxx); // transition density gradient - LR_Util::grad(rho, drho.data(), this->rho_basis_, this->tpiba_); - - std::vector vxc_tmp(nrxx, 0.0); - - // 1. the terms in grad_dot int f_uu - std::vector> gdot_terms(nrxx); - for (int ir = 0;ir < nrxx;++ir) + const double* const f = this->v2rho2_comb_.data(); + double* const v = v_eff.c; + const double fac = ModuleBase::e2 * prefac; +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ir = 0;ir < nrxx;++ir) { v[ir] += fac * f[ir] * rho[ir]; } + }; + break; + } + case SpinType::S2_updown: + funcs[s] = [this, &fxc](FXC_PARA_TYPE)->void { - // gdot terms in f_uu - gdot_terms[ir] = -(rho[ir] * fxc.v2rhosigma_drho_triplet.at(ir) - + (fxc.drho_gs.at(0).at(ir) * drho.at(ir)) * fxc.v2sigma2_drho_triplet.at(ir) - + drho.at(ir) * (fxc.vsigma.at(ir * 3) * 2. - fxc.vsigma.at(ir * 3 + 1))); - } - XC_Functional::grad_dot(gdot_terms.data(), vxc_tmp.data(), &this->rho_basis_, this->tpiba_); - - // 2. terms not in grad_dot - for (int ir = 0;ir < nrxx;++ir) + assert(ispin_op.size() >= 2); + const int is = ispin_op[0] + ispin_op[1]; + const double* const f = fxc.v2rho2.data(); + double* const v = v_eff.c; +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif + for (int ir = 0;ir < nrxx;++ir) { v[ir] += ModuleBase::e2 * f[3 * ir + is] * rho[ir]; } + }; + break; + default: + throw std::domain_error("SpinType =" + std::to_string(static_cast(s)) + + " unfinished in " + std::string(__FILE__) + " line " + std::to_string(__LINE__)); + break; + } + } + else if (xc == XCType::GGA || xc == XCType::HYB_GGA) { + switch (s) + { + case SpinType::S1: + case SpinType::S1_gs: + { + const double prefac = (s == SpinType::S1_gs) ? 1.0 : 2.0; // S1_gs is exactly half of S1. + funcs[s] = [this, &fxc, prefac](FXC_PARA_TYPE)->void { - vxc_tmp[ir] += rho[ir] * (fxc.v2rho2.at(ir * 3) - fxc.v2rho2.at(ir * 3 + 1)) - + drho.at(ir) * fxc.v2rhosigma_drho_triplet.at(ir); - } - BlasConnector::axpy(nrxx, ModuleBase::e2, vxc_tmp.data(), 1, v_eff.c, 1); - }; - break; - case SpinType::S2_updown: - funcs[s] = [this, &fxc](FXC_PARA_TYPE)->void - { - assert(ispin_op.size() >= 2); - std::vector> drho(nrxx); // transition density gradient - LR_Util::grad(rho, drho.data(), this->rho_basis_, this->tpiba_); + // transition density gradient, from the $\rho^X(G)$ `cal_v_eff` already built + Vec3* const drho = scratch().drho.data(); + XC_Functional::grad_rho(scratch().rhog.data(), drho, &this->rho_basis_, this->tpiba_); - std::vector vxc_tmp(nrxx, 0.0); + double* const vxc = scratch().vxc.data(); + Vec3* const e_drho = scratch().gdot.data(); + const Vec3* const rs2 = fxc.v2rhosigma_2drho.data(); + const Vec3* const ss4 = fxc.v2sigma2_4drho.data(); + const Vec3* const dgs = fxc.drho_gs.at(0).data(); + const double* const vsig = fxc.vsigma.data(); + const double* const rr = fxc.v2rho2.data(); - // 1. the terms in grad_dot int f_uu - std::vector> gdot_terms(nrxx); - switch (ispin_op[0] << 1 | ispin_op[1]) - { - case 0: // (0,0) - for (int ir = 0;ir < nrxx;++ir) - { - gdot_terms[ir] = -(rho[ir] * fxc.v2rhosigma_drho_uu.at(ir) - + (fxc.drho_gs.at(0).at(ir) * drho.at(ir)) * fxc.v2sigma2_drho_uu_u.at(ir) - + (fxc.drho_gs.at(1).at(ir) * drho.at(ir)) * fxc.v2sigma2_drho_uu_d.at(ir) - + drho.at(ir) * fxc.vsigma.at(ir * 3) * 2.); - } - break; - case 1: // (0,1) + //1. $\partial E/\partial\rho = 2f^{\rho\sigma}*\nabla\rho*\rho_1+4f^{\sigma\sigma}\nabla\rho(\nabla\rho\cdot\nabla\rho_1)+2v^\sigma\nabla\rho_1$ +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif for (int ir = 0;ir < nrxx;++ir) { - gdot_terms[ir] = -(rho[ir] * fxc.v2rhosigma_drho_du.at(ir) // rho_d, drho_u - + (fxc.drho_gs.at(0).at(ir) * drho.at(ir)) * fxc.v2sigma2_drho_ud_u.at(ir) - + (fxc.drho_gs.at(1).at(ir) * drho.at(ir)) * fxc.v2sigma2_drho_ud_d.at(ir) - + drho.at(ir) * fxc.vsigma.at(ir * 3 + 1)); + e_drho[ir] = -(rs2[ir] * rho[ir] + + ss4[ir] * (dgs[ir] * drho[ir]) + + drho[ir] * (vsig[ir] * 2.)); } - break; - case 2: // (1,0) + XC_Functional::grad_dot(e_drho, vxc, &this->rho_basis_, this->tpiba_); + + // 2. $f^{\rho\rho}\rho_1+2f^{\rho\sigma}\nabla\rho\cdot\nabla\rho_1$ +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif for (int ir = 0;ir < nrxx;++ir) { - gdot_terms[ir] = -(rho[ir] * fxc.v2rhosigma_drho_ud.at(ir) // rho_u, drho_d - + (fxc.drho_gs.at(0).at(ir) * drho.at(ir)) * fxc.v2sigma2_drho_du_u.at(ir) - + (fxc.drho_gs.at(1).at(ir) * drho.at(ir)) * fxc.v2sigma2_drho_du_d.at(ir) - + drho.at(ir) * fxc.vsigma.at(ir * 3 + 1)); + vxc[ir] += rr[ir] * rho[ir] + rs2[ir] * drho[ir]; } - break; - case 3: // (1,1) + BlasConnector::axpy(nrxx, ModuleBase::e2 * prefac, vxc, 1, v_eff.c, 1); + }; + break; + } + case SpinType::S2_singlet: + case SpinType::S2_gs: + case SpinType::S2_triplet: + { + // S2_gs is exactly half of S2_singlet; the whole expression is linear in the + // kernel, so scaling the final axpy is enough. Singlet and triplet differ only + // in which pre-contracted kernel arrays are read, so they share one lambda. + const double prefac = (s == SpinType::S2_gs) ? 0.5 : 1.0; + const bool triplet = (s == SpinType::S2_triplet); + funcs[s] = [this, &fxc, prefac, triplet](FXC_PARA_TYPE)->void + { + Vec3* const drho = scratch().drho.data(); + XC_Functional::grad_rho(scratch().rhog.data(), drho, &this->rho_basis_, this->tpiba_); + + double* const vxc = scratch().vxc.data(); + Vec3* const gdot = scratch().gdot.data(); + const Vec3* const rs = triplet ? fxc.v2rhosigma_drho_triplet.data() + : fxc.v2rhosigma_drho_singlet.data(); + const Vec3* const ss = triplet ? fxc.v2sigma2_drho_triplet.data() + : fxc.v2sigma2_drho_singlet.data(); + const Vec3* const dgs = fxc.drho_gs.at(0).data(); + const double* const vsig = this->vsigma_comb_.data(); // 2*vsigma_uu -+ vsigma_ud + const double* const rr = this->v2rho2_comb_.data(); // v2rho2_uu -+ v2rho2_ud + + // 1. the terms in grad_dot int f_uu +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif for (int ir = 0;ir < nrxx;++ir) { - gdot_terms[ir] = -(rho[ir] * fxc.v2rhosigma_drho_dd.at(ir) - + (fxc.drho_gs.at(0).at(ir) * drho.at(ir)) * fxc.v2sigma2_drho_dd_u.at(ir) - + (fxc.drho_gs.at(1).at(ir) * drho.at(ir)) * fxc.v2sigma2_drho_dd_d.at(ir) - + drho.at(ir) * fxc.vsigma.at(ir * 3 + 2) * 2.); + gdot[ir] = -(rs[ir] * rho[ir] + + ss[ir] * (dgs[ir] * drho[ir]) + + drho[ir] * vsig[ir]); } - break; - default: - throw std::runtime_error("Invalid ispin_op"); - } - XC_Functional::grad_dot(gdot_terms.data(), vxc_tmp.data(), &this->rho_basis_, this->tpiba_); + XC_Functional::grad_dot(gdot, vxc, &this->rho_basis_, this->tpiba_); - // 2. terms not in grad_dot - switch (ispin_op[0] << 1 | ispin_op[1]) - { - case 0: + // 2. terms not in grad_dot +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif for (int ir = 0;ir < nrxx;++ir) { - vxc_tmp[ir] += rho[ir] * fxc.v2rho2.at(ir * 3) + drho.at(ir) * fxc.v2rhosigma_drho_uu.at(ir); + vxc[ir] += rho[ir] * rr[ir] + drho[ir] * rs[ir]; } - break; - case 1: - for (int ir = 0;ir < nrxx;++ir) + BlasConnector::axpy(nrxx, ModuleBase::e2 * prefac, vxc, 1, v_eff.c, 1); + }; + break; + } + case SpinType::S2_updown: + funcs[s] = [this, &fxc](FXC_PARA_TYPE)->void + { + assert(ispin_op.size() >= 2); + Vec3* const drho = scratch().drho.data(); + XC_Functional::grad_rho(scratch().rhog.data(), drho, &this->rho_basis_, this->tpiba_); + + double* const vxc = scratch().vxc.data(); + Vec3* const gdot = scratch().gdot.data(); + const Vec3* const dgs_u = fxc.drho_gs.at(0).data(); + const Vec3* const dgs_d = fxc.drho_gs.at(1).data(); + const double* const vsig = fxc.vsigma.data(); + const double* const rr = fxc.v2rho2.data(); + + // the four (sigma, sigma') blocks: which pre-contracted arrays to read, + // which vsigma component multiplies drho^X, and which v2rho2 component. + const int blk = ispin_op[0] << 1 | ispin_op[1]; + const Vec3* rs = nullptr; const Vec3* su = nullptr; const Vec3* sd = nullptr; + int i_vsig = 0, i_rr = 0; double vsig_fac = 1.; + // `rs2` is the v2rhosigma array used again in the non-divergence term; + // for the off-diagonal blocks it is the transposed one (ud vs du). + const Vec3* rs2 = nullptr; + switch (blk) { - vxc_tmp[ir] += rho[ir] * fxc.v2rho2.at(ir * 3 + 1) + drho.at(ir) * fxc.v2rhosigma_drho_ud.at(ir); + case 0: // (0,0) + rs = fxc.v2rhosigma_drho_uu.data(); rs2 = rs; + su = fxc.v2sigma2_drho_uu_u.data(); sd = fxc.v2sigma2_drho_uu_d.data(); + i_vsig = 0; vsig_fac = 2.; i_rr = 0; + break; + case 1: // (0,1): rho_d, drho_u + rs = fxc.v2rhosigma_drho_du.data(); rs2 = fxc.v2rhosigma_drho_ud.data(); + su = fxc.v2sigma2_drho_ud_u.data(); sd = fxc.v2sigma2_drho_ud_d.data(); + i_vsig = 1; vsig_fac = 1.; i_rr = 1; + break; + case 2: // (1,0): rho_u, drho_d + rs = fxc.v2rhosigma_drho_ud.data(); rs2 = fxc.v2rhosigma_drho_du.data(); + su = fxc.v2sigma2_drho_du_u.data(); sd = fxc.v2sigma2_drho_du_d.data(); + i_vsig = 1; vsig_fac = 1.; i_rr = 1; + break; + case 3: // (1,1) + rs = fxc.v2rhosigma_drho_dd.data(); rs2 = rs; + su = fxc.v2sigma2_drho_dd_u.data(); sd = fxc.v2sigma2_drho_dd_d.data(); + i_vsig = 2; vsig_fac = 2.; i_rr = 2; + break; + default: + throw std::runtime_error("Invalid ispin_op"); } - break; - case 2: + // 1. the terms in grad_dot int f_uu +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif for (int ir = 0;ir < nrxx;++ir) { - vxc_tmp[ir] += rho[ir] * fxc.v2rho2.at(ir * 3 + 1) + drho.at(ir) * fxc.v2rhosigma_drho_du.at(ir); + gdot[ir] = -(rs[ir] * rho[ir] + + su[ir] * (dgs_u[ir] * drho[ir]) + + sd[ir] * (dgs_d[ir] * drho[ir]) + + drho[ir] * (vsig[ir * 3 + i_vsig] * vsig_fac)); } - break; - case 3: + XC_Functional::grad_dot(gdot, vxc, &this->rho_basis_, this->tpiba_); + + // 2. terms not in grad_dot +#ifdef _OPENMP +#pragma omp parallel for schedule(static) +#endif for (int ir = 0;ir < nrxx;++ir) { - vxc_tmp[ir] += rho[ir] * fxc.v2rho2.at(ir * 3 + 2) + drho.at(ir) * fxc.v2rhosigma_drho_dd.at(ir); + vxc[ir] += rho[ir] * rr[3 * ir + i_rr] + drho[ir] * rs2[ir]; } - break; - default: - throw std::runtime_error("Invalid ispin_op"); - } - BlasConnector::axpy(nrxx, ModuleBase::e2, vxc_tmp.data(), 1, v_eff.c, 1); - }; - break; - default: - throw std::domain_error("SpinType =" + std::to_string(static_cast(s)) + "for GGA or HYB_GGA is unfinished in " - + std::string(__FILE__) + " line " + std::to_string(__LINE__)); - break; + BlasConnector::axpy(nrxx, ModuleBase::e2, vxc, 1, v_eff.c, 1); + }; + break; + default: + throw std::domain_error("SpinType =" + std::to_string(static_cast(s)) + "for GGA or HYB_GGA is unfinished in " + + std::string(__FILE__) + " line " + std::to_string(__LINE__)); + break; + } } - } else + else { throw std::domain_error("GlobalV::XC_Functional::get_func_type() =" + std::to_string(XC_Functional::get_func_type()) + " unfinished in " + std::string(__FILE__) + " line " + std::to_string(__LINE__)); diff --git a/source/source_lcao/module_lr/potentials/pot_hxc_lrtd.h b/source/source_lcao/module_lr/potentials/pot_hxc_lrtd.h index 8a712253727..1a95c070e8b 100644 --- a/source/source_lcao/module_lr/potentials/pot_hxc_lrtd.h +++ b/source/source_lcao/module_lr/potentials/pot_hxc_lrtd.h @@ -3,40 +3,70 @@ #include "source_estate/module_pot/h_hartree_pw.h" #include "xc_kernel.h" +#include "pot_lr_base.h" #include #include namespace LR { - class PotHxcLR + class PotHxcLR : public PotLRBase { public: - /// S1: K^Hartree + K^xc + /// S1: the nspin=1 singlet LR kernel, 2*K^Hartree + 2*K^xc. + /// The factor 2 on *both* terms is what makes it equal to S2_singlet. + /// S1_gs: the nspin=1 counterpart of S2_gs, i.e. the *ground-state* Hxc kernel + /// K^Hartree + K^xc = S1 / 2. This is what `pot_hxc_gs` needs at nspin=1. /// S2_singlet: 2*K^Hartree + K^xc_{upup} + K^xc_{updown} /// S2_triplet: K^xc_{upup} - K^xc_{updown} /// S2_updown: K^Hartree + (K^xc_{upup}, K^xc_{updown}, K^xc_{downup} or K^xc_{downdown}), according to `ispin_op` (for spin-polarized systems) - enum SpinType { S1 = 0, S2_singlet = 1, S2_triplet = 2, S2_updown = 3 }; + /// S2_gs: the nspin=2 *ground-state* Hxc kernel K^Hartree + (K^xc_{upup} + K^xc_{updown})/2 = S2_singlet / 2. + /// Used for `pot_hxc_gs` in LR gradients, where the convention is `K_Hxc(singlet) = 2 * pot_hxc_gs`. + /// The 1/2 on the xc part is not a convention but the chain rule: the derivative is + /// taken w.r.t. the *total* density matrix, and $\partial v_u/\partial\rho = + /// (f_{uu}+f_{ud})/2$ because $\rho_u=\rho_d=\rho/2$. The Hartree part needs no halving. + /// Do NOT use S1 here when nspin=2: `KernelXC` is built with the input `nspin`, so the + /// kernel arrays carry 3 spin components per grid point while the S1 integrand indexes + /// them as if there were 1. + enum SpinType { S1 = 0, S2_singlet = 1, S2_triplet = 2, S2_updown = 3, S2_gs = 4, S1_gs = 5 }; /// XCType here is to determin the method of integration from kernel to potential, not the way calculating the kernel enum XCType { None = 0, LDA = 1, GGA = 2, HYB_GGA = 4 }; - /// constructor for exchange-correlation kernel + /// constructor building exchange-correlation kernel PotHxcLR(const std::string& xc_kernel, const ModulePW::PW_Basis& rho_basis, const UnitCell& ucell, const Charge& chg_gs/*ground state*/, const Parallel_Grid& pgrid, const SpinType& st = SpinType::S1, const std::vector& lr_init_xc_kernel = { "default" }); + // Explicit extension: existing callers keep the original constructor without new defaults. + PotHxcLR(const std::string& xc_kernel, const ModulePW::PW_Basis& rho_basis, + const UnitCell& ucell, const Charge& chg_gs, const Parallel_Grid& pgrid, + const SpinType& st, const std::vector& lr_init_xc_kernel, + const int gxc_spin); + /// Constructor taking an already-built kernel. Several `PotHxcLR` can share the same* $f^{xc}$ arrays. + /// The caller is responsible for the ordering: `KernelXC` calls `XC_Functional::set_xc_type`, + /// and this constructor reads the resulting global `get_func_type()`, so build the kernel + /// immediately before the potentials that use it. + PotHxcLR(std::shared_ptr kernel, const std::string& xc_kernel, + const ModulePW::PW_Basis& rho_basis, const UnitCell& ucell, const int nrxx, + const SpinType& st = SpinType::S1); ~PotHxcLR() {} - void cal_v_eff(double** rho, const UnitCell& ucell, ModuleBase::matrix& v_eff, const std::vector& ispin_op = { 0,0 }); - const int& nrxx = nrxx_; + virtual void cal_v_eff(double** rho, const UnitCell& ucell, ModuleBase::matrix& v_eff, const std::vector& ispin_op = { 0,0 }) const override; + + /// Build a kernel that can be shared by several `PotHxcLR` (see the constructor above). + /// `openshell` must match what every sharing potential would have passed, i.e. + /// `st == SpinType::S2_updown`; `gxc_spin` must cover every combination they will ask for. + static std::shared_ptr make_kernel(const std::string& xc_kernel, + const ModulePW::PW_Basis& rho_basis, const UnitCell& ucell, const Charge& chg_gs, + const Parallel_Grid& pgrid, const bool openshell, const int gxc_spin, + const std::vector& lr_init_xc_kernel = { "default" }); + + const KernelXC& xc_kernel_components() const { return *xc_kernel_components_; } private: - const ModulePW::PW_Basis& rho_basis_; - const int nspin_ = 1; - const int nrxx_ = 1; std::unique_ptr pot_hartree_; /// different components of local and semi-local xc kernels: /// LDA: v2rho2 /// GGA: v2rho2, v2rhosigma, v2sigma2 /// meta-GGA: v2rho2, v2rhosigma, v2sigma2, v2rholap, v2rhotau, v2sigmalap, v2sigmatau, v2laptau, v2lap2, v2tau2 - const KernelXC xc_kernel_components_; + /// To allow different potential objects sharing the same kernel. + const std::shared_ptr xc_kernel_components_; const std::string xc_kernel_; - const double& tpiba_; const SpinType spin_type_ = SpinType::S1; XCType xc_type_ = XCType::None; @@ -50,6 +80,45 @@ namespace LR std::map kernel_to_potential_; void set_integral_func(const SpinType& s, const XCType& xc); + + // ---- scratch buffers ------------------------------------------------------------ + // `cal_v_eff` used to allocate (and value-initialize) every temporary on every call: + // at a 200^3 grid that is ~450 MB of memset per call, all of it overwritten before it + // is read, plus the page faults of freshly mapped pages. They now come from a pool + // that grows on demand and is SHARED by every `PotHxcLR` (a calculation holds three of + // them -- singlet, triplet and the ground-state kernel -- and giving each its own copy + // cost ~0.9 GB on the bigger grids). Sharing is safe because `cal_v_eff` is only ever + // entered from a single thread: all the OpenMP lives inside the loops it calls. + struct Scratch + { + std::vector> drho; ///< $\nabla\rho^X$ + std::vector> gdot; ///< integrand of the divergence + std::vector vxc; ///< accumulated $v^{xc}$ + std::vector> rhog; ///< $\rho^X(G)$, npw + std::vector> vg; ///< G-space scratch, nmaxgr + void alloc(const int nrxx, const int npw, const int nmaxgr, const bool gga); + }; + static Scratch& scratch(); + /// Hartree potential of the transition density, built from `scratch().rhog` and + /// accumulated into `v_eff` scaled by `factor`. Replaces `H_Hartree_pw::v_hartree`, + /// which for the LR use case (a) re-did the forward FFT of $\rho^X$ that the GGA branch + /// needs anyway, (b) accumulated a Hartree "energy" of the transition density and pushed + /// it through `Parallel_Reduce::reduce_pool` on every call -- a per-call collective whose + /// result is never read, and which clobbers the global `H_Hartree_pw::hartree_energy` -- + /// and (c) returned a full `matrix` by value, to which `v_eff += 2 * (...)` then added a + /// second temporary of the same size. + void add_v_hartree(const UnitCell& ucell, ModuleBase::matrix& v_eff, const double factor) const; + // ---- pre-contracted spin combinations ------------------------------------------ + // At nspin=2 the singlet/triplet integrands read `v2rho2[3ir] +- v2rho2[3ir+1]` and + // `2*vsigma[3ir] +- vsigma[3ir+1]` at every grid point of every call. Both combinations + // depend only on the ground state, so they are formed once here. This also turns two + // strided reads into one contiguous one, which is where most of the gain is. + // Only the combination this potential's own `spin_type_` needs is built (one scalar + // array each, so 8 B/point, and nothing at all for nspin=1 or the open-shell branch), + // which is why they live here rather than in the shared `KernelXC`. + std::vector v2rho2_comb_; + std::vector vsigma_comb_; ///< GGA only + void build_spin_combos(const bool gga); }; } // namespace LR diff --git a/source/source_lcao/module_lr/potentials/pot_lr_base.h b/source/source_lcao/module_lr/potentials/pot_lr_base.h new file mode 100644 index 00000000000..85f45b462f2 --- /dev/null +++ b/source/source_lcao/module_lr/potentials/pot_lr_base.h @@ -0,0 +1,24 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_POTENTIALS_POT_LR_BASE_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_POTENTIALS_POT_LR_BASE_H +#include "source_cell/unitcell.h" +#include "source_basis/module_pw/pw_basis.h" + +namespace LR +{ + class PotLRBase + { + public: + PotLRBase(const ModulePW::PW_Basis& rho_basis, const int& nspin, const int& nrxx, const double& tpiba) : rho_basis_(rho_basis), nspin_(nspin), nrxx_(nrxx), tpiba_(tpiba) {} + virtual void cal_v_eff(double** rho, const UnitCell& ucell, ModuleBase::matrix& v_eff, const std::vector& ispin_op = { 0,0 }) const = 0; + // const references + const ModulePW::PW_Basis& get_rho_basis() const { return rho_basis_; } + const int& nrxx = nrxx_; + const int& nspin = nspin_; + protected: + const ModulePW::PW_Basis& rho_basis_; + const int nspin_ = 1; + const int nrxx_ = 1; + const double& tpiba_; + }; +} +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_POTENTIALS_POT_LR_BASE_H diff --git a/source/source_lcao/module_lr/potentials/xc_kernel.cpp b/source/source_lcao/module_lr/potentials/xc_kernel.cpp index de7146db670..4cd76221cc7 100644 --- a/source/source_lcao/module_lr/potentials/xc_kernel.cpp +++ b/source/source_lcao/module_lr/potentials/xc_kernel.cpp @@ -2,10 +2,15 @@ #include "source_hamilt/module_xc/xc_functional.h" #include "source_io/module_parameter/parameter.h" #include "source_base/timer.h" +#include "source_base/tool_quit.h" #include "source_lcao/module_lr/utils/lr_util.h" #include "source_lcao/module_lr/utils/lr_util_xc.hpp" #include #include +#include +#include +#include +#include #include "source_io/module_output/cube_io.h" #ifdef __LIBXC #include @@ -22,9 +27,10 @@ LR::KernelXC::KernelXC(const ModulePW::PW_Basis& rho_basis, const int& nspin, const std::string& kernel_name, const std::vector& lr_init_xc_kernel, - const bool openshell) :rho_basis_(rho_basis), openshell_(openshell) + const bool openshell, + const int gxc_spin) :rho_basis_(rho_basis), openshell_(openshell), gxc_spin_(gxc_spin) { - if (!std::set({ "lda", "pwlda", "pbe", "hse" }).count(kernel_name)) { return; } + if (!LR_Util::has_local_xc(kernel_name)) { return; } XC_Functional::set_xc_type(kernel_name); // for hse, (1-alpha) and omega are set here const int& nrxx = rho_basis.nrxx; @@ -96,7 +102,13 @@ inline void add_assign_op(const std::vector& src, std::vector& dst) template inline void cutoff_grid_data_spin2(std::vector& func, const std::vector& mask) { - const int& nrxx = mask.size() / 2; + const int nrxx = mask.size() / 2; + if (nrxx == 0) + { + ModuleBase::WARNING_QUIT("LR::cutoff_grid_data_spin2", + "An MPI rank has no local real-space grid points for the LR XC kernel. " + "Reduce the number of MPI processes (and increase OpenMP threads if needed)."); + } assert(func.size() % nrxx == 0 && func.size() / nrxx > 1); const int n_component = func.size() / nrxx; #ifdef _OPENMP @@ -120,95 +132,78 @@ void LR::KernelXC::f_xc_libxc(const int& nspin, const double& omega, const doubl assert(nspin == 1 || nspin == 2); - double hybrid_alpha = 0.0; - double hse_omega = 0.0; + // `hybrid_alpha` is neither $\alpha$ nor $\alpha+\beta$ of the range separation + // $v_1(r)=[\alpha+\beta\,\mathrm{erfc}(\omega r)]/r$: `input_conv` sets it to + // $\max(|\alpha|,|\beta|)$, the factor it divided `coulomb_param` by. + // A functional with both non-zero (CAM, LC, LRC) would get a meaningless + // value from it -- and does not use it: `in_built_xc_func_ext_params` reads + // $\alpha$ and $\beta$ for those straight from `exx_fock_alpha` / `exx_erfc_alpha`, and + // takes only $\omega$ from `hse_omega`. + const double hybrid_alpha = XC_Functional::get_hybrid_alpha(); + const double hse_omega = XC_Functional::get_hse_omega(); std::vector funcs = XC_Functional_Libxc::init_func( XC_Functional::get_func_id(), (1 == nspin) ? XC_UNPOLARIZED : XC_POLARIZED, hybrid_alpha, hse_omega); const int& nrxx = rho_basis_.nrxx; + const bool is_gga = std::any_of(funcs.begin(), funcs.end(), [](const xc_func_type& f) { return f.info->family == XC_FAMILY_GGA || f.info->family == XC_FAMILY_HYB_GGA; }); + // The third-order kernel exists for exactly one purpose: the $g^{xc}$ part of the LR + // gradient. If none was requested, skip it. Open shell needs the RAW arrays (see below) -- + // `build_gxc_coef` is a closed-shell-only pre-contraction. + const bool need_kxc = (this->gxc_spin_ != GxcSpin::NoGxc); - // converting rho (extract it as a subfuntion in the future) - // ----------------------------------------------------------------------------------- std::vector rho(nspin * nrxx); // r major / spin contigous - -#ifdef _OPENMP -#pragma omp parallel for collapse(2) schedule(static, 1024) -#endif - for (int is = 0; is < nspin; ++is) { for (int ir = 0; ir < nrxx; ++ir) { rho[ir * nspin + is] = rho_gs[is][ir]; } } - if (rho_core) - { - const double fac = 1.0 / nspin; - for (int is = 0; is < nspin; ++is) { for (int ir = 0; ir < nrxx; ++ir) { rho[ir * nspin + is] += fac * rho_core[ir]; } } - } - - // ----------------------------------------------------------------------------------- // for GGA - const bool is_gga = std::any_of(funcs.begin(), funcs.end(), [](const xc_func_type& f) { return f.info->family == XC_FAMILY_GGA || f.info->family == XC_FAMILY_HYB_GGA; }); - std::vector>> gradrho; // \nabla \rho std::vector sigma; // |\nabla\rho|^2 std::vector sgn; // sgn for threshold mask - if (is_gga) - { - // 0. set up sgn for threshold mask - // in the case of GGA correlation for polarized case, - // a cutoff for grho is required to ensure that libxc gives reasonable results + this->get_rho_drho_sigma(nspin, tpiba, rho_gs, rho_core, is_gga, rho, gradrho, sigma); - // 1. \nabla \rho - gradrho.resize(nspin); - for (int is = 0; is < nspin; ++is) - { - std::vector rhor(nrxx); -#ifdef _OPENMP -#pragma omp parallel for schedule(static, 1024) -#endif - for (int ir = 0; ir < nrxx; ++ir) { rhor[ir] = rho[ir * nspin + is]; -} - gradrho[is].resize(nrxx); - LR_Util::grad(rhor.data(), gradrho[is].data(), rho_basis_, tpiba); - } - // 2. |\nabla\rho|^2 - sigma.resize(nrxx * ((1 == nspin) ? 1 : 3)); - if (1 == nspin) - { -#ifdef _OPENMP -#pragma omp parallel for schedule(static, 1024) -#endif - for (int ir = 0; ir < nrxx; ++ir) { - sigma[ir] = gradrho[0][ir] * gradrho[0][ir]; -} - } - else - { -#ifdef _OPENMP -#pragma omp parallel for schedule(static, 256) -#endif - for (int ir = 0; ir < nrxx; ++ir) - { - sigma[ir * 3] = gradrho[0][ir] * gradrho[0][ir]; - sigma[ir * 3 + 1] = gradrho[0][ir] * gradrho[1][ir]; - sigma[ir * 3 + 2] = gradrho[1][ir] * gradrho[1][ir]; - } - } - } - // ----------------------------------------------------------------------------------- //==================== XC Kernels (f_xc)============================= this->vrho_.resize(nspin * nrxx, 0.); this->v2rho2_.resize(((1 == nspin) ? 1 : 3) * nrxx, 0.);//(nrxx* ((1 == nspin) ? 1 : 3)): 00, 01, 11 + if (need_kxc) + { + this->v3rho3_.resize(((1 == nspin) ? 1 : 4) * nrxx, 0.);//(nrxx* ((1 == nspin) ? 1 : 4)): 000, 001, 011, 111 + } if (is_gga) { this->vsigma_.resize(((1 == nspin) ? 1 : 3) * nrxx, 0.);//(nrxx*): 2 for rho * 3 for sigma: 00, 01, 02, 10, 11, 12 this->v2rhosigma_.resize(((1 == nspin) ? 1 : 6) * nrxx, 0.); //(nrxx*): 2 for rho * 3 for sigma: 00, 01, 02, 10, 11, 12 this->v2sigma2_.resize(((1 == nspin) ? 1 : 6) * nrxx, 0.); //(nrxx* ((1 == nspin) ? 1 : 6)): 00, 01, 02, 11, 12, 22 + if (need_kxc) + { + this->v3rho2sigma_.resize(((1 == nspin) ? 1 : 9) * nrxx, 0.); //000, 001, 002, 010, 011, 012, 110, 111, 112 + this->v3rhosigma2_.resize(((1 == nspin) ? 1 : 12) * nrxx, 0.); //000, 001, 002, 011, 012, 022, 100, 101, 102, 111, 112, 122 + this->v3sigma3_.resize(((1 == nspin) ? 1 : 10) * nrxx, 0.);//000, 001, 002, 011, 012, 022, 111, 112, 122, 222 + } } //MetaGGA ... for (xc_func_type& func : funcs) { - const double rho_threshold = 1E-6; - const double grho_threshold = 1E-10; + // These used to be 1E-6 / 1E-10, the values the ground-state SCF uses. That is far too + // aggressive for LR *gradients*: `cal_sgn` zeroes the kernel wherever rho < rho_threshold, + // and in a molecule-in-a-big-box most of the grid is below 1E-6, while the exact derivative + // the finite difference measures has no such truncation. + // + // Measured on H2/DZP TDRPA@LDA, where K^T = 0 makes the triplet gradient identical to + // d(eps_a - eps_i)/dx and the reference is therefore exact to ~13 digits: + // thresholds 1E-6 /1E-10 : max |err| = 0.0159 eV/Ang over 9 states + // thresholds 1E-14/1E-20 : max |err| = 0.0002 eV/Ang (at the force printout precision) + // and on H2/DZP TDLDA over 18 states: 0.038 -> 0.00095 eV/Ang (the latter is the finite + // difference's own resolution). H2/SZ was insensitive either way -- its compact 1s-only + // basis puts no T+D^Z density in the truncated region, which is why the problem only shows + // up once diffuse/p functions enter. + // + // These thresholds also change the shared second-order XC kernel and excitation energies, + // even with cal_force disabled, so excitation-energy references must use the same thresholds. + // CAVEAT: only LDA has been checked at these thresholds. The 1E-6/1E-10 pair exists because + // GGA *correlation* can misbehave at very low density; if PBE turns out to need protection, + // the fix is a separate threshold for the gradient path, not a return to 1E-6 everywhere. + const double rho_threshold = 1E-14; + const double grho_threshold = 1E-20; xc_func_set_dens_threshold(&func, rho_threshold); @@ -221,25 +216,98 @@ void LR::KernelXC::f_xc_libxc(const int& nspin, const double& omega, const doubl std::vector vsigma_tmp(this->vsigma_.size()); std::vector v2rhosigma_tmp(this->v2rhosigma_.size()); std::vector v2sigma2_tmp(this->v2sigma2_.size()); + // The kxc arrays used to be handed to libxc directly. Since the libxc interfaces *overwrite* + // their output (that is why every other component already goes through a temporary), only the + // last functional of a composite survived -- e.g. for PBE = XC_GGA_X_PBE + XC_GGA_C_PBE the + // exchange part of the third derivative was silently dropped. They are accumulated now too. + // Note these are left uncut by `sgn`: for nspin=1 the cutoff is a no-op anyway (a single + // spin component), and for nspin=2 `cutoff_grid_data_spin2` does not match the component + // layout of the third derivatives. + std::vector v3rho3_tmp(this->v3rho3_.size()); + std::vector v3rho2sigma_tmp(this->v3rho2sigma_.size()); + std::vector v3rhosigma2_tmp(this->v3rhosigma2_.size()); + std::vector v3sigma3_tmp(this->v3sigma3_.size()); switch (func.info->family) { case XC_FAMILY_LDA: xc_lda_vxc(&func, nrxx, rho.data(), vrho_tmp.data()); xc_lda_fxc(&func, nrxx, rho.data(), v2rho2_tmp.data()); + if (need_kxc) + { + xc_lda_kxc(&func, nrxx, rho.data(), v3rho3_tmp.data()); + } break; case XC_FAMILY_GGA: case XC_FAMILY_HYB_GGA: { - xc_gga_vxc(&func, nrxx, rho.data(), sigma.data(), vrho_tmp.data(), vsigma_tmp.data()); - xc_gga_fxc(&func, nrxx, rho.data(), sigma.data(), v2rho2_tmp.data(), v2rhosigma_tmp.data(), v2sigma2_tmp.data()); + // HSE06's short-range exchange (libxc's gga_x_wpbeh, folded into the single combined + // XC_HYB_GGA_XC_HSE06 functional evaluated here) has a closed-form whose derivatives + // are numerically unstable (catastrophic cancellation) near the reduced gradient s->0 + // (a symmetry-forced grad-rho=0 node at otherwise ordinary density, e.g. rho=0.415 in + // 06_N2/hse). The true derivative is FINITE there (the functional is smooth in s^2); + // libxc just loses precision evaluating the closed form at that point. Clamping sigma + // from below before the libxc call evaluates the same closed form at a point just off + // the instability, borrowing smoothness instead of removing the cancellation + // algebraically. This existing HSE06 stabilization is retained here; it is not + // a derivative-consistent regularization and needs a separate numerical fix. + // Other functionals receive the original sigma, including the signed mixed-spin + // invariant grad(rho_up).grad(rho_down), which can be smaller than -1. + // + // `cal_sgn` returns all-ones for exchange functionals, so `cutoff_grid_data_spin2` + // below is a no-op for wpbeh; vsigma (1st order, read by the LR GGA kernel formula in + // pot_hxc_lrtd.cpp as the coefficient of the transition-density gradient) and v2*/v3* + // are therefore all equally unprotected and must ALL be clamped with the same + // sigma_cut -- clamping only a subset is not just insufficient, it can be WORSE + // (breaks whatever partial cancellation existed between the orders at the raw, + // unclamped point). Confirmed on 09_CH4/hse: with 1st-order vsigma left unclamped (as + // it originally was here), max|vsigma| reaches 3.9e7 at a vacuum grid point + // (rho~0, sigma~3.3e-19), which propagates into a ~1e4-magnitude AO-basis V_Hxc + // element on the carbon 2nd-zeta s/p shell and three ~-19404 Ry ghost eigenvalues in + // the full (lr_solver=lapack) Casida spectrum. + // + // Threshold choice (1e-6): g(sigma) is smooth at sigma=0, so g(sigma_cut) = g(0) + + // O(sigma_cut) -- any sigma_cut small enough stays a good approximation to the true + // sigma->0 limit. The lower bound on "small enough" was set empirically, not derived: + // scanning 1e-10/1e-4/1e-2 with only ONE of fxc/kxc clamped left the result completely + // flat (06_N2/hse state 5 pinned at +1.887e7 across 6 orders of magnitude for kxc-only) + // -- the textbook sign that the knob isn't touching the actual cancellation. Clamping + // BOTH together at 1e-6/1e-4/1e-2 instead gives a smoothly-varying, converging sequence + // (-0.302/-0.299/-0.257), the behavior a well-posed limit should show. 1e-6 sits deep in + // that well-posed plateau (tightening it to 1e-8 or 1e-10 moves the clamped answer by + // far less than scanning between 1e-6/1e-4/1e-2 already does) while staying far above + // the sigma range where the raw, unclamped evaluation is already visibly corrupted + // (1e7-1e39 noise at the diagnosed grid points). Checked against archived + // finite-difference references for 06_N2/hse, 09_CH4/hse and 08_BeH2/rpa_at_hse: + // MAE/|F| dropped to 0.015%-0.556%, Max/|F| to 0.029%-1.279%, matching the error level + // of functionals that were never broken; also confirmed to resolve 08_BeH2/hse (both + // its DZP and TZDP bases, all excited states). + const bool is_hse06 = func.info->number == XC_HYB_GGA_XC_HSE06; + std::vector hse_sigma; + const std::vector& sigma_input = LR_Util::prepare_xc_sigma(sigma, is_hse06, hse_sigma); + xc_gga_vxc(&func, nrxx, rho.data(), sigma_input.data(), vrho_tmp.data(), vsigma_tmp.data()); + xc_gga_fxc(&func, nrxx, rho.data(), sigma_input.data(), v2rho2_tmp.data(), v2rhosigma_tmp.data(), v2sigma2_tmp.data()); // std::cout << "max element of v2sigma2_tmp: " << *std::max_element(v2sigma2_tmp.begin(), v2sigma2_tmp.end()) << std::endl; // std::cout << "rho corresponding to max element of v2sigma2_tmp: " << rho[(std::max_element(v2sigma2_tmp.begin(), v2sigma2_tmp.end()) - v2sigma2_tmp.begin()) / 6] << std::endl; - // cut off by sgn - cutoff_grid_data_spin2(vrho_tmp, sgn); - cutoff_grid_data_spin2(vsigma_tmp, sgn); - cutoff_grid_data_spin2(v2rho2_tmp, sgn); - cutoff_grid_data_spin2(v2rhosigma_tmp, sgn); - cutoff_grid_data_spin2(v2sigma2_tmp, sgn); + // cut off by sgn. nspin=2 only: `cutoff_grid_data_spin2` assumes >1 component per + // grid point (it asserts on it), and at nspin=1 there is exactly one, for which both + // of its `for_each` ranges are empty -- the cutoff is a no-op anyway. + if (nspin == 2) + { + cutoff_grid_data_spin2(vrho_tmp, sgn); + cutoff_grid_data_spin2(vsigma_tmp, sgn); + cutoff_grid_data_spin2(v2rho2_tmp, sgn); + cutoff_grid_data_spin2(v2rhosigma_tmp, sgn); + cutoff_grid_data_spin2(v2sigma2_tmp, sgn); + } + if (need_kxc) + { + // Use the same functional-specific input as vxc and fxc. + xc_gga_kxc(&func, nrxx, rho.data(), sigma_input.data(), + v3rho3_tmp.data(), + v3rho2sigma_tmp.data(), + v3rhosigma2_tmp.data(), + v3sigma3_tmp.data()); + } break; } default: @@ -254,6 +322,13 @@ void LR::KernelXC::f_xc_libxc(const int& nspin, const double& omega, const doubl add_assign_op(vsigma_tmp, this->vsigma_); add_assign_op(v2rhosigma_tmp, this->v2rhosigma_); add_assign_op(v2sigma2_tmp, this->v2sigma2_); + if (need_kxc) + { + add_assign_op(v3rho3_tmp, this->v3rho3_); + add_assign_op(v3rho2sigma_tmp, this->v3rho2sigma_); + add_assign_op(v3rhosigma2_tmp, this->v3rhosigma2_); + add_assign_op(v3sigma3_tmp, this->v3sigma3_); + } // auto end = std::chrono::high_resolution_clock::now(); // auto duration = std::chrono::duration_cast(end - start); // std::cout << "Time elapsed adding XC components: " << duration.count() << " ms\n"; @@ -291,6 +366,9 @@ void LR::KernelXC::f_xc_libxc(const int& nspin, const double& omega, const doubl { this->v2sigma2_4drho_[i] = gradrho[0][i] * v2s2[i] * 4.; } + + // The third-order kernels contracted with $\nabla\rho$ used to be pre-built here; + // they are now folded into `GxcCoef::e_*` (with $\nabla\rho$ factored back out) } else if (2 == nspin) //close-shell { @@ -352,16 +430,259 @@ void LR::KernelXC::f_xc_libxc(const int& nspin, const double& omega, const doubl throw std::domain_error("nspin =" + std::to_string(nspin) + " unfinished in " + std::string(__FILE__) + " line " + std::to_string(__LINE__)); } - this->drho_gs_ = std::move(gradrho); } - if (PARAM.inp.nspin == 1 || PARAM.inp.nspin == 2) { - ModuleBase::timer::end("XC_Functional", "f_xc_libxc"); + + // Build the $v^{(2)}$ coefficient sets. Must happen before `gradrho` is moved away below. + // Only the combinations the caller asked for: at nspin=2 each set costs 10 doubles per grid + // point for a GGA, so building an unused one is a third of the whole third-order kernel. + this->nspin_ = nspin; + // Open shell keeps the raw libxc third-order arrays and contracts them on the fly in + // `PotGradXCLR::cal_v_eff_openshell`: there is no single spin combination to pre-contract + // into (the free spin index $\tau$ stays open), and pre-contracting would cost ~117 doubles + // per grid point against the 35 the raw arrays already occupy. + if (!openshell_ && this->gxc_spin_ != GxcSpin::NoGxc) + { + // No `gradrho` here: the divergence coefficients are stored with $\nabla\rho$ factored + // out (see `GxcCoef`), so the spin algebra below is purely local. + // nspin=1 has no triplet: `gxc()` returns `gxc_s_` whatever is asked, so build it for any request. + if (nspin == 1 || (this->gxc_spin_ & GxcSpin::Singlet)) // Singlet or Bothspin + { + this->build_gxc_coef(this->gxc_s_, /*triplet=*/false, nspin, is_gga); + } + if (nspin == 2 && (this->gxc_spin_ & GxcSpin::Triplet)) // Triplet or Bothspin + { + this->build_gxc_coef(this->gxc_t_, /*triplet=*/true, nspin, is_gga); + } + } + if (is_gga) { this->drho_gs_ = std::move(gradrho); } + ModuleBase::timer::end("XC_Functional", "f_xc_libxc"); +} + +void LR::KernelXC::get_rho_drho_sigma(const int& nspin, + const double& tpiba, + const double* const* const rho_gs, + const double* const rho_core, + const bool& is_gga, + std::vector& rho, + std::vector>>& gradrho, + std::vector& sigma) +{ + const int nrxx = rho_basis_.nrxx; +#ifdef _OPENMP +#pragma omp parallel for collapse(2) schedule(static, 1024) +#endif + for (int is = 0; is < nspin; ++is) { for (int ir = 0; ir < nrxx; ++ir) { rho[ir * nspin + is] = rho_gs[is][ir]; } } + if (rho_core) + { + const double fac = 1.0 / nspin; + for (int is = 0; is < nspin; ++is) { for (int ir = 0; ir < nrxx; ++ir) { rho[ir * nspin + is] += fac * rho_core[ir]; } } + } + if (is_gga) + { + // 0. set up sgn for threshold mask + // in the case of GGA correlation for polarized case, + // a cutoff for grho is required to ensure that libxc gives reasonable results + + // 1. \nabla \rho + gradrho.resize(nspin); + for (int is = 0; is < nspin; ++is) + { + std::vector rhor(nrxx); +#ifdef _OPENMP +#pragma omp parallel for schedule(static, 1024) +#endif + for (int ir = 0; ir < nrxx; ++ir) { + rhor[ir] = rho[ir * nspin + is]; + } + gradrho[is].resize(nrxx); + LR_Util::grad(rhor.data(), gradrho[is].data(), rho_basis_, tpiba); + } + // 2. |\nabla\rho|^2 + sigma.resize(nrxx * ((1 == nspin) ? 1 : 3)); + if (1 == nspin) + { +#ifdef _OPENMP +#pragma omp parallel for schedule(static, 1024) +#endif + for (int ir = 0; ir < nrxx; ++ir) { + sigma[ir] = gradrho[0][ir] * gradrho[0][ir]; + } + } + else + { +#ifdef _OPENMP +#pragma omp parallel for schedule(static, 256) +#endif + for (int ir = 0; ir < nrxx; ++ir) + { + sigma[ir * 3] = gradrho[0][ir] * gradrho[0][ir]; + sigma[ir * 3 + 1] = gradrho[0][ir] * gradrho[1][ir]; + sigma[ir * 3 + 2] = gradrho[1][ir] * gradrho[1][ir]; + } + } + } +} + + +// --------------------------------------------------------------------------------------------- +// The spin algebra of $v^{(2)}$, for both the singlet and the triplet combination. +// +// Both combinations come from the SAME lambda-expansion; they differ only in five weight vectors, +// which is why they share this one routine: +// +// perturbation singlet: rho_u += L*s, rho_d += L*s triplet: rho_u += L*s, rho_d -= L*s +// eta_sigma (+1, +1) (+1, -1) d(rho_sigma)/dL / s +// mu_{ab} (+1, +1, +1) (+1, 0, -1) d(sigma_ab)/dL / (2t) +// nu_{ab} (+1, +1, +1) (+1, -1, +1) d2(sigma_ab)/dL2 / (2q) +// theta_{ab} ( 2, 1, 0) ( 2, 1, 0) from 2*e_{s_uu}*grad rho_u +// theta~_{ab} ( 2, 1, 0) ( 2, -1, 0) + e_{s_ud}*grad rho_d +// +// theta~ differs from theta only for the triplet, and only because grad rho_d there carries +// d(grad rho_d)/dL = -grad rho^1: W'' picks up 4u'_uu - 2u'_ud where W' picks up 2u'_uu + u'_ud. +// In the singlet the two coincide, which is why the singlet formula looked tidier than it is. +// Getting theta~ wrong is invisible in every singlet test. +// libxc component indices for nspin=2 now live in `xc_kernel.h` (namespace LR::libxc_idx). +// The open-shell g^xc code in `pot_grad_xc.cpp` needs the same tables. +using LR::libxc_idx::p2; +using LR::libxc_idx::p3; + +void LR::KernelXC::build_gxc_coef(GxcCoef& dst, const bool triplet, const int& nspin, const bool& is_gga) +{ + const int& nrxx = rho_basis_.nrxx; + const double eta[2] = { 1., triplet ? -1. : 1. }; + const double mu[3] = { 1., triplet ? 0. : 1., triplet ? -1. : 1. }; + const double nu[3] = { 1., triplet ? -1. : 1., 1. }; + const double th[3] = { 2., 1., 0. }; + const double tht[3] = { 2., triplet ? -1. : 1., 0. }; + + dst.a_s2.resize(nrxx, 0.); + if (is_gga) + { + dst.a_st.resize(nrxx, 0.); dst.a_t2.resize(nrxx, 0.); dst.a_q.resize(nrxx, 0.); + dst.c_s.resize(nrxx, 0.); dst.c_t.resize(nrxx, 0.); + dst.e_s2.resize(nrxx); dst.e_st.resize(nrxx); dst.e_t2.resize(nrxx); dst.e_q.resize(nrxx); + } + const std::vector& v2rs = this->v2rhosigma_; + const std::vector& v2s2 = this->v2sigma2_; + const std::vector& v3r3 = this->v3rho3_; + const std::vector& v3r2s = this->v3rho2sigma_; + const std::vector& v3rs2 = this->v3rhosigma2_; + const std::vector& v3s3 = this->v3sigma3_; + + if (nspin == 1) + { + // Single component everywhere; all the weight sums collapse to 1. + // + // ... and then the whole set is scaled by 4 to reach the *singlet* normalization, the same + // one `nspin=2` produces below and the only one the rest of the gradient code knows about. + // libxc's unpolarized derivatives are taken w.r.t. the TOTAL density, so for a closed shell + // (rho_u = rho_d = rho/2) an n-th derivative is 2^(n-1) smaller than the singlet spin + // combination: d^2E/drho^2 = (f_uu+f_ud)/2 -- which is why `SpinType::S1` carries a 2 -- + // and d^3E/drho^3 = (g_uuu + 3g_uud)/4, while the nspin=2 branch below builds + // a_s2 = g_uuu + 2g_uud + g_udd = g_uuu + 3g_uud. Hence 4 here, 2 there. + // + // Measured on H2/SZ/LDA: without it the GXC DMTRANS force was 0.08549 eV/Ang against the + // nspin=2 singlet's 0.17097 -- a factor 2, being 1/4 from this and 2 from the `dm_gs` + // channel convention (`gs_dm_channel_factor` in `lr_force.cpp`). + constexpr double to_singlet = 4.; +#ifdef _OPENMP +#pragma omp parallel for schedule(static, 4096) +#endif + for (int i = 0;i < nrxx;++i) + { + dst.a_s2[i] = to_singlet * v3r3[i]; + if (!is_gga) { continue; } + dst.a_st[i] = to_singlet * v3r2s[i] * 4.; + dst.a_t2[i] = to_singlet * v3rs2[i] * 4.; + dst.a_q[i] = to_singlet * v2rs[i] * 2.; + dst.c_s[i] = to_singlet * v2rs[i] * 4.; + dst.c_t[i] = to_singlet * v2s2[i] * 8.; + dst.e_s2[i] = to_singlet * v3r2s[i] * 2.; + dst.e_st[i] = to_singlet * v3rs2[i] * 8.; + dst.e_t2[i] = to_singlet * v3s3[i] * 8.; + dst.e_q[i] = to_singlet * v2s2[i] * 4.; + } return; - // else if (4 == PARAM.inp.nspin) - } else//NSPIN != 1,2,4 is not supported + } + + // nspin=2, close shell. Every sum below runs over ORDERED spin indices, while libxc only + // stores the unique combinations -- that is where the multiplicities come from. +#ifdef _OPENMP +#pragma omp parallel for schedule(static, 4096) +#endif + for (int i = 0;i < nrxx;++i) { - throw std::domain_error("PARAM.inp.nspin =" + std::to_string(PARAM.inp.nspin) - + " unfinished in " + std::string(__FILE__) + " line " + std::to_string(__LINE__)); + const int o4 = i * 4; + const int o6 = i * 6; + const int o9 = i * 9; + const int o10 = i * 10; + const int o12 = i * 12; + + // $a_{s^2}=\sum_{\sigma\sigma'}\eta_\sigma\eta_{\sigma'}g^{\rho_u\rho_\sigma\rho_{\sigma'}}$ + // v3rho3 = (uuu, uud, udd, ddd); with the first index pinned to u the component index is + // just the number of d's among (sigma, sigma'). + dst.a_s2[i] = eta[0] * eta[0] * v3r3[o4] + + 2. * eta[0] * eta[1] * v3r3[o4 + 1] + + eta[1] * eta[1] * v3r3[o4 + 2]; + if (!is_gga) { continue; } + + // ---- the local part $A$ ---- + // $a_{st}=4\sum_\sigma\eta_\sigma\sum_{\alpha\beta}\mu_{\alpha\beta}g^{\rho_u\rho_\sigma\sigma_{\alpha\beta}}$ + // v3rho2sigma = [rho-pair uu,ud,dd] x [sigma uu,ud,dd]; first rho index pinned to u means + // rho-pair block = (sigma==d). + double a_st = 0.; + for (int sg = 0;sg < 2;++sg) { + for (int b = 0;b < 3;++b) { a_st += eta[sg] * mu[b] * v3r2s[o9 + 3 * sg + b]; } } + dst.a_st[i] = a_st * 4.; + + // $a_{t^2}=4\sum_{\alpha\beta,\gamma\delta}\mu\mu\,g^{\rho_u\sigma_{\alpha\beta}\sigma_{\gamma\delta}}$ + double a_t2 = 0.; + for (int b = 0;b < 3;++b) { + for (int c = 0;c < 3;++c) { a_t2 += mu[b] * mu[c] * v3rs2[o12 + p2[b][c]]; } } + dst.a_t2[i] = a_t2 * 4.; + + // $a_q=2F$, $F=\sum_{\alpha\beta}\nu_{\alpha\beta}f^{\rho_u\sigma_{\alpha\beta}}$ + double F = 0.; + for (int b = 0;b < 3;++b) { F += nu[b] * v2rs[o6 + b]; } + dst.a_q[i] = F * 2.; + + // ---- the divergence part $\boldsymbol{E}$ ---- + // $P=\sum_{\alpha\beta}\theta_{\alpha\beta}\sum_{\sigma\sigma'}\eta\eta\,g^{\rho_\sigma\rho_{\sigma'}\sigma_{\alpha\beta}}$ + // rho-pair block index = number of d's in (sigma, sigma'), with multiplicity 2 for ud. + double P = 0.; + double Q = 0.; + double R = 0.; + double S = 0.; + double T = 0.; + double St = 0.; + for (int a = 0;a < 3;++a) + { + const double w = th[a]; + const double wt = tht[a]; + if (w != 0.) + { + P += w * (eta[0] * eta[0] * v3r2s[o9 + 0 + a] + + 2. * eta[0] * eta[1] * v3r2s[o9 + 3 + a] + + eta[1] * eta[1] * v3r2s[o9 + 6 + a]); + for (int sg = 0;sg < 2;++sg) { + for (int c = 0;c < 3;++c) { Q += w * eta[sg] * mu[c] * v3rs2[o12 + 6 * sg + p2[a][c]]; } } + for (int c = 0;c < 3;++c) { + for (int d = 0;d < 3;++d) { R += w * mu[c] * mu[d] * v3s3[o10 + p3[a][c][d]]; } } + for (int c = 0;c < 3;++c) { S += w * nu[c] * v2s2[o6 + p2[a][c]]; } + } + if (wt != 0.) + { + for (int sg = 0;sg < 2;++sg) { T += wt * eta[sg] * v2rs[o6 + 3 * sg + a]; } + for (int c = 0;c < 3;++c) { St += wt * mu[c] * v2s2[o6 + p2[a][c]]; } + } + } + dst.c_s[i] = T * 2.; + dst.c_t[i] = St * 4.; + dst.e_s2[i] = P; + dst.e_st[i] = Q * 4.; + dst.e_t2[i] = R * 4.; + dst.e_q[i] = S * 2.; } } -#endif \ No newline at end of file + +#endif diff --git a/source/source_lcao/module_lr/potentials/xc_kernel.h b/source/source_lcao/module_lr/potentials/xc_kernel.h index 96f4e42bc18..93af73e3671 100644 --- a/source/source_lcao/module_lr/potentials/xc_kernel.h +++ b/source/source_lcao/module_lr/potentials/xc_kernel.h @@ -5,16 +5,51 @@ #include "source_cell/unitcell.h" #include "source_base/parallel_grid.h" #include "source_estate/module_charge/charge.h" +#include #define CREF(x) const std::vector& x = x##_ #define CREF3(x) const std::vector>& x = x##_ namespace LR { + /// libxc component layout for the spin-POLARIZED case, shared by the closed- and open-shell + /// $g^{xc}$ code. Spin: 0=u, 1=d. Sigma: 0=uu, 1=ud, 2=dd. + /// libxc stores only the unique (unordered) index combinations, which is where all the + /// multiplicities in the contractions come from. + namespace libxc_idx + { + /// unordered sigma PAIR -> v2sigma2 (6) and the per-rho block of v3rhosigma2 + constexpr int p2[3][3] = { {0,1,2},{1,3,4},{2,4,5} }; + /// unordered sigma TRIPLE -> v3sigma3 (10): + /// (000)(001)(002)(011)(012)(022)(111)(112)(122)(222) + constexpr int p3[3][3][3] = { + { {0,1,2},{1,3,4},{2,4,5} }, + { {1,3,4},{3,6,7},{4,7,8} }, + { {2,4,5},{4,7,8},{5,8,9} } }; + /// v3rho3 (4): (uuu,uud,udd,ddd) -- indexed by the number of d's + inline constexpr int r3(const int s0, const int s1, const int s2) { return s0 + s1 + s2; } + /// v2rho2 (3): (uu,ud,dd) + inline constexpr int r2(const int s0, const int s1) { return s0 + s1; } + /// v2rhosigma (6): [rho u,d] x [sigma uu,ud,dd] + inline constexpr int rs(const int s, const int a) { return 3 * s + a; } + /// v3rho2sigma (9): [rho-pair uu,ud,dd] x [sigma uu,ud,dd] + inline constexpr int r2s(const int s0, const int s1, const int a) { return 3 * (s0 + s1) + a; } + /// v3rhosigma2 (12): [rho u,d] x [unordered sigma pair] + inline constexpr int rs2(const int s, const int a, const int b) { return 6 * s + p2[a][b]; } + /// $\partial\sigma_a/\partial\nabla\rho_\tau = \theta^\tau_a\,\nabla\rho_{c(\tau,a)}$ + /// tau=u: (uu -> 2 grad rho_u, ud -> 1 grad rho_d, dd -> 0) + /// tau=d: (uu -> 0, ud -> 1 grad rho_u, dd -> 2 grad rho_d) + constexpr double theta[2][3] = { {2., 1., 0.}, {0., 1., 2.} }; + /// which density-gradient channel goes with (tau, a); -1 where theta vanishes + constexpr int chan[2][3] = { {0, 1, -1}, {-1, 0, 1} }; + } + /// @brief Calculate the exchange-correlation (XC) kernel ($f_{xc}=\delta^2E_xc/\delta\rho^2$) and store its components. class KernelXC { using Tvec = std::vector; using Tvec3 = std::vector>; public: + /// Which spin combinations of the $g^{xc}$ coefficient set (`GxcCoef`) to build. + enum GxcSpin { NoGxc = 0, Singlet = 1, Triplet = 2, BothSpins = 3 }; KernelXC(const ModulePW::PW_Basis& rho_basis, const UnitCell& ucell, const Charge& chg_gs, @@ -22,7 +57,8 @@ namespace LR const int& nspin, const std::string& kernel_name, const std::vector& lr_init_xc_kernel, - const bool openshell = false); + const bool openshell = false, + const int gxc_spin = GxcSpin::NoGxc); ~KernelXC() {} // const references @@ -32,14 +68,70 @@ namespace LR CREF3(v2rhosigma_drho_uu); CREF3(v2rhosigma_drho_ud); CREF3(v2rhosigma_drho_du); CREF3(v2rhosigma_drho_dd); CREF3(v2sigma2_drho_uu_u); CREF3(v2sigma2_drho_uu_d); CREF3(v2sigma2_drho_ud_u); CREF3(v2sigma2_drho_ud_d); CREF3(v2sigma2_drho_du_u); CREF3(v2sigma2_drho_du_d); CREF3(v2sigma2_drho_dd_u); CREF3(v2sigma2_drho_dd_d); + CREF(v3rho3); CREF(v3rho2sigma); CREF(v3rhosigma2); CREF(v3sigma3); + /// @brief The coefficients of + /// $v^{(2)}(r)=\iint dr'dr''\,g^{xc}(r,r',r'')\rho^1(r')\rho^1(r'')$ + /// for ONE spin combination, stored exactly as they appear in the final formula: + /// $v^{(2)} = a_{s^2}s^2 + a_{st}\,s\,t + a_{t^2}t^2 + a_q\,q + /// - \nabla\cdot[\,(e_{s^2}s^2 + e_{st}\,s\,t + e_{t^2}t^2 + e_q\,q)\,\nabla\rho + /// + (c_s\,s + c_t\,t)\,\nabla\rho^1\,]$ + /// with $s=\rho^1$, $t=\nabla\rho\cdot\nabla\rho^1$, $q=\nabla\rho^1\cdot\nabla\rho^1$. + /// + /// Every numeric factor and every spin sum is folded in here, so `PotGradXCLR::cal_v_eff` + /// is a literal transcription of the formula with no arithmetic of its own, and nspin=1, + /// singlet and triplet all run through the same code. For LDA only `a_s2` is filled. + /// $v^{(2)} = A - \nabla\cdot E$ + /// + /// NOTE the four $e$ are *scalars*, with the common $\nabla\rho^{gs}$ factored out of the sum. + /// Should an open-shell version ever need $\nabla\rho_u\ne\nabla\rho_d$ under the same + /// divergence, this factorization no longer holds and they must go back to `Vector3`. + struct GxcCoef + { + std::vector a_s2, a_st, a_t2, a_q; ///< the local part $A$ + std::vector c_s, c_t; ///< the two $\nabla\rho^1$-weighted scalars + std::vector e_s2, e_st, e_t2, e_q; ///< under the divergence, all times $\nabla\rho^{gs}$ + }; + /// nspin=1 has no singlet/triplet distinction, so it always returns the one set that is built. + /// Throws instead of handing back an empty set when the requested combination was not + /// requested at construction -- silently returning zeros would look like a physics bug. + const GxcCoef& gxc(const bool triplet) const + { + const GxcCoef& ret = (nspin_ == 1 || !triplet) ? gxc_s_ : gxc_t_; + if (ret.a_s2.empty()) + { + throw std::runtime_error("KernelXC: the " + std::string(triplet ? "triplet" : "singlet") + + " g^xc coefficients were not built; pass the matching `GxcSpin` flag to the constructor."); + } + return ret; + } + + const bool& openshell = openshell_; const std::vector>>& drho_gs = drho_gs_; + /// Whether THIS kernel (built for its own functional name, which may differ from the + /// ground state's `dft_functional` in a cross-functional run such as TDLDA@PBE) needs + /// the GGA gradient terms. `drho_gs_` is only ever filled when this kernel's own `is_gga` + /// was true at construction (see `f_xc_libxc`), so its emptiness is a reliable per-kernel + /// proxy -- unlike the global `XC_Functional::get_func_type()`, which reflects the + /// ground state's functional and disagrees with this kernel whenever the two differ. + bool is_gga() const { return !this->drho_gs_.empty(); } private: #ifdef __LIBXC /// @brief Calculate the XC kernel using libxc. void f_xc_libxc(const int& nspin, const double& omega, const double& tpiba, const double* const* const rho_gs, const double* const rho_core = nullptr); + /// calculate the input rho, grad rho, and sigma for libxc + void get_rho_drho_sigma(const int& nspin, + const double& tpiba, + const double* const* const rho_gs, + const double* const rho_core, + const bool& is_gga, + std::vector& rho, + std::vector>>& gradrho, + std::vector& sigma); #endif // See https://libxc.gitlab.io/manual/libxc-5.1.x/ for the naming convention of the following members. // std::map> kernel_set_; // [kernel_type][nrxx][nspin] + + // ================================== XC kernels ============================================ std::vector vrho_; std::vector vsigma_; std::vector v2rho2_; @@ -70,6 +162,23 @@ namespace LR Tvec3 v2sigma2_drho_du_d_; /// $2f^{\sigma_{ud}\sigma_{dd}}\nabla\rho_d+f^{\sigma_{ud}\sigma_{ud}}\nabla\rho_u$ Tvec3 v2sigma2_drho_dd_u_; /// $2f^{\sigma_{ud}\sigma_{dd}}\nabla\rho_d+f^{\sigma_{ud}\sigma_{ud}}\nabla\rho_u$ Tvec3 v2sigma2_drho_dd_d_; /// $4f^{\sigma_{dd}\sigma_{dd}}\nabla\rho_d+2f^{\sigma_{ud}\sigma_{dd}\nabla\rho_u$ + // ================================== XC kernels ============================================ + // ================================== XC kernel Gradiants ==================================== + Tvec v3rho3_; + Tvec v3rho2sigma_; + Tvec v3rhosigma2_; + Tvec v3sigma3_; + + // The two spin combinations of $v^{(2)}$'s coefficients (see `GxcCoef` above). + // `gxc_t_` stays empty for nspin=1, where there is no triplet. + GxcCoef gxc_s_; + GxcCoef gxc_t_; + /// @brief Fill `dst` for one spin combination. All the spin algebralives here, driven by the weight + /// vectors that distinguish singlet from triplet -- the two differ only in those weights. + void build_gxc_coef(GxcCoef& dst, const bool triplet, const int& nspin, const bool& is_gga); + int nspin_ = 1; + const int gxc_spin_ = GxcSpin::NoGxc; ///< which `GxcCoef` sets to build, see `GxcSpin` + // ================================== XC kernel Gradiants ==================================== const ModulePW::PW_Basis& rho_basis_; const bool openshell_ = false; }; diff --git a/source/source_lcao/module_lr/pulay_hc.h b/source/source_lcao/module_lr/pulay_hc.h new file mode 100644 index 00000000000..ffa188f5528 --- /dev/null +++ b/source/source_lcao/module_lr/pulay_hc.h @@ -0,0 +1,155 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_PULAY_HC_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_PULAY_HC_H +#include "source_basis/module_nao/two_center_bundle.h" +#include "source_estate/module_dm/density_matrix.h" +#include "source_cell/unitcell.h" +#include "source_lcao/module_lr/utils/lr_util.h" +#include "source_lcao/module_lr/potentials/pot_lr_base.h" +#include "source_hamilt/module_gint/gint_interface.h" +#include "source_base/parallel_reduce.h" + +namespace PulayForceStress +{ +// add stess later +/// for 2-center-integration terms, provided HS derivatives + template +ModuleBase::matrix cal_pulay_fs( + const module_dm::DensityMatrix& dm, ///< [in] density matrix or energy density matrix + const UnitCell& ucell, ///< [in] unit cell + const std::vector>& dHS, ///< [in] dHS x, y, z, for force + const double& factor_force = 1.0) +{ + ModuleBase::matrix f(ucell.nat, 3); + const Parallel_Orbitals& pv = *dHS[0].get_paraV(); + const int& npol = ucell.get_npol(); + const int nspin_dmr = dm.get_dmr_vec().size(); + for (int ixyz = 0;ixyz < 3;++ixyz) + { + for (int iat0 = 0;iat0 < ucell.nat;++iat0) + { + for (int iat1 = 0;iat1 < ucell.nat;++iat1) + { + hamilt::AtomPair* ap = dHS[ixyz].find_pair(iat0, iat1); + if (ap) + { + for (int iR = 0;iR < ap->get_R_size();++iR) + { + const ModuleBase::Vector3& R = ap->get_R_index(iR); + hamilt::BaseMatrix& mat_dhs = ap->get_HR_values(R.x, R.y, R.z); + std::vector*> mat_dmr; + for (int is = 0; is < nspin_dmr; ++is) + { + mat_dmr.push_back(dm.get_dmr_ptr(is + 1)->find_matrix(iat0, iat1, R.x, R.y, R.z)); + } + + for (int mu = 0; mu < pv.get_nrow_atom(iat0); mu += npol) + { + for (int nu = 0; nu < pv.get_ncol_atom(iat1); nu += npol) + { + double dm2d = 0.0; + for (int is = 0; is < nspin_dmr; ++is) { dm2d += mat_dmr[is]->get_value(mu, nu); } + f(iat0, ixyz) += dm2d * factor_force * 2.0 * mat_dhs.get_value(mu, nu); + } + } + } + } + } + } + } + Parallel_Reduce::reduce_all(f.c, f.nr * f.nc); + return f; +} + +/// for grid-integration terms +template +ModuleBase::matrix cal_pulay_fs( + const module_dm::DensityMatrix& dm, ///< [in] density matrix or energy density matrix + const UnitCell& ucell, ///< [in] unit cell + const LR::PotLRBase* pot ///< [in] potential on grid +) +{ + ModuleBase::matrix force(ucell.nat, 3); + ModuleBase::matrix stress_tmp(3, 3); + + // The LR kernel yields a single potential channel, so the grid integrals run with + // nspin=1 on the first density-matrix channel -- the same thing the old + // `Gint_inout(is=0, ...)` calls did. `dm` may still carry two channels (the + // `test_force` path feeds in the ground-state DM and doubles the result afterwards); + // ModuleGint reads only the leading `nspin` entries of `dm_vec`, but `vr_eff` is + // indexed up to `nspin`, so passing the DM's spin count here would read past + // `p_vr_hxc` and segfault. + constexpr int nspin_gint = 1; + + // 1. dm->rho + double** rho; + const int& nrxx = pot->nrxx; + LR_Util::_allocate_2order_nested_ptr(rho, nspin_gint, nrxx); + ModuleBase::GlobalFunc::ZEROS(rho[0], nrxx); + ModuleGint::cal_gint_rho(dm.get_dmr_vec(), nspin_gint, rho, false); + + // 2. v_hxc = f_hxc * rho + ModuleBase::matrix vr_hxc(1, nrxx); //grid + pot->cal_v_eff(rho, ucell, vr_hxc); + LR_Util::_deallocate_2order_nested_ptr(rho, nspin_gint); + + // 3. v(r) -> force + // An empty local FFT slab has no (0, 0) element. Pass the storage pointer + // directly; the grid integrator has no local points to read on that rank. + const std::vector p_vr_hxc(nspin_gint, vr_hxc.c); + ModuleGint::cal_gint_fvl(nspin_gint, p_vr_hxc, dm.get_dmr_vec(), /*isforce=*/true, /*isstress=*/false, &force, &stress_tmp); + // `cal_gint_fvl` only sums the grid points (and their atom pairs) this rank's share of the + // real-space FFT box touches; core ABACUS always follows it with this same reduction (see + // `force_stress_lcao.cpp`'s `Parallel_Reduce::reduce_pool` right after its own grid-based + // `cal_pulay_fs` call) -- omitted here, so this was silently wrong for nprocs>1. + Parallel_Reduce::reduce_pool(force.c, force.nr * force.nc); + return force; +} + +/// @brief Open-shell (spin-unrestricted) counterpart of the grid `cal_pulay_fs` above. +/// +/// The LR kernel mixes the two channels, +/// $v_\sigma=\sum_{\sigma'}f^{\sigma\sigma'}\rho^1_{\sigma'}$ (+ the full Hartree), +/// so the density has to be built per channel, the potential accumulated over the inner +/// spin, and only then contracted with the density matrix of the outer spin. `PotHxcLR` +/// with `SpinType::S2_updown` selects the $(\sigma,\sigma')$ component via `ispin_op`. +template +ModuleBase::matrix cal_pulay_fs_openshell( + const module_dm::DensityMatrix& dm, ///< [in] 2-channel density matrix + const UnitCell& ucell, + const LR::PotLRBase* pot) +{ + ModuleBase::matrix force(ucell.nat, 3); + ModuleBase::matrix stress_tmp(3, 3); + constexpr int nspin_dm = 2; + assert(dm.get_dmr_vec().size() == nspin_dm); + + // 1. dm -> rho, one channel each + double** rho; + const int& nrxx = pot->nrxx; + LR_Util::_allocate_2order_nested_ptr(rho, nspin_dm, nrxx); + for (int is = 0; is < nspin_dm; ++is) { ModuleBase::GlobalFunc::ZEROS(rho[is], nrxx); } + ModuleGint::cal_gint_rho(dm.get_dmr_vec(), nspin_dm, rho, false); + + // 2. $v_\sigma=\sum_{\sigma'}f^{\sigma\sigma'}\rho_{\sigma'}$ + std::vector vr_hxc(nspin_dm, ModuleBase::matrix(1, nrxx)); + for (int sl = 0; sl < nspin_dm; ++sl) + { + for (int sr = 0; sr < nspin_dm; ++sr) + { + double* rho_in[1] = { rho[sr] }; + pot->cal_v_eff(rho_in, ucell, vr_hxc[sl], { sl, sr }); + } + } + LR_Util::_deallocate_2order_nested_ptr(rho, nspin_dm); + + // 3. v(r) -> force, summed over the outer spin by `cal_gint_fvl` + std::vector p_vr_hxc(nspin_dm); + // Keep empty-grid ranks in the integration and subsequent pool reduction. + for (int is = 0; is < nspin_dm; ++is) { p_vr_hxc[is] = vr_hxc[is].c; } + ModuleGint::cal_gint_fvl(nspin_dm, p_vr_hxc, dm.get_dmr_vec(), /*isforce=*/true, false, &force, &stress_tmp); + Parallel_Reduce::reduce_pool(force.c, force.nr * force.nc); // see the closed-shell overload above + return force; +} +} + +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_PULAY_HC_H diff --git a/source/source_lcao/module_lr/root_ovlp.cpp b/source/source_lcao/module_lr/root_ovlp.cpp new file mode 100644 index 00000000000..21f096a352f --- /dev/null +++ b/source/source_lcao/module_lr/root_ovlp.cpp @@ -0,0 +1,244 @@ +#include "root_ovlp.h" +#include "source_base/matrix3.h" +#include "source_base/module_external/blas_connector.h" +#include +#include +#include +#include +#include +#include "source_base/parallel_2d.h" +#include "source_basis/module_ao/parallel_orbitals.h" +#include "source_cell/unitcell.h" +#include "source_lcao/module_operator_lcao/ovlp_block.h" +#include "source_hamilt/module_hcontainer/hcontainer_funcs.h" +#ifdef __MPI +#include "source_base/module_external/scalapack_connector.h" +#endif + +namespace LR +{ +std::vector> cross_image_indices( + const ModuleBase::Vector3& displacement, + const ModuleBase::Matrix3& lattice, const double cutoff) +{ + const double volume = lattice.Det(); + if (cutoff <= 0.0 || std::abs(volume) < 1e-14) + { + throw std::invalid_argument("LR root overlap: invalid lattice or orbital cutoff"); + } + const ModuleBase::Matrix3 inverse = lattice.Inverse(); + const ModuleBase::Vector3 fractional = displacement * inverse; + const ModuleBase::Vector3 reciprocal_x(inverse.e11, inverse.e21, inverse.e31); + const ModuleBase::Vector3 reciprocal_y(inverse.e12, inverse.e22, inverse.e32); + const ModuleBase::Vector3 reciprocal_z(inverse.e13, inverse.e23, inverse.e33); + const double extent_x = cutoff * reciprocal_x.norm(); + const double extent_y = cutoff * reciprocal_y.norm(); + const double extent_z = cutoff * reciprocal_z.norm(); + const double lower_x = -fractional.x - extent_x; + const double upper_x = -fractional.x + extent_x; + const double lower_y = -fractional.y - extent_y; + const double upper_y = -fractional.y + extent_y; + const double lower_z = -fractional.z - extent_z; + const double upper_z = -fractional.z + extent_z; + const int xmin = static_cast(std::ceil(lower_x)); + const int xmax = static_cast(std::floor(upper_x)); + const int ymin = static_cast(std::ceil(lower_y)); + const int ymax = static_cast(std::floor(upper_y)); + const int zmin = static_cast(std::ceil(lower_z)); + const int zmax = static_cast(std::floor(upper_z)); + std::vector> images; + for (int x = xmin; x <= xmax; ++x) + { + for (int y = ymin; y <= ymax; ++y) + { + for (int z = zmin; z <= zmax; ++z) + { + const ModuleBase::Vector3 image(x, y, z); + const ModuleBase::Vector3 shifted = displacement + image * lattice; + if (shifted.norm() < cutoff) { images.push_back(image); } + } + } + } + return images; +} + +// Build ordered old-bra/new-ket S(R) blocks on the existing AO 2D distribution. +// Fourier folding preserves their nonsymmetry; Gamma is simply k = 0. +std::vector cross_ao_overlap(const UnitCell& cell, + const std::vector>& old_positions, + const std::vector& cutoff, const TwoCenterIntegrator& integrator, + const Parallel_Orbitals& pmat) +{ + if (cell.get_npol() != 1) { throw std::invalid_argument("LR root overlap requires collinear spin"); } + ModuleBase::Matrix3 lattice = cell.latvec; + lattice *= cell.lat0; + hamilt::HContainer overlap_r(&pmat); + for (int i = 0; i < cell.nat; ++i) + { + const int type_i = cell.iat2it[i]; + for (int j = 0; j < cell.nat; ++j) + { + if (pmat.is_invalid_atom_pair(i, j)) { continue; } + const int type_j = cell.iat2it[j]; + const auto displacement = (cell.get_tau(j) - old_positions[i]) * cell.lat0; + const double radius = cutoff[type_i] + cutoff[type_j]; + const auto indices = cross_image_indices(displacement, lattice, radius); + for (const auto& index : indices) + { + const hamilt::AtomPair pair(i, j, index, &pmat); + overlap_r.insert_pair(pair); + } + } + } + overlap_r.allocate(nullptr, true); + for (int pair_index = 0; pair_index < overlap_r.size_atom_pairs(); ++pair_index) + { + auto& pair = overlap_r.get_atom_pair(pair_index); + const int i = pair.get_atom_i(); + const int j = pair.get_atom_j(); + const auto displacement = (cell.get_tau(j) - old_positions[i]) * cell.lat0; + for (int ir = 0; ir < pair.get_R_size(); ++ir) + { + const auto index = pair.get_R_index(ir); + const auto shifted = displacement + index * lattice; + double* const values = pair.get_pointer(ir); + hamilt::cal_overlap_block(cell, integrator, i, j, pmat, shifted, values); + } + } + // Use the complex Fourier overload even at Gamma: it folds every retained R + // without fix_gamma(), so the same S(R) can later be folded at nonzero k. + const int local_count = pmat.get_local_size(); + const std::complex zero(0.0, 0.0); + std::vector> local_overlap(local_count, zero); + const ModuleBase::Vector3 gamma(0.0, 0.0, 0.0); + const int leading_dimension = pmat.get_row_size(); + if (local_count > 0) + { + hamilt::folding_HR(overlap_r, local_overlap.data(), gamma, leading_dimension, 1); + } + std::vector overlap(local_count, 0.0); + for (int index = 0; index < local_count; ++index) + { + overlap[index] = local_overlap[index].real(); + } + return overlap; +} + +namespace +{ +void check_grid(const Parallel_2D& a, const Parallel_2D& b) +{ +#ifdef __MPI + if (a.blacs_ctxt != b.blacs_ctxt || a.get_block_size() != b.get_block_size()) + { + throw std::invalid_argument("Root overlap matrices require the same BLACS grid and block size"); + } +#endif +} +} + +template +std::vector mo_overlap_dist(const std::vector& old_coefficients, + const std::vector& ao, const std::vector& current_coefficients, + const Parallel_2D& pc, const Parallel_2D& ps, const Parallel_2D& po) +{ + check_grid(pc, ps); + check_grid(pc, po); + const int naos = pc.get_global_row_size(); + const int bands = pc.get_global_col_size(); + if (ps.get_global_row_size() != naos || ps.get_global_col_size() != naos + || po.get_global_row_size() != bands || po.get_global_col_size() != bands + || old_coefficients.size() < pc.get_local_size() + || current_coefficients.size() < pc.get_local_size() || ao.size() < ps.get_local_size()) + { + throw std::invalid_argument("Root overlap matrix dimensions do not match their distributions"); + } + const T zero = T(0); + const T one = T(1); + const int coefficient_count = std::max(1, pc.get_local_size()); + const int overlap_count = std::max(1, po.get_local_size()); + std::vector intermediate(coefficient_count, zero); + std::vector overlap(overlap_count, zero); + const T* const old_data = old_coefficients.empty() ? &zero : old_coefficients.data(); + const T* const new_data = current_coefficients.empty() ? &zero : current_coefficients.data(); + const T* const ao_data = ao.empty() ? &zero : ao.data(); + const char adjoint = std::is_same::value ? 'T' : 'C'; +#ifdef __MPI + const int first = 1; + // Even ranks with empty local blocks join both collective PBLAS calls. + ScalapackConnector::gemm('N', 'N', naos, bands, naos, one, + ao_data, first, first, ps.desc, new_data, first, first, pc.desc, + zero, intermediate.data(), first, first, pc.desc); + ScalapackConnector::gemm(adjoint, 'N', bands, bands, naos, one, + old_data, first, first, pc.desc, intermediate.data(), first, first, pc.desc, + zero, overlap.data(), first, first, po.desc); +#else + BlasConnector::gemm_cm('N', 'N', naos, bands, naos, one, + ao_data, naos, new_data, naos, zero, intermediate.data(), naos); + BlasConnector::gemm_cm(adjoint, 'N', bands, bands, naos, one, + old_data, naos, intermediate.data(), naos, zero, overlap.data(), bands); +#endif + overlap.resize(po.get_local_size()); + return overlap; +} + +template +std::vector project_reference_dist(const std::vector& previous, + const std::vector& overlap, const Parallel_2D& po, + const Parallel_2D& px, const int nocc, const int nvirt) +{ + check_grid(po, px); + const int bands = nocc + nvirt; + if (po.get_global_row_size() != bands || po.get_global_col_size() != bands + || px.get_global_row_size() != nvirt || px.get_global_col_size() != nocc + || previous.size() != px.get_local_size() || overlap.size() != po.get_local_size() + || po.get_block_size() != 1) + { + throw std::invalid_argument("Root projection requires matching windows and unit block size"); + } + const T zero = T(0); + const T one = T(1); + const int local_count = std::max(1, px.get_local_size()); + std::vector intermediate(local_count, zero); + std::vector projected(local_count, zero); + const T* const old_data = previous.empty() ? &zero : previous.data(); + const T* const mo_data = overlap.empty() ? &zero : overlap.data(); + const char adjoint = std::is_same::value ? 'T' : 'C'; +#ifdef __MPI + const int first = 1; + const int virtual_start = nocc + 1; + // O's virtual/occupied blocks are submatrices; unit blocks keep their + // origins aligned with px, including ranks that own no electron-hole pairs. + ScalapackConnector::gemm(adjoint, 'N', nvirt, nocc, nvirt, one, + mo_data, virtual_start, virtual_start, po.desc, + old_data, first, first, px.desc, zero, + intermediate.data(), first, first, px.desc); + ScalapackConnector::gemm('N', 'N', nvirt, nocc, nocc, one, + intermediate.data(), first, first, px.desc, + mo_data, first, first, po.desc, zero, + projected.data(), first, first, px.desc); +#else + const int virtual_offset = nocc * bands + nocc; + const T* const virtual_data = mo_data + virtual_offset; + BlasConnector::gemm_cm(adjoint, 'N', nvirt, nocc, nvirt, one, + virtual_data, bands, old_data, nvirt, zero, intermediate.data(), nvirt); + BlasConnector::gemm_cm('N', 'N', nvirt, nocc, nocc, one, + intermediate.data(), nvirt, mo_data, bands, zero, projected.data(), nvirt); +#endif + projected.resize(px.get_local_size()); + return projected; +} + +template std::vector mo_overlap_dist(const std::vector&, + const std::vector&, const std::vector&, + const Parallel_2D&, const Parallel_2D&, const Parallel_2D&); +template std::vector> mo_overlap_dist( + const std::vector>&, const std::vector>&, + const std::vector>&, const Parallel_2D&, const Parallel_2D&, const Parallel_2D&); +template std::vector project_reference_dist(const std::vector&, + const std::vector&, const Parallel_2D&, const Parallel_2D&, int, int); +template std::vector> project_reference_dist( + const std::vector>&, const std::vector>&, + const Parallel_2D&, const Parallel_2D&, int, int); + +} diff --git a/source/source_lcao/module_lr/root_ovlp.h b/source/source_lcao/module_lr/root_ovlp.h new file mode 100644 index 00000000000..f993796cae5 --- /dev/null +++ b/source/source_lcao/module_lr/root_ovlp.h @@ -0,0 +1,49 @@ +#ifndef ABACUS_LR_ROOT_OVERLAP_H +#define ABACUS_LR_ROOT_OVERLAP_H +#include "source_base/vector3.h" +#include + +// Cross-geometry AO/MO overlaps and reference projection, without tracking state. +// Distributed buffers are local column-major blocks; no matrices are gathered here. +namespace ModuleBase { class Matrix3; } +class UnitCell; +class Parallel_Orbitals; +class Parallel_2D; +class TwoCenterIntegrator; +namespace LR +{ +// Construct ordered old-bra/new-ket S(R), then fold at Gamma. Return this rank's +// column-major AO block; old_positions are Cartesian in cell.lat0 units, cutoffs in Bohr. +// Uses independent cross-geometry candidates and preserves nonsymmetry. +std::vector cross_ao_overlap(const UnitCell& cell, + const std::vector>& old_positions, + const std::vector& cutoff, const TwoCenterIntegrator& integrator, + const Parallel_Orbitals& pmat); + +// Local column-major matrices on one BLACS grid: O = C_old^dagger S C_current. +// pc describes AO-by-band coefficients; ps describes S; po describes band-by-band O. +// Row/column MO indices refer to the old/current step respectively; S need not be Hermitian. +// All layouts share a BLACS grid and block size, including ranks with empty local blocks. +template +std::vector mo_overlap_dist(const std::vector& old_coefficients, + const std::vector& ao, const std::vector& current_coefficients, + const Parallel_2D& pc, const Parallel_2D& ps, const Parallel_2D& po); +// X is stored as A = X^T (virtual rows, occupied columns). Return local +// A_projected = O_virt^dagger A_previous O_occ, without normalization. +// po describes the full MO overlap; px describes the electron-hole block. +// previous is the old-step local reference (possibly a JT mixture), not the current roots. +// Unit blocks align the occupied/virtual submatrix origins with px. +template +std::vector project_reference_dist(const std::vector& previous, + const std::vector& overlap, const Parallel_2D& po, + const Parallel_2D& px, int nocc, int nvirt); + + +// Integer lattice translations selected by the same cross-geometry cutoff. +// Retain these indices for HContainer storage and exp(+i k.R) Fourier phases. +std::vector> cross_image_indices( + const ModuleBase::Vector3& displacement, + const ModuleBase::Matrix3& lattice, double cutoff); + +} +#endif diff --git a/source/source_lcao/module_lr/root_track.cpp b/source/source_lcao/module_lr/root_track.cpp new file mode 100644 index 00000000000..f27bfe6e88d --- /dev/null +++ b/source/source_lcao/module_lr/root_track.cpp @@ -0,0 +1,122 @@ +#include "root_track.h" +#include "root_ovlp.h" +#include "source_psi/psi.h" +#include "lr_amp.h" +#include "source_base/parallel_2d.h" +#include "source_cell/unitcell.h" +#include "source_basis/module_ao/parallel_orbitals.h" +#include +#include + +namespace LR +{ +namespace +{ +template +RootBasis capture_basis(const RootInputs& inputs) +{ + RootBasis basis; + basis.lattice = inputs.cell.latvec; + basis.lat0 = inputs.cell.lat0; + basis.nbasis = inputs.pc.get_global_row_size(); + basis.nbands = inputs.nocc[0] + inputs.nvirt[0]; + basis.nocc = inputs.nocc; + basis.nvirt = inputs.nvirt; + basis.layout = {inputs.pc.get_block_size(), inputs.pc.get_dim0(), inputs.pc.get_dim1(), + inputs.pc.get_coord_row(), inputs.pc.get_coord_col(), inputs.pc.get_row_size(), inputs.pc.get_col_size()}; + for (const auto& px : inputs.px) + { + basis.layout.push_back(px.get_block_size()); + basis.layout.push_back(px.get_row_size()); + basis.layout.push_back(px.get_col_size()); + } + + for (int atom = 0; atom < inputs.cell.nat; ++atom) + { + const auto position = inputs.cell.get_tau(atom); + basis.positions.push_back(position); + } + // Snapshot only this rank's column-major coefficient block; no AO-by-band gather. + const int coefficient_count = inputs.pc.get_local_size(); + const T zero = T(0); + for (int spin = 0; spin < inputs.orbitals.get_nk(); ++spin) + { + std::vector coefficients(coefficient_count, zero); + for (int band = 0; band < inputs.pc.get_col_size(); ++band) + { + for (int ao = 0; ao < inputs.pc.get_row_size(); ++ao) + { + coefficients[band * inputs.pc.get_row_size() + ao] = inputs.orbitals(spin, band, ao); + } + } + basis.coefficients.push_back(std::move(coefficients)); + } + return basis; +} + + +template +void project_local_reference(const RootInputs& inputs, const RootBasis& old, + const RootBasis& current, const std::vector& ao, std::vector& previous) +{ + Parallel_2D po; +#ifdef __MPI + po.set(current.nbands, current.nbands, inputs.pc.get_block_size(), inputs.pc.blacs_ctxt); +#else + po.set_serial(current.nbands, current.nbands); +#endif + const std::vector typed_ao(ao.begin(), ao.end()); + int offset = 0; + const std::vector spins = inputs.openshell ? std::vector{0, 1} : std::vector{inputs.channel}; + for (const int spin : spins) + { + const Parallel_2D& px = inputs.px[spin]; + const int no = inputs.nocc[spin]; + const int nv = inputs.nvirt[spin]; + const int local_count = px.get_local_size(); + std::vector local_reference(local_count); + std::copy_n(previous.begin() + offset, local_count, local_reference.begin()); + const auto local_mo = mo_overlap_dist(old.coefficients[spin], typed_ao, + current.coefficients[spin], inputs.pc, inputs.pmat, po); + const auto projected = project_reference_dist(local_reference, local_mo, po, px, no, nv); + std::copy(projected.begin(), projected.end(), previous.begin() + offset); + offset += px.get_local_size(); + } +} +} + +template +void follow_cross_root(const RootInputs& inputs, const T* amplitudes, + const int nlocal, const int nstates, const int initial_state, int& target, + std::vector& previous, RootBasis& basis, std::ostream& log) +{ + RootBasis current = capture_basis(inputs); + if (!basis.positions.empty()) + { + const auto difference = current.lattice - basis.lattice; + const double cell_change = std::abs(difference.e11) + std::abs(difference.e12) + std::abs(difference.e13) + + std::abs(difference.e21) + std::abs(difference.e22) + std::abs(difference.e23) + + std::abs(difference.e31) + std::abs(difference.e32) + std::abs(difference.e33); + if (cell_change > 1e-12 || std::abs(current.lat0 - basis.lat0) > 1e-12 + || current.nbasis != basis.nbasis || current.nbands != basis.nbands + || current.nocc != basis.nocc || current.nvirt != basis.nvirt + || current.layout != basis.layout + || current.positions.size() != basis.positions.size() + || current.coefficients.size() != basis.coefficients.size()) + { + throw std::invalid_argument("LR root overlap: cell, orbital window or MPI layout changed between steps"); + } + const auto ao = cross_ao_overlap(inputs.cell, basis.positions, inputs.cutoff, inputs.integrator, inputs.pmat); + project_local_reference(inputs, basis, current, ao, previous); + log << " EXCITED-STATE RELAX: root overlap uses old-new AO and MO bases." << std::endl; + } + follow_root(amplitudes, nlocal, nstates, initial_state, target, previous, log); + basis = std::move(current); +} + +template void follow_cross_root(const RootInputs&, const double*, int, int, int, + int&, std::vector&, RootBasis&, std::ostream&); +template void follow_cross_root>(const RootInputs>&, + const std::complex*, int, int, int, int&, std::vector>&, + RootBasis>&, std::ostream&); +} diff --git a/source/source_lcao/module_lr/root_track.h b/source/source_lcao/module_lr/root_track.h new file mode 100644 index 00000000000..ccc5aca18b4 --- /dev/null +++ b/source/source_lcao/module_lr/root_track.h @@ -0,0 +1,57 @@ +#ifndef ABACUS_LR_ROOT_TRACK_H +#define ABACUS_LR_ROOT_TRACK_H +#include "source_base/matrix3.h" +#include +#include +class UnitCell; +class Parallel_2D; +class Parallel_Orbitals; +class TwoCenterIntegrator; +namespace base_device { struct DEVICE_CPU; } +namespace psi { template class Psi; } +namespace LR +{ +// Snapshot of the previous geometry and MO basis; reference amplitudes are stored separately. +// Lattice and Cartesian positions use lat0 units; coefficients are local per-spin blocks. +template +struct RootBasis +{ + ModuleBase::Matrix3 lattice; + double lat0 = 0.0; + int nbasis = 0; + int nbands = 0; + std::vector> positions; + std::vector nocc; + std::vector nvirt; + // Local column-major AO-by-band blocks; layout must stay fixed between steps. + std::vector layout; + std::vector> coefficients; +}; + +// Borrow current geometry, orbitals and AO/MO/electron-hole distributions for one tracking step. +// Cutoffs are in Bohr; pc/pmat/px describe coefficient, AO and amplitude matrices respectively. +template +struct RootInputs +{ + const UnitCell& cell; + const std::vector& cutoff; + const TwoCenterIntegrator& integrator; + const psi::Psi& orbitals; + const Parallel_2D& pc; + const Parallel_Orbitals& pmat; + const std::vector& px; + const std::vector& nocc; + const std::vector& nvirt; + bool openshell; + int channel; +}; + +// Form old-new AO/MO overlaps and project the old reference into the current basis, +// then select the current root with largest overlap. Update target, previous and basis; +// current amplitudes stay unchanged. Requires Gamma, fixed cell/window and local layout. +template +void follow_cross_root(const RootInputs& inputs, const T* amplitudes, + int nlocal, int nstates, int initial_state, int& target, + std::vector& previous, RootBasis& basis, std::ostream& log); +} +#endif diff --git a/source/source_lcao/module_lr/test/CMakeLists.txt b/source/source_lcao/module_lr/test/CMakeLists.txt new file mode 100644 index 00000000000..b6596cd6a9f --- /dev/null +++ b/source/source_lcao/module_lr/test/CMakeLists.txt @@ -0,0 +1,53 @@ +AddTest( + TARGET MODULE_LR_exx_projection + LIBS base parameter ${math_libs} container device + SOURCES test_exx_proj.cpp ../exx_proj.cpp +) + +AddTest( + TARGET MODULE_LR_grad_degen + LIBS base parameter ${math_libs} container device + SOURCES test_grad_degen.cpp test_grad_jt.cpp ../grad_degen.cpp ../grad_jt.cpp +) + +if(ENABLE_MPI) + if(TARGET ELPA::ELPA) + AddTest( + TARGET MODULE_LR_zeqlin_solv + LIBS base parameter ${math_libs} container device psi ELPA::ELPA MPI::MPI_CXX + SOURCES test_zeqlin.cpp ../zeqlin_solv.cpp ../utils/lr_util.cpp + ) + else() + AddTest( + TARGET MODULE_LR_zeqlin_solv + LIBS base parameter ${math_libs} container device psi MPI::MPI_CXX + SOURCES test_zeqlin.cpp ../zeqlin_solv.cpp ../utils/lr_util.cpp + ) + endif() +endif() + +AddTest( + TARGET MODULE_LR_gradient_amplitudes + LIBS base parameter ${math_libs} container device + SOURCES test_lr_amp.cpp +) + +# AO construction shares the existing two-center mock fixture with operator tests. +set(root_overlap_test_sources + ../root_ovlp.cpp + ../../module_operator_lcao/ovlp_block.cpp + ../../module_operator_lcao/test/tmp_mocks.cpp + ../../../source_basis/module_ao/parallel_orbitals.cpp + ../../../source_basis/module_ao/orb_atomic_lm.cpp + ../../../source_hamilt/operator.cpp + ../../../source_hamilt/module_hcontainer/base_matrix.cpp + ../../../source_hamilt/module_hcontainer/hcontainer.cpp + ../../../source_hamilt/module_hcontainer/atom_pair.cpp + ../../../source_hamilt/module_hcontainer/func_folding.cpp +) + +AddTest( + TARGET MODULE_LR_root_ovlp + LIBS base parameter ${math_libs} container device psi + SOURCES test_root_ovlp.cpp ${root_overlap_test_sources} +) diff --git a/source/source_lcao/module_lr/test/test_exx_proj.cpp b/source/source_lcao/module_lr/test/test_exx_proj.cpp new file mode 100644 index 00000000000..d999252108a --- /dev/null +++ b/source/source_lcao/module_lr/test/test_exx_proj.cpp @@ -0,0 +1,314 @@ +#include "../exx_proj.h" + +#include +#include + +namespace +{ +const int naos = 5; +const int nocc = 2; +const int nvirt = 3; + +std::vector coefficients(const int columns, const double offset) +{ + std::vector c(naos * columns); + for (int col = 0; col < columns; ++col) + { + for (int row = 0; row < naos; ++row) + { + c[col * naos + row] = offset + 0.13 * row - 0.21 * col; + } + } + return c; +} + +// Independent reference: construct every rank-one probe density and contract +// it with a nonsymmetric AO matrix, as DMBand + cal_energy did. +std::vector reference(const std::vector& h, + const std::vector& left, + const std::vector& right, + const int nright, + const double factor) +{ + std::vector result(nocc * nright, 0.0); + for (int io = 0; io < nocc; ++io) + { + for (int iv = 0; iv < nright; ++iv) + { + for (int mu = 0; mu < naos; ++mu) + { + for (int nu = 0; nu < naos; ++nu) + { + const double density = left[io * naos + mu] * right[iv * naos + nu]; + result[io * nright + iv] += factor * density * h[nu * naos + mu]; + } + } + } + } + return result; +} + +TEST(ExxProjection, AllFourModesPreserveNonsymmetricProbeContraction) +{ + std::vector h(naos * naos); + for (int col = 0; col < naos; ++col) + { + for (int row = 0; row < naos; ++row) + { + h[col * naos + row] = 0.17 * row + 0.32 * col + 0.09 * row * col; + } + } + const auto co = coefficients(nocc, 0.3); + const auto cv = coefficients(nvirt, -0.2); + const auto cvx = coefficients(nocc, 0.7); + const auto coxt = coefficients(nvirt, -0.5); + const double factor = 0.5; + const double minus_factor = -factor; + for (int mode = 0; mode < 4; ++mode) + { + const int nright = mode == 1 ? nocc : nvirt; + std::vector result(nocc * nright, 1.25); + std::vector scratch(naos * nocc); + std::vector expected; + if (mode == 0 || mode == 1) + { + const auto& right = mode == 1 ? co : cv; + LR::project_exx(h.data(), co.data(), right.data(), naos, nocc, nright, + factor, scratch.data(), result.data()); + expected = reference(h, co, right, nright, factor); + } + else if (mode == 2) + { + LR::project_exx(h.data(), cvx.data(), cv.data(), naos, nocc, nvirt, + factor, scratch.data(), result.data()); + LR::project_exx(h.data(), co.data(), coxt.data(), naos, nocc, nvirt, + minus_factor, scratch.data(), result.data()); + expected = reference(h, cvx, cv, nvirt, factor); + const auto occ = reference(h, co, coxt, nvirt, minus_factor); + for (std::size_t i = 0; i < expected.size(); ++i) { expected[i] += occ[i]; } + } + else + { + LR::project_exx(h.data(), co.data(), coxt.data(), naos, nocc, nvirt, + factor, scratch.data(), result.data()); + expected = reference(h, co, coxt, nvirt, factor); + } + for (std::size_t i = 0; i < expected.size(); ++i) + { + EXPECT_NEAR(result[i], 1.25 + expected[i], 1e-12) << "mode=" << mode << " element=" << i; + } + } +} + +TEST(ExxProjection, EmptyProjectionDoesNotDereferenceInputs) +{ + double result = 3.0; + LR::project_exx(nullptr, nullptr, nullptr, naos, 0, nvirt, 1.0, nullptr, &result); + EXPECT_DOUBLE_EQ(result, 3.0); +} + +TEST(ExxProjection, ComplexBlochPhasesAndAllFourProbeModes) +{ + typedef std::complex Complex; + const double factor = 0.5; + const double minus_factor = -factor; + const double angles[] = {0.0, 0.37, -0.61}; + std::vector> blocks(3, std::vector(naos * naos)); + std::vector co(naos * nocc); + std::vector cv(naos * nvirt); + std::vector cvx(naos * nocc); + std::vector coxt(naos * nvirt); + for (int i = 0; i < naos * nvirt; ++i) + { + cv[i] = Complex(0.2 - 0.03 * i, 0.13 + 0.07 * i); + coxt[i] = Complex(-0.1 + 0.02 * i, 0.04 - 0.05 * i); + } + for (int i = 0; i < naos * nocc; ++i) + { + co[i] = Complex(0.3 + 0.04 * i, -0.11 + 0.09 * i); + cvx[i] = Complex(0.1 - 0.07 * i, 0.23 + 0.03 * i); + } + for (int r = 0; r < 3; ++r) + { + for (int i = 0; i < naos * naos; ++i) + { + blocks[r][i] = Complex(0.17 * i + 0.2 * r, 0.32 - 0.03 * i + 0.11 * r); + } + } + // Multiple distinct k points and real-space cells. The reference forms + // the original negative-phase probe and conjugates it in dotc(D,H). + for (int ik = 0; ik < 3; ++ik) + { + std::vector h(naos * naos, Complex(0)); + for (int r = 0; r < 3; ++r) + { + const Complex phase = std::exp(Complex(0, (ik + 1) * angles[r])); + for (int i = 0; i < naos * naos; ++i) { h[i] += phase * blocks[r][i]; } + } + for (int mode = 0; mode < 4; ++mode) + { + const int nright = mode == 1 ? nocc : nvirt; + const auto& left = mode == 2 ? cvx : co; + const auto& right = mode == 1 ? co : (mode == 3 ? coxt : cv); + const Complex initial(1.25, -0.3); + std::vector result(nocc * nright, initial); + std::vector scratch(naos * nocc); + LR::project_exx(h.data(), left.data(), right.data(), naos, nocc, nright, + factor, scratch.data(), result.data()); + if (mode == 2) + { + LR::project_exx(h.data(), co.data(), coxt.data(), naos, nocc, nvirt, + minus_factor, scratch.data(), result.data()); + } + for (int io = 0; io < nocc; ++io) + { + for (int iv = 0; iv < nright; ++iv) + { + Complex expected = initial; + for (int r = 0; r < 3; ++r) + { + const Complex phase = std::exp(Complex(0, -(ik + 1) * angles[r])); + for (int mu = 0; mu < naos; ++mu) + { + for (int nu = 0; nu < naos; ++nu) + { + Complex density = phase * right[iv * naos + mu] + * std::conj(left[io * naos + nu]); + if (mode == 1) + { + density = phase * co[io * naos + mu] * std::conj(co[iv * naos + nu]); + } + if (mode == 2) + { + density -= phase * coxt[iv * naos + mu] * std::conj(co[io * naos + nu]); + } + expected += factor * std::conj(density) * blocks[r][nu * naos + mu]; + } + } + } + const int index = mode == 1 ? iv * nocc + io : io * nright + iv; + EXPECT_NEAR(std::abs(result[index] - expected), 0.0, 1e-11) + << "k=" << ik << " mode=" << mode << " io=" << io << " iv=" << iv; + } + } + } + } +} + +TEST(ExxProjection, ComplexEmptyProjectionDoesNotDereferenceInputs) +{ + std::complex result(3.0, 2.0); + const std::complex* input = nullptr; + LR::project_exx(input, input, input, naos, nocc, 0, 1.0, nullptr, &result); + EXPECT_EQ(result, std::complex(3.0, 2.0)); +} + + +TEST(ExxProjection, BenchmarkMappedIndicesPreserveProbeContraction) +{ + typedef std::complex Complex; + // Two atom types with 2 and 3 orbitals. The historical benchmark probe + // uses those counts as coefficient indices for every orbital of a type. + const int basis_index[] = {2, 2, 3, 3, 3}; + std::vector original(naos * naos); + std::vector folded(naos * naos, 0.0); + std::vector original_complex(naos * naos); + std::vector folded_complex(naos * naos, Complex(0)); + for (int mu = 0; mu < naos; ++mu) + { + for (int nu = 0; nu < naos; ++nu) + { + const int index = nu * naos + mu; + original[index] = 0.11 * mu - 0.23 * nu + 0.07 * mu * nu; + original_complex[index] = Complex(original[index], 0.03 * mu + 0.13 * nu); + const int mapped = basis_index[nu] * naos + basis_index[mu]; + folded[mapped] += original[index]; + folded_complex[mapped] += original_complex[index]; + } + } + const auto co = coefficients(nocc, 0.3); + const auto cv = coefficients(nvirt, -0.2); + std::vector co_complex(co.begin(), co.end()); + std::vector cv_complex(cv.begin(), cv.end()); + for (int i = 0; i < naos * nocc; ++i) { co_complex[i] += Complex(0, 0.05 * i); } + for (int i = 0; i < naos * nvirt; ++i) { cv_complex[i] += Complex(0, -0.09 * i); } + std::vector result(nocc * nvirt, 0.0); + std::vector scratch(naos * nocc); + std::vector result_complex(nocc * nvirt, Complex(0)); + std::vector scratch_complex(naos * nocc); + const double factor = 0.5; + LR::project_exx(folded.data(), co.data(), cv.data(), naos, nocc, nvirt, + factor, scratch.data(), result.data()); + LR::project_exx(folded_complex.data(), co_complex.data(), cv_complex.data(), + naos, nocc, nvirt, factor, scratch_complex.data(), result_complex.data()); + for (int io = 0; io < nocc; ++io) + { + for (int iv = 0; iv < nvirt; ++iv) + { + double expected = 0.0; + Complex expected_complex(0); + for (int mu = 0; mu < naos; ++mu) + { + for (int nu = 0; nu < naos; ++nu) + { + const int row = basis_index[mu]; + const int col = basis_index[nu]; + expected += factor * co[io * naos + row] * cv[iv * naos + col] + * original[nu * naos + mu]; + const Complex probe = cv_complex[iv * naos + row] + * std::conj(co_complex[io * naos + col]); + expected_complex += factor * std::conj(probe) * original_complex[nu * naos + mu]; + } + } + const int index = io * nvirt + iv; + EXPECT_NEAR(result[index], expected, 1e-12); + EXPECT_NEAR(std::abs(result_complex[index] - expected_complex), 0.0, 1e-12); + } + } +} + + +TEST(ExxProjection, TransposedResponseMatchesSwappedCxcOProbe) +{ + std::vector h(naos * naos); + std::vector transposed(naos * naos); + for (int mu = 0; mu < naos; ++mu) + { + for (int nu = 0; nu < naos; ++nu) + { + const double value = 0.17 * mu - 0.32 * nu + 0.09 * mu * nu; + h[nu * naos + mu] = value; + transposed[mu * naos + nu] = value; + } + } + const auto co = coefficients(nocc, 0.3); + const auto coxt = coefficients(nvirt, -0.5); + const double factor = 0.5; + std::vector result(nocc * nvirt, 0.0); + std::vector scratch(naos * nocc); + LR::project_exx(transposed.data(), co.data(), coxt.data(), naos, nocc, nvirt, + factor, scratch.data(), result.data()); + const auto normal = reference(h, co, coxt, nvirt, factor); + double difference = 0.0; + for (int io = 0; io < nocc; ++io) + { + for (int iv = 0; iv < nvirt; ++iv) + { + double expected = 0.0; + for (int mu = 0; mu < naos; ++mu) + { + for (int nu = 0; nu < naos; ++nu) + { + expected += factor * coxt[iv * naos + mu] * co[io * naos + nu] + * h[nu * naos + mu]; + } + } + const int index = io * nvirt + iv; + EXPECT_NEAR(result[index], expected, 1e-12); + difference += std::abs(result[index] - normal[index]); + } + } + EXPECT_GT(difference, 1e-6); +} + +} diff --git a/source/source_lcao/module_lr/test/test_grad_degen.cpp b/source/source_lcao/module_lr/test/test_grad_degen.cpp new file mode 100644 index 00000000000..95f14546d38 --- /dev/null +++ b/source/source_lcao/module_lr/test/test_grad_degen.cpp @@ -0,0 +1,295 @@ +#include + +#include +#include + +#include "source_base/matrix.h" +#include "../grad_degen.h" + +/// Tests for the degenerate-subspace gradient matrix algebra; `../grad_matrix_degenerate.h` +/// states the identity being exercised. +/// +/// The interesting test is not the arithmetic of `assemble_grad_matrix` but whether the route it +/// implements really recovers the bilinear form: `QuadraticForm` below plays the part of the +/// analytic-gradient pipeline (a genuine quadratic form in X), and the tests check that feeding it +/// the normalized combinations reproduces $G_{kl}=X_k^\top B X_l$ and that the result is covariant +/// under a rotation of the basis of the degenerate subspace -- the two self-checks the document +/// lists in section 5.4.5 as the ones needing no finite differences. + +namespace +{ + constexpr int nov = 5; ///< ambient particle-hole space, stands in for nocc*nvirt + + /// $\mathcal F[X]=X^\top B X$ with B symmetric: the same shape the real force map has on a + /// multiplet. Scalar-valued, which is enough -- the real one is just 3N independent copies. + struct QuadraticForm + { + std::vector b; ///< nov*nov, symmetric + + QuadraticForm() : b(nov * nov, 0.0) + { + // arbitrary but fixed and definitely not diagonal, so the off-diagonal elements of G + // are nonzero and the test can fail + const double raw[nov][nov] = { { 1.3, -0.7, 0.4, 0.9, -0.2 }, + { 0.0, 2.1, -1.1, 0.3, 0.6 }, + { 0.0, 0.0, -0.5, 0.8, 1.4 }, + { 0.0, 0.0, 0.0, 0.7, -0.9 }, + { 0.0, 0.0, 0.0, 0.0, 1.9 } }; + for (int i = 0; i < nov; ++i) + { + for (int j = i; j < nov; ++j) + { + b[i * nov + j] = raw[i][j]; + b[j * nov + i] = raw[i][j]; + } + } + } + + /// the bilinear form B behind it, i.e. the exact answer + double bilinear(const std::vector& x, const std::vector& y) const + { + double s = 0.0; + for (int i = 0; i < nov; ++i) + { + for (int j = 0; j < nov; ++j) { s += x[i] * b[i * nov + j] * y[j]; } + } + return s; + } + + /// what "the code" returns for one state + double operator()(const std::vector& x) const { return bilinear(x, x); } + }; + + /// three orthonormal vectors spanning the stand-in degenerate subspace + std::vector> subspace_basis() + { + std::vector> x; + x.push_back({ 1.0, 0.0, 0.0, 0.0, 0.0 }); + x.push_back({ 0.0, 0.6, 0.0, 0.8, 0.0 }); + x.push_back({ 0.0, 0.8, 0.0, -0.6, 0.0 }); + return x; + } + + /// run route A2 over a subspace basis and return G + std::vector> route_a2(const QuadraticForm& f, + const std::vector>& x) + { + const int d = static_cast(x.size()); + const std::vector> pairs = LR::degenerate_pairs(d); + std::vector diag(d); + for (int k = 0; k < d; ++k) { diag[k] = f(x[k]); } + std::vector plus(pairs.size()); + for (size_t ip = 0; ip < pairs.size(); ++ip) + { + std::vector xp(nov); + LR::combine_normalized(x[pairs[ip].first].data(), x[pairs[ip].second].data(), nov, + xp.data()); + plus[ip] = f(xp); + } + return LR::assemble_grad_matrix(diag, plus, pairs); + } +} + +/// The whole point: route A2 reproduces the bilinear form exactly, off-diagonal included. +TEST(GradMatrixDegenerate, PolarizationRecoversBilinearForm) +{ + const QuadraticForm f; + const std::vector> x = subspace_basis(); + const std::vector> g = route_a2(f, x); + + const int d = static_cast(x.size()); + bool any_offdiag = false; + for (int k = 0; k < d; ++k) + { + for (int l = 0; l < d; ++l) + { + EXPECT_NEAR(g[k][l], f.bilinear(x[k], x[l]), 1e-12) << "k=" << k << " l=" << l; + if (k != l && std::abs(g[k][l]) > 1e-6) { any_offdiag = true; } + } + } + // guard against a degenerate test case that would pass with G assembled as zero off-diagonal + EXPECT_TRUE(any_offdiag); +} + +/// G is symmetric by construction; assert it so a future refactor cannot lose it silently. +TEST(GradMatrixDegenerate, IsSymmetric) +{ + const QuadraticForm f; + const std::vector> g = route_a2(f, subspace_basis()); + for (size_t k = 0; k < g.size(); ++k) + { + for (size_t l = 0; l < g.size(); ++l) { EXPECT_DOUBLE_EQ(g[k][l], g[l][k]); } + } +} + +/// Section 5.4.5(ii), the strongest check: rotating the basis of the degenerate subspace must +/// rotate G, $G'=U^\top G U$. This is what would fail if the pipeline were not a pure quadratic +/// form, and it needs no finite differences. +TEST(GradMatrixDegenerate, RotationCovariance) +{ + const QuadraticForm f; + const std::vector> x = subspace_basis(); + const std::vector> g = route_a2(f, x); + const int d = static_cast(x.size()); + + // an orthogonal U mixing all three members (rotation by theta in (0,1), then by phi in (1,2)) + const double ct = std::cos(0.7); + const double st = std::sin(0.7); + const double cp = std::cos(0.4); + const double sp = std::sin(0.4); + std::vector> u(d, std::vector(d, 0.0)); + u[0][0] = ct; + u[0][1] = -st * cp; + u[0][2] = st * sp; + u[1][0] = st; + u[1][1] = ct * cp; + u[1][2] = -ct * sp; + u[2][0] = 0.0; + u[2][1] = sp; + u[2][2] = cp; + + // rotated basis $X'_k=\sum_l U_{lk}X_l$ + std::vector> xr(d, std::vector(nov, 0.0)); + for (int k = 0; k < d; ++k) + { + for (int l = 0; l < d; ++l) + { + for (int i = 0; i < nov; ++i) { xr[k][i] += u[l][k] * x[l][i]; } + } + } + const std::vector> gr = route_a2(f, xr); + + for (int k = 0; k < d; ++k) + { + for (int l = 0; l < d; ++l) + { + double expect = 0.0; + for (int p = 0; p < d; ++p) + { + for (int q = 0; q < d; ++q) { expect += u[p][k] * g[p][q] * u[q][l]; } + } + EXPECT_NEAR(gr[k][l], expect, 1e-12) << "k=" << k << " l=" << l; + } + } +} + +/// The trace is basis-invariant while the individual diagonal elements are not -- the reason the +/// validation tables compare group means (section 1.1). +TEST(GradMatrixDegenerate, TraceIsBasisInvariantButDiagonalIsNot) +{ + const QuadraticForm f; + const std::vector> x = subspace_basis(); + // swap-and-mix the last two members only + std::vector> xr = x; + const double r = 1.0 / std::sqrt(2.0); + for (int i = 0; i < nov; ++i) + { + xr[1][i] = r * (x[1][i] + x[2][i]); + xr[2][i] = r * (x[1][i] - x[2][i]); + } + const std::vector> g = route_a2(f, x); + const std::vector> gr = route_a2(f, xr); + + double tr = 0.0; + double trr = 0.0; + for (size_t k = 0; k < g.size(); ++k) + { + tr += g[k][k]; + trr += gr[k][k]; + } + EXPECT_NEAR(tr, trr, 1e-12); + // and the mixing really did change the individual entries + EXPECT_GT(std::abs(g[1][1] - gr[1][1]), 1e-6); +} + +TEST(GradMatrixDegenerate, CombineNormalizedKeepsNorm) +{ + const std::vector a = { 1.0, 0.0, 0.0, 0.0, 0.0 }; + const std::vector b = { 0.0, 1.0, 0.0, 0.0, 0.0 }; + std::vector p(nov); + LR::combine_normalized(a.data(), b.data(), nov, p.data()); + double n = 0.0; + for (int i = 0; i < nov; ++i) { n += p[i] * p[i]; } + EXPECT_NEAR(n, 1.0, 1e-14); + EXPECT_NEAR(p[0], 1.0 / std::sqrt(2.0), 1e-14); + EXPECT_NEAR(p[1], 1.0 / std::sqrt(2.0), 1e-14); +} + +TEST(GradMatrixDegenerate, PairOrder) +{ + const std::vector> p = LR::degenerate_pairs(3); + ASSERT_EQ(p.size(), 3u); + EXPECT_EQ(p[0], std::make_pair(0, 1)); + EXPECT_EQ(p[1], std::make_pair(0, 2)); + EXPECT_EQ(p[2], std::make_pair(1, 2)); + EXPECT_TRUE(LR::degenerate_pairs(1).empty()); + EXPECT_TRUE(LR::degenerate_pairs(0).empty()); + // d + d(d-1)/2 = d(d+1)/2 evaluations in total, the number of independent components + for (int d = 1; d < 8; ++d) + { + EXPECT_EQ(static_cast(LR::degenerate_pairs(d).size()) + d, d * (d + 1) / 2); + } +} + +TEST(GradMatrixDegenerate, GroupingCollectsMultiplets) +{ + // a triplet, then a singlet, then a doublet + const std::vector omega = { 0.5000000, 0.5000001, 0.5000002, 0.9, 1.30, 1.3000005 }; + const std::vector> g = LR::group_degenerate_states(omega, 2e-3); + ASSERT_EQ(g.size(), 3u); + EXPECT_EQ(g[0], std::vector({ 0, 1, 2 })); + EXPECT_EQ(g[1], std::vector({ 3 })); + EXPECT_EQ(g[2], std::vector({ 4, 5 })); +} + +TEST(GradMatrixDegenerate, GroupingSortsAndHandlesEdgeCases) +{ + // unsorted input must still group correctly, and the returned indices are into `omega` + const std::vector omega = { 1.3, 0.5, 1.3000005, 0.5000001 }; + const std::vector> g = LR::group_degenerate_states(omega, 2e-3); + ASSERT_EQ(g.size(), 2u); + EXPECT_EQ(g[0], std::vector({ 1, 3 })); + EXPECT_EQ(g[1], std::vector({ 0, 2 })); + + // thr <= 0 disables grouping entirely (the default: current per-state behaviour) + const std::vector> none = LR::group_degenerate_states(omega, 0.0); + EXPECT_EQ(none.size(), omega.size()); + for (const std::vector& grp : none) { EXPECT_EQ(grp.size(), 1u); } + + EXPECT_TRUE(LR::group_degenerate_states(std::vector(), 2e-3).empty()); +} + +/// Anchoring to the group's first member, not the predecessor: a ladder of steps each below `thr` +/// must NOT chain into one group whose total spread exceeds it. +TEST(GradMatrixDegenerate, GroupingDoesNotChain) +{ + const double thr = 1e-3; + std::vector omega; + for (int i = 0; i < 6; ++i) { omega.push_back(0.5 + i * 0.9e-3); } + const std::vector> g = LR::group_degenerate_states(omega, thr); + for (const std::vector& grp : g) + { + const double lo = omega[grp.front()]; + const double hi = omega[grp.back()]; + EXPECT_LT(hi - lo, thr); + } + EXPECT_GT(g.size(), 1u); +} + +/// The template has to work on the type the driver actually uses. +TEST(GradMatrixDegenerate, AssemblesModuleBaseMatrix) +{ + constexpr int nat = 2; + std::vector diag(2, ModuleBase::matrix(nat, 3)); + diag[0](0, 0) = 1.0; + diag[1](0, 0) = 3.0; + std::vector plus(1, ModuleBase::matrix(nat, 3)); + plus[0](0, 0) = 5.0; // -> off-diagonal 5 - (1+3)/2 = 3 + const std::vector> g + = LR::assemble_grad_matrix(diag, plus, LR::degenerate_pairs(2)); + ASSERT_EQ(g.size(), 2u); + EXPECT_DOUBLE_EQ(g[0][0](0, 0), 1.0); + EXPECT_DOUBLE_EQ(g[1][1](0, 0), 3.0); + EXPECT_DOUBLE_EQ(g[0][1](0, 0), 3.0); + EXPECT_DOUBLE_EQ(g[1][0](0, 0), 3.0); +} diff --git a/source/source_lcao/module_lr/test/test_grad_jt.cpp b/source/source_lcao/module_lr/test/test_grad_jt.cpp new file mode 100644 index 00000000000..0e70128fd13 --- /dev/null +++ b/source/source_lcao/module_lr/test/test_grad_jt.cpp @@ -0,0 +1,253 @@ +#include +#include +#include +#include "../grad_degen.h" + +// ----------------------------- the Jahn-Teller direction search ----------------------------- + +namespace +{ + /// A fixed, deliberately unsymmetric G tensor: `ncoord` blocks of d x d, symmetric in (k, l). + std::vector sample_gflat(const int ncoord, const int d, const unsigned seed) + { + std::vector g(static_cast(ncoord) * d * d, 0.0); + unsigned x = seed; + const auto next = [&x]() { + x = x * 1664525u + 1013904223u; // deterministic; a test must not depend on rand() + return static_cast(static_cast(x >> 8) % 2000 - 1000) / 500.0; + }; + for (int a = 0; a < ncoord; ++a) + { + double* const b = g.data() + static_cast(a) * d * d; + for (int k = 0; k < d; ++k) + { + for (int l = k; l < d; ++l) + { + const double v = next(); + b[k * d + l] = v; + b[l * d + k] = v; + } + } + } + return g; + } + + double branch_grad_norm(const std::vector& g, const int ncoord, const int d, + const std::vector& v) + { + double s = 0.0; + for (int a = 0; a < ncoord; ++a) + { + const double* const b = g.data() + static_cast(a) * d * d; + double q = 0.0; + for (int k = 0; k < d; ++k) + { + for (int l = 0; l < d; ++l) { q += v[k] * b[k * d + l] * v[l]; } + } + s += q * q; + } + return std::sqrt(s); + } + + /// lambda_min of sum_a u_a G^(a), by brute force over the 2x2 / 3x3 block. + double lambda_min_along(const std::vector& g, const int ncoord, const int d, + const std::vector& u) + { + std::vector m(static_cast(d) * d, 0.0); + for (int a = 0; a < ncoord; ++a) + { + const double* const b = g.data() + static_cast(a) * d * d; + for (int i = 0; i < d * d; ++i) { m[i] += u[a] * b[i]; } + } + // smallest Rayleigh quotient, sampled densely; d is 2 or 3 in these tests + double lo = 1e300; + const int n = 2000; + if (d == 2) + { + for (int i = 0; i <= n; ++i) + { + const double t = M_PI * i / n; + const double v[2] = { std::cos(t), std::sin(t) }; + const double r = v[0] * v[0] * m[0] + 2.0 * v[0] * v[1] * m[1] + v[1] * v[1] * m[3]; + lo = std::min(lo, r); + } + } + else + { + for (int i = 0; i <= 200; ++i) + { + for (int j = 0; j <= 400; ++j) + { + const double th = M_PI * i / 200; + const double ph = 2.0 * M_PI * j / 400; + const double v[3] = { std::sin(th) * std::cos(ph), std::sin(th) * std::sin(ph), + std::cos(th) }; + double r = 0.0; + for (int p = 0; p < 3; ++p) + { + for (int q = 0; q < 3; ++q) { r += v[p] * m[p * 3 + q] * v[q]; } + } + lo = std::min(lo, r); + } + } + } + return lo; + } +} + +/// The identity the whole search rests on: +/// min_{|u|=1} lambda_min(M(u)) = -max_{|v|=1} |q(v)|. +/// Checked against a brute-force scan of the left-hand side over directions built from the +/// returned optimum, so a wrong reduction cannot pass. +TEST(JTDirection, ReductionIdentityHolds) +{ + for (const int d : { 2, 3 }) + { + const int ncoord = 6; + const std::vector g = sample_gflat(ncoord, d, 7u + d); + const LR::JTDirection jt = LR::find_jt_direction(g, ncoord, d); + ASSERT_EQ(static_cast(jt.displacement.size()), ncoord) << "d=" << d; + ASSERT_EQ(static_cast(jt.mixing.size()), d); + EXPECT_GT(jt.slope, 0.0); + + // |q(v*)| must equal the reported slope + EXPECT_NEAR(branch_grad_norm(g, ncoord, d, jt.mixing), jt.slope, 1e-9 * jt.slope); + // the displacement is the normalized branch gradient + double n = 0.0; + for (int a = 0; a < ncoord; ++a) { n += jt.displacement[a] * jt.displacement[a]; } + EXPECT_NEAR(std::sqrt(n), 1.0, 1e-12); + // and along MINUS it the smallest eigenvalue is -slope: the two sides of the identity + std::vector u(ncoord); + for (int a = 0; a < ncoord; ++a) { u[a] = -jt.displacement[a]; } + EXPECT_NEAR(lambda_min_along(g, ncoord, d, u), -jt.slope, 1e-4 * jt.slope) << "d=" << d; + } +} + +/// Optimality: no other unit mixing may give a larger branch-gradient norm. Brute-forced over the +/// subspace sphere, which is the independent check that the alternating iteration converged to the +/// global optimum and not merely to a stationary point. +TEST(JTDirection, IsGlobalOverTheSubspace) +{ + const int ncoord = 6; + { + const int d = 2; + const std::vector g = sample_gflat(ncoord, d, 9u); + const LR::JTDirection jt = LR::find_jt_direction(g, ncoord, d); + double best = 0.0; + for (int i = 0; i <= 4000; ++i) + { + const double t = M_PI * i / 4000; + best = std::max(best, branch_grad_norm(g, ncoord, d, { std::cos(t), std::sin(t) })); + } + EXPECT_NEAR(jt.slope, best, 1e-6 * best); + } + { + const int d = 3; + const std::vector g = sample_gflat(ncoord, d, 11u); + const LR::JTDirection jt = LR::find_jt_direction(g, ncoord, d); + double best = 0.0; + for (int i = 0; i <= 300; ++i) + { + for (int j = 0; j <= 600; ++j) + { + const double th = M_PI * i / 300; + const double ph = 2.0 * M_PI * j / 600; + best = std::max(best, branch_grad_norm(g, ncoord, d, + { std::sin(th) * std::cos(ph), std::sin(th) * std::sin(ph), std::cos(th) })); + } + } + EXPECT_NEAR(jt.slope, best, 1e-4 * best); + } +} + +/// A multiplet whose every G block is a multiple of the identity has no Jahn-Teller direction to +/// find: all branches share one gradient, so |q(v)| is the same for every v and nothing splits. +/// The search must still return that common direction rather than something arbitrary. +TEST(JTDirection, DegenerateCaseGivesTheCommonGradient) +{ + const int ncoord = 6; + const int d = 2; + std::vector g(static_cast(ncoord) * d * d, 0.0); + const double comm[6] = { 0.3, -0.7, 1.1, 0.0, 0.5, -0.2 }; + for (int a = 0; a < ncoord; ++a) + { + g[static_cast(a) * 4 + 0] = comm[a]; + g[static_cast(a) * 4 + 3] = comm[a]; + } + const LR::JTDirection jt = LR::find_jt_direction(g, ncoord, d); + double n = 0.0; + for (int a = 0; a < ncoord; ++a) { n += comm[a] * comm[a]; } + n = std::sqrt(n); + EXPECT_NEAR(jt.slope, n, 1e-10); + for (int a = 0; a < ncoord; ++a) { EXPECT_NEAR(jt.displacement[a], comm[a] / n, 1e-10); } + // every start reaches the same value, since the objective is constant on the sphere + EXPECT_GE(jt.restarts_agreeing, 2); +} + +/// Homogeneity: scaling G scales the slope and leaves the direction alone. +TEST(JTDirection, ScalesLinearly) +{ + const int ncoord = 6; + const int d = 3; + const std::vector g = sample_gflat(ncoord, d, 13u); + std::vector g3 = g; + for (size_t i = 0; i < g3.size(); ++i) { g3[i] *= 3.0; } + const LR::JTDirection a = LR::find_jt_direction(g, ncoord, d); + const LR::JTDirection b = LR::find_jt_direction(g3, ncoord, d); + EXPECT_NEAR(b.slope, 3.0 * a.slope, 1e-8 * a.slope); + for (int i = 0; i < ncoord; ++i) { EXPECT_NEAR(b.displacement[i], a.displacement[i], 1e-8); } +} + +/// The symmetric / Jahn-Teller split must reconstruct the branch gradient, and the symmetric part +/// must be the multiplet average -- independent of which branch was chosen. +TEST(JTDirection, SymmetricSplitReconstructs) +{ + const int ncoord = 6; + const int d = 3; + const std::vector g = sample_gflat(ncoord, d, 17u); + const LR::JTDirection jt = LR::find_jt_direction(g, ncoord, d); + std::vector jt_part; + const std::vector sym = LR::split_symmetric_part(g, ncoord, d, jt.mixing, jt_part); + for (int a = 0; a < ncoord; ++a) + { + const double* const b = g.data() + static_cast(a) * d * d; + double q = 0.0; + for (int k = 0; k < d; ++k) + { + for (int l = 0; l < d; ++l) { q += jt.mixing[k] * b[k * d + l] * jt.mixing[l]; } + } + EXPECT_NEAR(sym[a] + jt_part[a], q, 1e-12); + double tr = 0.0; + for (int k = 0; k < d; ++k) { tr += b[k * d + k]; } + EXPECT_NEAR(sym[a], tr / d, 1e-12); + } + // a different mixing keeps the same symmetric part + std::vector other(d, 0.0); + other[0] = 1.0; + std::vector jt2; + const std::vector sym2 = LR::split_symmetric_part(g, ncoord, d, other, jt2); + for (int a = 0; a < ncoord; ++a) { EXPECT_NEAR(sym2[a], sym[a], 1e-14); } +} + +TEST(JTDirection, ZeroGradientReturnsAValidStationaryBranch) +{ + for (const int d : {2, 3}) + { + const int ncoord = 3; + const size_t size = static_cast(ncoord) * d * d; + const std::vector zero(size, 0.0); + const LR::JTDirection jt = LR::find_jt_direction(zero, ncoord, d); + ASSERT_EQ(jt.mixing.size(), d); + ASSERT_EQ(jt.displacement.size(), ncoord); + EXPECT_DOUBLE_EQ(jt.slope, 0.0); + EXPECT_DOUBLE_EQ(jt.mixing[0], 1.0); + double norm_squared = 0.0; + for (const double value : jt.mixing) { norm_squared += value * value; } + EXPECT_DOUBLE_EQ(norm_squared, 1.0); + std::vector jt_part; + const std::vector sym = LR::split_symmetric_part(zero, ncoord, d, jt.mixing, jt_part); + EXPECT_EQ(sym, (std::vector(ncoord, 0.0))); + EXPECT_EQ(jt_part, sym); + EXPECT_EQ(jt.displacement, sym); + } +} diff --git a/source/source_lcao/module_lr/test/test_lr_amp.cpp b/source/source_lcao/module_lr/test/test_lr_amp.cpp new file mode 100644 index 00000000000..4ba15aa52e8 --- /dev/null +++ b/source/source_lcao/module_lr/test/test_lr_amp.cpp @@ -0,0 +1,146 @@ +#include +#include "../lr_amp.h" +#include "source_base/parallel_global.h" +#include +#include + +namespace { int test_rank = 0; } + +TEST(GradientAmplitudes, InitialRootSeedsTheReference) +{ + const double roots[] = {1.0, 0.0, 0.0, 1.0}; + int target = -1; + std::vector previous; + std::ostringstream log; + LR::follow_root(roots, 2, 2, 1, target, previous, log); + EXPECT_EQ(target, 1); + EXPECT_EQ(previous, (std::vector{0.0, 1.0})); +} + +TEST(GradientAmplitudes, FollowsACrossingIndependentOfSign) +{ + const double roots[] = {0.0, 1.0, -1.0, 0.0}; + int target = 0; + std::vector previous{1.0, 0.0}; + std::ostringstream log; + LR::follow_root(roots, 2, 2, 0, target, previous, log); + EXPECT_EQ(target, 1); + EXPECT_EQ(previous, (std::vector{-1.0, 0.0})); + EXPECT_NE(log.str().find("root moved"), std::string::npos); +} + +TEST(GradientAmplitudes, ComplexPhaseDoesNotChangeTheFollowedRoot) +{ + using Complex = std::complex; + const Complex roots[] = {Complex(0.0), Complex(1.0), Complex(0.0, 1.0), Complex(0.0)}; + int target = 0; + std::vector previous{Complex(1.0), Complex(0.0)}; + std::ostringstream log; + LR::follow_root(roots, 2, 2, 0, target, previous, log); + EXPECT_EQ(target, 1); + EXPECT_EQ(previous[0], Complex(0.0, 1.0)); +} + +TEST(GradientAmplitudes, OpenShellDensitiesReadTheEntireDownSpinKBlock) +{ + // Two roots, two k points; up has two pairs per k and down has one. + const double roots[] = {10, 11, 12, 13, 20, 21, 30, 31, 32, 33, 40, 41}; + const int first_down = LR::electron_hole_offset(0, 2, 2, 1, 1, true); + const int second_down = LR::electron_hole_offset(1, 2, 2, 1, 1, true); + const int second_up = LR::electron_hole_offset(1, 2, 2, 1, 0, true); + EXPECT_DOUBLE_EQ(roots[first_down], 20); + EXPECT_DOUBLE_EQ(roots[first_down + 1], 21); + EXPECT_DOUBLE_EQ(roots[second_down], 40); + EXPECT_DOUBLE_EQ(roots[second_down + 1], 41); + EXPECT_DOUBLE_EQ(roots[second_up], 30); + const int closed = LR::electron_hole_offset(1, 2, 2, 1, 1, false); + EXPECT_EQ(closed, 2); + const int gamma = LR::electron_hole_offset(1, 1, 2, 1, 1, true); + EXPECT_EQ(gamma, 5); +} + +TEST(GradientAmplitudes, FollowsTheSelectedJTBranchAfterSplitting) +{ + const double degenerate_roots[] = {1.0, 0.0, 0.0, 1.0}; + const std::vector group{0, 1}; + const std::vector mixing{0.0, 1.0}; + std::vector previous{1.0, 0.0}; + LR::save_mixed_root(degenerate_roots, 2, group, mixing, previous); + // The selected second component becomes the first root at the next geometry. + const double split_roots[] = {0.0, 1.0, 1.0, 0.0}; + int target = 0; + std::ostringstream log; + LR::follow_root(split_roots, 2, 2, 0, target, previous, log); + EXPECT_EQ(target, 0); + EXPECT_EQ(previous, (std::vector{0.0, 1.0})); +} + +TEST(GradientAmplitudes, MixedReferenceHasGlobalUnitNorm) +{ + using Complex = std::complex; + const Complex roots[] = {Complex(0, 1), Complex(0), Complex(0), Complex(1)}; + const std::vector group{0, 1}; + const std::vector mixing{0.6, 0.8}; + std::vector previous; + LR::save_mixed_root(roots, 2, group, mixing, previous); + ASSERT_EQ(previous.size(), 2); + double norm_squared = std::norm(previous[0]) + std::norm(previous[1]); + Parallel_Reduce::reduce_all(norm_squared); + EXPECT_NEAR(norm_squared, 1.0, 1e-14); + EXPECT_NEAR(previous[0].imag() / previous[1].real(), 0.75, 1e-14); +} + +TEST(GradientAmplitudes, KeepsJTMixtureAfterIndividualOrbitalSignChange) +{ + const double old_roots[] = {1, 0, 0, 1}; + const std::vector group{0, 1}; + const std::vector mixing{0.6, 0.8}; + std::vector previous; + LR::save_mixed_root(old_roots, 2, group, mixing, previous); + // The sign-flipped second virtual orbital changes that reference component. + // The full projection is tested against its formula in test_root_ovlp.cpp. + previous[1] = -previous[1]; + const double current_roots[] = {0.6, -0.8, 0.8, 0.6}; + int target = 0; + std::ostringstream log; + LR::follow_root(current_roots, 2, 2, 0, target, previous, log); + EXPECT_EQ(target, 0); +} + +TEST(GradientAmplitudes, EmptyLocalRanksParticipateInRootSelection) +{ + const int local_size = test_rank == 0 ? 2 : 0; + const double first[] = {1, 0, 0, 1}; + const double second[] = {0, 1, 1, 0}; + int target = 0; + std::vector previous; + std::ostringstream log; + LR::follow_root(first, local_size, 2, 0, target, previous, log); + LR::follow_root(second, local_size, 2, 0, target, previous, log); + EXPECT_EQ(target, 1); + EXPECT_EQ(previous.size(), static_cast(local_size)); +} + +int main(int argc, char** argv) +{ + int processes = 1; + int threads = 1; + int rank = 0; + Parallel_Global::read_pal_param(argc, argv, processes, threads, rank); + test_rank = rank; +#ifdef __MPI + // This focused test has no pool/grid setup; the cleanup wrapper needs null handles. + POOL_WORLD = MPI_COMM_NULL; + KP_WORLD = MPI_COMM_NULL; + INT_BGROUP = MPI_COMM_NULL; + BP_WORLD = MPI_COMM_NULL; + GRID_WORLD = MPI_COMM_NULL; + DIAG_WORLD = MPI_COMM_NULL; +#endif + ::testing::InitGoogleTest(&argc, argv); + const int result = RUN_ALL_TESTS(); +#ifdef __MPI + Parallel_Global::finalize_mpi(); +#endif + return result; +} diff --git a/source/source_lcao/module_lr/test/test_root_ovlp.cpp b/source/source_lcao/module_lr/test/test_root_ovlp.cpp new file mode 100644 index 00000000000..8fa1fdc6b86 --- /dev/null +++ b/source/source_lcao/module_lr/test/test_root_ovlp.cpp @@ -0,0 +1,440 @@ +#include "source_base/parallel_2d.h" +#include "source_basis/module_ao/parallel_orbitals.h" +#include "source_basis/module_nao/two_center_integrator.h" +#include "source_cell/unitcell.h" +#include +#include "source_base/parallel_global.h" +#include "source_base/parallel_reduce.h" +#include +#include +#include "source_lcao/module_lr/root_ovlp.h" +#include "source_base/matrix3.h" +#include +#include + +namespace +{ +template +std::vector transform_for_test(const std::vector& old, const std::vector& ao, + const std::vector& current, int naos, int bands); +template +std::vector project(const std::vector& previous, const std::vector& overlap, int no, int nv); +} + +TEST(RootOverlap, IncludesPairsAbsentFromBothSameGeometryLists) +{ + // Old centers 0, 1.1; new centers 0.2, 1.3. Both intra-step distances + // exceed 1, but old second/new first are only 0.9 apart. + const ModuleBase::Matrix3 lattice(10, 0, 0, 0, 10, 0, 0, 0, 10); + const ModuleBase::Vector3 displacement(-0.9, 0, 0); + const auto images = LR::cross_image_indices(displacement, lattice, 1.0); + ASSERT_EQ(images.size(), 1u); + const auto shifted = displacement + images[0] * lattice; + EXPECT_NEAR(shifted.x, -0.9, 1e-14); +} + +TEST(RootOverlap, IncludesImagesAfterLargeTranslationsInSkewCell) +{ + const ModuleBase::Matrix3 lattice(2, 0, 0, 1.7, 1.1, 0, 0.4, 0.3, 2.2); + const ModuleBase::Vector3 translation(4, -3, 2); + const ModuleBase::Vector3 residual(0.2, 0.1, 0.3); + const ModuleBase::Vector3 displacement = translation * lattice + residual; + const auto images = LR::cross_image_indices(displacement, lattice, 1.6); + std::vector actual; + std::vector expected; + for (const auto& image : images) + { + const auto shifted = displacement + image * lattice; + const double distance = shifted.norm(); + actual.push_back(distance); + } + for (int x = -8; x <= 8; ++x) + { + for (int y = -8; y <= 8; ++y) + { + for (int z = -8; z <= 8; ++z) + { + const ModuleBase::Vector3 image(x, y, z); + const auto shifted = displacement + image * lattice; + if (shifted.norm() < 1.6) { expected.push_back(shifted.norm()); } + } + } + } + std::sort(actual.begin(), actual.end()); + std::sort(expected.begin(), expected.end()); + ASSERT_EQ(actual.size(), expected.size()); + for (std::size_t i = 0; i < actual.size(); ++i) { EXPECT_NEAR(actual[i], expected[i], 1e-13); } +} + +TEST(RootOverlap, RetainsIntegerTranslationsForFourierPhases) +{ + const ModuleBase::Matrix3 lattice(2, 0, 0, 0, 3, 0, 0, 0, 4); + const ModuleBase::Vector3 displacement(2.2, 0.0, 0.0); + const auto indices = LR::cross_image_indices(displacement, lattice, 0.5); + ASSERT_EQ(indices.size(), 1u); + EXPECT_EQ(indices[0].x, -1); + EXPECT_EQ(indices[0].y, 0); + EXPECT_EQ(indices[0].z, 0); + const auto shifted = displacement + indices[0] * lattice; + EXPECT_NEAR(shifted.x, 0.2, 1e-14); +} + +TEST(RootOverlap, RemovesOrbitalSignGaugeFromRootSelection) +{ + const double a = 1.0 / std::sqrt(2.0); + const std::vector previous{a, a}; + const std::vector overlap{1, 0, 0, 0, 1, 0, 0, 0, -1}; + const auto projected = project(previous, overlap, 1, 2); + EXPECT_DOUBLE_EQ(projected[0], a); + EXPECT_DOUBLE_EQ(projected[1], -a); + const double physical_root = projected[0] * a - projected[1] * a; + const double other_root = projected[0] * a + projected[1] * a; + EXPECT_NEAR(physical_root, 1.0, 1e-14); + EXPECT_NEAR(other_root, 0.0, 1e-14); +} + +TEST(RootOverlap, PreservesComplexConjugationAndBothOrbitalRotations) +{ + using C = std::complex; + const std::vector previous{C(1), C(0), C(0), C(0)}; + // Swap occupied orbitals with a phase and independently rotate virtuals. + const double a = 1.0 / std::sqrt(2.0); + const std::vector overlap{C(0), C(0, 1), C(0), C(0), + C(1), C(0), C(0), C(0), + C(0), C(0), C(a), C(0, a), + C(0), C(0), C(0, a), C(a)}; + const auto projected = project(previous, overlap, 2, 2); + EXPECT_EQ(projected[0], C(0)); + EXPECT_EQ(projected[1], C(0)); + EXPECT_EQ(projected[2], C(0, a)); + EXPECT_EQ(projected[3], C(a)); +} + +TEST(RootOverlap, RetainsProjectionLossInsteadOfRenormalizing) +{ + const std::vector previous{1}; + const std::vector overlap{0.5, 0.7, -0.3, 0.8}; + const auto projected = project(previous, overlap, 1, 1); + EXPECT_DOUBLE_EQ(projected[0], 0.4); +} + + +TEST(RootOverlap, UsesOrderedNonsymmetricAOOverlap) +{ + const std::vector coefficients{1, 0, 0, 1}; + const std::vector ao{1, 0.2, -0.4, 0.9}; + const auto mo = transform_for_test(coefficients, ao, coefficients, 2, 2); + EXPECT_EQ(mo, ao); +} + +TEST(RootOverlap, MOTransformConjugatesOldOrbitals) +{ + using C = std::complex; + const std::vector old{C(0, 1), C(0), C(0), C(1)}; + const std::vector current{C(1), C(0), C(0), C(0, 1)}; + const std::vector ao{1, 0.2, -0.4, 0.9}; + const auto mo = transform_for_test(old, ao, current, 2, 2); + EXPECT_EQ(mo[0], C(0, -1)); + EXPECT_EQ(mo[1], C(0.2)); + EXPECT_EQ(mo[2], C(-0.4)); + EXPECT_EQ(mo[3], C(0, 0.9)); +} + + +namespace +{ +void setup(Parallel_2D& target, int rows, int cols, const Parallel_2D& grid) +{ +#ifdef __MPI + target.set(rows, cols, 1, grid.blacs_ctxt); +#else + target.set_serial(rows, cols); +#endif +} + +template +std::vector scatter(const std::vector& full, const Parallel_2D& layout) +{ + std::vector local(layout.get_local_size(), T(0)); + const int columns = layout.get_global_col_size(); + for (int col = 0; col < layout.get_col_size(); ++col) + { + const int global_col = layout.local2global_col(col); + for (int row = 0; row < layout.get_row_size(); ++row) + { + const int global_row = layout.local2global_row(row); + local[col * layout.get_row_size() + row] = full[global_row * columns + global_col]; + } + } + return local; +} + +// Independent test oracles: direct formula contractions, with no BLAS/PBLAS calls. +template T conjugate(T value) { return value; } +template <> std::complex conjugate(std::complex value) { return std::conj(value); } + +template +std::vector reference_mo(const std::vector& old, const std::vector& ao, + const std::vector& current, int naos, int bands) +{ + std::vector result(bands * bands, T(0)); + for (int i = 0; i < bands; ++i) + { + for (int j = 0; j < bands; ++j) + { + for (int mu = 0; mu < naos; ++mu) + { + for (int nu = 0; nu < naos; ++nu) + { + result[i * bands + j] += conjugate(old[i * naos + mu]) + * ao[mu * naos + nu] * current[j * naos + nu]; + } + } + } + } + return result; +} + +template +std::vector reference_projection(const std::vector& previous, + const std::vector& overlap, int no, int nv) +{ + const int bands = no + nv; + std::vector result(no * nv, T(0)); + for (int j = 0; j < no; ++j) + { + for (int b = 0; b < nv; ++b) + { + for (int i = 0; i < no; ++i) + { + for (int a = 0; a < nv; ++a) + { + result[j * nv + b] += overlap[i * bands + j] * previous[i * nv + a] + * conjugate(overlap[(no + a) * bands + no + b]); + } + } + } + } + return result; +} + +template +std::vector gather(const std::vector& local, const Parallel_2D& layout) +{ + const int columns = layout.get_global_col_size(); + const int count = layout.get_global_row_size() * columns; + std::vector result(count, T(0)); + EXPECT_EQ(local.size(), layout.get_local_size()); + for (int col = 0; col < layout.get_col_size(); ++col) + { + const int global_col = layout.local2global_col(col); + for (int row = 0; row < layout.get_row_size(); ++row) + { + const int global_row = layout.local2global_row(row); + result[global_row * columns + global_col] = local[col * layout.get_row_size() + row]; + } + } + Parallel_Reduce::reduce_all(result.data(), count); + return result; +} + +template +std::vector transform_for_test(const std::vector& old, const std::vector& ao, + const std::vector& current, int naos, int bands) +{ + Parallel_2D pc; +#ifdef __MPI + pc.init(naos, bands, 1, MPI_COMM_WORLD); +#else + pc.set_serial(naos, bands); +#endif + Parallel_2D ps; + Parallel_2D po; + setup(ps, naos, naos, pc); + setup(po, bands, bands, pc); + std::vector old_rows(old.size()); + std::vector current_rows(current.size()); + for (int i = 0; i < naos; ++i) + { + for (int band = 0; band < bands; ++band) + { + old_rows[i * bands + band] = old[band * naos + i]; + current_rows[i * bands + band] = current[band * naos + i]; + } + } + const std::vector typed_ao(ao.begin(), ao.end()); + const auto local_old = scatter(old_rows, pc); + const auto local_current = scatter(current_rows, pc); + const auto local_ao = scatter(typed_ao, ps); + const auto actual = LR::mo_overlap_dist(local_old, local_ao, local_current, pc, ps, po); + return gather(actual, po); +} + +template +std::vector project(const std::vector& previous, const std::vector& overlap, int no, int nv) +{ + const int bands = no + nv; + Parallel_2D po; +#ifdef __MPI + po.init(bands, bands, 1, MPI_COMM_WORLD); +#else + po.set_serial(bands, bands); +#endif + Parallel_2D px; + setup(px, nv, no, po); + std::vector previous_rows(previous.size()); + for (int i = 0; i < no; ++i) + { + for (int a = 0; a < nv; ++a) { previous_rows[a * no + i] = previous[i * nv + a]; } + } + const auto local_previous = scatter(previous_rows, px); + const auto local_overlap = scatter(overlap, po); + const auto actual = LR::project_reference_dist(local_previous, local_overlap, po, px, no, nv); + const auto rows = gather(actual, px); + std::vector result(previous.size()); + for (int i = 0; i < no; ++i) + { + for (int a = 0; a < nv; ++a) { result[i * nv + a] = rows[a * no + i]; } + } + return result; +} + +template +void check_mo(int naos, int bands, T phase) +{ + std::vector ao(naos * naos); + std::vector old(naos * bands); + std::vector current(naos * bands); + for (int i = 0; i < naos; ++i) + { + for (int j = 0; j < naos; ++j) { ao[i * naos + j] = (2 * i - j + 1) / 7.0; } + for (int band = 0; band < bands; ++band) + { + old[band * naos + i] = phase * ((i + band + 1) / 11.0); + current[band * naos + i] = T((i - 2 * band + 2) / 13.0); + } + } + const auto actual = transform_for_test(old, ao, current, naos, bands); + const auto expected = reference_mo(old, ao, current, naos, bands); + ASSERT_EQ(actual.size(), expected.size()); + for (std::size_t i = 0; i < actual.size(); ++i) + { + EXPECT_NEAR(std::abs(actual[i] - expected[i]), 0.0, 1e-13); + } +} +template +void check_projection(int no, int nv, T phase) +{ + const int bands = no + nv; + Parallel_2D po; +#ifdef __MPI + po.init(bands, bands, 1, MPI_COMM_WORLD); +#else + po.set_serial(bands, bands); +#endif + Parallel_2D px; + setup(px, nv, no, po); + std::vector overlap(bands * bands); + std::vector previous(no * nv); + for (int i = 0; i < bands; ++i) + { + for (int j = 0; j < bands; ++j) + { + overlap[i * bands + j] = T((i - 2 * j + 1) / 17.0) + phase * ((i + j + 1) / 19.0); + } + } + for (int i = 0; i < no * nv; ++i) { previous[i] = phase * ((i + 1) / 23.0); } + // Convert occupied-major X into row-major A=X^T before scattering into px. + std::vector previous_rows(previous.size()); + for (int i = 0; i < no; ++i) + { + for (int a = 0; a < nv; ++a) { previous_rows[a * no + i] = previous[i * nv + a]; } + } + const auto local_previous = scatter(previous_rows, px); + const auto local_overlap = scatter(overlap, po); + const auto actual = LR::project_reference_dist(local_previous, local_overlap, po, px, no, nv); + const auto expected = reference_projection(previous, overlap, no, nv); + std::vector expected_rows(expected.size()); + for (int i = 0; i < no; ++i) + { + for (int a = 0; a < nv; ++a) { expected_rows[a * no + i] = expected[i * nv + a]; } + } + const auto local_expected = scatter(expected_rows, px); + ASSERT_EQ(actual.size(), local_expected.size()); + for (std::size_t i = 0; i < actual.size(); ++i) + { + EXPECT_NEAR(std::abs(actual[i] - local_expected[i]), 0.0, 1e-13); + } +} + +} + +TEST(RootDistributed, RealNonsymmetricRectangularTransformation) { check_mo(5, 3, 1.0); } +TEST(RootDistributed, ComplexAdjointTransformation) +{ + const std::complex phase(0.6, 0.8); + check_mo(5, 3, phase); +} +TEST(RootDistributed, EmptyCoefficientBlocksParticipate) { check_mo(2, 1, 1.0); } + +TEST(RootDistributed, RealProjectionPreservesWindowLoss) { check_projection(2, 3, 0.4); } +TEST(RootDistributed, ComplexProjectionUsesVirtualAdjoint) +{ + const std::complex phase(0.3, 0.4); + check_projection(2, 3, phase); +} +TEST(RootDistributed, EmptyReferenceBlocksParticipate) { check_projection(1, 1, 0.4); } + +TEST(RootOverlap, CrossGeometryAOOverlapIsOrderedAndNonsymmetric) +{ + UnitCell cell; + std::unique_ptr atoms(new Atom); + cell.atoms = atoms.get(); + cell.ntype = 1; + cell.nat = 2; + cell.lat0 = 1.0; + cell.latvec = ModuleBase::Matrix3(10, 0, 0, 0, 10, 0, 0, 0, 10); + cell.iat2it = new int[2]{0, 0}; + cell.iat2ia = new int[2]{0, 1}; + atoms->nw = 1; + atoms->na = 2; + atoms->iw2l = {0}; + atoms->iw2n = {0}; + atoms->iw2m = {0}; + atoms->tau = {{0.2, 0.0, 0.0}, {1.3, 0.0, 0.0}}; + const std::vector offsets{0, 1}; + cell.set_iat2iwt_for_test(offsets, 1); + Parallel_Orbitals distribution; + distribution.set_serial(2, 2); + distribution.set_atomic_trace(cell.get_iat2iwt(), cell.nat, 2); + TwoCenterIntegrator integrator; + const std::vector> old_positions{{0.0, 0.0, 0.0}, {1.1, 0.0, 0.0}}; + const std::vector cutoff{0.5}; + const auto actual = LR::cross_ao_overlap(cell, old_positions, cutoff, integrator, distribution); + // Existing integral mock returns one: old second/new first is inside the cutoff, + // whereas old first/new second is outside. Output is column-major. + const std::vector expected{1.0, 1.0, 0.0, 1.0}; + EXPECT_EQ(actual, expected); +} + +int main(int argc, char** argv) +{ + int processes = 1; + int threads = 1; + int rank = 0; + Parallel_Global::read_pal_param(argc, argv, processes, threads, rank); +#ifdef __MPI + POOL_WORLD = MPI_COMM_NULL; + KP_WORLD = MPI_COMM_NULL; + INT_BGROUP = MPI_COMM_NULL; + BP_WORLD = MPI_COMM_NULL; + GRID_WORLD = MPI_COMM_NULL; + DIAG_WORLD = MPI_COMM_NULL; +#endif + ::testing::InitGoogleTest(&argc, argv); + const int result = RUN_ALL_TESTS(); +#ifdef __MPI + Parallel_Global::finalize_mpi(); +#endif + return result; +} diff --git a/source/source_lcao/module_lr/test/test_zeqlin.cpp b/source/source_lcao/module_lr/test/test_zeqlin.cpp new file mode 100644 index 00000000000..69b831ba715 --- /dev/null +++ b/source/source_lcao/module_lr/test/test_zeqlin.cpp @@ -0,0 +1,310 @@ +#include +#include "../zeq_solver.hpp" +#include "../cal_edm.h" +#include "../cal_edm.h" +#include "mpi.h" +#include "../zeqlin_solv.h" + +#include +#include +#include +#include + +#include "source_lcao/module_lr/utils/lr_util.h" + +// The distributed Z-vector solvers are checked against the replicated LAPACK solve +// (`LR_Util::lapack_linear_solver`) on matrices generated identically on every rank. +namespace +{ + // deterministic pseudo-random entry in [-1, 1], identical on every rank + double entry(const int i, const int j, const int salt) + { + return std::sin(0.7 * i + 1.3 * j + 0.37 * salt + 0.1 * i * j); + } + double conj_if_complex(const double v) { return v; } + std::complex conj_if_complex(const std::complex& v) { return std::conj(v); } + template T make_value(const double re, const double im); + template <> double make_value(const double re, const double im) { return re; } + template <> std::complex make_value>(const double re, const double im) + { + return std::complex(re, im); + } + + /// column-major n x n Hermitian matrix; positive definite when `shift` > n + template + std::vector hermitian_matrix(const int n, const double shift) + { + std::vector a(static_cast(n) * n); + for (int j = 0; j < n; ++j) + { + for (int i = 0; i <= j; ++i) + { + const T v = make_value(entry(i, j, 0), entry(i, j, 1)); + a[j * n + i] = v; + a[i * n + j] = conj_if_complex(v); + } + a[j * n + j] = T(entry(j, j, 2) + shift); + } + return a; + } + + template + std::vector rhs_matrix(const int n, const int nrhs) + { + std::vector b(static_cast(n) * nrhs); + for (int j = 0; j < nrhs; ++j) + { + for (int i = 0; i < n; ++i) { b[j * n + i] = make_value(entry(i, j, 3), entry(i, j, 4)); } + } + return b; + } + + template + std::vector to_local(const std::vector& full, const Parallel_2D& pv) + { + const int nrow = pv.get_global_row_size(); + std::vector loc(pv.get_local_size()); + for (int lc = 0; lc < pv.get_col_size(); ++lc) + { + for (int lr = 0; lr < pv.get_row_size(); ++lr) + { + loc[lc * pv.get_row_size() + lr] = full[pv.local2global_col(lc) * nrow + pv.local2global_row(lr)]; + } + } + return loc; + } + + void expect_near(const double a, const double b) { EXPECT_NEAR(a, b, 1e-10); } + void expect_near(const std::complex& a, const std::complex& b) + { + EXPECT_NEAR(a.real(), b.real(), 1e-10); + EXPECT_NEAR(a.imag(), b.imag(), 1e-10); + } + + /// solve the distributed system with `solver` and compare with LAPACK on `a_ref` + template + void check_against_lapack(void (*solver)(T*, T*, const Parallel_2D&, const Parallel_2D&), + const std::vector& a_in, const std::vector& a_ref, const int n, const int nrhs, const int nb) + { + const std::vector b = rhs_matrix(n, nrhs); + std::vector x_ref(b.size()); + LR_Util::lapack_linear_solver(a_ref.data(), x_ref.data(), b.data(), n, nrhs); + + Parallel_2D pa; + LR_Util::setup_2d_division(pa, nb, n, n); + Parallel_2D pb; + LR_Util::setup_2d_division(pb, nb, n, nrhs, pa.blacs_ctxt); + std::vector a_loc = to_local(a_in, pa); + std::vector x_loc = to_local(b, pb); + solver(a_loc.data(), x_loc.data(), pa, pb); + + std::vector x(b.size(), T(0)); + LR_Util::gather_2d_to_full(pb, x_loc.data(), x.data(), false, n, nrhs); + for (std::size_t i = 0; i < x.size(); ++i) { expect_near(x[i], x_ref[i]); } + } +} + +class ZeqLinearSolverTest : public testing::Test +{ + public: + // (n, nrhs, nb): n not a multiple of nb, a single rhs, and nb = 1 + const std::vector> sizes{ {37, 3, 4}, {20, 1, 3}, {9, 2, 1} }; +}; + +TEST_F(ZeqLinearSolverTest, ScalapackDouble) +{ + for (const auto& s : sizes) + { + const std::vector a = hermitian_matrix(s[0], s[0] + 1.0); + check_against_lapack(&LR::scalapack_linear_solver, a, a, s[0], s[1], s[2]); + } +} + +TEST_F(ZeqLinearSolverTest, ScalapackComplex) +{ + typedef std::complex C; + for (const auto& s : sizes) + { + const std::vector a = hermitian_matrix(s[0], s[0] + 1.0); + check_against_lapack(&LR::scalapack_linear_solver, a, a, s[0], s[1], s[2]); + } +} + +TEST_F(ZeqLinearSolverTest, ScalapackNonSymmetric) +{ + // LU must not assume symmetry: perturb with an antisymmetric part + for (const auto& s : sizes) + { + const int n = s[0]; + std::vector a = hermitian_matrix(n, n + 1.0); + for (int j = 0; j < n; ++j) + { + for (int i = 0; i < j; ++i) + { + a[j * n + i] += 0.3 * entry(i, j, 5); + a[i * n + j] -= 0.3 * entry(i, j, 5); + } + } + check_against_lapack(&LR::scalapack_linear_solver, a, a, n, s[1], s[2]); + } +} + +TEST_F(ZeqLinearSolverTest, ScalapackCholDouble) +{ + for (const auto& s : sizes) + { + const std::vector a = hermitian_matrix(s[0], s[0] + 1.0); + check_against_lapack(&LR::scalapack_cholesky_linear_solver, a, a, s[0], s[1], s[2]); + } +} + +TEST_F(ZeqLinearSolverTest, ScalapackCholComplex) +{ + typedef std::complex C; + for (const auto& s : sizes) + { + const std::vector a = hermitian_matrix(s[0], s[0] + 1.0); + check_against_lapack(&LR::scalapack_cholesky_linear_solver, a, a, s[0], s[1], s[2]); + } +} + +TEST_F(ZeqLinearSolverTest, ScalapackCholSolvesHermitianPart) +{ + // as for ELPA: a slightly non-symmetric input is solved as (A + A^T)/2 + for (const auto& s : sizes) + { + const int n = s[0]; + const std::vector a_sym = hermitian_matrix(n, n + 1.0); + std::vector a = a_sym; + for (int j = 0; j < n; ++j) + { + for (int i = 0; i < j; ++i) + { + a[j * n + i] += 1e-2 * entry(i, j, 5); + a[i * n + j] -= 1e-2 * entry(i, j, 5); + } + } + check_against_lapack(&LR::scalapack_cholesky_linear_solver, a, a_sym, n, s[1], s[2]); + } +} + +TEST_F(ZeqLinearSolverTest, ScalapackCholRejectsIndefinite) +{ + // every rank: p?potrf's INFO is global, so unlike ELPA this must throw everywhere, not hang + const int n = 20; + const int nb = 3; + const std::vector a = hermitian_matrix(n, -(n + 1.0)); // negative definite + Parallel_2D pa; + LR_Util::setup_2d_division(pa, nb, n, n); + Parallel_2D pb; + LR_Util::setup_2d_division(pb, nb, n, 1, pa.blacs_ctxt); + std::vector a_loc = to_local(a, pa); + std::vector b_loc = to_local(rhs_matrix(n, 1), pb); + EXPECT_THROW(LR::scalapack_cholesky_linear_solver(a_loc.data(), b_loc.data(), pa, pb), + std::runtime_error); +} + +#ifdef __ELPA +TEST_F(ZeqLinearSolverTest, ElpaDouble) +{ + for (const auto& s : sizes) + { + const std::vector a = hermitian_matrix(s[0], s[0] + 1.0); + check_against_lapack(&LR::elpa_linear_solver, a, a, s[0], s[1], s[2]); + } +} + +TEST_F(ZeqLinearSolverTest, ElpaComplex) +{ + typedef std::complex C; + for (const auto& s : sizes) + { + const std::vector a = hermitian_matrix(s[0], s[0] + 1.0); + check_against_lapack(&LR::elpa_linear_solver, a, a, s[0], s[1], s[2]); + } +} + +TEST_F(ZeqLinearSolverTest, ElpaSolvesHermitianPart) +{ + // a slightly non-symmetric input must be solved as its symmetric part (A + A^T)/2, + // not as whatever its upper triangle happens to be + for (const auto& s : sizes) + { + const int n = s[0]; + const std::vector a_sym = hermitian_matrix(n, n + 1.0); + std::vector a = a_sym; + for (int j = 0; j < n; ++j) + { + for (int i = 0; i < j; ++i) + { + a[j * n + i] += 1e-2 * entry(i, j, 5); + a[i * n + j] -= 1e-2 * entry(i, j, 5); + } + } + check_against_lapack(&LR::elpa_linear_solver, a, a_sym, n, s[1], s[2]); + } +} + +TEST_F(ZeqLinearSolverTest, ElpaRejectsIndefinite) +{ + // single rank only: with several ranks ELPA's Cholesky deadlocks on an indefinite matrix + // (only the rank owning the failing block returns), see `elpa_linear_solver` + int nproc = 1; + MPI_Comm_size(MPI_COMM_WORLD, &nproc); + if (nproc > 1) { return; } + const int n = 20; + const int nb = 3; + const std::vector a = hermitian_matrix(n, -(n + 1.0)); // negative definite + Parallel_2D pa; + LR_Util::setup_2d_division(pa, nb, n, n); + Parallel_2D pb; + LR_Util::setup_2d_division(pb, nb, n, 1, pa.blacs_ctxt); + std::vector a_loc = to_local(a, pa); + std::vector b_loc = to_local(rhs_matrix(n, 1), pb); + EXPECT_THROW(LR::elpa_linear_solver(a_loc.data(), b_loc.data(), pa, pb), std::runtime_error); +} +#endif + + +TEST(ZeqCGReview, SolvesIdentityWithoutDumpingProductionVector) +{ + double rhs[] = {2.0, -3.0}; + double z[] = {0.0, 0.0}; + const std::function identity = + [](const double* x, double* y) { y[0] = x[0]; y[1] = x[1]; }; + testing::internal::CaptureStdout(); + LR::solve_Z_CG(z, rhs, 2, 1, identity, false); + const std::string output = testing::internal::GetCapturedStdout(); + EXPECT_NEAR(z[0], rhs[0], 1e-12); + EXPECT_NEAR(z[1], rhs[1], 1e-12); + EXPECT_EQ(output.find("Final Z-vector:"), std::string::npos); +} + +TEST(ZeqCGReview, ZeroRhsIsConverged) +{ + double rhs[] = {0.0, 0.0}; + double z[] = {9.0, -1.0}; + const std::function identity = + [](const double* x, double* y) { y[0] = x[0]; y[1] = x[1]; }; + LR::solve_Z_CG(z, rhs, 2, 1, identity, false); + EXPECT_DOUBLE_EQ(z[0], 0.0); + EXPECT_DOUBLE_EQ(z[1], 0.0); +} + +TEST(ZeqCGReview, RejectsAnUnsolvableSystem) +{ + double rhs[] = {1.0, 2.0}; + double z[] = {0.0, 0.0}; + const std::function zero = + [](const double*, double* y) { y[0] = 0.0; y[1] = 0.0; }; + EXPECT_THROW(LR::solve_Z_CG(z, rhs, 2, 1, zero, false), std::runtime_error); +} + +int main(int argc, char** argv) +{ + MPI_Init(&argc, &argv); + testing::InitGoogleTest(&argc, argv); + const int result = RUN_ALL_TESTS(); + MPI_Finalize(); + return result; +} diff --git a/source/source_lcao/module_lr/utils/lr_util.cpp b/source/source_lcao/module_lr/utils/lr_util.cpp index f720012413a..987aa9f972a 100644 --- a/source/source_lcao/module_lr/utils/lr_util.cpp +++ b/source/source_lcao/module_lr/utils/lr_util.cpp @@ -1,12 +1,72 @@ #include "source_base/constants.h" #include "lr_util.h" +#include "source_cell/unitcell.h" +#include +#include +#include #include "source_base/module_external/lapack_connector.h" #include "source_base/module_external/scalapack_connector.h" +#include "source_base/module_container/base/third_party/lapack.h" namespace LR_Util { + void check_force_pp(const UnitCell& cell, const bool cal_force, const std::string& calculation) + { + const bool relaxation = calculation == "relax" || calculation == "cell-relax"; + if (!cal_force && !relaxation) { return; } + for (int type = 0; type < cell.ntype; ++type) + { + if (cell.atoms[type].ncpp.nlcc) + { + throw std::invalid_argument( + "LR excited-state forces and relaxation do not support NLCC pseudopotentials: element " + + cell.atoms[type].label + ", file " + cell.pseudo_fn.at(type) + + ". Core-density response-force and XC-kernel derivative terms are not implemented." + " LR spectra remain available with cal_force=0."); + } + } + } + + const std::vector& prepare_xc_sigma(const std::vector& sigma, + const bool is_hse06, + std::vector& hse_buffer) + { + if (!is_hse06) { return sigma; } + hse_buffer.resize(sigma.size()); + for (std::size_t index = 0; index < sigma.size(); ++index) + { + hse_buffer[index] = std::max(sigma[index], 1e-6); + } + return hse_buffer; + } + /// =================PHYSICS==================== int cal_nocc(int nelec) { return nelec / ModuleBase::DEGSPIN + nelec % static_cast(ModuleBase::DEGSPIN); } + int cal_nocc(double nelec, int nspin, int nupdown) + { + const double polarization = nspin == 2 ? std::abs(nupdown) : 0.0; + if (!std::isfinite(nelec) || nelec <= 0.0 || polarization > nelec) + { + throw std::invalid_argument("LR: invalid electron number or spin population"); + } + // Use ceil to include a partially occupied frontier orbital, without promoting + // a numerically noisy integer population to the next orbital. + const double population = (nelec + polarization) / 2.0; + const double nearest_integer = std::round(population); + const double stable_population = std::abs(population - nearest_integer) < 1e-8 + ? nearest_integer : population; + return static_cast(std::ceil(stable_population)); + } + + int cal_nocc_window(int requested_nocc, int nocc_max) + { + if (nocc_max <= 0) + { + throw std::invalid_argument("LR: no occupied orbitals"); + } + return requested_nocc <= 0 ? nocc_max : std::min(requested_nocc, nocc_max); + } + std::pair>> set_ix_map_diagonal(bool mode, int nocc, int nvirt) { @@ -93,6 +153,64 @@ namespace LR_Util const int i1 = 1; pztranc_(&n, &n, &alpha, tmp.data(), &i1, &i1, pmat.desc, &beta, inout, &i1, &i1, pmat.desc); } + + template<> + void mattrans(const double* in, const int n, const Parallel_2D& pmat, double* out) + { + std::copy(in, in + pmat.get_local_size(), out); + const double alpha = 1.0; + const double beta = 0.0; + const int i1 = 1; + pdtran_(&n, &n, &alpha, in, &i1, &i1, pmat.desc, &beta, out, &i1, &i1, pmat.desc); + } + template<> + void mattrans(double* inout, const int n, const Parallel_2D& pmat) + { + std::vector tmp(pmat.get_local_size()); + std::copy(inout, inout + pmat.get_local_size(), tmp.begin()); + const double alpha = 1.0; + const double beta = 0.0; + const int i1 = 1; + pdtran_(&n, &n, &alpha, tmp.data(), &i1, &i1, pmat.desc, &beta, inout, &i1, &i1, pmat.desc); + } + template<> + void mattrans>(const std::complex* in, const int n, const Parallel_2D& pmat, std::complex* out) + { + std::copy(in, in + pmat.get_local_size(), out); + const std::complex alpha(1.0, 0.0), beta(0.0, 0.0); + const int i1 = 1; + pztranc_(&n, &n, &alpha, in, &i1, &i1, pmat.desc, &beta, out, &i1, &i1, pmat.desc); + } + template<> + void mattrans>(std::complex* inout, const int n, const Parallel_2D& pmat) + { + std::vector> tmp(pmat.get_local_size()); + std::copy(inout, inout + pmat.get_local_size(), tmp.begin()); + const std::complex alpha(1.0, 0.0), beta(0.0, 0.0); + const int i1 = 1; + pztranc_(&n, &n, &alpha, tmp.data(), &i1, &i1, pmat.desc, &beta, inout, &i1, &i1, pmat.desc); + } + + + template<> + void matantisym(double* inout, const int n, const Parallel_2D& pmat) + { + std::vector tmp(pmat.get_local_size()); + std::copy(inout, inout + pmat.get_local_size(), tmp.begin()); + const double alpha = -0.5; + const double beta = 0.5; + const int i1 = 1; + pdtran_(&n, &n, &alpha, tmp.data(), &i1, &i1, pmat.desc, &beta, inout, &i1, &i1, pmat.desc); + } + template<> + void matantisym>(std::complex* inout, const int n, const Parallel_2D& pmat) + { + std::vector> tmp(pmat.get_local_size()); + std::copy(inout, inout + pmat.get_local_size(), tmp.begin()); + const std::complex alpha(-0.5, 0.0), beta(0.5, 0.0); + const int i1 = 1; + pztranc_(&n, &n, &alpha, tmp.data(), &i1, &i1, pmat.desc, &beta, inout, &i1, &i1, pmat.desc); + } #endif // for the first matrix in the commutator @@ -115,7 +233,8 @@ namespace LR_Util } #endif - void diag_lapack(const int& n, double* mat, double* eig) + template<> + void diag_lapack(const int& n, double* mat, double* eig) { ModuleBase::TITLE("LR_Util", "diag_lapack"); int info = 0; @@ -129,8 +248,8 @@ namespace LR_Util if (info) { std::cout << "ERROR: Lapack solver, info=" << info << std::endl; } delete[] work2; } - - void diag_lapack(const int& n, std::complex* mat, double* eig) + template<> + void diag_lapack>(const int& n, std::complex* mat, double* eig) { ModuleBase::TITLE("LR_Util", "diag_lapack >"); int lwork = 2 * n; @@ -143,8 +262,8 @@ namespace LR_Util delete[] rwork; delete[] work2; } - - void diag_lapack_nh(const int& n, double* mat, std::complex* eig) + template<> + void diag_lapack_nh(const int& n, double* mat, std::complex* eig) { ModuleBase::TITLE("LR_Util", "diag_lapack_nh"); int info = 0; @@ -164,8 +283,8 @@ namespace LR_Util if (info) { std::cout << "ERROR: Lapack solver dgeev, info=" << info << std::endl; } for (int i = 0;i < n;++i) { eig[i] = std::complex(eig_real[i], eig_imag[i]); } } - - void diag_lapack_nh(const int& n, std::complex* mat, std::complex* eig) + template<> + void diag_lapack_nh>(const int& n, std::complex* mat, std::complex* eig) { ModuleBase::TITLE("LR_Util", "diag_lapack_nh >"); int lwork = 2 * n; @@ -180,6 +299,46 @@ namespace LR_Util if (info) { std::cout << "ERROR: Lapack solver zgeev, info=" << info << std::endl; } } + template<> + int lapack_linear_solver(const double* A, double* x, const double* b, const int n, const int nrhs) + { + ModuleBase::TITLE("LR_Util", "lapack_linear_solver"); + // 1. copy A to a mutable array + std::vector A_copy(A, A + n * n); + // copy b to x + std::copy(b, b + n * nrhs, x); + // 2. LU decomposition: A->LU + std::vector ipiv(n); // pivot indices + int info = 0; + dgetrf_(&n, &n, A_copy.data(), &n, ipiv.data(), &info); + if (info) { std::cout << "ERROR: Lapack solver dgetrf, info=" << info << std::endl; } + // 3. Solve Ax=b + const char trans = 'N'; + dgetrs_(&trans, &n, &nrhs, A_copy.data(), &n, ipiv.data(), x, &n, &info); + if (info) { std::cout << "ERROR: Lapack solver dgetrs, info=" << info << std::endl; } + return info; + } + + template<> + int lapack_linear_solver>(const std::complex* A, std::complex* x, const std::complex* b, const int n, const int nrhs) + { + ModuleBase::TITLE("LR_Util", "lapack_linear_solver>"); + // 1. copy A to a mutable array + std::vector> A_copy(A, A + n * n); + // copy b to x + std::copy(b, b + n * nrhs, x); + // 2. LU decomposition: A->LU + std::vector ipiv(n); // pivot indices, for + int info = 0; + zgetrf_(&n, &n, A_copy.data(), &n, ipiv.data(), &info); + if (info) { std::cout << "ERROR: Lapack solver zgetrf, info=" << info << std::endl; } + // 3. Solve Ax=b + const char trans = 'N'; + zgetrs_(&trans, &n, &nrhs, A_copy.data(), &n, ipiv.data(), x, &n, &info); + if (info) { std::cout << "ERROR: Lapack solver zgetrs, info=" << info << std::endl; } + return info; + } + std::string tolower(const std::string& str) { std::string str_lower = str; @@ -192,4 +351,4 @@ namespace LR_Util std::transform(str_upper.begin(), str_upper.end(), str_upper.begin(), ::toupper); return str_upper; } -} \ No newline at end of file +} diff --git a/source/source_lcao/module_lr/utils/lr_util.h b/source/source_lcao/module_lr/utils/lr_util.h index c8c716c7402..fecf6f094bc 100644 --- a/source/source_lcao/module_lr/utils/lr_util.h +++ b/source/source_lcao/module_lr/utils/lr_util.h @@ -10,6 +10,10 @@ #include "source_base/parallel_2d.h" #include "source_psi/psi.h" #include +#include +#include + +class UnitCell; using DAT = container::DataType; using DEV = container::DeviceType; @@ -25,7 +29,41 @@ template <> struct ToComplex> { using type = std::complex& prepare_xc_sigma(const std::vector& sigma, + bool is_hse06, + std::vector& hse_buffer); + /// =====================PHYSICS==================== + /// @brief the xc kernels that carry an exact-exchange (EXX) term + /// + /// Membership here says only *that* the kernel has an EXX part, never *which* Coulomb + /// operator that part uses. The operator is the general range-separated form + /// $v_1(r)=[\alpha+\beta\,\mathrm{erfc}(\mu r)]/r$, and it is built from + /// `coulomb_param`, which `input_conv` fills from `dft_functional` plus + /// `exx_fock_alpha` ($\alpha$), `exx_erfc_alpha` ($\beta$) and `exx_erfc_omega` ($\mu$). + /// Its overall weight reaches this module as `exx_info.info_global.hybrid_alpha` + /// ($=\max(|\alpha|,|\beta|)$, the factor by which `coulomb_param` was normalized). + inline const std::set& hybrid_xc_list() + { + static const std::set l = { "hf", "hse", "pbe0", "b3lyp", + "cam_pbeh", "lc_pbe", "lc_wpbe", "lrc_wpbe", "lrc_wpbeh" }; + return l; + } + + /// @brief check if the xc functional has local xc kernel + inline bool has_local_xc(const std::string& name) + { + if (std::set({ "lda", "pwlda", "pbe" }).count(name)) { return true; } + // Every hybrid but pure HF keeps a semilocal remainder: the KS exchange left over after + // the EXX part is taken out, $(1-\alpha)E_x^\text{KS-LR}+[1-(\alpha+\beta)]E_x^ + // \text{KS-SR}$. libxc returns exactly that once `f_xc_libxc` hands the functional its + // external parameters, so no per-functional code is needed here -- only the name. + return name != "hf" && hybrid_xc_list().count(name); + } /// @brief calculate the number of electrons /// @tparam TCell @@ -36,6 +74,12 @@ namespace LR_Util /// @brief calculate the number of occupied orbitals /// @param nelec int cal_nocc(int nelec); + + /// Largest occupied spin channel; nelec already includes the charge correction. + int cal_nocc(double nelec, int nspin, int nupdown); + + /// Retain a positive user window, otherwise use all occupied orbitals. + int cal_nocc_window(int requested_nocc, int nocc_max); /// @brief set the index map: ix to (ic, iv) and vice versa /// by diagonal traverse the c-v pairs @@ -97,6 +141,14 @@ namespace LR_Util void matsym(const T* in, const int n, const Parallel_2D& pmat, T* out); template void matsym(T* inout, const int n, const Parallel_2D& pmat); + template + void mattrans(const T* in, const int n, const Parallel_2D& pmat, T* out); + template + void mattrans(T* inout, const int n, const Parallel_2D& pmat); + + // calculate (A-A^T)/2 (in-place version) + template + void matantisym(T* inout, const int n, const Parallel_2D& pmat); #endif template bool is_hermitian(const T* mat, const Parallel_2D& pmat, const double threshold, const int my_rank); @@ -156,20 +208,39 @@ namespace LR_Util template void gather_2d_to_full(const Parallel_2D& pv, const T* submat, T* fullmat, const bool row_major, const std::size_t global_nrow, const std::size_t global_ncol); + + /// @brief scatter full matrix to 2d block-cyclic distributed matrix + template + void scatter_full_to_2d(const Parallel_2D& pv, const T* fullmat, T* submat, const bool col_first = false); #endif ///=================diago-lapack==================== /// @brief diagonalize a hermitian matrix - void diag_lapack(const int& n, double* mat, double* eig); - void diag_lapack(const int& n, std::complex* mat, double* eig); - /// @brief diagonalize a general matrix - void diag_lapack_nh(const int& n, double* mat, std::complex* eig); - void diag_lapack_nh(const int& n, std::complex* mat, std::complex* eig); + template + void diag_lapack(const int& n, T* mat, double* eig); + /// @brief diagonalize a general matrix + template + void diag_lapack_nh(const int& n, T* mat, std::complex* eig); + ///================linear-solver-lapack============== + /// @brief solve linear equations Ax=b using LAPACK + template + int lapack_linear_solver(const T* A, T* x, const T* b, const int n, const int nrhs); ///=================string option==================== std::string tolower(const std::string& str); std::string toupper(const std::string& str); } +///=================operators======================= (should ot in namespace LR_Util) +template +std::vector operator+(const std::vector& a, const std::vector& b) +{ + const int maxsize = std::max(a.size(), b.size()); + const int minsize = std::min(a.size(), b.size()); + std::vector c(maxsize); + for (int i = 0;i < minsize;++i) { c[i] = a[i] + b[i]; } + for (int i = minsize;i < maxsize;++i) { c[i] = (a.size() > b.size() ? a[i] : b[i]); } + return c; +} #include "lr_util.hpp" #endif // ABACUS_SOURCE_LCAO_MODULE_LR_UTILS_LR_UTIL_H diff --git a/source/source_lcao/module_lr/utils/lr_util.hpp b/source/source_lcao/module_lr/utils/lr_util.hpp index c5bd9440d49..e233b5e450b 100644 --- a/source/source_lcao/module_lr/utils/lr_util.hpp +++ b/source/source_lcao/module_lr/utils/lr_util.hpp @@ -582,8 +582,21 @@ namespace LR_Util } MPI_Allreduce(MPI_IN_PLACE, fullmat, global_nrow * global_ncol, LR_Util::MPIType::value(), MPI_SUM, pv.comm()); }; -#endif + template + void scatter_full_to_2d(const Parallel_2D& pv, const T* fullmat, T* submat, const bool col_first) + { + ModuleBase::TITLE("LR_Util", "scatter_full_to_2d"); + const int global_nrow = pv.get_global_row_size(); + const int global_ncol = pv.get_global_col_size(); + for (int i = 0;i < pv.get_row_size();++i) + for (int j = 0;j < pv.get_col_size();++j) + if (col_first) + submat[i * pv.get_col_size() + j] = fullmat[pv.local2global_row(i) * global_ncol + pv.local2global_col(j)]; + else + submat[j * pv.get_row_size() + i] = fullmat[pv.local2global_col(j) * global_nrow + pv.local2global_row(i)]; + } +#endif } #endif // ABACUS_SOURCE_LCAO_MODULE_LR_UTILS_LR_UTIL_HPP diff --git a/source/source_lcao/module_lr/utils/lr_util_hcontainer.cpp b/source/source_lcao/module_lr/utils/lr_util_hcontainer.cpp deleted file mode 100644 index b655f00c99a..00000000000 --- a/source/source_lcao/module_lr/utils/lr_util_hcontainer.cpp +++ /dev/null @@ -1,96 +0,0 @@ -#include "lr_util_hcontainer.h" -namespace LR_Util -{ - void get_DMR_real_imag_part(const module_dm::DensityMatrix, std::complex>& DMR, - module_dm::DensityMatrix, double>& DMR_real, - const int& nat, - const char& type) - { - assert(DMR.get_dmr_vec().size() == DMR_real.get_dmr_vec().size()); - bool get_imag = (type == 'I' || type == 'i'); - for (int is = 0;is < DMR.get_dmr_vec().size();++is) - { - auto dr = DMR.get_dmr_vec()[is]; //get_dmr_ptr() has bug when is=0 - auto dr_real = DMR_real.get_dmr_vec()[is]; - assert(dr != nullptr); - assert(dr_real != nullptr); - for (int ia = 0;ia < nat;ia++) { - for (int ja = 0;ja < nat;ja++) - { - auto ap = dr->find_pair(ia, ja); - auto ap_real = dr_real->find_pair(ia, ja); - // under MPI-parallel (2D block-cyclic) HContainer, an atom pair not owned by this rank - // is absent from find_pair() and returns nullptr here; skip it - if (!ap || !ap_real) { continue; } - for (int iR = 0;iR < ap->get_R_size();++iR) - { - // R index may be different between the two HContainers, find by R value instead of R-index - auto dR = ap->get_R_index(iR); - auto ptr = ap->get_HR_values(iR).get_pointer(); - auto ptr_real = ap_real->get_HR_values(dR.x, dR.y, dR.z).get_pointer(); - for (int i = 0;i < ap->get_size();++i) { ptr_real[i] = (get_imag ? ptr[i].imag() : ptr[i].real()); } - } - } - } - } - } - - void get_DMR_real_imag_part(const module_dm::DensityMatrix, std::complex>& DMR, - module_dm::DensityMatrix, double>& DMR_real, - const int& nat, - const int& is, - const char& type) - { - assert(is < DMR.get_dmr_vec().size()); - assert(DMR_real.get_dmr_vec().size() == 1); - bool get_imag = (type == 'I' || type == 'i'); - auto dr = DMR.get_dmr_vec()[is]; //get_dmr_ptr() has bug when is=0 - auto dr_real = DMR_real.get_dmr_vec()[0]; - assert(dr != nullptr); - assert(dr_real != nullptr); - for (int ia = 0;ia < nat;ia++) { - for (int ja = 0;ja < nat;ja++) - { - auto ap = dr->find_pair(ia, ja); - auto ap_real = dr_real->find_pair(ia, ja); - // under MPI-parallel (2D block-cyclic) HContainer, an atom pair not owned by this rank - // is absent from find_pair() and returns nullptr here; skip it - if (!ap || !ap_real) { continue; } - for (int iR = 0;iR < ap->get_R_size();++iR) - { - // R index may be different between the two HContainers, find by R value instead of R-index - auto dR = ap->get_R_index(iR); - auto ptr = ap->get_HR_values(iR).get_pointer(); - auto ptr_real = ap_real->get_HR_values(dR.x, dR.y, dR.z).get_pointer(); - for (int i = 0;i < ap->get_size();++i) { ptr_real[i] = (get_imag ? ptr[i].imag() : ptr[i].real()); } - } - } - } - } - - void set_HR_real_imag_part(const hamilt::HContainer& HR_real, - hamilt::HContainer>& HR, - const int& nat, - const char& type) - { - bool get_imag = (type == 'I' || type == 'i'); - for (int ia = 0;ia < nat;ia++) { - for (int ja = 0;ja < nat;ja++) - { - auto ap = HR.find_pair(ia, ja); - auto ap_real = HR_real.find_pair(ia, ja); - // under MPI-parallel (2D block-cyclic) HContainer, an atom pair not owned by this rank - // is absent from find_pair() and returns nullptr here; skip it - if (!ap || !ap_real) { continue; } - for (int iR = 0;iR < ap->get_R_size();++iR) - { - // R index may be different between the two HContainers, find by R value instead of R-index - auto dR = ap->get_R_index(iR); - auto ptr = ap->get_HR_values(iR).get_pointer(); - auto ptr_real = ap_real->get_HR_values(dR.x, dR.y, dR.z).get_pointer(); - for (int i = 0;i < ap->get_size();++i) { get_imag ? ptr[i].imag(ptr_real[i]) : ptr[i].real(ptr_real[i]); } - } - } - } - } -} \ No newline at end of file diff --git a/source/source_lcao/module_lr/utils/lr_util_hcontainer.h b/source/source_lcao/module_lr/utils/lr_util_hcontainer.h index 2c5fa6bb7f9..e98726cd8e5 100644 --- a/source/source_lcao/module_lr/utils/lr_util_hcontainer.h +++ b/source/source_lcao/module_lr/utils/lr_util_hcontainer.h @@ -4,57 +4,139 @@ #include "source_estate/module_dm/density_matrix.h" #include #include "source_base/parallel_reduce.h" +#include "source_base/macros.h" +#include +#include "source_io/module_parameter/parameter.h" +#include "source_io/module_hs/lat_r_csr.h" +#include "source_lcao/module_lr/utils/lr_util.h" +#ifdef __EXX +#include "source_lcao/module_ri/abfs_vector3_order.h" +#include "source_lcao/module_ri/ri_2d_comm.h" +#endif namespace LR_Util { + template + using Real = typename GetTypeReal::type; template - void print_HR(const hamilt::HContainer& HR, const int& nat, const std::string& label, const double& threshold = 1e-10) + void print_HR(const hamilt::HContainer& HR, const std::string& label, const double& threshold = 1e-10) { std::cout << label << "\n"; - for (int ia = 0;ia < nat;ia++) - for (int ja = 0;ja < nat;ja++) + for (int iap = 0; iap < HR.size_atom_pairs(); ++iap) + { + auto ap = HR.get_atom_pair(iap); + const int ia = ap.get_atom_i(); + const int ja = ap.get_atom_j(); + for (int iR = 0;iR < ap.get_R_size();++iR) { - auto ap = HR.find_pair(ia, ja); - for (int iR = 0;iR < ap->get_R_size();++iR) + std::cout << "atom pair (" << ia << ", " << ja << "), " + << "R=(" << ap.get_R_index(iR)[0] << ", " << ap.get_R_index(iR)[1] << ", " << ap.get_R_index(iR)[2] << "): \n"; + auto& mat = ap.get_HR_values(iR); + std::cout << "rowsize=" << ap.get_row_size() << ", colsize=" << ap.get_col_size() << "\n"; + for (int i = 0;i < ap.get_row_size();++i) { - std::cout << "atom pair (" << ia << ", " << ja << "), " - << "R=(" << ap->get_R_index(iR)[0] << ", " << ap->get_R_index(iR)[1] << ", " << ap->get_R_index(iR)[2] << "): \n"; - auto& mat = ap->get_HR_values(iR); - std::cout << "rowsize=" << ap->get_row_size() << ", colsize=" << ap->get_col_size() << "\n"; - for (int i = 0;i < ap->get_row_size();++i) + for (int j = 0;j < ap.get_col_size();++j) { - for (int j = 0;j < ap->get_col_size();++j) - { - auto& v = mat.get_value(i, j); - std::cout << (std::abs(v) > threshold ? v : 0) << " "; - } - std::cout << "\n"; + auto& v = mat.get_value(i, j); + std::cout << (std::abs(v) > threshold ? v : 0) << " "; } + std::cout << "\n"; } } + } } + template - void print_DMR(const module_dm::DensityMatrix& DMR, const int& nat, const std::string& label, const double& threshold = 1e-10) + void print_DMR(const module_dm::DensityMatrix& DMR, const std::string& label, const double& threshold = 1e-10) { std::cout << label << "\n"; int is = 0; for (auto& dr : DMR.get_dmr_vec()) - print_HR(*dr, nat, "DMR[" + std::to_string(is++) + "]", threshold); - } - void get_DMR_real_imag_part(const module_dm::DensityMatrix, std::complex>& DMR, - module_dm::DensityMatrix, double>& DMR_real, - const int& nat, - const char& type = 'R'); - /// overload: only copy the `is`-th spin channel of DMR (source) into the (single-channel) DMR_real, - /// to avoid mixing/overlapping spin channels when DMR has more than one spin channel - void get_DMR_real_imag_part(const module_dm::DensityMatrix, std::complex>& DMR, - module_dm::DensityMatrix, double>& DMR_real, - const int& nat, + print_HR(*dr, "DMR[ispin=s" + std::to_string(is++) + "]", threshold); + } + + /// copy one atom pair's HContainer values from a complex DMR into a real-valued one, by + /// matching atom-pair/R rather than by nat-based index iteration (so it needs no `nat`). + template + inline void copy_hcontainer_real_imag_part(const hamilt::HContainer& HR, + hamilt::HContainer& HR_real, const bool get_imag) + { + for (int iap = 0; iap < HR.size_atom_pairs(); ++iap) + { + auto ap = &const_cast&>(HR).get_atom_pair(iap); + const int ia = ap->get_atom_i(); + const int ja = ap->get_atom_j(); + auto ap_real = HR_real.find_pair(ia, ja); + // under MPI-parallel (2D block-cyclic) HContainer, an atom pair not owned by this + // rank is absent from find_pair() and returns nullptr here; skip it + if (!ap_real) { continue; } + for (int iR = 0;iR < ap->get_R_size();++iR) + { + // R index may be different between the two HContainers, find by R value instead of R-index + auto dR = ap->get_R_index(iR); + auto ptr = ap->get_HR_values(iR).get_pointer(); + auto ptr_real = ap_real->get_HR_values(dR.x, dR.y, dR.z).get_pointer(); + for (int i = 0;i < ap->get_size();++i) { ptr_real[i] = (get_imag ? std::imag(ptr[i]) : std::real(ptr[i])); } + } + } + } + + template + void get_DMR_real_imag_part(const module_dm::DensityMatrix& DMR, + module_dm::DensityMatrix>& DMR_real, + const char& type = 'R') + { + assert(DMR.get_dmr_vec().size() == DMR_real.get_dmr_vec().size()); + const bool get_imag = (type == 'I' || type == 'i'); + for (size_t is = 0;is < DMR.get_dmr_vec().size();++is) + { + auto dr = DMR.get_dmr_vec()[is]; //get_dmr_ptr() has bug when is=0 + auto dr_real = DMR_real.get_dmr_vec()[is]; + assert(dr != nullptr); + assert(dr_real != nullptr); + copy_hcontainer_real_imag_part(*dr, *dr_real, get_imag); + } + } + + /// overload: only copy the `is`-th spin channel of DMR (source) into the (single-channel) + /// DMR_real, to avoid mixing/overlapping spin channels when DMR has more than one spin channel + template + void get_DMR_real_imag_part(const module_dm::DensityMatrix& DMR, + module_dm::DensityMatrix>& DMR_real, const int& is, - const char& type = 'R'); - void set_HR_real_imag_part(const hamilt::HContainer& HR_real, + const char& type = 'R') + { + assert(static_cast(is) < DMR.get_dmr_vec().size()); + assert(DMR_real.get_dmr_vec().size() == 1); + const bool get_imag = (type == 'I' || type == 'i'); + auto dr = DMR.get_dmr_vec()[is]; //get_dmr_ptr() has bug when is=0 + auto dr_real = DMR_real.get_dmr_vec()[0]; + assert(dr != nullptr); + assert(dr_real != nullptr); + copy_hcontainer_real_imag_part(*dr, *dr_real, get_imag); + } + + inline void set_HR_real_imag_part(const hamilt::HContainer& HR_real, hamilt::HContainer>& HR, - const int& nat, - const char& type = 'R'); + const char& type) + { + bool get_imag = (type == 'I' || type == 'i'); + for (int iap = 0; iap < HR.size_atom_pairs(); ++iap) + { + auto ap = &HR.get_atom_pair(iap); + const int ia = ap->get_atom_i(); + const int ja = ap->get_atom_j(); + auto ap_real = HR_real.find_pair(ia, ja); + assert(ap_real != nullptr); + for (int iR = 0;iR < ap->get_R_size();++iR) + { + // R index may be different between the two HContainers, find by R value instead of R-index + auto dR = ap->get_R_index(iR); + auto ptr = ap->get_HR_values(iR).get_pointer(); + auto ptr_real = ap_real->get_HR_values(dR.x, dR.y, dR.z).get_pointer(); + for (int i = 0;i < ap->get_size();++i) { get_imag ? ptr[i].imag(ptr_real[i]) : ptr[i].real(ptr_real[i]); } + } + } + } template void initialize_HR(hamilt::HContainer& hR, @@ -100,33 +182,304 @@ namespace LR_Util /// $\sum_{uvR} H1_{uv}(R) H2_{uv}(R)$ template - TR1 dot_R_matrix(const hamilt::HContainer& h1, const hamilt::HContainer& h2, const int& nat) + TR1 dot_R_matrix(const hamilt::HContainer& h1, const hamilt::HContainer& h2) { - const auto& pmat = *h1.get_paraV(); TR1 sum = 0; - // in case of the different order of atom pair and R-index in h1 and h2, we search by value instead of index - for (int iat1 = 0;iat1 < nat;++iat1) + // Match by atom/R values: containers may store pairs and images in different orders. + for (int iap = 0; iap < h1.size_atom_pairs(); ++iap) { - for (int iat2 = 0;iat2 < nat;++iat2) + const auto& ap1 = h1.get_atom_pair(iap); + const int iat1 = ap1.get_atom_i(); + const int iat2 = ap1.get_atom_j(); + const auto* ap2 = h2.find_pair(iat1, iat2); + assert(ap2); + for (int iR = 0; iR < ap1.get_R_size(); ++iR) { - auto ap1 = h1.find_pair(iat1, iat2); - if (!ap1) { continue; } - auto ap2 = h2.find_pair(iat1, iat2); - assert(ap2); - for (int iR = 0;iR < ap1->get_R_size();++iR) - { - const ModuleBase::Vector3& R = ap1->get_R_index(iR); - // std::cout<< "dot_R_matrix: iat1=" << iat1 << ", iat2=" << iat2 << ", R=(" << R.x << ", " << R.y << ", " << R.z << ")"<get_HR_values(R.x, R.y, R.z); - auto mat2 = ap2->get_HR_values(R.x, R.y, R.z); - sum += std::inner_product(mat1.get_pointer(), mat1.get_pointer() + mat1.get_col_size()*mat1.get_row_size(), mat2.get_pointer(), (TR1)0.0); - } + const auto& R = ap1.get_R_index(iR); + const auto& mat1 = ap1.get_HR_values(R.x, R.y, R.z); + const auto& mat2 = ap2->get_HR_values(R.x, R.y, R.z); + const int size = mat1.get_col_size() * mat1.get_row_size(); + const TR1* begin = mat1.get_pointer(); + const TR1* end = begin + size; + const TR2* values = mat2.get_pointer(); + const TR1 zero = 0; + sum += std::inner_product(begin, end, values, zero); } } // Parallel_Reduce::reduce_all(sum); // not needed, since it will be reduced outside return sum; } + + // cal_dmr converts column-major DMK into row-major DMR without changing AO indices. + // It already uses the ground-state DMR convention, including exp(+ik.R) for multi-k. + // A physical transpose must therefore be applied to DMK, never by swapping atom pairs. + template + void transpose_DMR(module_dm::DensityMatrix& dm, const Parallel_Orbitals& pv) + { + // 1. transpose dm(k) + for (auto& dk : dm.get_dmk_vec()) + { +#ifdef __MPI + // dm(k) is 2D-block-cyclic distributed, so the transpose needs the PBLAS routine + // `mattrans` (pdtran_/pztranc_) rather than a plain serial swap. + LR_Util::mattrans(dk.data(), pv.get_global_row_size(), pv); +#else + // In a serial build the entire square AO matrix is local, in column-major order. + const int n = pv.get_global_row_size(); + std::vector transposed(dk.size()); + for (int j = 0; j < n; ++j) + { + for (int i = 0; i < n; ++i) + { + transposed[j * n + i] = LR_Util::get_conj(dk[i * n + j]); + } + } + dk.swap(transposed); +#endif + } + + // 2. FT + dm.cal_dmr(-1); + } + template + void transpose_DMR(module_dm::DensityMatrix>& dm, const Parallel_Orbitals& pv) + { + throw std::runtime_error("transpose_DMR is not implemented for complex DMR, due to the lack of minus-sign FT."); + // 1. dm(k) dagger + for (auto& dk : dm.get_dmk_vec()) + { +#ifdef __MPI + LR_Util::mattrans(dk.data(), pv.get_global_row_size(), pv); +#else + throw std::runtime_error("transpose_DMR requires MPI (PBLAS mattrans) for the 2D-block-cyclic dm(k) transpose."); +#endif + } + + // 2. FT with the minus sign in the exponent (TO DO) + dm.cal_dmr(-1); + } + + template + module_dm::DensityMatrix build_dm_from_dmk(const std::vector& dmk, + const Parallel_Orbitals& pmat, + const int& nk, + const std::vector>& kvec_d, + const UnitCell& ucell, + const Grid_Driver& gd, + const std::vector& orb_cutoff, + const bool symmetrize = false, + const bool cal_dmr = true, + const bool transpose = false) + { + module_dm::DensityMatrix dm(&pmat, 1, kvec_d, nk); + initialize_DMR(dm, pmat, ucell, gd, orb_cutoff); + + if (symmetrize) + for (int ik = 0; ik < nk; ++ik) +#ifdef __MPI + LR_Util::matsym(dmk[ik].data(), pmat.get_global_row_size(), pmat); +#else + LR_Util::matsym(dmk[ik].data(), pmat.get_global_row_size()); +#endif + + for (int ik = 0; ik < nk; ++ik) + dm.set_dmk_ptr(ik, dmk[ik].data()); + + if (cal_dmr) + { + dm.cal_dmr(-1); + } + return dm; + } + + /// @brief Spin-resolved counterpart of `build_dm_from_dmk`: one `dmk` list per spin channel. + /// + /// Needed by the open-shell gradient, where $D^X$, $T$, $D^Z$ and the energy-weighted + /// density matrix all have two independent channels. `DensityMatrix` stores DMK as a flat + /// `[nspin][nk]` array, so channel `is` starts at `is * nk`. + template + module_dm::DensityMatrix build_dm_from_dmk_spin(const std::vector>& dmk, + const Parallel_Orbitals& pmat, + const int& nk, + const std::vector>& kvec_d, + const UnitCell& ucell, + const Grid_Driver& gd, + const std::vector& orb_cutoff, + const bool symmetrize = false, + const bool cal_dmr = true) + { + const int nspin_dm = static_cast(dmk.size()); + module_dm::DensityMatrix dm(&pmat, nspin_dm, kvec_d, nk); + initialize_DMR(dm, pmat, ucell, gd, orb_cutoff); + for (int is = 0; is < nspin_dm; ++is) + { + assert(static_cast(dmk[is].size()) >= nk); + if (symmetrize) + { + for (int ik = 0; ik < nk; ++ik) + { +#ifdef __MPI + LR_Util::matsym(dmk[is][ik].data(), pmat.get_global_row_size(), pmat); +#else + LR_Util::matsym(dmk[is][ik].data(), pmat.get_global_row_size()); +#endif + } + } + for (int ik = 0; ik < nk; ++ik) { dm.set_dmk_ptr(is * nk + ik, dmk[is][ik].data()); } + } + if (cal_dmr) + { + dm.cal_dmr(-1); + } + return dm; + } + + namespace sparse_format + { + // ref: sparse_format::cal_HContainer_d/cd and sparse_format::cal_HSR + // but more general(not depend on LCAO_HS_Arrays) + template + std::map, std::map>> + get_sparse_format( + const hamilt::HContainer& hR, + const Parallel_Orbitals& pv, + const double& sparse_thr = 1e-10) + { + std::map, std::map>> target; + auto row_indexes = pv.get_indexes_row(); + auto col_indexes = pv.get_indexes_col(); + for (int iap = 0; iap < hR.size_atom_pairs(); ++iap) { + int atom_i = hR.get_atom_pair(iap).get_atom_i(); + int atom_j = hR.get_atom_pair(iap).get_atom_j(); + int start_i = pv.atom_begin_row[atom_i]; + int start_j = pv.atom_begin_col[atom_j]; + int row_size = pv.get_nrow_atom(atom_i); + int col_size = pv.get_ncol_atom(atom_j); + for (int iR = 0; iR < hR.get_atom_pair(iap).get_R_size(); ++iR) { + auto& matrix = hR.get_atom_pair(iap).get_HR_values(iR); + const ModuleBase::Vector3 r_index + = hR.get_atom_pair(iap).get_R_index(iR); + Abfs::Vector3_Order dR(r_index.x, r_index.y, r_index.z); + for (int i = 0; i < row_size; ++i) { + int mu = row_indexes[start_i + i]; + for (int j = 0; j < col_size; ++j) { + int nu = col_indexes[start_j + j]; + const auto& value_tmp = matrix.get_value(i, j); + if (std::abs(value_tmp) > sparse_thr) { + target[dR][mu][nu] = value_tmp; + } + } + } + } + } + return target; + } + + //a more general version of save_HSR_sparse (not depend on LCAO_HS_Arrays) + template + void save_sparse( + const std::map, std::map>>& smat, + // const std::set>& all_R_coor, + // const bool& binary, + const std::string& filename, + const Parallel_Orbitals& pv, + const std::string& out_dir, + const int& nlocal, + const int& my_rank, + const double& sparse_thr = 1e-10) + { + // calculate the total number of non-zero elements of the (nbasis, nbasis) matrix for each R + std::vector non_zero_counts(smat.size(), 0); // number of Rs + int i = 0; + for (const auto& Rij : smat) + non_zero_counts[i++] = std::accumulate(Rij.second.begin(), Rij.second.end(), 0, + [](int sum, const std::pair>& line) { return sum + line.second.size(); }); + Parallel_Reduce::reduce_all(non_zero_counts.data(), non_zero_counts.size()); + + + std::string out_file = out_dir + filename; + std::ofstream ofs; + if (my_rank == 0) + { + ofs.open(out_file); + // if (binary) ofs.open(out_file, std::ios::binary); + ofs << "STEP: 0" << std::endl; + ofs << "Matrix Dimension: " << nlocal << std::endl; + ofs << "Matrix number: " << non_zero_counts.size() << std::endl; + } + i = 0; + for (const auto& Rij : smat) + { + const auto& R = Rij.first; + ofs << R.x << " " << R.y << " " << R.z << " " << non_zero_counts[i++] << std::endl; + ModuleIO::SparseWriteOptions single_R_options; + single_R_options.threshold = sparse_thr; + single_R_options.binary = false; + ModuleIO::save_lat_r(ofs, Rij.second, pv, single_R_options); + } + if (my_rank == 0) { ofs.close(); } + } + } + + template + void save_HR( + const hamilt::HContainer& hR, + // const std::set>& all_R_coor, + // const bool& binary, + const std::string& filename, + const Parallel_Orbitals& pv, + const std::string& out_dir, + const int& nlocal, + const int& my_rank, + const double& sparse_thr = 1e-10) + { + sparse_format::save_sparse(sparse_format::get_sparse_format(hR, pv, sparse_thr), + filename, pv, out_dir, nlocal, my_rank, sparse_thr); + } + + template + void save_DMR(const module_dm::DensityMatrix& DMR, + const std::string& filename, + const Parallel_Orbitals& pv, + const std::string& out_dir, + const int& nlocal, + const int& my_rank, + const double& sparse_thr = 1e-10) + { + int is = 0; + for (auto& dr : DMR.get_dmr_vec()) + save_HR(*dr, filename + "_s" + std::to_string(is), pv, out_dir, nlocal, my_rank, sparse_thr); + } + +#ifdef __EXX + // convert DensityMatrix to maps of RI::Tensors + // return 0.5*D[0] + template + auto get_exx_Ds_spin1(const module_dm::DensityMatrix& dm, + const UnitCell& ucell, const K_Vectors& kv, const Parallel_Orbitals& pmat) + -> std::map>, RI::Tensor>> + { + const int& nk = dm.get_DMK_nks(); // nks/nspin + std::vector*> DMk_trans_pointer(nk); + for (int ik = 0;ik < nk;++ik) { DMk_trans_pointer[ik] = &dm.get_dmk_vec()[ik]; } + return RI_2D_Comm::split_m2D_ktoR(ucell, kv, DMk_trans_pointer, pmat, /*nspin=*/1)[0]; + } + // return SPIN_multiple*D[0] as implemented in split_m2D_ktoR + // SPIN_multiple = map({ {1,0.5}, {2,1}, {4,1} }).at(nspin) + template + auto get_exx_Ds_gs(const module_dm::DensityMatrix& dm, + const UnitCell& ucell, const K_Vectors& kv, const Parallel_Orbitals& pmat) + -> std::vector>, RI::Tensor>>> + { + const int& nspin = dm.get_dmr_vec().size(); + const int nks = dm.get_DMK_nks(); + std::vector*> DMk_trans_pointer(nks); + for (int iks = 0;iks < dm.get_DMK_nks();++iks) + DMk_trans_pointer[iks] = &dm.get_dmk_vec()[iks]; + return RI_2D_Comm::split_m2D_ktoR(ucell, kv, DMk_trans_pointer, pmat, nspin); + } +#endif } #endif // ABACUS_SOURCE_LCAO_MODULE_LR_UTILS_LR_UTIL_HCONTAINER_H diff --git a/source/source_lcao/module_lr/utils/lr_util_xc.hpp b/source/source_lcao/module_lr/utils/lr_util_xc.hpp index e8992bfd429..98a2b75ee61 100644 --- a/source/source_lcao/module_lr/utils/lr_util_xc.hpp +++ b/source/source_lcao/module_lr/utils/lr_util_xc.hpp @@ -2,8 +2,18 @@ #define ABACUS_SOURCE_LCAO_MODULE_LR_UTILS_LR_UTIL_XC_HPP #include "lr_util.h" +#include "source_io/module_parameter/parameter.h" namespace LR_Util { + /// the `nspin` `KernelXC` (and the LR Hxc/xc potential built on it) is constructed with: + /// 1 for a non-magnetic calculation, 2 otherwise. A single shared definition so the several + /// consumers of this value cannot drift apart. + inline int kernel_nspin() + { + return (PARAM.inp.nspin == 1 + || (PARAM.inp.nspin == 4 && !PARAM.globalv.domag && !PARAM.globalv.domag_z)) ? 1 : 2; + } + template void grad(const T* rhor, ModuleBase::Vector3* gradrho, diff --git a/source/source_lcao/module_lr/utils/spectrum_mo.hpp b/source/source_lcao/module_lr/utils/spectrum_mo.hpp index ae3818331a5..552dd596c18 100644 --- a/source/source_lcao/module_lr/utils/spectrum_mo.hpp +++ b/source/source_lcao/module_lr/utils/spectrum_mo.hpp @@ -2,6 +2,7 @@ #define ABACUS_SOURCE_LCAO_MODULE_LR_UTILS_SPECTRUM_MO_HPP #include "source_base/tool_title.h" +#include "source_base/parallel_device.h" #include "source_basis/module_nao/two_center_bundle.h" #include "source_cell/klist.h" #include "source_io/module_parameter/parameter.h" @@ -104,7 +105,7 @@ std::vector> cal_velocity_mo(const UnitCell& ucell, for (int ik = 0; ik < nk; ++ik) { int glb_offset = (is * 3 * nk + id * nk + ik) * KS_num * KS_num; - int loc_offset = ik * pmo.get_local_size(); + int loc_offset = (is * nk + ik) * pmo.get_local_size(); for (int j = 0; j < pmo.get_col_size(); ++j){ for (int i = 0; i < pmo.get_row_size(); ++i){ velocity_mo[glb_offset + pmo.local2global_col(j) * KS_num + pmo.local2global_row(i)] @@ -115,7 +116,10 @@ std::vector> cal_velocity_mo(const UnitCell& ucell, } }//id #ifdef __MPI - MPI_Allreduce(MPI_IN_PLACE, velocity_mo.data(), velocity_mo.size(), LR_Util::MPIType::value(), MPI_SUM, pmo.comm()); + // velocity_mo is always complex, so MPIType cannot be used here + const int velocity_size = static_cast(velocity_mo.size()); + const MPI_Comm velocity_comm = pmo.comm(); + Parallel_Common::reduce_data(velocity_mo.data(), velocity_size, velocity_comm); #endif ModuleBase::GlobalFunc::DONE(GlobalV::ofs_running, "Finish velocity matrix in KS presentation."); ModuleBase::timer::end("LR_Util", "cal_velocity_mo"); diff --git a/source/source_lcao/module_lr/utils/test/CMakeLists.txt b/source/source_lcao/module_lr/utils/test/CMakeLists.txt index c5a71d7c3e1..75d763b7d02 100644 --- a/source/source_lcao/module_lr/utils/test/CMakeLists.txt +++ b/source/source_lcao/module_lr/utils/test/CMakeLists.txt @@ -10,4 +10,10 @@ AddTest( TARGET MODULE_LR_lr_util_algo_test LIBS parameter base device psi container planewave #for FFT SOURCES lr_util_algo_test.cpp ../lr_util.cpp -) \ No newline at end of file +) + +AddTest( + TARGET MODULE_LR_lr_util_pp + LIBS parameter base device container planewave cell_info symmetry psi + SOURCES test_lr_util.cpp ../lr_util.cpp ../../../../source_cell/magnetism.cpp +) diff --git a/source/source_lcao/module_lr/utils/test/lr_util_physics_test.cpp b/source/source_lcao/module_lr/utils/test/lr_util_physics_test.cpp index 7b0f547f151..bbddff19c7e 100644 --- a/source/source_lcao/module_lr/utils/test/lr_util_physics_test.cpp +++ b/source/source_lcao/module_lr/utils/test/lr_util_physics_test.cpp @@ -1,4 +1,5 @@ #include +#include #include "../lr_util.h" struct Atom_pseudo_Test @@ -36,6 +37,28 @@ TEST(LR_Util, cal_nocc) EXPECT_EQ(nocc, 3); } +TEST(LR_Util, ChargedSpinOccupiedWindow) +{ + // The effective electron number includes nelec_delta exactly once. + EXPECT_EQ(LR_Util::cal_nocc(254.0, 2, 2), 128); // NV-: 128 up, 126 down + EXPECT_EQ(LR_Util::cal_nocc(254.0, 2, -2), 128); + EXPECT_EQ(LR_Util::cal_nocc(252.0, 2, 2), 127); + EXPECT_EQ(LR_Util::cal_nocc(7.0, 2, 1), 4); // CH3 + EXPECT_EQ(LR_Util::cal_nocc(8.0, 2, 0), 4); + EXPECT_EQ(LR_Util::cal_nocc(8.0, 1, 0), 4); + EXPECT_EQ(LR_Util::cal_nocc(7.5, 2, 1), 5); // retain a partial frontier orbital + EXPECT_EQ(LR_Util::cal_nocc(8.0 + 1e-10, 1, 0), 4); + EXPECT_THROW(LR_Util::cal_nocc(2.0, 2, 3), std::invalid_argument); + const int full = LR_Util::cal_nocc(254.0, 2, 2); + EXPECT_EQ(LR_Util::cal_nocc_window(-1, full), 128); + EXPECT_EQ(LR_Util::cal_nocc_window(0, full), 128); + EXPECT_EQ(LR_Util::cal_nocc_window(819, full), 128); + const int selected = LR_Util::cal_nocc_window(4, full); + EXPECT_EQ(selected, 4); + EXPECT_EQ(full - selected, 124); // shared core prefix, up 4 and down 2 + EXPECT_EQ(selected - 2, 2); +} + TEST(LR_Util, set_ix_map_diagonal) { diff --git a/source/source_lcao/module_lr/utils/test/test_lr_util.cpp b/source/source_lcao/module_lr/utils/test/test_lr_util.cpp new file mode 100644 index 00000000000..6f712e84760 --- /dev/null +++ b/source/source_lcao/module_lr/utils/test/test_lr_util.cpp @@ -0,0 +1,86 @@ +#include "../lr_util.h" +#include "source_cell/unitcell.h" +#include +#include + +TEST(LRForcePP, AllowsForcesAndRelaxationWithoutNLCC) +{ + UnitCell cell; + cell.ntype = 1; + cell.atoms = new Atom[1]; + cell.set_atom_flag = true; + EXPECT_NO_THROW(LR_Util::check_force_pp(cell, true, "scf")); + EXPECT_NO_THROW(LR_Util::check_force_pp(cell, true, "relax")); +} + +TEST(LRForcePP, AllowsNLCCTDDFTSpectra) +{ + UnitCell cell; + cell.ntype = 2; + cell.atoms = new Atom[2]; + cell.set_atom_flag = true; + cell.pseudo_fn = {"H.upf", "Si_NLCC.upf"}; + cell.atoms[1].ncpp.nlcc = true; + EXPECT_NO_THROW(LR_Util::check_force_pp(cell, false, "scf")); + EXPECT_NO_THROW(LR_Util::check_force_pp(cell, false, "nscf")); +} + +TEST(LRForcePP, RejectsNLCCForcesAndIdentifiesThePseudopotential) +{ + UnitCell cell; + cell.ntype = 2; + cell.atoms = new Atom[2]; + cell.set_atom_flag = true; + cell.pseudo_fn = {"H.upf", "Si_NLCC.upf"}; + cell.atoms[1].ncpp.nlcc = true; + cell.atoms[1].label = "Si"; + try + { + LR_Util::check_force_pp(cell, true, "scf"); + FAIL() << "NLCC forces were accepted"; + } + catch (const std::invalid_argument& error) + { + const std::string message = error.what(); + EXPECT_NE(message.find("element Si"), std::string::npos); + EXPECT_NE(message.find("Si_NLCC.upf"), std::string::npos); + EXPECT_NE(message.find("not implemented"), std::string::npos); + } +} + +TEST(LRForcePP, RejectsNLCCRelaxationEvenWithoutExplicitForceFlag) +{ + UnitCell cell; + cell.ntype = 2; + cell.atoms = new Atom[2]; + cell.set_atom_flag = true; + cell.pseudo_fn = {"H.upf", "Si_NLCC.upf"}; + cell.atoms[1].ncpp.nlcc = true; + EXPECT_THROW(LR_Util::check_force_pp(cell, false, "relax"), std::invalid_argument); + EXPECT_THROW(LR_Util::check_force_pp(cell, true, "relax"), std::invalid_argument); +} + +TEST(LR_XCSigma, PreservesAntiparallelSpinGradientsWithoutHSE) +{ + // grad(up)=(2,0,0), grad(down)=(-1,0,0): a valid Gram matrix with sigma_ud < -1. + const std::vector sigma{4.0, -2.0, 1.0}; + std::vector hse_buffer; + const bool is_hse06 = false; + const auto& input = LR_Util::prepare_xc_sigma(sigma, is_hse06, hse_buffer); + EXPECT_EQ(input.data(), sigma.data()); + EXPECT_EQ(input[1], -2.0); + EXPECT_TRUE(hse_buffer.empty()); +} + +TEST(LR_XCSigma, RetainsExistingHSEStabilizationWithoutMutatingTheDensity) +{ + const std::vector sigma{0.01, -0.008, 0.0064}; + std::vector hse_buffer; + const bool is_hse06 = true; + const auto& input = LR_Util::prepare_xc_sigma(sigma, is_hse06, hse_buffer); + EXPECT_NE(input.data(), sigma.data()); + EXPECT_EQ(sigma[1], -0.008); + EXPECT_EQ(input[0], sigma[0]); + EXPECT_EQ(input[1], 1e-6); + EXPECT_EQ(input[2], sigma[2]); +} diff --git a/source/source_lcao/module_lr/zeq_solver.h b/source/source_lcao/module_lr/zeq_solver.h new file mode 100644 index 00000000000..64e13b79be5 --- /dev/null +++ b/source/source_lcao/module_lr/zeq_solver.h @@ -0,0 +1,11 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_ZEQ_SOLVER_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_ZEQ_SOLVER_H +#include "hamilt_zeq_l.h" +#include "hamilt_zeq_r.h" + +// `Z_vector_equation` is defined in zeq_solver.hpp. A declaration used to sit here with an older +// signature (no pot_hxc_gs / openshell / zvec_solver) and no definition anywhere, so any call that +// happened to match it would have failed at link time. + +#include "zeq_solver.hpp" +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_ZEQ_SOLVER_H diff --git a/source/source_lcao/module_lr/zeq_solver.hpp b/source/source_lcao/module_lr/zeq_solver.hpp new file mode 100644 index 00000000000..f92f3713a25 --- /dev/null +++ b/source/source_lcao/module_lr/zeq_solver.hpp @@ -0,0 +1,453 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_ZEQ_SOLVER_HPP +#define ABACUS_SOURCE_LCAO_MODULE_LR_ZEQ_SOLVER_HPP +#include "gradient_inputs.h" +#include +#include "zeq_solver.h" +#include +#include +#include +#include +#include +#include "source_base/opt_cg.h" +#include "source_lcao/module_lr/utils/lr_util.h" +#include "source_lcao/module_lr/utils/lr_util_print.h" +#include "zeqlin_solv.h" + +namespace LR +{ + // inline void solve_Z_CG(double* const Z, const double* const R, const int& ld, const int& nstates, + // std::function f_LZ, const bool test_force) + // Opt_CG's interfaces have no const qualifier for the pointer + inline void solve_Z_CG(double* const Z, double* R, const int& ld, const int& nstates, + std::function f_LZ, const bool test_force) + { + ModuleBase::TITLE("Z_vector", "solve_Z_CG"); + ModuleBase::timer::start("Z_vector", "solve_Z_CG"); + const int maxiter = 100; + double tol = 1e-6; + double residual = 10.; + const int size = nstates * ld; + container::Tensor P = LR_Util::newTensor({ nstates, ld }); // step length + container::Tensor LP = LR_Util::newTensor({ nstates, ld }); //f_LZ(P) + ModuleBase::zeros(Z, size); + + ModuleBase::Opt_CG cg; + cg.allocate(size); + cg.init_b(R); + std::cout << "Start solving Z-vector equaiton with CG method ..." << std::endl; + for (int iter = 0; iter < maxiter; ++iter) + { + if (residual < tol) + { + break; + } + cg.next_direct(LP.data(), 0, P.data()); + std::cout << "iter=" << iter << " residual=" << cg.get_residual() << std::endl; + // std::cout << "Z=" << std::endl; + // LR_Util::print_value(P.data(), nstates, ld); + // std::cout << "LZ before=" << std::endl; + // LR_Util::print_value(LP.data(), nstates, ld); + f_LZ(P.data(), LP.data()); // L: act each operators on P + // std::cout << "LZ=" << std::endl; + // LR_Util::print_value(LP.data(), nstates, ld); + int ifPD = 0; //??? + double step = cg.step_length(LP.data(), P.data(), ifPD); + // for (int i = 0; i < size; ++i) Z[i] += step * P[i]; + std::transform(Z, Z + size, P.data(), Z, [step](const double& z, const double& p) { return z + step * p; }); + residual = cg.get_residual(); + } + if (!(residual < tol)) + { + // Opt_CG's residual precedes its last step; check the returned vector itself. + f_LZ(Z, LP.data()); + double residual_squared = 0.0; + for (int i = 0; i < size; ++i) + { + const double difference = LP.data()[i] - R[i]; + residual_squared += difference * difference; + } + Parallel_Reduce::reduce_all(residual_squared); + residual = std::sqrt(residual_squared); + if (!(residual < tol)) + { + std::ostringstream message; + message << "Z-vector CG did not converge after " << maxiter + << " iterations (residual=" << residual << ", tolerance=" << tol + << "). Excited-state forces are not valid."; + throw std::runtime_error(message.str()); + } + } + if (test_force) + { + std::cout << "Final Z-vector:" << std::endl; + LR_Util::print_value(Z, nstates, ld); + } + ModuleBase::timer::end("Z_vector", "solve_Z_CG"); + } + + inline void solve_Z_CG(std::complex* const Z, std::complex* R, const int& ld, const int& nstates, + std::function* const, std::complex* const)> f_LZ, const bool test_force) + { + throw std::runtime_error("complex Z-vector solver is not implemented yet"); + } + + /// @brief Global length of one state's Z-vector: $\sum_\sigma n_k n_{occ,\sigma} n_{virt,\sigma}$. + /// `THam` only has to expose `nk`, `nocc` and `nvirt`. + template + inline int zvec_global_dim(THam& hm, const int nspin_x) + { + int n_global = 0; + for (int is = 0;is < nspin_x;++is) { n_global += hm.nk * hm.nocc[is] * hm.nvirt[is]; } + return n_global; + } + +#ifdef __MPI + /// @brief Gather one state's Z-vector from the `hm.pX` layout (`local`, length `ld`) into the + /// global vector `full` (length `zvec_global_dim`), replicated on every rank. Collective. + template + inline void zvec_local_to_full(THam& hm, const int nspin_x, const T* const local, T* const full) + { + const int n_global = zvec_global_dim(hm, nspin_x); + // `gather_2d_to_full` sums over the ranks, so the entries a rank does not own must be zero + std::fill(full, full + n_global, T(0)); + int loffset = 0; + int goffset = 0; + for (int is = 0;is < nspin_x;++is) + { + const int npairs = hm.nocc[is] * hm.nvirt[is]; + const int lsize = hm.pX[is].get_local_size(); + for (int ik = 0;ik < hm.nk;++ik) + { + LR_Util::gather_2d_to_full(hm.pX[is], local + loffset + ik * lsize, + full + goffset + ik * npairs, false, hm.nvirt[is], hm.nocc[is]); + } + loffset += hm.nk * lsize; + goffset += hm.nk * npairs; + } + } + + /// @brief Inverse of `zvec_local_to_full`: pick this rank's entries of the global vector + /// `full` into the `hm.pX` layout `local`. No communication. + template + inline void zvec_full_to_local(THam& hm, const int nspin_x, const T* const full, T* const local) + { + int loffset = 0; + int goffset = 0; + for (int is = 0;is < nspin_x;++is) + { + const int npairs = hm.nocc[is] * hm.nvirt[is]; + const int lsize = hm.pX[is].get_local_size(); + for (int ik = 0;ik < hm.nk;++ik) + { + LR_Util::scatter_full_to_2d(hm.pX[is], full + goffset + ik * npairs, + local + loffset + ik * lsize, false); + } + loffset += hm.nk * lsize; + goffset += hm.nk * npairs; + } + } +#endif + + /// @brief Solve the Z-vector equation with a dense LAPACK solve. + /// + /// Works for both the closed-shell (`nspin_x == 1`) and the open-shell (`nspin_x == 2`, + /// X = [up | down]) layouts: everything is expressed through per-spin segment sizes, which + /// collapse to the single-block case when `nspin_x == 1`. + /// `THam` only has to expose `matrix()`, `nk`, `nocc`, `nvirt` and `pX`. + /// @attention Every rank builds the full Hessian and solves the same system redundantly; + /// see `solve_Z_scalapack` / `solve_Z_elpa` for the distributed solves. + template + inline void solve_Z_lapack(T* const Z, const T* const R, const int& ld, const int& nstates, + THam& hm, const int nspin_x, const bool test_force) + { + ModuleBase::TITLE("Z_vector", "solve_Z_lapack"); + const int n_global = zvec_global_dim(hm, nspin_x); + int ld_expect = 0; + for (int is = 0;is < nspin_x;++is) { ld_expect += hm.nk * hm.pX[is].get_local_size(); } + assert(ld == ld_expect); + + std::vector hessian_full = hm.matrix(); // MO-hessian, A+B + std::vector Z_full(static_cast(n_global) * nstates, T(0.0)); + + // `hessian_full` is replicated, so the right-hand side must be global too. + // `R` is distributed over `pX` (length `ld` per state), so gather it first -- + // the mirror image of the scatter of `Z_full` below. Reading `n_global` entries + // straight out of `R` would be an out-of-bounds read as soon as + // ld < n_global, and a plain segfault on a rank whose local size is 0. + std::vector R_full(static_cast(n_global) * nstates, T(0.0)); +#ifdef __MPI + for (int istate = 0; istate < nstates; ++istate) + { + zvec_local_to_full(hm, nspin_x, R + istate * ld, R_full.data() + istate * n_global); + } +#else + std::copy(R, R + static_cast(n_global) * nstates, R_full.begin()); +#endif + + // use lapack to solve the linear equation + ModuleBase::timer::start("Z_vector", "lapack_solver"); + LR_Util::lapack_linear_solver(hessian_full.data(), Z_full.data(), R_full.data(), n_global, nstates); + ModuleBase::timer::end("Z_vector", "lapack_solver"); + + // test: print full Z + if (test_force) + { + std::cout << "The full Z-vector solved by LAPACK:" << std::endl; + LR_Util::print_value(Z_full.data(), nstates, n_global); + } + + // copy the local part of Z_full to Z +#ifdef __MPI + for (int istate = 0; istate < nstates; ++istate) + { + zvec_full_to_local(hm, nspin_x, Z_full.data() + istate * n_global, Z + istate * ld); + } +#else + std::copy(Z_full.begin(), Z_full.end(), Z); +#endif + if (test_force) + { + std::cout << "The local Z-vector solved by LAPACK:" << std::endl; + LR_Util::print_value(Z, nstates, ld); + } + } + +#ifdef __MPI + /// @brief Build this rank's block of the Z-vector Hessian on the 2D block-cyclic layout `ph` + /// (n_global x n_global), one column at a time: column g is `hm.hPsi` of the g-th unit vector, + /// i.e. exactly the operator the CG solve applies. Only O(n_global) is replicated per rank + /// (one column in flight), instead of the O(n_global^2) of `hm.matrix()`. + /// Collective: every rank walks every column, since `hPsi` and the gather communicate. + template + std::vector zvec_hessian_2d(THam& hm, const int nspin_x, const int ld, const Parallel_2D& ph) + { + ModuleBase::TITLE("Z_vector", "zvec_hessian_2d"); + ModuleBase::timer::start("Z_vector", "zvec_hessian_2d"); + const int n_global = ph.get_global_row_size(); + std::vector h_loc(ph.get_local_size(), T(0)); + std::vector e_full(n_global, T(0)); + std::vector e_loc(ld, T(0)); + std::vector he_loc(ld, T(0)); + std::vector he_full(n_global, T(0)); + for (int gcol = 0; gcol < n_global; ++gcol) + { + e_full[gcol] = T(1); + zvec_full_to_local(hm, nspin_x, e_full.data(), e_loc.data()); + e_full[gcol] = T(0); + hm.hPsi(e_loc.data(), he_loc.data(), ld, 1); + zvec_local_to_full(hm, nspin_x, he_loc.data(), he_full.data()); + const int lcol = ph.global2local_col(gcol); + if (lcol < 0) { continue; } + T* const h_col = h_loc.data() + static_cast(lcol) * ph.get_row_size(); + for (int lrow = 0; lrow < ph.get_row_size(); ++lrow) { h_col[lrow] = he_full[ph.local2global_row(lrow)]; } + } + ModuleBase::timer::end("Z_vector", "zvec_hessian_2d"); + return h_loc; + } + + /// @brief Distributed dense Z-vector solve: the Hessian (`zvec_hessian_2d`) and the + /// right-hand side are laid out 2D block-cyclically on the BLACS grid of `hm.pX[0]`, and + /// `linear_solver` (`scalapack_linear_solver`, `scalapack_cholesky_linear_solver` or + /// `elpa_linear_solver`) solves them in place. + template + void solve_Z_2d(T* const Z, const T* const R, const int ld, const int nstates, + THam& hm, const int nspin_x, + void (*linear_solver)(T*, T*, const Parallel_2D&, const Parallel_2D&)) + { + ModuleBase::TITLE("Z_vector", "solve_Z_2d"); + ModuleBase::timer::start("Z_vector", "solve_Z_2d"); + const int n_global = zvec_global_dim(hm, nspin_x); + const Parallel_2D& px0 = hm.pX[0]; + // 32 suits both ScaLAPACK and ELPA; shrink it for small systems so no rank is left empty + const int nproc_dim = std::max(px0.get_dim0(), px0.get_dim1()); + const int nb = std::max(1, std::min(32, n_global / nproc_dim)); + Parallel_2D ph; + LR_Util::setup_2d_division(ph, nb, n_global, n_global, px0.blacs_ctxt); + Parallel_2D pz; + LR_Util::setup_2d_division(pz, nb, n_global, nstates, px0.blacs_ctxt); + + std::vector h_loc = zvec_hessian_2d(hm, nspin_x, ld, ph); + + // right-hand side: pX layout -> 2D block-cyclic on pz + std::vector z_loc(pz.get_local_size(), T(0)); + std::vector r_full(n_global, T(0)); + for (int istate = 0; istate < nstates; ++istate) + { + zvec_local_to_full(hm, nspin_x, R + istate * ld, r_full.data()); + const int lcol = pz.global2local_col(istate); + if (lcol < 0) { continue; } + T* const z_col = z_loc.data() + static_cast(lcol) * pz.get_row_size(); + for (int lrow = 0; lrow < pz.get_row_size(); ++lrow) { z_col[lrow] = r_full[pz.local2global_row(lrow)]; } + } + + linear_solver(h_loc.data(), z_loc.data(), ph, pz); + + // solution: 2D block-cyclic on pz -> pX layout + std::vector z_full(static_cast(n_global) * nstates, T(0)); + LR_Util::gather_2d_to_full(pz, z_loc.data(), z_full.data(), false, n_global, nstates); + for (int istate = 0; istate < nstates; ++istate) + { + zvec_full_to_local(hm, nspin_x, z_full.data() + static_cast(istate) * n_global, Z + istate * ld); + } + ModuleBase::timer::end("Z_vector", "solve_Z_2d"); + } +#endif + + /// @brief Distributed dense Z-vector solve with ScaLAPACK LU (p?gesv). Same system as + /// `solve_Z_lapack`, but neither the Hessian nor the factorization is replicated. + template + inline void solve_Z_scalapack(T* const Z, const T* const R, const int ld, const int nstates, + THam& hm, const int nspin_x) + { +#ifdef __MPI + solve_Z_2d(Z, R, ld, nstates, hm, nspin_x, &scalapack_linear_solver); +#else + throw std::runtime_error("Z-vector solver 'scalapack' needs an MPI build; use 'lapack' or 'cg'"); +#endif + } + + /// @brief Distributed dense Z-vector solve with a ScaLAPACK Cholesky factorization (p?potrf), + /// about half the flops of the LU in `solve_Z_scalapack`. Like `solve_Z_elpa` it needs the + /// orbital Hessian to be positive definite, but a failure is an error on every rank, not a hang. + template + inline void solve_Z_scalapack_chol(T* const Z, const T* const R, const int ld, const int nstates, + THam& hm, const int nspin_x) + { +#ifdef __MPI + solve_Z_2d(Z, R, ld, nstates, hm, nspin_x, &scalapack_cholesky_linear_solver); +#else + throw std::runtime_error("Z-vector solver 'scalapack_chol' needs an MPI build; use 'lapack' or 'cg'"); +#endif + } + + /// @brief Distributed dense Z-vector solve with an ELPA Cholesky factorization: the orbital + /// Hessian A+B is symmetric positive definite at a stable ground state, so this needs about + /// half the flops of the LU in `solve_Z_scalapack`. Fails loudly if the Hessian is not positive + /// definite (an unstable ground state). + template + inline void solve_Z_elpa(T* const Z, const T* const R, const int ld, const int nstates, + THam& hm, const int nspin_x) + { +#ifdef __MPI + solve_Z_2d(Z, R, ld, nstates, hm, nspin_x, &elpa_linear_solver); +#else + throw std::runtime_error("Z-vector solver 'elpa' needs an MPI build; use 'lapack' or 'cg'"); +#endif + } + + /// @brief Run the configured solver (and then, for testing, every supported one). + /// Shared by the closed- and open-shell paths; `THamL` only needs `hPsi` plus what + /// `solve_Z_lapack` reads. + template + inline void solve_zeq_with(T* const Z, T* const R, const int ld, const int nstates, + THamL& ops_L, const int nspin_x, const std::string& zvec_solver, const bool test_force) + { + for (int i = 0; i < nstates * ld; ++i) { Z[i] = T(0.0); } // clear Z + if (zvec_solver == "cg") + { + solve_Z_CG(Z, R, ld, nstates, + [&ops_L, ld, nstates](const T* const in, T* const out) + { ops_L.hPsi(in, out, ld, nstates); }, test_force); + } + else if (zvec_solver == "lapack") { solve_Z_lapack(Z, R, ld, nstates, ops_L, nspin_x, test_force); } + else if (zvec_solver == "scalapack") { solve_Z_scalapack(Z, R, ld, nstates, ops_L, nspin_x); } + else if (zvec_solver == "scalapack_chol") { solve_Z_scalapack_chol(Z, R, ld, nstates, ops_L, nspin_x); } + else if (zvec_solver == "elpa") { solve_Z_elpa(Z, R, ld, nstates, ops_L, nspin_x); } + else { throw std::runtime_error("Unsupported Z-vector solver: " + zvec_solver); } + } + + /// Builds the Z-vector equation's RHS from `ops_R` and solves it with `ops_L`. Extracted to + /// a template function (rather than a generic lambda, which needs C++14) so the closed- and + /// open-shell call sites in `Z_vector_equation` -- which pass different `Z_vector_R`/`UR` and + /// `Z_vector_L`/`UL` types -- can share this body under the repository's C++11 baseline. + template + void build_and_solve_zeq(TOpsR& ops_R, TOpsL& ops_L, const int nspin_x, + const T* const X, container::Tensor& R, T* const Z, + const int nloc_per_band, const int nstates, const std::string& zvec_solver, const bool test_force) + { + ModuleBase::timer::start("Z_vector", "Z_vector_R"); + ops_R.hPsi(X, R.template data(), nloc_per_band, nstates); // act each operator on X + ModuleBase::timer::end("Z_vector", "Z_vector_R"); + // std::cout << "The right side of the Z-vector equation:" << std::endl; + // LR_Util::print_value(R.template data(), nstates, nloc_per_band); + solve_zeq_with(Z, R.template data(), nloc_per_band, nstates, ops_L, nspin_x, zvec_solver, test_force); + } + + template + void Z_vector_equation(const GradientInputs& inputs, + const T* const X, + T* const Z, + const int& nstates, + std::weak_ptr pot, + const std::string& spin_type, + const std::string& in_dir, + const bool openshell, + const std::string& zvec_solver) + { + const int& nspin = inputs.nspin; + const int& naos = inputs.nbasis; + const std::vector& nocc = inputs.nocc; + const std::vector& nvirt = inputs.nvirt; + const UnitCell& ucell = inputs.ucell; + const std::vector& orb_cutoff = inputs.orb_cutoff; + const Grid_Driver& gd = inputs.gd; + const K_Vectors& kv = inputs.kv; + const std::vector& px = inputs.px; + const Parallel_2D& pc = inputs.pc; + const Parallel_Orbitals& pmat = inputs.pmat; + const std::string& dft_functional = inputs.dft_functional; + std::weak_ptr pot_hxc_gs = inputs.pot_hxc_gs; +#ifdef __EXX + std::weak_ptr> exx_lri = inputs.exx_lri; + const double& exx_alpha = inputs.hybrid_alpha; +#endif + const std::string& xc_kernel = inputs.xc_kernel; + const psi::Psi& psi_ks = inputs.psi_ks; + const ModuleBase::matrix& eig_ks = inputs.eig_ks; + const std::string& out_dir = inputs.out_dir; + const std::string& ks_solver = inputs.ks_solver; + const bool test_force = inputs.test_force; + ModuleBase::TITLE("Z_vector", "Z_vector"); + const int nk = kv.get_nks() / nspin; + const int nloc_per_band = openshell + ? nk * (px[0].get_local_size() + px[1].get_local_size()) + : nk * px[0].get_local_size(); + container::Tensor R = LR_Util::newTensor({ nstates, nloc_per_band }); + R.zero(); + + if (openshell) + { + Z_vector_UR ops_R(xc_kernel, nspin, naos, nocc, nvirt, + ucell, orb_cutoff, gd, psi_ks, eig_ks, +#ifdef __EXX + exx_lri, exx_alpha, +#endif + pot, pot_hxc_gs, kv, px, pc, pmat, ks_solver, dft_functional); + Z_vector_UL ops_L(xc_kernel, nspin, naos, nocc, nvirt, + ucell, orb_cutoff, gd, psi_ks, eig_ks, +#ifdef __EXX + exx_lri, exx_alpha, +#endif + pot_hxc_gs, kv, px, pc, pmat, dft_functional); + build_and_solve_zeq(ops_R, ops_L, /*nspin_x=*/2, X, R, Z, nloc_per_band, nstates, zvec_solver, test_force); + } + else + { + Z_vector_R ops_R(xc_kernel, nspin, naos, nocc, nvirt, + ucell, orb_cutoff, gd, psi_ks, eig_ks, +#ifdef __EXX + exx_lri, exx_alpha, +#endif + pot, pot_hxc_gs, kv, px, pc, pmat, in_dir, out_dir, dft_functional, spin_type); + Z_vector_L ops_L(xc_kernel, nspin, naos, nocc, nvirt, + ucell, orb_cutoff, gd, psi_ks, eig_ks, +#ifdef __EXX + exx_lri, exx_alpha, +#endif + pot_hxc_gs, kv, px, pc, pmat, spin_type, in_dir, out_dir, dft_functional); + build_and_solve_zeq(ops_R, ops_L, /*nspin_x=*/1, X, R, Z, nloc_per_band, nstates, zvec_solver, test_force); + } + } +} + +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_ZEQ_SOLVER_HPP diff --git a/source/source_lcao/module_lr/zeqlin_solv.cpp b/source/source_lcao/module_lr/zeqlin_solv.cpp new file mode 100644 index 00000000000..11531828c9f --- /dev/null +++ b/source/source_lcao/module_lr/zeqlin_solv.cpp @@ -0,0 +1,199 @@ +#include "zeqlin_solv.h" + +#ifdef __MPI +#include +#include +#include +#include +#include +#include +#include + +#include "source_base/module_external/blacs_connector.h" +#include "source_base/module_external/scalapack_connector.h" +#include "source_base/timer.h" +#include "source_base/tool_title.h" +#ifdef _OPENMP +#include +#endif +#ifdef __ELPA +#include "source_hsolver/module_genelpa/elpa_new.h" +#ifdef I // avoid conflict with the macro defined by ELPA +#undef I +#endif +#endif + +namespace LR +{ + /// A <- (A + A^H) / 2 on the 2D layout `pA`. The Cholesky solvers read only the upper + /// triangle, so any numerical asymmetry of A would otherwise be dropped silently instead of + /// averaged. + template + void hermitize_2d(T* A, const Parallel_2D& pA) + { + const int n = pA.get_global_row_size(); + std::vector AH(pA.get_local_size(), T(0)); + const T one(1.0); + const T zero(0.0); + ScalapackConnector::tranc(n, n, one, A, 1, 1, pA.desc, zero, AH.data(), 1, 1, pA.desc); + for (std::size_t i = 0; i < AH.size(); ++i) { A[i] = 0.5 * (A[i] + AH[i]); } + } + +#ifdef __ELPA + /// Number of diagonal entries of the distributed Cholesky factor U that are not real, positive + /// and finite, summed over the BLACS grid of `pA`. A failed potrf leaves its non-positive pivot + /// on the diagonal, so this is nonzero exactly when the factorization broke down. + template + int count_bad_cholesky_pivots(const T* U, const Parallel_2D& pA) + { + int nbad = 0; + for (int lc = 0; lc < pA.get_col_size(); ++lc) + { + const int lr = pA.global2local_row(pA.local2global_col(lc)); + if (lr < 0) { continue; } + const T u = U[static_cast(lc) * pA.get_row_size() + lr]; + const double re = std::real(u); + const bool good = std::isfinite(re) && re > 0.0 && std::abs(std::imag(u)) <= 1e-12 * re; + if (!good) { ++nbad; } + } + char scope[] = "All"; + char top[] = " "; + Cigsum2d(pA.blacs_ctxt, scope, top, 1, 1, &nbad, 1, -1, -1); + return nbad; + } +#endif + + template + void scalapack_linear_solver(T* A, T* B, const Parallel_2D& pA, const Parallel_2D& pB) + { + ModuleBase::TITLE("LR", "scalapack_linear_solver"); + ModuleBase::timer::start("LR", "scalapack_linear_solver"); + const int n = pA.get_global_row_size(); + const int nrhs = pB.get_global_col_size(); + assert(pA.get_global_col_size() == n); + assert(pB.get_global_row_size() == n); + assert(pA.blacs_ctxt == pB.blacs_ctxt); + assert(pA.get_block_size() == pB.get_block_size()); + // p?gesv needs LOCr(M_A) + MB_A pivot entries + std::vector ipiv(pA.get_row_size() + pA.get_block_size()); + int info = 0; + ScalapackConnector::gesv(n, nrhs, A, 1, 1, pA.desc, ipiv.data(), B, 1, 1, pB.desc, &info); + if (info != 0) + { + throw std::runtime_error("scalapack_linear_solver: p?gesv failed, info=" + std::to_string(info)); + } + ModuleBase::timer::end("LR", "scalapack_linear_solver"); + } + + template + void scalapack_cholesky_linear_solver(T* A, T* B, const Parallel_2D& pA, const Parallel_2D& pB) + { + ModuleBase::TITLE("LR", "scalapack_cholesky_linear_solver"); + ModuleBase::timer::start("LR", "scalapack_cholesky_linear_solver"); + const int n = pA.get_global_row_size(); + const int nrhs = pB.get_global_col_size(); + assert(pA.get_global_col_size() == n); + assert(pB.get_global_row_size() == n); + assert(pA.blacs_ctxt == pB.blacs_ctxt); + assert(pA.get_block_size() == pB.get_block_size()); + + hermitize_2d(A, pA); + + // A = U^H U, U in the upper triangle. Unlike ELPA's, p?potrf's INFO is global output, so a + // non-positive-definite A is reported identically on every rank. + int desc_a[9]; + std::copy(pA.desc, pA.desc + 9, desc_a); + int info = ScalapackConnector::potrf('U', n, A, desc_a); + if (info != 0) + { + throw std::runtime_error("scalapack_cholesky_linear_solver: p?potrf failed, info=" + + std::to_string(info) + " -- the matrix is not positive definite; " + "use the LU-based 'scalapack' solver instead"); + } + ScalapackConnector::potrs('U', n, nrhs, A, 1, 1, pA.desc, B, 1, 1, pB.desc, &info); + if (info != 0) + { + throw std::runtime_error("scalapack_cholesky_linear_solver: p?potrs failed, info=" + + std::to_string(info)); + } + ModuleBase::timer::end("LR", "scalapack_cholesky_linear_solver"); + } + + template + void elpa_linear_solver(T* A, T* B, const Parallel_2D& pA, const Parallel_2D& pB) + { + ModuleBase::TITLE("LR", "elpa_linear_solver"); + ModuleBase::timer::start("LR", "elpa_linear_solver"); +#ifdef __ELPA + const int n = pA.get_global_row_size(); + const int nrhs = pB.get_global_col_size(); + assert(pA.get_global_col_size() == n); + assert(pB.get_global_row_size() == n); + assert(pA.blacs_ctxt == pB.blacs_ctxt); + assert(pA.get_block_size() == pB.get_block_size()); + + hermitize_2d(A, pA); + + int status = 0; + if (elpa_init(20210430) != ELPA_OK) + { + throw std::runtime_error("elpa_linear_solver: ELPA API version not supported"); + } + elpa_t handle = elpa_allocate(&status); + if (status != ELPA_OK) { throw std::runtime_error("elpa_linear_solver: elpa_allocate failed"); } + elpa_set(handle, "na", n, &status); + elpa_set(handle, "nev", n, &status); + elpa_set(handle, "local_nrows", pA.get_row_size(), &status); + elpa_set(handle, "local_ncols", pA.get_col_size(), &status); + elpa_set(handle, "nblk", pA.get_block_size(), &status); + elpa_set(handle, "mpi_comm_parent", MPI_Comm_c2f(pA.comm()), &status); + elpa_set(handle, "process_row", pA.get_coord_row(), &status); + elpa_set(handle, "process_col", pA.get_coord_col(), &status); +#ifdef _OPENMP + const int num_threads = omp_get_max_threads(); +#else + const int num_threads = 1; +#endif + elpa_set(handle, "omp_threads", num_threads, &status); + if (elpa_setup(handle) != ELPA_OK) { throw std::runtime_error("elpa_linear_solver: elpa_setup failed"); } + + // A = U^H U, U in the upper triangle. + // The returned status cannot be trusted: ELPA 2022.11 declares the C binding's `error` as + // intent(in), so a failed Cholesky still reports ELPA_OK. Check the pivots instead. + // On a non-positive-definite A only the rank owning the failing diagonal block returns + // (elpa_cholesky_template.F90) while the others wait in a broadcast, so with several + // ranks the check below is never reached and the run hangs. + elpa_cholesky(handle, A, &status); + // No `elpa_uninit` here: the KS solver may still hold ELPA handles of its own. + elpa_deallocate(handle, &status); + if (count_bad_cholesky_pivots(A, pA) > 0) + { + throw std::runtime_error("elpa_linear_solver: ELPA Cholesky failed -- the matrix is not " + "positive definite; use the LU-based 'scalapack' solver instead"); + } + + int info = 0; + ScalapackConnector::potrs('U', n, nrhs, A, 1, 1, pA.desc, B, 1, 1, pB.desc, &info); + if (info != 0) + { + throw std::runtime_error("elpa_linear_solver: p?potrs failed, info=" + std::to_string(info)); + } +#else + throw std::runtime_error("elpa_linear_solver: ABACUS was built without ELPA; rebuild with " + "ENABLE_ELPA=ON or use the 'scalapack' solver"); +#endif + ModuleBase::timer::end("LR", "elpa_linear_solver"); + } + + template void scalapack_linear_solver(double*, double*, const Parallel_2D&, const Parallel_2D&); + template void scalapack_linear_solver>(std::complex*, std::complex*, + const Parallel_2D&, const Parallel_2D&); + template void scalapack_cholesky_linear_solver(double*, double*, const Parallel_2D&, + const Parallel_2D&); + template void scalapack_cholesky_linear_solver>(std::complex*, + std::complex*, const Parallel_2D&, const Parallel_2D&); + template void elpa_linear_solver(double*, double*, const Parallel_2D&, const Parallel_2D&); + template void elpa_linear_solver>(std::complex*, std::complex*, + const Parallel_2D&, const Parallel_2D&); +} +#endif diff --git a/source/source_lcao/module_lr/zeqlin_solv.h b/source/source_lcao/module_lr/zeqlin_solv.h new file mode 100644 index 00000000000..75c546f3413 --- /dev/null +++ b/source/source_lcao/module_lr/zeqlin_solv.h @@ -0,0 +1,37 @@ +#ifndef ABACUS_SOURCE_LCAO_MODULE_LR_ZEQLIN_SOLV_H +#define ABACUS_SOURCE_LCAO_MODULE_LR_ZEQLIN_SOLV_H +#include "source_base/parallel_2d.h" + +namespace LR +{ +#ifdef __MPI + /// @brief Solve A Z = B with ScaLAPACK LU (p?gesv), the distributed counterpart of + /// `LR_Util::lapack_linear_solver`. + /// @param A [in/out] local block of the n x n matrix on `pA`; overwritten by its LU factors + /// @param B [in/out] local block of the n x nrhs right-hand side on `pB`; overwritten by Z + /// @attention `pA` and `pB` must share the BLACS context and the block size. + template + void scalapack_linear_solver(T* A, T* B, const Parallel_2D& pA, const Parallel_2D& pB); + + /// @brief Solve A Z = B for a Hermitian positive-definite A with ScaLAPACK alone: Cholesky + /// p?potrf A = U^H U, then p?potrs. A is Hermitized as (A + A^H)/2 first, as in + /// `elpa_linear_solver`. Arguments as in `scalapack_linear_solver`; A is overwritten by U. + /// Throws on every rank if A is not positive definite (p?potrf's INFO is global). + template + void scalapack_cholesky_linear_solver(T* A, T* B, const Parallel_2D& pA, const Parallel_2D& pB); + + /// @brief Solve A Z = B for a Hermitian positive-definite A: ELPA Cholesky A = U^H U, + /// then the two triangular solves with ScaLAPACK p?potrs. + /// A is Hermitized as (A + A^H)/2 first, since the Cholesky reads only the upper triangle. + /// Arguments as in `scalapack_linear_solver`; A is overwritten by U. + /// @attention requires an ELPA build (`__ELPA`); throws otherwise. A non-positive-definite A + /// throws only on a single rank: ELPA's Cholesky returns early on the rank owning the failing + /// diagonal block and leaves the others in a broadcast, so with several ranks it HANGS. + /// (Its error code is no help either: ELPA 2022.11 reports ELPA_OK even then, so the failure + /// is detected from the pivots of U.) + template + void elpa_linear_solver(T* A, T* B, const Parallel_2D& pA, const Parallel_2D& pB); +#endif +} + +#endif // ABACUS_SOURCE_LCAO_MODULE_LR_ZEQLIN_SOLV_H diff --git a/source/source_lcao/module_operator_lcao/CMakeLists.txt b/source/source_lcao/module_operator_lcao/CMakeLists.txt index 8a6ddaeeb50..439a20fa6cc 100644 --- a/source/source_lcao/module_operator_lcao/CMakeLists.txt +++ b/source/source_lcao/module_operator_lcao/CMakeLists.txt @@ -7,6 +7,7 @@ add_library( veff_dh.cpp deepks_lcao.cpp overlap.cpp + ovlp_block.cpp overlap_fs.cpp ekinetic.cpp ekinetic_fs.cpp diff --git a/source/source_lcao/module_operator_lcao/operator_lcao.cpp b/source/source_lcao/module_operator_lcao/operator_lcao.cpp index d98ce868d4e..a623243bdc2 100644 --- a/source/source_lcao/module_operator_lcao/operator_lcao.cpp +++ b/source/source_lcao/module_operator_lcao/operator_lcao.cpp @@ -185,6 +185,11 @@ void OperatorLCAO::init(const int ik_in) { break; } case calculation_type::lcao_exx: + case calculation_type::lr_dmtrans_hxc: + case calculation_type::lr_dmtrans_exx: + case calculation_type::lr_dmdiff_hxc: + case calculation_type::lr_dmdiff_exx: + case calculation_type::lr_dmtrans_gxc: { // EXX is accumulated in H(R); the last operator-chain node folds // the complete H(R) into H(k), including the TD gauge phase. diff --git a/source/source_lcao/module_operator_lcao/overlap.cpp b/source/source_lcao/module_operator_lcao/overlap.cpp index 6b206c60a01..3c7228580f7 100644 --- a/source/source_lcao/module_operator_lcao/overlap.cpp +++ b/source/source_lcao/module_operator_lcao/overlap.cpp @@ -1,4 +1,5 @@ #include "overlap.h" +#include "ovlp_block.h" #include "source_base/timer.h" #include "source_base/tool_title.h" @@ -154,7 +155,8 @@ void Overlap>::calculate_SR() const ModuleBase::Vector3 R_index = tmp.get_R_index(iR); auto dtau = ucell->cal_dtau(iat1, iat2, R_index); TR* data_pointer = tmp.get_pointer(iR); - this->cal_SR_IJR(iat1, iat2, paraV, dtau, data_pointer); + const auto displacement = dtau * this->ucell->lat0; + cal_overlap_block(*this->ucell, *this->intor_, iat1, iat2, *paraV, displacement, data_pointer); } } // if TK == double, then SR should be fixed to gamma case @@ -166,75 +168,6 @@ void Overlap>::calculate_SR() ModuleBase::timer::end("Overlap", "calculate_SR"); } -// cal_SR_IJR() -template -void Overlap>::cal_SR_IJR(const int& iat1, - const int& iat2, - const Parallel_Orbitals* paraV, - const ModuleBase::Vector3& dtau, - TR* data_pointer) -{ - // --------------------------------------------- - // get info of orbitals of atom1 and atom2 from ucell - // --------------------------------------------- - int T1=0; - int I1=0; - this->ucell->iat2iait(iat1, &I1, &T1); - int T2=0; - int I2=0; - this->ucell->iat2iait(iat2, &I2, &T2); - Atom& atom1 = this->ucell->atoms[T1]; - Atom& atom2 = this->ucell->atoms[T2]; - - // npol is the number of polarizations, - // 1 for non-magnetic (one Hamiltonian matrix only has spin-up or spin-down), - // 2 for magnetic (one Hamiltonian matrix has both spin-up and spin-down) - const int npol = this->ucell->get_npol(); - - const int* iw2l1 = atom1.iw2l.data(); - const int* iw2n1 = atom1.iw2n.data(); - const int* iw2m1 = atom1.iw2m.data(); - const int* iw2l2 = atom2.iw2l.data(); - const int* iw2n2 = atom2.iw2n.data(); - const int* iw2m2 = atom2.iw2m.data(); - - // --------------------------------------------- - // calculate the overlap matrix for each pair of orbitals - // --------------------------------------------- - double olm[3] = {0, 0, 0}; - auto row_indexes = paraV->get_indexes_row(iat1); - auto col_indexes = paraV->get_indexes_col(iat2); - const int step_trace = col_indexes.size() + 1; - for (int iw1l = 0; iw1l < row_indexes.size(); iw1l += npol) - { - const int iw1 = row_indexes[iw1l] / npol; - const int L1 = iw2l1[iw1]; - const int N1 = iw2n1[iw1]; - const int m1 = iw2m1[iw1]; - - // convert m (0,1,...2l) to M (-l, -l+1, ..., l-1, l) - int M1 = (m1 % 2 == 0) ? -m1 / 2 : (m1 + 1) / 2; - - for (int iw2l = 0; iw2l < col_indexes.size(); iw2l += npol) - { - const int iw2 = col_indexes[iw2l] / npol; - const int L2 = iw2l2[iw2]; - const int N2 = iw2n2[iw2]; - const int m2 = iw2m2[iw2]; - - // convert m (0,1,...2l) to M (-l, -l+1, ..., l-1, l) - int M2 = (m2 % 2 == 0) ? -m2 / 2 : (m2 + 1) / 2; - intor_->calculate(T1, L1, N1, M1, T2, L2, N2, M2, dtau * this->ucell->lat0, olm); - for (int ipol = 0; ipol < npol; ipol++) - { - data_pointer[ipol * step_trace] += olm[0]; - } - data_pointer += npol; - } - data_pointer += (npol - 1) * col_indexes.size(); - } -} - // contributeHR() template void Overlap>::contributeHR() @@ -387,7 +320,8 @@ HContainer* Overlap>::calculate_SR_async(const UnitCell } TR* data_pointer = atom_pair.get_pointer(iR); - this->cal_SR_IJR(iat1, iat2, paraV_local, dtau, data_pointer); + const auto displacement = dtau * this->ucell->lat0; + cal_overlap_block(*this->ucell, *this->intor_, iat1, iat2, *paraV_local, displacement, data_pointer); } } diff --git a/source/source_lcao/module_operator_lcao/overlap.h b/source/source_lcao/module_operator_lcao/overlap.h index c9df271c2b0..222928d29d4 100644 --- a/source/source_lcao/module_operator_lcao/overlap.h +++ b/source/source_lcao/module_operator_lcao/overlap.h @@ -95,15 +95,6 @@ class Overlap> : public OperatorLCAO */ void calculate_SR(); - /** - * @brief calculate the SR local matrix of atom pair - */ - void cal_SR_IJR(const int& iat1, - const int& iat2, - const Parallel_Orbitals* paraV, - const ModuleBase::Vector3& dtau, - TR* data_pointer); - /** * @brief calculate force contribution for atom pair */ diff --git a/source/source_lcao/module_operator_lcao/ovlp_block.cpp b/source/source_lcao/module_operator_lcao/ovlp_block.cpp new file mode 100644 index 00000000000..89b132b56a4 --- /dev/null +++ b/source/source_lcao/module_operator_lcao/ovlp_block.cpp @@ -0,0 +1,52 @@ +#include "ovlp_block.h" +#include "source_basis/module_ao/parallel_orbitals.h" +#include "source_basis/module_nao/two_center_integrator.h" +#include "source_cell/unitcell.h" +#include + +namespace hamilt +{ +template +void cal_overlap_block(const UnitCell& cell, const TwoCenterIntegrator& integrator, + const int atom_i, const int atom_j, const Parallel_Orbitals& distribution, + const ModuleBase::Vector3& displacement, TR* values) +{ + const int type_i = cell.iat2it[atom_i]; + const int type_j = cell.iat2it[atom_j]; + const Atom& bra = cell.atoms[type_i]; + const Atom& ket = cell.atoms[type_j]; + const int npol = cell.get_npol(); + const auto rows = distribution.get_indexes_row(atom_i); + const auto cols = distribution.get_indexes_col(atom_j); + const int column_count = cols.size(); + const int spin_stride = column_count + 1; + for (std::size_t row = 0; row < rows.size(); row += npol) + { + const int mu = rows[row] / npol; + const int l1 = bra.iw2l[mu]; + const int n1 = bra.iw2n[mu]; + const int bra_m = bra.iw2m[mu]; + const int m1 = bra_m % 2 == 0 ? -bra_m / 2 : (bra_m + 1) / 2; + for (std::size_t col = 0; col < cols.size(); col += npol) + { + const int nu = cols[col] / npol; + const int l2 = ket.iw2l[nu]; + const int n2 = ket.iw2n[nu]; + const int ket_m = ket.iw2m[nu]; + const int m2 = ket_m % 2 == 0 ? -ket_m / 2 : (ket_m + 1) / 2; + double integral = 0.0; + integrator.calculate(type_i, l1, n1, m1, type_j, l2, n2, m2, displacement, &integral); + const std::size_t offset = row * column_count + col; + for (int spin = 0; spin < npol; ++spin) + { + values[offset + spin * spin_stride] += integral; + } + } + } +} + +template void cal_overlap_block(const UnitCell&, const TwoCenterIntegrator&, int, int, + const Parallel_Orbitals&, const ModuleBase::Vector3&, double*); +template void cal_overlap_block(const UnitCell&, const TwoCenterIntegrator&, int, int, + const Parallel_Orbitals&, const ModuleBase::Vector3&, std::complex*); +} diff --git a/source/source_lcao/module_operator_lcao/ovlp_block.h b/source/source_lcao/module_operator_lcao/ovlp_block.h new file mode 100644 index 00000000000..f21ce8a883f --- /dev/null +++ b/source/source_lcao/module_operator_lcao/ovlp_block.h @@ -0,0 +1,18 @@ +#ifndef ABACUS_LCAO_OVLP_BLOCK_H +#define ABACUS_LCAO_OVLP_BLOCK_H +namespace ModuleBase { template class Vector3; } +class UnitCell; +class Parallel_Orbitals; +class TwoCenterIntegrator; +namespace hamilt +{ +// Add one ordered atom-pair overlap block in local row-major AO layout. +// displacement is ket-center minus bra-center in Bohr, including any lattice image +// or cross-geometry shift. No neighbor search, coordinate reconstruction or symmetrization. +// For npol=2, add the same spatial integral to the two spin-diagonal entries. +template +void cal_overlap_block(const UnitCell& cell, const TwoCenterIntegrator& integrator, + int atom_i, int atom_j, const Parallel_Orbitals& distribution, + const ModuleBase::Vector3& displacement, TR* values); +} +#endif diff --git a/source/source_lcao/module_operator_lcao/test/CMakeLists.txt b/source/source_lcao/module_operator_lcao/test/CMakeLists.txt index b93390dfed8..dbd321da8f7 100644 --- a/source/source_lcao/module_operator_lcao/test/CMakeLists.txt +++ b/source/source_lcao/module_operator_lcao/test/CMakeLists.txt @@ -1,10 +1,23 @@ if(ENABLE_LCAO) abacus_disable_feature_definitions(__FFT_TWO_CENTER) +AddTest( + TARGET MODULE_LCAO_ovlp_block + LIBS parameter psi base device container + SOURCES test_ovlp_block.cpp ../ovlp_block.cpp tmp_mocks.cpp + ../../../source_basis/module_ao/parallel_orbitals.cpp + ../../../source_basis/module_ao/orb_atomic_lm.cpp + ../../../source_hamilt/operator.cpp + ../../../source_hamilt/module_hcontainer/base_matrix.cpp + ../../../source_hamilt/module_hcontainer/hcontainer.cpp + ../../../source_hamilt/module_hcontainer/atom_pair.cpp + ../../../source_hamilt/module_hcontainer/func_folding.cpp +) + AddTest( TARGET MODULE_LCAO_operator_overlap_test LIBS parameter psi base device container - SOURCES test_overlap.cpp ../overlap.cpp ../overlap_fs.cpp ../operator_fs_utils.cpp ../../../source_hamilt/module_hcontainer/func_folding.cpp + SOURCES test_overlap.cpp ../overlap.cpp ../ovlp_block.cpp ../overlap_fs.cpp ../operator_fs_utils.cpp ../../../source_hamilt/module_hcontainer/func_folding.cpp ../../../source_hamilt/module_hcontainer/base_matrix.cpp ../../../source_hamilt/module_hcontainer/hcontainer.cpp ../../../source_hamilt/module_hcontainer/atom_pair.cpp ../../../source_hamilt/module_hcontainer/func_transfer.cpp ../../../source_hamilt/module_hcontainer/output_hcontainer.cpp ../../../source_hamilt/module_hcontainer/transfer.cpp @@ -17,7 +30,7 @@ AddTest( AddTest( TARGET MODULE_LCAO_operator_overlap_serial_test LIBS parameter psi base device container - SOURCES test_overlap_serial.cpp ../overlap.cpp ../overlap_fs.cpp ../operator_fs_utils.cpp ../../../source_hamilt/module_hcontainer/func_folding.cpp + SOURCES test_overlap_serial.cpp ../overlap.cpp ../ovlp_block.cpp ../overlap_fs.cpp ../operator_fs_utils.cpp ../../../source_hamilt/module_hcontainer/func_folding.cpp ../../../source_hamilt/module_hcontainer/base_matrix.cpp ../../../source_hamilt/module_hcontainer/hcontainer.cpp ../../../source_hamilt/module_hcontainer/atom_pair.cpp ../../../source_hamilt/module_hcontainer/func_transfer.cpp ../../../source_hamilt/module_hcontainer/output_hcontainer.cpp ../../../source_hamilt/module_hcontainer/transfer.cpp @@ -30,7 +43,7 @@ AddTest( AddTest( TARGET MODULE_LCAO_operator_overlap_cd_test LIBS parameter psi base device container - SOURCES test_overlap_cd.cpp ../overlap.cpp ../overlap_fs.cpp ../operator_fs_utils.cpp ../../../source_hamilt/module_hcontainer/func_folding.cpp + SOURCES test_overlap_cd.cpp ../overlap.cpp ../ovlp_block.cpp ../overlap_fs.cpp ../operator_fs_utils.cpp ../../../source_hamilt/module_hcontainer/func_folding.cpp ../../../source_hamilt/module_hcontainer/base_matrix.cpp ../../../source_hamilt/module_hcontainer/hcontainer.cpp ../../../source_hamilt/module_hcontainer/atom_pair.cpp ../../../source_hamilt/module_hcontainer/func_transfer.cpp ../../../source_hamilt/module_hcontainer/output_hcontainer.cpp ../../../source_hamilt/module_hcontainer/transfer.cpp diff --git a/source/source_lcao/module_operator_lcao/test/test_overlap.cpp b/source/source_lcao/module_operator_lcao/test/test_overlap.cpp index e3c017857b2..3af696a7ab1 100644 --- a/source/source_lcao/module_operator_lcao/test/test_overlap.cpp +++ b/source/source_lcao/module_operator_lcao/test/test_overlap.cpp @@ -1,6 +1,7 @@ #include "../overlap.h" #include "gtest/gtest.h" +#include //--------------------------------------- // Unit test of Overlap class @@ -172,6 +173,33 @@ TEST_F(OverlapTest, constructHRd2cd) } } +TEST_F(OverlapTest, AsyncOverlapKeepsSeparateContainer) +{ + ucell.lat0 = 1.0; + const ModuleBase::Vector3 velocity(0.001, 0.0, 0.0); + ucell.atoms[0].vel.assign(ucell.nat, velocity); + const ModuleBase::Vector3 gamma(0.0, 0.0, 0.0); + const std::vector> kpoints{gamma}; + const std::vector cutoff{1.0}; + hamilt::HS_Matrix_K hsk(paraV); + Grid_Driver neighbors(0, 0); + hamilt::Overlap> op( + &hsk, kpoints, nullptr, SR, &ucell, cutoff, &neighbors, &intor_); + auto* const async_container = op.calculate_SR_async(ucell, 0.1, paraV); + std::unique_ptr> async(async_container); + ASSERT_EQ(async->size_atom_pairs(), SR->size_atom_pairs()); + for (int i = 0; i < async->size_atom_pairs(); ++i) + { + const auto& pair = async->get_atom_pair(i); + const int count = pair.get_row_size() * pair.get_col_size(); + for (int j = 0; j < count; ++j) + { + EXPECT_DOUBLE_EQ(pair.get_pointer(0)[j], 1.0); + EXPECT_DOUBLE_EQ(SR->get_atom_pair(i).get_pointer(0)[j], 0.0); + } + } +} + int main(int argc, char** argv) { #ifdef __MPI diff --git a/source/source_lcao/module_operator_lcao/test/test_ovlp_block.cpp b/source/source_lcao/module_operator_lcao/test/test_ovlp_block.cpp new file mode 100644 index 00000000000..1f57e7ce918 --- /dev/null +++ b/source/source_lcao/module_operator_lcao/test/test_ovlp_block.cpp @@ -0,0 +1,67 @@ +#include "../ovlp_block.h" +#include "source_basis/module_ao/parallel_orbitals.h" +#include "source_basis/module_nao/two_center_integrator.h" +#include "source_cell/unitcell.h" +#include +#include +#include + +class OverlapBlockTest : public ::testing::Test +{ + protected: + void SetUp() override + { + atom.reset(new Atom); + cell.ntype = 1; + cell.nat = 1; + cell.atoms = atom.get(); + // Statistics owns these maps and releases them with UnitCell. + cell.iat2it = new int[1]{0}; + cell.iat2ia = new int[1]{0}; + atom->nw = 2; + atom->iw2l = {0, 0}; + atom->iw2n = {0, 0}; + atom->iw2m = {0, 0}; + } + void set_polarization(int npol) + { + const std::vector offsets{0}; + cell.set_iat2iwt_for_test(offsets, npol); + const int size = atom->nw * npol; + distribution.set_serial(size, size); + distribution.set_atomic_trace(cell.get_iat2iwt(), cell.nat, size); + } + UnitCell cell; + std::unique_ptr atom; + Parallel_Orbitals distribution; + TwoCenterIntegrator integrator; +}; + +TEST_F(OverlapBlockTest, AddsToExistingRealBlock) +{ + set_polarization(1); + const ModuleBase::Vector3 displacement(0.3, -0.2, 0.1); + std::vector values(4, 2.0); + // The existing operator fixture mocks each two-center integral as one. + hamilt::cal_overlap_block(cell, integrator, 0, 0, distribution, displacement, values.data()); + hamilt::cal_overlap_block(cell, integrator, 0, 0, distribution, displacement, values.data()); + for (double value : values) { EXPECT_DOUBLE_EQ(value, 4.0); } +} + +TEST_F(OverlapBlockTest, PreservesComplexSpinOffDiagonalEntries) +{ + set_polarization(2); + const ModuleBase::Vector3 displacement(-0.3, 0.2, -0.1); + const std::complex initial(2.0, 3.0); + std::vector> values(16, initial); + hamilt::cal_overlap_block(cell, integrator, 0, 0, distribution, displacement, values.data()); + for (int row = 0; row < 4; ++row) + { + for (int col = 0; col < 4; ++col) + { + const double increment = row % 2 == col % 2 ? 1.0 : 0.0; + const std::complex expected(2.0 + increment, 3.0); + EXPECT_EQ(values[row * 4 + col], expected); + } + } +} diff --git a/source/source_lcao/module_ri/CMakeLists.txt b/source/source_lcao/module_ri/CMakeLists.txt index c45ca9751a4..b3ce16f3e58 100644 --- a/source/source_lcao/module_ri/CMakeLists.txt +++ b/source/source_lcao/module_ri/CMakeLists.txt @@ -22,6 +22,7 @@ if (ENABLE_LIBRI) gaussian_abfs.cpp singular_value.cpp exx_lri_detail.cpp + exx_lr_ws.cpp ) endif() add_library( diff --git a/source/source_lcao/module_ri/exx_lr_ws.cpp b/source/source_lcao/module_ri/exx_lr_ws.cpp new file mode 100644 index 00000000000..9670c6412ef --- /dev/null +++ b/source/source_lcao/module_ri/exx_lr_ws.cpp @@ -0,0 +1,78 @@ +#include "exx_lr_ws.h" +#include "exx_lri.h" + +namespace +{ +template +void copy_pack(const RI::Exx& source, + RI::Exx& destination, + const std::string& name) +{ + destination.lri.data_pool.emplace(name, source.lri.data_pool.at(name)); +} +} + +template +void share_exx_geometry(const RI::Exx& source, + RI::Exx& destination) +{ + // Only geometry packs may cross the KS/LR boundary. In particular, do not copy + // Ds, Hs, cvc, label bindings, symmetry filters or parallel/post-processing objects. + if (source.flag_finish.Cs) { copy_pack(source, destination, "Cs_"); } + if (source.flag_finish.Vs) { copy_pack(source, destination, "Vs_"); } + for (int axis = 0; axis < 3; ++axis) + { + const std::string suffix = std::to_string(axis) + "_"; + const std::string dc_name = "dCs_" + suffix; + const std::string dv_name = "dVs_" + suffix; + if (source.flag_finish.dCs) { copy_pack(source, destination, dc_name); } + if (source.flag_finish.dVs) { copy_pack(source, destination, dv_name); } + for (int second = 0; second < 3; ++second) + { + const std::string stress_suffix = suffix + std::to_string(second) + "_"; + const std::string dcr_name = "dCRs_" + stress_suffix; + const std::string dvr_name = "dVRs_" + stress_suffix; + if (source.flag_finish.dCRs) { copy_pack(source, destination, dcr_name); } + if (source.flag_finish.dVRs) { copy_pack(source, destination, dvr_name); } + } + } + destination.flag_finish.Cs = source.flag_finish.Cs; + destination.flag_finish.Vs = source.flag_finish.Vs; + destination.flag_finish.dCs = source.flag_finish.dCs; + destination.flag_finish.dVs = source.flag_finish.dVs; + destination.flag_finish.dCRs = source.flag_finish.dCRs; + destination.flag_finish.dVRs = source.flag_finish.dVRs; +} + +template +std::shared_ptr> Exx_LRI::make_lr_workspace(const UnitCell& ucell, + const K_Vectors& kv) const +{ + auto workspace = std::make_shared>(this->info); + workspace->mpi_comm = this->mpi_comm; + workspace->p_kv = &kv; + workspace->abfs_Lmax_ = this->abfs_Lmax_; + std::map atoms_pos; + for (int atom = 0; atom < ucell.nat; ++atom) + { + const int type = ucell.iat2it[atom]; + const int index = ucell.iat2ia[atom]; + atoms_pos[atom] = RI_Util::Vector3_to_array3(ucell.atoms[type].tau[index]); + } + const std::array lattice = {RI_Util::Vector3_to_array3(ucell.a1), + RI_Util::Vector3_to_array3(ucell.a2), + RI_Util::Vector3_to_array3(ucell.a3)}; + const std::array period = {kv.nmp[0], kv.nmp[1], kv.nmp[2]}; + workspace->exx_lri.set_parallel(this->mpi_comm, atoms_pos, lattice, period); + share_exx_geometry(this->exx_lri, workspace->exx_lri); + return workspace; +} + +template void share_exx_geometry(const RI::Exx&, + RI::Exx&); +template void share_exx_geometry(const RI::Exx>&, + RI::Exx>&); +template std::shared_ptr> Exx_LRI::make_lr_workspace(const UnitCell&, + const K_Vectors&) const; +template std::shared_ptr>> +Exx_LRI>::make_lr_workspace(const UnitCell&, const K_Vectors&) const; diff --git a/source/source_lcao/module_ri/exx_lr_ws.h b/source/source_lcao/module_ri/exx_lr_ws.h new file mode 100644 index 00000000000..4f0090b6ff4 --- /dev/null +++ b/source/source_lcao/module_ri/exx_lr_ws.h @@ -0,0 +1,18 @@ +#ifndef EXX_LR_WS_H +#define EXX_LR_WS_H + +#include + +namespace RI +{ +template +class Exx; +} + +// The destination must be fresh, with its own parallel context for the same geometry. +// Tensor storage is shared read-only; map containers and electronic state are independent. +template +void share_exx_geometry(const RI::Exx& source, + RI::Exx& destination); + +#endif diff --git a/source/source_lcao/module_ri/exx_lri.h b/source/source_lcao/module_ri/exx_lri.h index 67eaa321c28..f734d36472b 100644 --- a/source/source_lcao/module_ri/exx_lri.h +++ b/source/source_lcao/module_ri/exx_lri.h @@ -61,6 +61,15 @@ class Exx_LRI Exx_LRI operator=(const Exx_LRI&) = delete; Exx_LRI operator=(Exx_LRI&&); + // accessors used by the LR-TDDFT analytical-gradient module + RI::Exx& get() { return this->exx_lri; } + auto& get_info() const { return this->info; } + auto& get_mpi_comm() const { return this->mpi_comm; } + + // Share read-only geometry tensors, but keep all electronic work buffers independent. + std::shared_ptr> make_lr_workspace(const UnitCell& ucell, + const K_Vectors& kv) const; + void init( const MPI_Comm &mpi_comm_in, const UnitCell &ucell, @@ -111,6 +120,9 @@ class Exx_LRI ModuleBase::matrix force_exx; ModuleBase::matrix stress_exx; + void post_process_Hexx(std::map>>& Hexxs_io) const; + double post_process_Eexx(const double& Eexx_in) const; + int abfs_Lmax() const { return abfs_Lmax_; } const Exx_Info_RI& get_info_ri() const { return info; } @@ -133,9 +145,6 @@ class Exx_LRI std::map>>>> coulomb_settings; - void post_process_Hexx( std::map>> &Hexxs_io ) const; - double post_process_Eexx(const double& Eexx_in) const; - friend class RPA_LRI; friend class RPA_LRI, Tdata>; friend class Exx_LRI_Interface; diff --git a/source/source_lcao/module_ri/exx_lri.hpp b/source/source_lcao/module_ri/exx_lri.hpp index 04a134a7748..e9b6ef2874d 100644 --- a/source/source_lcao/module_ri/exx_lri.hpp +++ b/source/source_lcao/module_ri/exx_lri.hpp @@ -924,7 +924,8 @@ void Exx_LRI::cal_exx_force(const int& nat) ModuleBase::timer::start("Exx_LRI", "cal_exx_force"); this->force_exx.create(nat, Ndim); - for(int is=0; isexx_lri.cal_force({"","",std::to_string(is),"",""}); for(std::size_t idim=0; idim::cal_exx_force(const int& nat) this->force_exx(force_item.first, idim) += std::real(force_item.second); } } } - - const double SPIN_multiple = std::map{{1,2}, {2,1}, {4,1}}.at(PARAM.inp.nspin); // why? - const double frac = -2 * SPIN_multiple; // why? + // SPIN_multiple cancels the one in `split_m2D_ktoR`, which are 0.5*0.5 at nspin=1. + // but only u-u and d-d pairs of Ds has contribution, so here's 2 instead of 4. + // And -2 is the same as post_process_Hexx (which didn't act on Hs) + const double SPIN_multiple = std::map{{1,2}, {2,1}, {4,1}}.at(nspin); + const double frac = -2 * SPIN_multiple; this->force_exx *= frac; ModuleBase::timer::end("Exx_LRI", "cal_exx_force"); } diff --git a/source/source_lcao/module_ri/test/CMakeLists.txt b/source/source_lcao/module_ri/test/CMakeLists.txt index ff75006e5b6..817017e6fe3 100644 --- a/source/source_lcao/module_ri/test/CMakeLists.txt +++ b/source/source_lcao/module_ri/test/CMakeLists.txt @@ -1,6 +1,13 @@ abacus_disable_feature_definitions(__MLALGO) abacus_disable_feature_definitions(__CUDA) abacus_disable_feature_definitions(__ROCM) +if(ENABLE_LCAO) + AddTest( + TARGET MODULE_RI_exx_lr_ws + LIBS base parameter device ${math_libs} + SOURCES test_exx_lr_ws.cpp ../exx_lr_ws.cpp ../../../source_basis/module_ao/orb_atomic_lm.cpp + ) +endif() AddTest( TARGET MODULE_RI_dm_mixing_test LIBS parameter base device diff --git a/source/source_lcao/module_ri/test/test_exx_lr_ws.cpp b/source/source_lcao/module_ri/test/test_exx_lr_ws.cpp new file mode 100644 index 00000000000..a7e11d45f92 --- /dev/null +++ b/source/source_lcao/module_ri/test/test_exx_lr_ws.cpp @@ -0,0 +1,155 @@ +#include "../exx_lr_ws.h" +#include "source_base/parallel_global.h" +#include +#include + +namespace +{ +template +void check_workspace() +{ + using Engine = RI::Exx; + using Cell = std::array; + using Pair = std::pair; + using TensorMap = std::map>>; + Engine ks; + Engine lr; + const std::map> positions{{0, {0.0, 0.0, 0.0}}}; + const std::array, 3> lattice{{{1.0, 0.0, 0.0}, + {0.0, 1.0, 0.0}, + {0.0, 0.0, 1.0}}}; + const Cell period{1, 1, 1}; + ks.set_parallel(MPI_COMM_WORLD, positions, lattice, period); + lr.set_parallel(MPI_COMM_WORLD, positions, lattice, period); + RI::Tensor c({1, 1, 1}); + RI::Tensor v({1, 1}); + c.ptr()[0] = T(2.0); + v.ptr()[0] = T(3.0); + const Pair pair{0, {0, 0, 0}}; + const TensorMap cs{{0, {{pair, c}}}}; + const TensorMap vs{{0, {{pair, v}}}}; + ks.set_Cs(cs, 0.0); + ks.set_Vs(vs, 0.0); + const std::array dcs{{cs, cs, cs}}; + const std::array dvs{{vs, vs, vs}}; + ks.set_dCs(dcs, 0.0); + ks.set_dVs(dvs, 0.0); + const std::array, 3> dcrs{{dcs, dcs, dcs}}; + const std::array, 3> dvrs{{dvs, dvs, dvs}}; + ks.set_dCRs(dcrs, 0.0); + ks.set_dVRs(dvrs, 0.0); + ks.set_Ds(vs, 0.0, "0"); + ks.set_Ds(vs, 0.0, "1"); + const std::array gs_names{{"", "", "1"}}; + ks.cal_Hs(gs_names); + TensorMap gs_hs; + for (const auto& atom : ks.Hs) + { + for (const auto& entry : atom.second) { gs_hs[atom.first][entry.first] = entry.second.copy(); } + } + std::map original_packs; + for (const auto& pack : ks.lri.data_pool) + { + for (const auto& atom : pack.second.Ds_ab) + { + for (const auto& entry : atom.second) + { + original_packs[pack.first][atom.first][entry.first] = entry.second.copy(); + } + } + } + const auto gs_bindings = ks.lri.data_ab_name; + const auto gs_pool_size = ks.lri.data_pool.size(); + const auto gs_energy = ks.energy; + share_exx_geometry(ks, lr); + EXPECT_NE(ks.lri.parallel.get(), lr.lri.parallel.get()); + EXPECT_NE(ks.lri.filter_atom.get(), lr.lri.filter_atom.get()); + EXPECT_FALSE(lr.flag_finish.Ds); + EXPECT_TRUE(lr.Hs.empty()); + EXPECT_EQ(lr.lri.data_pool.count("Ds_0"), 0); + EXPECT_EQ(lr.lri.data_pool.count("Ds_1"), 0); + EXPECT_EQ(lr.lri.data_pool.size(), 26); + for (const auto& pack : lr.lri.data_pool) + { + for (const auto& atom : pack.second.Ds_ab) + { + for (const auto& entry : atom.second) + { + const auto& original = ks.lri.data_pool.at(pack.first).Ds_ab.at(atom.first).at(entry.first); + EXPECT_EQ(original.ptr(), entry.second.ptr()); + } + } + } + // Exercise the operations that previously replaced the KS density and Hamiltonian. + RI::Tensor response({1, 1}); + response.ptr()[0] = T(0.5); + const TensorMap dx{{0, {{pair, response}}}}; + lr.set_Ds(dx, 0.0); + lr.cal_Hs(); + lr.cal_force(); + EXPECT_EQ(ks.lri.data_pool.size(), gs_pool_size); + EXPECT_EQ(ks.lri.data_ab_name, gs_bindings); + EXPECT_EQ(ks.energy, gs_energy); + for (const auto& pack : original_packs) + { + for (const auto& atom : pack.second) + { + for (const auto& entry : atom.second) + { + const auto& unchanged = ks.lri.data_pool.at(pack.first).Ds_ab.at(atom.first).at(entry.first); + EXPECT_EQ(unchanged.ptr()[0], entry.second.ptr()[0]); + } + } + } + EXPECT_EQ(ks.Hs.size(), gs_hs.size()); + for (const auto& atom : gs_hs) + { + for (const auto& entry : atom.second) + { + const auto& unchanged = ks.Hs.at(atom.first).at(entry.first); + EXPECT_EQ(unchanged.ptr()[0], entry.second.ptr()[0]); + } + } + EXPECT_EQ(c.ptr()[0], T(2.0)); + EXPECT_EQ(v.ptr()[0], T(3.0)); + // A new ionic-step workspace must see the newly prepared geometry, not old packs. + RI::Tensor next_c({1, 1, 1}); + next_c.ptr()[0] = T(4.0); + const TensorMap next_cs{{0, {{pair, next_c}}}}; + ks.set_Cs(next_cs, 0.0); + Engine next_lr; + next_lr.set_parallel(MPI_COMM_WORLD, positions, lattice, period); + share_exx_geometry(ks, next_lr); + for (const auto& atom : next_lr.lri.data_pool.at("Cs_").Ds_ab) + { + for (const auto& entry : atom.second) + { + const auto& updated = ks.lri.data_pool.at("Cs_").Ds_ab.at(atom.first).at(entry.first); + const auto& previous = lr.lri.data_pool.at("Cs_").Ds_ab.at(atom.first).at(entry.first); + EXPECT_EQ(entry.second.ptr()[0], updated.ptr()[0]); + EXPECT_NE(entry.second.ptr()[0], previous.ptr()[0]); + } + } +} +} + +TEST(ExxLRWorkspace, RealGeometryAndElectronicIsolation) { check_workspace(); } +TEST(ExxLRWorkspace, ComplexGeometryAndElectronicIsolation) { check_workspace>(); } + +int main(int argc, char** argv) +{ + int processes = 1; + int threads = 1; + int rank = 0; + Parallel_Global::read_pal_param(argc, argv, processes, threads, rank); + POOL_WORLD = MPI_COMM_NULL; + KP_WORLD = MPI_COMM_NULL; + INT_BGROUP = MPI_COMM_NULL; + BP_WORLD = MPI_COMM_NULL; + GRID_WORLD = MPI_COMM_NULL; + DIAG_WORLD = MPI_COMM_NULL; + ::testing::InitGoogleTest(&argc, argv); + const int result = RUN_ALL_TESTS(); + Parallel_Global::finalize_mpi(); + return result; +} diff --git a/source/source_lcao/pulay_fs.h b/source/source_lcao/pulay_fs.h index 3a8091df392..1a369a3dc49 100644 --- a/source/source_lcao/pulay_fs.h +++ b/source/source_lcao/pulay_fs.h @@ -46,7 +46,7 @@ namespace PulayForceStress /// for grid-integration terms template - void cal_pulay_fs( + void cal_pulay_fs(const int nspin, ModuleBase::matrix& f, ///< [out] force ModuleBase::matrix& s, ///< [out] stress const module_dm::DensityMatrix& dm, ///< [in] density matrix or energy density matrix diff --git a/source/source_lcao/pulay_fs_gint.h b/source/source_lcao/pulay_fs_gint.h index a2eab535178..26d935aa079 100644 --- a/source/source_lcao/pulay_fs_gint.h +++ b/source/source_lcao/pulay_fs_gint.h @@ -9,7 +9,7 @@ namespace PulayForceStress { template - void cal_pulay_fs( + void cal_pulay_fs(const int nspin, ModuleBase::matrix& f, ///< [out] force ModuleBase::matrix& s, ///< [out] stress const module_dm::DensityMatrix& dm, ///< [in] density matrix @@ -19,7 +19,6 @@ namespace PulayForceStress const bool& isstress, const bool& set_dmr_gint) { - const int nspin = PARAM.inp.nspin; std::vector vr_eff(nspin, nullptr); std::vector vofk_eff(nspin, nullptr); if (XC_Functional::get_func_type() == 3 || XC_Functional::get_func_type() == 5) diff --git a/source/source_pw/module_pwdft/force_pw.h b/source/source_pw/module_pwdft/force_pw.h index 6aeaff6069e..0cc4b56cdeb 100644 --- a/source/source_pw/module_pwdft/force_pw.h +++ b/source/source_pw/module_pwdft/force_pw.h @@ -32,6 +32,8 @@ class Forces friend class Force_Stress_LCAO; template friend class hamilt::Veff; + template + friend class ForcePWTerms; /* This routine is a driver routine which compute the forces * acting on the atoms, the complete forces in plane waves * is computed from 4 main parts diff --git a/tests/08_RI/CASES_CPU.txt b/tests/08_RI/CASES_CPU.txt index d94befe08e8..c970e418247 100644 --- a/tests/08_RI/CASES_CPU.txt +++ b/tests/08_RI/CASES_CPU.txt @@ -13,6 +13,7 @@ scf_out_xc_multik scf_campbeh_gamma scf_hse_soc_symm_multik lr_tddft_lda_gamma +lr_tddft_lda_s1_gamma lr_tddft_pbe_gamma lr_tddft_hf_gamma lr_tddft_hf_ulr_gamma diff --git a/tests/08_RI/lr_tddft_hf_gamma/INPUT b/tests/08_RI/lr_tddft_hf_gamma/INPUT index 5fbb2944db2..c52491b00d6 100644 --- a/tests/08_RI/lr_tddft_hf_gamma/INPUT +++ b/tests/08_RI/lr_tddft_hf_gamma/INPUT @@ -46,3 +46,4 @@ nocc 4 nvirt 2 abs_wavelen_range 40 180 abs_broadening 0.01 +cal_force 1 diff --git a/tests/08_RI/lr_tddft_hf_gamma/result.ref b/tests/08_RI/lr_tddft_hf_gamma/result.ref index 17fd842289c..7b6c6232562 100644 --- a/tests/08_RI/lr_tddft_hf_gamma/result.ref +++ b/tests/08_RI/lr_tddft_hf_gamma/result.ref @@ -1,5 +1,6 @@ +totallrforceref 121.6163457489 excitationenergyref1 1.531730 excitationenergyref2 1.532420 excitationenergyref3 1.353650 excitationenergyref4 1.354670 -totaltimeref 1.94 +totaltimeref 3.05 diff --git a/tests/08_RI/lr_tddft_hf_ulr_gamma/INPUT b/tests/08_RI/lr_tddft_hf_ulr_gamma/INPUT index 6c3da3549ac..790bcc8b757 100644 --- a/tests/08_RI/lr_tddft_hf_ulr_gamma/INPUT +++ b/tests/08_RI/lr_tddft_hf_ulr_gamma/INPUT @@ -42,3 +42,4 @@ esolver_type ks-lr nvirt 2 nocc 2 +cal_force 1 diff --git a/tests/08_RI/lr_tddft_hf_ulr_gamma/result.ref b/tests/08_RI/lr_tddft_hf_ulr_gamma/result.ref index e8aae3c7069..5b7164b15be 100644 --- a/tests/08_RI/lr_tddft_hf_ulr_gamma/result.ref +++ b/tests/08_RI/lr_tddft_hf_ulr_gamma/result.ref @@ -1,4 +1,5 @@ +totallrforceref 100.1957279900 excitationenergyref1 -0.981255 excitationenergyref2 -0.977372 excitationenergyref3 -0.765054 -totaltimeref 1.77 +totaltimeref 5.68 diff --git a/tests/08_RI/lr_tddft_lda_gamma/INPUT b/tests/08_RI/lr_tddft_lda_gamma/INPUT index df3c352eb4c..976722a548b 100644 --- a/tests/08_RI/lr_tddft_lda_gamma/INPUT +++ b/tests/08_RI/lr_tddft_lda_gamma/INPUT @@ -37,3 +37,5 @@ esolver_type ks-lr nvirt 2 abs_wavelen_range 40 180 abs_broadening 0.01 +# CI uses distribution libxc without kxc; gradient coverage is kept in HF cases. +cal_force 0 diff --git a/tests/08_RI/lr_tddft_lda_gamma/result.ref b/tests/08_RI/lr_tddft_lda_gamma/result.ref index 38e23a0cc97..a8fbd04efa2 100644 --- a/tests/08_RI/lr_tddft_lda_gamma/result.ref +++ b/tests/08_RI/lr_tddft_lda_gamma/result.ref @@ -1,5 +1,5 @@ -excitationenergyref1 0.587373 -excitationenergyref2 0.727934 -excitationenergyref3 0.531918 -excitationenergyref4 0.663441 -totaltimeref 1.74 +excitationenergyref1 0.587347 +excitationenergyref2 0.727911 +excitationenergyref3 0.531913 +excitationenergyref4 0.663432 +totaltimeref 2.55 diff --git a/tests/08_RI/lr_tddft_lda_s1_gamma/INPUT b/tests/08_RI/lr_tddft_lda_s1_gamma/INPUT new file mode 100644 index 00000000000..50ba3d57cc4 --- /dev/null +++ b/tests/08_RI/lr_tddft_lda_s1_gamma/INPUT @@ -0,0 +1,41 @@ +INPUT_PARAMETERS +#Parameters (1.General) +suffix autotest +pseudo_dir ../../../tests/PP_ORB +orbital_dir ../../../tests/PP_ORB +calculation scf +nbands 6 +symmetry -1 +nspin 1 + +#Parameters (2.Iteration) +ecutwfc 10 +scf_thr 1e-8 +scf_nmax 100 + +#Parameters (3.Basis) +basis_type lcao +gamma_only 1 + +#Parameters (4.Smearing) +smearing_method gaussian +smearing_sigma 0.02 + +#Parameters (5.Mixing) +mixing_type pulay +mixing_beta 0.4 +mixing_gg0 0 + +lr_nstates 2 +xc_kernel lda +lr_solver lapack +lr_thr 1e-8 +pw_diag_ndim 2 + +esolver_type ks-lr + +nvirt 2 +abs_wavelen_range 40 180 +abs_broadening 0.01 +# CI uses distribution libxc without kxc; gradient coverage is kept in HF cases. +cal_force 0 diff --git a/tests/08_RI/lr_tddft_lda_s1_gamma/KPT b/tests/08_RI/lr_tddft_lda_s1_gamma/KPT new file mode 100644 index 00000000000..c289c0158aa --- /dev/null +++ b/tests/08_RI/lr_tddft_lda_s1_gamma/KPT @@ -0,0 +1,4 @@ +K_POINTS +0 +Gamma +1 1 1 0 0 0 diff --git a/tests/08_RI/lr_tddft_lda_s1_gamma/README b/tests/08_RI/lr_tddft_lda_s1_gamma/README new file mode 100644 index 00000000000..7fd7114ec16 --- /dev/null +++ b/tests/08_RI/lr_tddft_lda_s1_gamma/README @@ -0,0 +1,2 @@ +LR-TDDFT LDA on H2O, gamma_only, nspin=1, lr_nstates=2, lr_solver=lapack. +Regression for the factor-two closed-shell singlet kernel; no kxc-dependent forces. diff --git a/tests/08_RI/lr_tddft_lda_s1_gamma/STRU b/tests/08_RI/lr_tddft_lda_s1_gamma/STRU new file mode 100644 index 00000000000..ceb8aaa84cb --- /dev/null +++ b/tests/08_RI/lr_tddft_lda_s1_gamma/STRU @@ -0,0 +1,29 @@ +ATOMIC_SPECIES +H 1.008 H_ONCV_PBE-1.0.upf +O 15.9994 O_ONCV_PBE-1.0.upf + +NUMERICAL_ORBITAL +H_gga_8au_60Ry_2s1p.orb +O_gga_7au_60Ry_2s2p1d.orb + + +LATTICE_CONSTANT +1 + +LATTICE_VECTORS +28 0 0 +0 28 0 +0 0 28 + +ATOMIC_POSITIONS +Cartesian + +H +0 +2 +-12.046787058887078 18.76558614676448 8.395247471328744 1 1 1 +-14.228868795885418 20.61549300274637 7.611989524516571 1 1 1 +O +0 +1 +-13.486789117423204 19.684192208418636 8.958321352749174 1 1 1 diff --git a/tests/08_RI/lr_tddft_lda_s1_gamma/result.ref b/tests/08_RI/lr_tddft_lda_s1_gamma/result.ref new file mode 100644 index 00000000000..c6718ab3cf9 --- /dev/null +++ b/tests/08_RI/lr_tddft_lda_s1_gamma/result.ref @@ -0,0 +1,3 @@ +excitationenergyref1 0.587550 +excitationenergyref2 0.728069 +totaltimeref 5.18 diff --git a/tests/08_RI/lr_tddft_pbe_gamma/INPUT b/tests/08_RI/lr_tddft_pbe_gamma/INPUT index 9641538f48b..fdb39f7499f 100644 --- a/tests/08_RI/lr_tddft_pbe_gamma/INPUT +++ b/tests/08_RI/lr_tddft_pbe_gamma/INPUT @@ -37,3 +37,5 @@ esolver_type ks-lr nvirt 2 abs_wavelen_range 40 180 abs_broadening 0.01 +# CI uses distribution libxc without kxc; gradient coverage is kept in HF cases. +cal_force 0 diff --git a/tests/08_RI/lr_tddft_pbe_gamma/result.ref b/tests/08_RI/lr_tddft_pbe_gamma/result.ref index a6e3b42a47b..b7e1521f9e3 100644 --- a/tests/08_RI/lr_tddft_pbe_gamma/result.ref +++ b/tests/08_RI/lr_tddft_pbe_gamma/result.ref @@ -1,5 +1,5 @@ -excitationenergyref1 0.589637 -excitationenergyref2 0.731349 -excitationenergyref3 0.526037 -excitationenergyref4 0.657779 -totaltimeref 1.82 +excitationenergyref1 0.589604 +excitationenergyref2 0.731319 +excitationenergyref3 0.526004 +excitationenergyref4 0.657739 +totaltimeref 2.81 diff --git a/tests/integrate/Autotest.sh b/tests/integrate/Autotest.sh index 48b6d2308f7..70f5f5fea93 100755 --- a/tests/integrate/Autotest.sh +++ b/tests/integrate/Autotest.sh @@ -13,6 +13,8 @@ njobs=1 # threshold with unit: eV threshold=0.0000001 force_threshold=0.0001 +# Absolute tolerance for the sum of excited-state force components. +lr_force_threshold=0.000001 stress_threshold=0.001 # descriptor mean threshold descriptor_threshold=0.00001 @@ -30,6 +32,7 @@ threshold_file="threshold" # threshold file example: # threshold 0.0000001 # force_threshold 0.0001 +# lr_force_threshold 0.000001 # stress_threshold 0.001 # fatal_threshold 1 @@ -116,6 +119,7 @@ echo "Number of threads: $nt" echo "Concurrent test cases: $njobs" echo "Test accuracy totenergy: $threshold eV" echo "Test accuracy force: $force_threshold" +echo "Test accuracy excited-state force: $lr_force_threshold" echo "Test accuracy stress: $stress_threshold" echo "Test accuracy descriptor mean: $descriptor_threshold" echo "Check accuaracy: $ca" @@ -148,6 +152,7 @@ check_out(){ stress_thr=$4 fatal_thr=$5 descriptor_thr=$6 + lr_force_thr=$7 #------------------------------------------------------ # outfile = result.out @@ -205,7 +210,9 @@ check_out(){ break else compare_thr=$thr - if [[ $key == ml_desc_mean_* ]]; then + if [ "$key" == "totallrforceref" ]; then + compare_thr=$lr_force_thr + elif [[ $key == ml_desc_mean_* ]]; then compare_thr=$descriptor_thr fi if [ $(check_deviation_pass $deviation $compare_thr) = 0 ]; then @@ -350,7 +357,8 @@ run_case() my_stress_threshold=$(get_threshold $threshold_file "stress_threshold" $stress_threshold) my_fatal_threshold=$(get_threshold $threshold_file "fatal_threshold" $fatal_threshold) my_descriptor_threshold=$(get_threshold $threshold_file "descriptor_threshold" $descriptor_threshold) - check_out result.out $my_threshold $my_force_threshold $my_stress_threshold $my_fatal_threshold $my_descriptor_threshold + my_lr_force_threshold=$(get_threshold $threshold_file "lr_force_threshold" $lr_force_threshold) + check_out result.out $my_threshold $my_force_threshold $my_stress_threshold $my_fatal_threshold $my_descriptor_threshold $my_lr_force_threshold fi else bash -e ../../integrate/tools/catch_properties.sh result.ref diff --git a/tests/integrate/tools/catch_properties.sh b/tests/integrate/tools/catch_properties.sh index 94b4f526119..fbae5b837cb 100755 --- a/tests/integrate/tools/catch_properties.sh +++ b/tests/integrate/tools/catch_properties.sh @@ -170,7 +170,7 @@ fi # force information # echo "hasforce:"$has_force #---------------------------- -if ! test -z "$has_force" && [ $has_force == 1 ]; then +if ! test -z "$has_force" && [ $has_force == 1 ] && [ $is_lr == 0 ]; then nn3=`echo "$natom + 3" |bc` # echo "nn3=$nn3" # check the last step result @@ -180,6 +180,22 @@ if ! test -z "$has_force" && [ $has_force == 1 ]; then echo "totalforceref $total_force" >>$1 fi +#---------------------------- +# excited-state force (LR-TDDFT analytic gradients) +# ESolver_LR prints its own "Forces (-gradients) of each excited +# state" table instead of the ground-state TOTAL-FORCE one, so it +# needs a separate extraction: pull the 3 numbers that follow every +# literal "force" token, for every state (and every spin channel, for +# open-shell runs where the table is printed once per channel). +#---------------------------- +if [ $is_lr == 1 ] && ! test -z "$has_force" && [ $has_force == 1 ]; then + awk '/Forces \(-gradients\) of each excited state/{flag=1; next} flag{for(i=1;i<=NF;i++) if($i=="force"){print $(i+1),$(i+2),$(i+3)}}' $running_path > lr_force.txt + # Accumulate all printed components before rounding the LR force total. + total_lr_force=$(awk '{for (i=1; i<=NF; ++i) sum += sqrt($i*$i)} END {printf "%.10f\n", sum}' lr_force.txt) + rm lr_force.txt + echo "totallrforceref $total_lr_force" >>$1 +fi + #------------------------------- # stress information # echo "has_stress:"$has_stress