Skip to content

Julia 1.12's LLVM SLP vectorizer miscompiles calc_forces! for Dual{10} on AVX-512 Zen 4/5 under --check-bounds=yes (JuliaLang/julia#62368) #379

Description

@1-Bort-1

On Julia 1.12 with --check-bounds=yes (what Pkg.test uses), linearize's default AutoForwardDiff() gives a Jacobian about 5% off on AMD Zen 4/5 CPUs that expose AVX-512. The cause is LLVM 18's SLP vectorizer, which miscompiles calc_forces!'s per-panel force assembly for 10-partial Duals. That is JuliaLang/julia#62368: the LLVM fix (llvm/llvm-project@5d7cf504) is identified there but not yet backported to Julia 1.12's LLVM. This bug is what #360's flake was.

Measured in four diagnostic CI rounds on Julia 1.12.7, 16 runners each, running repro379.jl below against main (a324968):

runner --check-bounds=yes default bounds checks --check-bounds=yes, JULIA_LLVM_ARGS=-vectorize-slp=false
EPYC 9V74 (znver4) / 9V45 (znver5) with +avx512f rel. 0.042 or 0.050: 7 of 7 in the rounds that logged CPU features 0 3e-19
EPYC 9V74 with AVX-512 masked (-avx512f) 3e-19 0 3e-19
EPYC 7763 (znver3), Intel Sapphire Rapids / Granite Rapids / Ice Lake 3e-19 0 3e-19
  • It happens only with --check-bounds=yes: coverage alone and default bounds checking give correct Jacobians. Ordinary runs are not affected.
  • Julia 1.13 (LLVM 20) was correct on the same CPUs.
  • The miscompile needs the whole loop. Leaving out either the moment_dist[i] line or the m_body_3D store fixes it, and so does moving the loop into its own function or into a standalone script without VortexStepMethod's types. So there is no package-free reproducer yet.
repro379.jl (run with julia --check-bounds=yes --project=. in a VortexStepMethod checkout)
using VortexStepMethod, DifferentiationInterface, LinearAlgebra

function affine_matrix_wing()
    alphas = deg2rad.(-5:5:25)
    deltas = deg2rad.(-3:3:3)
    cl = [0.2 + 5.5alpha + 1.5delta for alpha in alphas, delta in deltas]
    cd = [0.03 + 0.2alpha + 0.05delta for alpha in alphas, delta in deltas]
    cm = [-0.05 - 0.1alpha - 0.3delta for alpha in alphas, delta in deltas]
    wing = Wing(8, spanwise_distribution=LINEAR)
    for phi in deg2rad.((50, 17, -17, -50))
        le = [0.0, 3.0 * sin(phi), 3.0 * (cos(phi) - 1)]
        add_section!(wing, le, le .+ [1.0, 0.0, 0.0], POLAR_MATRICES,
            (alphas, deltas, cl, cd, cm))
    end
    refine!(wing)
    return wing
end

wing = affine_matrix_wing()
body = BodyAerodynamics([wing])
solver = Solver(wing.n_panels, wing.n_unrefined_sections; aerodynamic_model_type=VSM,
    rtol=1e-11, solver_type=LOOP, use_gamma_prev=false)
aoa = deg2rad(7.5)
y = [zeros(4); [cos(aoa), 0.0, sin(aoa)] * 15.0; zeros(3)]
jacobian(chunk) = first(VortexStepMethod.linearize(solver, body, y;
    theta_idxs=1:4, va_vec_idxs=5:7, omega_idxs=8:10, aero_coeffs=true,
    backend=AutoForwardDiff(chunksize=chunk)))
jac10 = jacobian(10)
jac5 = jacobian(5)
println("rel=", maximum(abs.(jac10 .- jac5)) / maximum(abs, jac5))

#378 works around it: the per-panel loads come from a @noinline panel_body_loads, and the POLAR_MATRICES testset checks that the Jacobian does not depend on the chunk size. With it, repro379.jl under --check-bounds=yes gives rel. 0 on three AVX-512 Zen runners (2× 9V74, 1× 9V45). Without it, main failed on 12 Zen runners across the four rounds, including all 7 in the two rounds that logged CPU features. This issue can close once Julia 1.12 ships the LLVM backport, and the @noinline can go then.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions