On Julia 1.12 with --check-bounds=yes (what Pkg.test uses), linearize's default AutoForwardDiff() gives a Jacobian about 5% off on AMD Zen 4/5 CPUs that expose AVX-512. The cause is LLVM 18's SLP vectorizer, which miscompiles calc_forces!'s per-panel force assembly for 10-partial Duals. That is JuliaLang/julia#62368: the LLVM fix (llvm/llvm-project@5d7cf504) is identified there but not yet backported to Julia 1.12's LLVM. This bug is what #360's flake was.
Measured in four diagnostic CI rounds on Julia 1.12.7, 16 runners each, running repro379.jl below against main (a324968):
| runner |
--check-bounds=yes |
default bounds checks |
--check-bounds=yes, JULIA_LLVM_ARGS=-vectorize-slp=false |
EPYC 9V74 (znver4) / 9V45 (znver5) with +avx512f |
rel. 0.042 or 0.050: 7 of 7 in the rounds that logged CPU features |
0 |
3e-19 |
EPYC 9V74 with AVX-512 masked (-avx512f) |
3e-19 |
0 |
3e-19 |
| EPYC 7763 (znver3), Intel Sapphire Rapids / Granite Rapids / Ice Lake |
3e-19 |
0 |
3e-19 |
- It happens only with
--check-bounds=yes: coverage alone and default bounds checking give correct Jacobians. Ordinary runs are not affected.
- Julia 1.13 (LLVM 20) was correct on the same CPUs.
- The miscompile needs the whole loop. Leaving out either the
moment_dist[i] line or the m_body_3D store fixes it, and so does moving the loop into its own function or into a standalone script without VortexStepMethod's types. So there is no package-free reproducer yet.
repro379.jl (run with julia --check-bounds=yes --project=. in a VortexStepMethod checkout)
using VortexStepMethod, DifferentiationInterface, LinearAlgebra
function affine_matrix_wing()
alphas = deg2rad.(-5:5:25)
deltas = deg2rad.(-3:3:3)
cl = [0.2 + 5.5alpha + 1.5delta for alpha in alphas, delta in deltas]
cd = [0.03 + 0.2alpha + 0.05delta for alpha in alphas, delta in deltas]
cm = [-0.05 - 0.1alpha - 0.3delta for alpha in alphas, delta in deltas]
wing = Wing(8, spanwise_distribution=LINEAR)
for phi in deg2rad.((50, 17, -17, -50))
le = [0.0, 3.0 * sin(phi), 3.0 * (cos(phi) - 1)]
add_section!(wing, le, le .+ [1.0, 0.0, 0.0], POLAR_MATRICES,
(alphas, deltas, cl, cd, cm))
end
refine!(wing)
return wing
end
wing = affine_matrix_wing()
body = BodyAerodynamics([wing])
solver = Solver(wing.n_panels, wing.n_unrefined_sections; aerodynamic_model_type=VSM,
rtol=1e-11, solver_type=LOOP, use_gamma_prev=false)
aoa = deg2rad(7.5)
y = [zeros(4); [cos(aoa), 0.0, sin(aoa)] * 15.0; zeros(3)]
jacobian(chunk) = first(VortexStepMethod.linearize(solver, body, y;
theta_idxs=1:4, va_vec_idxs=5:7, omega_idxs=8:10, aero_coeffs=true,
backend=AutoForwardDiff(chunksize=chunk)))
jac10 = jacobian(10)
jac5 = jacobian(5)
println("rel=", maximum(abs.(jac10 .- jac5)) / maximum(abs, jac5))
#378 works around it: the per-panel loads come from a @noinline panel_body_loads, and the POLAR_MATRICES testset checks that the Jacobian does not depend on the chunk size. With it, repro379.jl under --check-bounds=yes gives rel. 0 on three AVX-512 Zen runners (2× 9V74, 1× 9V45). Without it, main failed on 12 Zen runners across the four rounds, including all 7 in the two rounds that logged CPU features. This issue can close once Julia 1.12 ships the LLVM backport, and the @noinline can go then.
On Julia 1.12 with
--check-bounds=yes(whatPkg.testuses),linearize's defaultAutoForwardDiff()gives a Jacobian about 5% off on AMD Zen 4/5 CPUs that expose AVX-512. The cause is LLVM 18's SLP vectorizer, which miscompilescalc_forces!'s per-panel force assembly for 10-partial Duals. That is JuliaLang/julia#62368: the LLVM fix (llvm/llvm-project@5d7cf504) is identified there but not yet backported to Julia 1.12's LLVM. This bug is what #360's flake was.Measured in four diagnostic CI rounds on Julia 1.12.7, 16 runners each, running
repro379.jlbelow againstmain(a324968):--check-bounds=yes--check-bounds=yes,JULIA_LLVM_ARGS=-vectorize-slp=false+avx512f-avx512f)--check-bounds=yes: coverage alone and default bounds checking give correct Jacobians. Ordinary runs are not affected.moment_dist[i]line or them_body_3Dstore fixes it, and so does moving the loop into its own function or into a standalone script without VortexStepMethod's types. So there is no package-free reproducer yet.repro379.jl (run with
julia --check-bounds=yes --project=.in a VortexStepMethod checkout)#378 works around it: the per-panel loads come from a
@noinlinepanel_body_loads, and the POLAR_MATRICES testset checks that the Jacobian does not depend on the chunk size. With it,repro379.jlunder--check-bounds=yesgives rel. 0 on three AVX-512 Zen runners (2× 9V74, 1× 9V45). Without it,mainfailed on 12 Zen runners across the four rounds, including all 7 in the two rounds that logged CPU features. This issue can close once Julia 1.12 ships the LLVM backport, and the@noinlinecan go then.