Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
806 commits
Select commit Hold shift + click to select a range
718d16f
Ledger + constitution: P1 gate results (np=8 first completion, W8 hol…
sbryngelson Aug 22, 2026
190647c
Gather-batching design: S0 np=8 reverses the level-1/level-2 priority…
sbryngelson Aug 22, 2026
3a6f592
Derive the rebuild gather message set up front and assert the per-box…
sbryngelson Aug 22, 2026
a7a0d63
Gather-batching step-2 design: chunked pre-posted exchange, both fami…
sbryngelson Aug 22, 2026
1056e5c
Step-2 design review: fatal same-chunk parent-pack defect found by bo…
sbryngelson Aug 22, 2026
01cc431
Chunk the rebuild gather: pre-posted recvs, plan-driven sends, one wa…
sbryngelson Aug 22, 2026
3de4724
Arm the parent send-size assert on the chunked gather path
sbryngelson Aug 22, 2026
2505d9c
Record the step-2 verdict and the expert-audit re-aim in the AMR ledgers
sbryngelson Aug 22, 2026
7d8cc4a
Add hcid 306: the AMR benchmark blob as a hardcoded IC
sbryngelson Aug 22, 2026
7b7ae5c
Add the [amr-cov] dead-word counters for the gather families
sbryngelson Aug 22, 2026
47460b3
Record the [amr-cov] verdict: clipping promoted ahead of T1
sbryngelson Aug 22, 2026
81bea2e
Close G-B: the AMReX S0 weak-scaling bar is 1.20x/1.15x per np-doubling
sbryngelson Aug 22, 2026
93e5a7d
Ring-clip design for the runtime fill gathers, adversarially reviewed…
sbryngelson Aug 22, 2026
e53db27
Close ring-clip finding F5 by evidence: boundary patches are unreachable
sbryngelson Aug 22, 2026
dc6d412
Ring-clip the runtime fill gathers to the hollow shell
sbryngelson Aug 22, 2026
bd85c79
Stage ring-clip slab metadata through a device-resident buffer
sbryngelson Aug 22, 2026
a797074
Use full-column pool slices in the clipped packs
sbryngelson Aug 22, 2026
dcb1995
Ledger: ring clip landed and bit-correct, wall regression open
sbryngelson Aug 23, 2026
43ba113
Revert the ring clip: amdflang whole-image codegen regression
sbryngelson Aug 23, 2026
873cce6
Add the regrid-cadence containment audit ([amr-cad])
sbryngelson Aug 23, 2026
58c0b61
Right-size the migration pack pools and request array (T1/I4a)
sbryngelson Aug 23, 2026
de12b8d
Split rg:move with mg:slot/pack/unpk/push brackets (I4b pricing)
sbryngelson Aug 23, 2026
180cdf1
Pre-reserve the migration wave's stash slots in one exact-target grow…
sbryngelson Aug 23, 2026
c7de5b6
Ledger: I4b priced (growth, not pack), I4b-a landed, I4b-b deferred b…
sbryngelson Aug 23, 2026
0b36c14
Cadence move validated (int=20, -41% wall); ladder 1.392x/1.993x; rea…
sbryngelson Aug 23, 2026
b554aa2
Ledger: T1 re-priced at int=20 (I2 first); I1b-gather implementation …
sbryngelson Aug 23, 2026
e531d35
Add per-xfer identity headers to the gather trio (T1/I1b-gather)
sbryngelson Aug 23, 2026
bdb00d5
Convert the level-1 stage fill to plan-based waves (T1/I2a)
sbryngelson Aug 23, 2026
cdf78c1
Convert level>=2 parent gathers to per-level F2 waves (T1/I3)
sbryngelson Aug 24, 2026
5537999
Convert the seam-halo cross-rank pairs to per-peer waves (T1/I5-F6)
sbryngelson Aug 24, 2026
37a4d9b
Convert the reflux-face and freg exchanges to single waves (T1/I5-F5)
sbryngelson Aug 24, 2026
083c782
Ledger: the post-wave measurement (np8 -17.7 pct, top rung 1.99x -> 1…
sbryngelson Aug 24, 2026
5ac7b6f
Batch the reflux apply: one kernel per face direction over all level-…
sbryngelson Aug 24, 2026
62bf3e1
Move the regrid stash chain device-side (stash copy, migration pack/u…
sbryngelson Aug 24, 2026
a1ea597
Move the prolongation kernels device-side; delete the per-box slot pu…
sbryngelson Aug 24, 2026
eeee912
Grow the flat store device-natively (no PCIe round trip)
sbryngelson Aug 25, 2026
3c3bbc5
Instrument the stage-fill wave: gw:plan/gw:pack/gw:wait sub-brackets
sbryngelson Aug 25, 2026
c738918
Fix three review findings: IB merge separation, bounded grow transien…
sbryngelson Aug 25, 2026
34296de
Re-land the stepfill ring clip on the wave plan walks (F1 wire -61 pct)
sbryngelson Aug 25, 2026
0bd9222
Adopt the amdflang attributor-cap workaround at the offload link (mir…
sbryngelson Aug 25, 2026
1e3bc7f
Ledger (18): overnight pairs priced the attributor cliff; workaround …
sbryngelson Aug 25, 2026
1414285
Grow the reflux registers device-natively below a transient threshold…
sbryngelson Aug 25, 2026
55e391c
Ring-clip the parent-fill wave (F2): the largest wire family -54 pct,…
sbryngelson Aug 25, 2026
d08dd54
Ledger (21): first inter-node rung 1.594x vs bar 1.192x; restr and th…
sbryngelson Aug 25, 2026
3731eb6
Seam-clip the freg wave (F5b) and sub-bracket restr: sibling-seam fac…
sbryngelson Aug 25, 2026
570f299
Face-selective reflux multicast (F5a): each participant is shipped ex…
sbryngelson Aug 25, 2026
ee155dc
Ledger (25): third rung 1.368x; restr growth is the F7 per-box chain,…
sbryngelson Aug 26, 2026
bc91de5
Run the lock-step restrict fold as per-level waves (F7): the per-box …
sbryngelson Aug 26, 2026
b5194a6
Disable Euler-Euler bubbles under AMR (user decision, simplifies the …
sbryngelson Aug 26, 2026
969a7c4
Ledger (26): fourth rung 1.343x; the F7 chain is dead (restr -74% at …
sbryngelson Aug 26, 2026
3ed9f57
Phase 2a: batch the AMR fine-block cons->prim conversion to one launc…
sbryngelson Aug 26, 2026
70fea5e
2a verdict: default the batched conversion OFF; keep the machinery fo…
sbryngelson Aug 26, 2026
44edcc5
Defer the coexist L0 coarse RHS past the fine advance: its rhs values…
sbryngelson Aug 26, 2026
f30cc1e
Ledger (27): 2a priced and gated off (bridge-loads are the cost, 2b's…
sbryngelson Aug 26, 2026
c37dea4
Use associated(), not allocated(), on the scalar_field pointer member…
sbryngelson Aug 26, 2026
f81239a
Guard every wave-scratch deallocate individually: spsz/rpsz are sized…
sbryngelson Aug 26, 2026
c5abe1b
Ledger (28): fifth rung 1.241x - within 4% of the AMReX bar; coarse d…
sbryngelson Aug 26, 2026
5312e83
Merge upstream master d74cc378 into up/mega
sbryngelson Aug 26, 2026
1f38c02
Exclude Doxygen's dangling *_8md.html auto-links from the lychee check
sbryngelson Aug 26, 2026
ed74885
Ledger (29): master merged into up/mega, PR 1628 CI-live; hook-env li…
sbryngelson Aug 26, 2026
0820567
Merge branch 'master' into up/mega
sbryngelson Aug 27, 2026
0acad7b
Participation-local flux registers: dense register slots replace glob…
sbryngelson Aug 27, 2026
60b931f
Default template: pass --oversubscribe to Open MPI's mpirun
sbryngelson Aug 26, 2026
d9e0228
Ledger (30): participation-local registers landed; the one red CI tes…
sbryngelson Aug 27, 2026
4cff2ad
Ledger (31): sixth rung 1.274x vs the AMReX np32 bar 1.278x - at SOTA…
sbryngelson Aug 27, 2026
c5b461f
Fix nine confirmed findings from the full-PR review
sbryngelson Aug 27, 2026
24d0300
Ledger (32): full-PR review - nine confirmed findings fixed and gated…
sbryngelson Aug 27, 2026
67d92c1
Gate each side's hypoelastic interface energy on its own damage state
sbryngelson Aug 27, 2026
7e2428d
Retract the HLLD-hypoelasticity CBC prohibit
sbryngelson Aug 27, 2026
87c13a2
Ledger (33): elastic-gate per-side fix landed; one review prohibit re…
sbryngelson Aug 27, 2026
9c9e0a7
Merge branch 'master' into up/mega
sbryngelson Aug 27, 2026
98c5279
Ledger (34): merge audit and strategy - the performance program is co…
sbryngelson Aug 27, 2026
caabd06
Remove the last stage ifdefs from src/common
sbryngelson Aug 27, 2026
49dfb86
Ledger (35): the ladder is not the scorecard - W1/W2/W4/W7 unmet, D-l…
sbryngelson Aug 27, 2026
be33eac
Ledger (35) corrections: S1 is a coefficient fix, not the W4 fix; W4'…
sbryngelson Aug 27, 2026
82dd12f
Remove the checker prohibits the master merge resurrected
sbryngelson Aug 27, 2026
4e21d0f
Move the O(global boxes) regrid work arrays off the stack
sbryngelson Aug 27, 2026
f702fb2
Move the AMR input constraints into the Python validator
sbryngelson Aug 27, 2026
5ec798e
Build the seam-pair list by lookup instead of an all-pairs scan
sbryngelson Aug 27, 2026
619998b
Ledger (36) and constitution refresh: record the hygiene batch, resol…
sbryngelson Aug 27, 2026
42f2105
Ledger (37): the migration gate only tested one direction, and S3's d…
sbryngelson Aug 27, 2026
d4edbce
Measure the Berger-Rigoutsos tree shape and its reduction cost (S3.0a…
sbryngelson Aug 27, 2026
55f4001
Fail formatting when ffmt reports unmatched Fortran block structure
sbryngelson Aug 28, 2026
2b4da7d
Cluster from per-rank tag lists with a fused signature reduction (S3.1)
sbryngelson Aug 28, 2026
c468496
Bring the AMR plan documents current: ledger (38), verified invariant…
sbryngelson Aug 28, 2026
320fe10
Add a minimum box size (B0) and measure how much of the clustering tr…
sbryngelson Aug 28, 2026
e6a007b
Record the S3.2 design contract, the B1 merge-order prerequisite, and…
sbryngelson Aug 28, 2026
bdfc64d
Merge branch 'master' into up/mega
sbryngelson Aug 28, 2026
4f6ce85
Merge remote-tracking branch 'upstream/master' into up/mega
sbryngelson Aug 28, 2026
dc80eb6
Merge remote-tracking branch 'origin/up/mega' into up/mega
sbryngelson Aug 28, 2026
dc27e4a
Ledger (39): B0/S3.2a golden gate closed 69/69, and the np32 question…
sbryngelson Aug 28, 2026
eb6ba24
B1: canonicalise the clusterer's merge order, plus per-rank scope ins…
sbryngelson Aug 28, 2026
7251139
Ledger (40-41): the forest is the only O(P) term, and 'level-1 is fla…
sbryngelson Aug 28, 2026
58aa086
S3.3: cluster the level->=2 forest per owner, and exchange its tags p…
sbryngelson Aug 28, 2026
ee7758b
Ledger (42): level 1 is O(P) too, W1 is ~48 global block scans per ST…
sbryngelson Aug 28, 2026
263730c
W1a: iterate this rank's own blocks instead of scanning the global bl…
sbryngelson Aug 28, 2026
3cd90f2
Ledger (43): the MPI_TAG_UB wall is an assumed number, not a measured…
sbryngelson Aug 28, 2026
cefd317
Restart format v2: store the per-block owner and extents, not a per-R…
sbryngelson Aug 28, 2026
e20576c
Correct the MPI_TAG_UB claim in the source, and record the S3.2b/W5 f…
sbryngelson Aug 28, 2026
8bc027b
Ledger (44): AMR aborts on Frontier under CCE, pre-existing, and it o…
sbryngelson Aug 28, 2026
b055b0d
Give the Cray debug build bounds checking, which every other compiler…
sbryngelson Aug 28, 2026
8e10b44
Device-declare the module allocatables that @:ALLOCATE maps to the de…
sbryngelson Aug 28, 2026
fa20cba
Fix NVHPC build: a GPU_DECLARE must follow every symbol it names
sbryngelson Aug 28, 2026
4ade13c
Work around the CCE descriptor defect for the AMR scratch and batch a…
sbryngelson Aug 29, 2026
a72c125
Default amr_blocking_factor to 4 so the clusterer stops splitting int…
sbryngelson Aug 29, 2026
79fe684
Loosen the two churn-growth goldens to 1e-11, which CCE reassociation…
sbryngelson Aug 29, 2026
c09797a
Ledger: current state after B1, S3.3, W1a, restart v2 and the CCE fix
sbryngelson Aug 29, 2026
e2cb607
Let a rank-local clustering node skip its global reduction
sbryngelson Aug 29, 2026
299c368
Build the level-2 nesting coverage from owned blocks and one reduction
sbryngelson Aug 29, 2026
4a2edfb
Walk the clustering tree in level order so one depth needs one reduction
sbryngelson Aug 29, 2026
b865a9a
Reduce a narrow clustering node among the ranks it spans, not the who…
sbryngelson Aug 29, 2026
f0cde87
Ledger: S3.2b-2 landed, the W4 gate defect, item P restated, the box …
sbryngelson Aug 29, 2026
2fedf22
Poison a reflux register only when this rank holds one
sbryngelson Aug 29, 2026
708db3c
Loosen the churn goldens to 1e-9, the scale of their numerically-zero…
sbryngelson Aug 29, 2026
0187c68
Drop the GPU declares added for a CCE abort they did not fix
sbryngelson Aug 29, 2026
0fb9752
Count the bytes each rank receives from the two box-list gathers
sbryngelson Aug 29, 2026
00467d9
Iterate the recorded receive list instead of rescanning every block
sbryngelson Aug 29, 2026
9fd9592
Revert "Iterate the recorded receive list instead of rescanning every…
sbryngelson Aug 29, 2026
89f1b94
Reapply "Iterate the recorded receive list instead of rescanning ever…
sbryngelson Aug 29, 2026
9e5bdea
Drop the churn tolerance override: the failure is a flipped tag, not …
sbryngelson Aug 29, 2026
8f5d950
Drop the GPU_DECLARE on the move_alloc'd AMR device arrays (fixes Fro…
sbryngelson Aug 30, 2026
97eff75
Bracket the base-grid halo exchange as its own phase
sbryngelson Aug 30, 2026
007dd8c
Add the halo, memory and grid-efficiency scaling probes
sbryngelson Aug 30, 2026
d2accee
Bound the churn test's amplification window instead of chasing the th…
sbryngelson Aug 30, 2026
06f32d8
Match the dual-pass flux teardown guard to its allocation guard
sbryngelson Aug 30, 2026
45ba759
Zero flux_gsrc_hatR_rsx_vf at allocation
sbryngelson Aug 30, 2026
bf3de26
Zero pc_iter_count on host and device at allocation
sbryngelson Aug 30, 2026
15f0ddf
Fail closed in post_process on v2 AMR restart files
sbryngelson Aug 30, 2026
c2f3213
Drop the dead v2 guard from the serial post reader
sbryngelson Aug 30, 2026
68bcaf1
Check post_process's exit code in the test suite
sbryngelson Aug 30, 2026
c64351a
Read v2 AMR restart files in post_process
sbryngelson Aug 30, 2026
441d844
Reduce the grid-efficiency numerator across ranks
sbryngelson Aug 30, 2026
358ac1c
Skip L0 tiles in the post_process AMR reader instead of aborting
sbryngelson Aug 30, 2026
d705abb
Add the post-pad footprint counter and warn on silent box-set truncation
sbryngelson Aug 30, 2026
1b56e95
Correct three stale subcycle claims
sbryngelson Aug 30, 2026
1afba26
Ledger (44): two expert reviews reorder the program
sbryngelson Aug 30, 2026
51778c6
Skip skipped save indices in post_process instead of aborting
sbryngelson Aug 30, 2026
e860a5f
Exempt the CCE-only tracer-bubble failure from --test-all, tracked in…
sbryngelson Aug 30, 2026
b5c7ee4
Bracket the subcycle advance path with the lock-step phase ids
sbryngelson Aug 31, 2026
1301a91
Ledger (45): the cadence result was a broken control; multi-node was …
sbryngelson Aug 31, 2026
abb8574
Ledger (46): subcycle parity at matched fidelity; binned merge in, bi…
sbryngelson Aug 31, 2026
4d5f6ee
Use single-precision post output on --single builds in the test suite
sbryngelson Aug 31, 2026
c881e4e
Binned candidate merge: O(1)-removal survivor list + spatial prune, b…
sbryngelson Aug 31, 2026
08a8f49
Budget the store-growth device transient in bytes, not columns
sbryngelson Aug 31, 2026
805f26e
Fix single-precision post output at the site the test path actually uses
sbryngelson Aug 31, 2026
1e57e79
Count amr_slots in the [amr-mem] replicated-footprint report; ledger …
sbryngelson Aug 31, 2026
2f401b0
SUM-reduce the migration counters globally; ledger (48): four scaling…
sbryngelson Aug 31, 2026
3eeb080
Dirty-box merge continuation: eliminate the per-fusion pass restart
sbryngelson Sep 1, 2026
5f57b03
Revert the private+map(alloc) clause overlap (fixes the NVHPC gpu-omp…
sbryngelson Sep 1, 2026
17db501
Replace the cluster sort's insertion sort with a stable bottom-up mer…
sbryngelson Sep 1, 2026
2c4b2f4
Keyed-tags M0: an always-on order oracle over the F5 wave exchanges
sbryngelson Sep 1, 2026
c8c4f2c
Ledger (49): four landings, the oracle, and the phase transition
sbryngelson Sep 1, 2026
acaae35
W1 batch 1: epoch-keyed level-1 receive list replaces the fill-wave r…
sbryngelson Sep 1, 2026
dcfce95
W1 batch 2: padded receive list + two owned-list conversions
sbryngelson Sep 1, 2026
75a1220
Ledger (50): W1 underway; P-prime interim and the misanchored bands
sbryngelson Sep 1, 2026
1212778
W1 batch 3: the last six per-stage scans, and the parent-index cache
sbryngelson Sep 1, 2026
2a57b32
Ledger (51): W1 batch 3 lands; np1024 postmortem memory-gates the rung
sbryngelson Sep 1, 2026
0f6e9f0
M1 on family F5: keyed wave tags with plan-derived sequences
sbryngelson Sep 1, 2026
025da79
Ledger (52): the proper two-code tax test; MFC AMR overhead below AMR…
sbryngelson Sep 2, 2026
9d0ec27
Ledger (53): np512 A/B closed at -5.0%; gain is the dirty-box merge, …
sbryngelson Sep 2, 2026
7a5fbd1
Ledger (53): probe verdict, k002-005 is the slow node
sbryngelson Sep 2, 2026
4bd1747
Ledger (53): the clean np512 rung closes the constant-density ladder …
sbryngelson Sep 2, 2026
3368b16
Ledger (53): name the 256->512 growth term (regrid rebuild)
sbryngelson Sep 2, 2026
fd9f242
Compile only the Riemann solver the case selects under case optimization
sbryngelson Sep 2, 2026
e9e80c9
Ledger (54): expert review overturns ledger 52's inference; regrid is…
sbryngelson Sep 2, 2026
948298e
Ledger (55): AMReX at its own GPU-sane grids taxes 7.7-8.5x, not 20.5x
sbryngelson Sep 2, 2026
1c1658f
AMR: abort at init when a refined block's fine extent exceeds the ran…
sbryngelson Sep 2, 2026
49473b0
Ledger (56): fixture pointers
sbryngelson Sep 2, 2026
52d2469
Regrid rebuild: print the xchg-flag arrival skew and collective time …
sbryngelson Sep 3, 2026
594fcaa
Ledger (57): tax replication 12.4x +/-2% on k004-003, w1 control 24.1…
sbryngelson Sep 3, 2026
7934718
Ledger (58): Task 3 review corrections + steady-state AMR-arm profile…
sbryngelson Sep 3, 2026
5651257
Ledger (58) correction: coarse phase is the level-0 rhs; exchanges 0.…
sbryngelson Sep 3, 2026
7bb359a
AMR: fill the WENO coefficient tail past the coarse subdomain (fixes …
sbryngelson Sep 3, 2026
6b8d631
AMR: size the dual-pass and NC-interface flux scratch to idwbuff_alloc
sbryngelson Sep 3, 2026
d6c73c9
AMR: delete the pinned-cap init guard; the widened scratch is now com…
sbryngelson Sep 3, 2026
3afefa3
AMR: 3D np=8 golden with the box cap pinned above a rank's coarse ext…
sbryngelson Sep 3, 2026
889b2d6
AMR: trim the WENO tail-fill comment
sbryngelson Sep 3, 2026
d4c50b8
Ledger (59): describe the committed 20-step golden
sbryngelson Sep 3, 2026
a94607d
W1 batch 4: reflux-to-parent walks the owned/foreign-child union, not…
sbryngelson Sep 3, 2026
1905d65
W1 batch 4: level>=2 relax loop walks the owned list
sbryngelson Sep 3, 2026
17706eb
W1 batch 4: level-1 relax loop walks the owned list
sbryngelson Sep 3, 2026
bb197c0
Ledger (60): rdma_mpi under OpenMP (-4%), device pools falsified, the…
sbryngelson Sep 3, 2026
03b6c27
Bracket-free MPI-wait instrument: [mpiwait] table under rank_time_wrt
sbryngelson Sep 4, 2026
f123623
Allow rdma_mpi under OpenMP offload: the checker gate predates the OM…
sbryngelson Sep 3, 2026
62e45e7
Ledger (61): the analytic-IC pre_process trap, the bracket-free MPI-w…
sbryngelson Sep 4, 2026
ffdfd18
Ledger (62): Task 4 concluded (regrid arrival skew 1.9x/doubling), th…
sbryngelson Sep 4, 2026
ed2d06f
Ledger (63): Task 9 count gate met (regrid doubling 1.94x -> 1.32x, r…
sbryngelson Sep 4, 2026
56138f8
Audit: MFC_XA_SEED_FAM aims the seeded fold at one family's first key…
sbryngelson Sep 4, 2026
feacbe3
M1 on family F2W: keyed parent-fill wave tags
sbryngelson Sep 4, 2026
11184f8
M1 on families F1W/F3W: keyed stage-fill wave tags (bands 3 and 4)
sbryngelson Sep 4, 2026
8731e19
M1 on family F6W: keyed fine-fine halo wave tags
sbryngelson Sep 4, 2026
bbd566c
M1 on family F7W: keyed level-1 restrict wave tags
sbryngelson Sep 4, 2026
d8633d8
M1 on family F7BW: keyed parent restrict wave tags
sbryngelson Sep 4, 2026
081ab44
M1 comments: amr_tag_base survives only for regrid migration; freg wa…
sbryngelson Sep 4, 2026
f0db26e
Ledger (64): np8 redone (5.69 s/step), the np16 rung never ran and it…
sbryngelson Sep 4, 2026
78c4b60
Merge upstream/master into up/mega: adopt the centralized Riemann EOS…
sbryngelson Sep 4, 2026
307b62a
Merge remote-tracking branch 'origin/merge/master-0904' into up/mega
sbryngelson Sep 4, 2026
f17a9aa
Merge remote-tracking branch 'upstream/master' into up/mega
sbryngelson Sep 4, 2026
43234ac
Ledger (65): master merged and gated on the combination; the batched …
sbryngelson Sep 4, 2026
c1f859b
Ledger (66): retract the restart-metadata padding finding; the compar…
sbryngelson Sep 4, 2026
ce37c52
AMR: the lock-step fine advance walks the owned-block list (W1 leftov…
sbryngelson Sep 4, 2026
8806561
AMR: batched fine advance behind amr_batched_advance (stacked bridge,…
sbryngelson Sep 4, 2026
525d4b7
AMR: L0 tile migration marks the owned-block list dirty after writing…
sbryngelson Sep 4, 2026
b0f601c
AMR: amr_batched_advance review fixes (abort on the coefficient-recom…
sbryngelson Sep 4, 2026
c60ce8a
Ledger (67): the batched fine advance is merged behind a default-off …
sbryngelson Sep 5, 2026
f1f8410
AMR: report the batched-advance batch population as [amr-bat] under r…
sbryngelson Sep 5, 2026
32e66e6
Ledger (68): the MI210 GPU ladder, its MPI-wait split, and the mechan…
sbryngelson Sep 5, 2026
4799a4c
Ledger (69): goal v2 -- gated increments are pushed the session they …
sbryngelson Sep 5, 2026
13a18f0
Merge branch 'task10/batchcount' into up/mega
sbryngelson Sep 5, 2026
704582d
Walk the rebuild box loop over this rank's participants, not every bo…
sbryngelson Sep 4, 2026
b188e75
Walk the rebuild's old-block loops over the stashes this rank holds
sbryngelson Sep 4, 2026
560f21c
Check seam topology from this rank's owned blocks, not all pairs
sbryngelson Sep 4, 2026
e2fc388
Print per-rank seconds for the regrid sub-phases in the phase-rank table
sbryngelson Sep 4, 2026
1a4344d
Drop the participant role array: the consumers' own predicates alread…
sbryngelson Sep 4, 2026
83e484a
Ledger (70): load balance replayed offline at zero cost -- do not imp…
sbryngelson Sep 5, 2026
082f65f
Merge branch 'task9/rebuild' into up/mega
sbryngelson Sep 5, 2026
5f2e184
Ledger (71): Task 9 merged -- the regrid rebuild's O(P) rows fall fro…
sbryngelson Sep 5, 2026
3ed5bae
Add the amr_device_pack case flag (default F): requires amr, excludes…
sbryngelson Sep 5, 2026
2f650c3
F1/F2 coarse-patch gather: one fused pack/unpack kernel per family pe…
sbryngelson Sep 5, 2026
f64fea7
Ledger (72): the controlled ladder and the MI250X A/B, reviewed -- ba…
sbryngelson Sep 5, 2026
5e0f5ea
Ledger (73): the steady AMR excess is 1.44 s/step (2.1x target) with …
sbryngelson Sep 5, 2026
80225d4
Ledger (74): the step is only ~26% MPI wait (a lower bound), and the …
sbryngelson Sep 5, 2026
d238236
Merge branch 'up/mega' into task10/fusedpack
sbryngelson Sep 5, 2026
b126ee8
Ledger (75, 76): the fused gather packs are worth 0.14 s/step with 12…
sbryngelson Sep 5, 2026
23800cf
AMR: correct the amr_device_pack description (sends fuse per family, …
sbryngelson Sep 5, 2026
a11b4fe
Merge task10/fusedpack: fused F1/F2 gather packs behind the default-o…
sbryngelson Sep 5, 2026
3e208d3
Ledger (75): record the gates the fused-pack merge passed, and the 3 …
sbryngelson Sep 5, 2026
2d3381c
Ledger (76): the 8-GPU/node first doubling is 1.30x -- the 21 percent…
sbryngelson Sep 5, 2026
a4618f6
Ledger (77): my own hypothesis falsified -- the allocator setting rec…
sbryngelson Sep 5, 2026
e36a680
Ledger (75): all 70 AMR goldens pass on the merged tree -- the 3 chem…
sbryngelson Sep 5, 2026
4b534eb
Ledger (78): this node's intra-node MPI wait degraded 4.4x during the…
sbryngelson Sep 5, 2026
cede444
Ledger (78): name the canary script and record that its first design …
sbryngelson Sep 5, 2026
7babe17
AMR: delete amr_rg_gather and its 35 unreachable sites -- a flag noth…
sbryngelson Sep 5, 2026
97eedbb
Ledger (80): per-block cost is ~13 ms/block/step across six phases, p…
sbryngelson Sep 5, 2026
04fb2f0
Ledger (81): negative, pre-registered, falsifier fired -- pooling the…
sbryngelson Sep 5, 2026
9002f3c
AMR instrument: five bracket-free host-time rows (h:slot/shell/own/un…
sbryngelson Sep 5, 2026
8dc669a
AMD OpenMP lane: per-file opt-in defaultmap(present:allocatable) (MFC…
sbryngelson Sep 5, 2026
00caa29
Ledger (82): the per-block AMR cost is amdflang's per-launch re-map o…
sbryngelson Sep 5, 2026
256355c
Ledger (82) correction: unallocated module arrays abort under present…
sbryngelson Sep 5, 2026
a235b5a
AMD OpenMP lane: m_amr_registers.fpp opts in to defaultmap(present:al…
sbryngelson Sep 5, 2026
22b4fba
m_amr_registers opt-in header: fypp comments only (the formatter had …
sbryngelson Sep 5, 2026
11f4a77
Ledger (83): m_amr_registers opted in -- at most 2% at cap 32, nothin…
sbryngelson Sep 5, 2026
b5b1782
AMR: one device-resident slab table + one GPU_UPDATE per launch for t…
sbryngelson Sep 5, 2026
daaa80c
Ledger (84): one device slab table + one update per launch, pre-regis…
sbryngelson Sep 5, 2026
55c735d
AMR: amr_batched_gather (default F) -- pool the gathered coarse patch…
sbryngelson Sep 5, 2026
5ee8e1d
AMR: amr_batched_gather -- wire the F2 parent-fill wave to the pooled…
sbryngelson Sep 5, 2026
f920995
AMR: amr_batched_gather -- keep the per-member tables host-only so co…
sbryngelson Sep 5, 2026
f1510c7
AMR: amr_batched_gather -- the pooled unpack copies in only this wave…
sbryngelson Sep 5, 2026
ab091b8
amr_batched_gather rebase: restore the four preprocessor directives t…
sbryngelson Sep 6, 2026
6ddd8f1
Ledger (85): ledger 81 re-tested under the clause -- the pooled gathe…
sbryngelson Sep 6, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
38 changes: 38 additions & 0 deletions .claude/rules/common-pitfalls.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,8 +19,37 @@ covered in `docs/documentation/contributing.md`.
`contxb`/`momxb` shorthands are gone. Index positions depend on `model_eqns` and
enabled features — changing either moves ALL indices; never hard-code one.

## AMR levels (silent-index traps)

- **A level-`l` block's fine extent is `amr_ref_ratio**l * (coarse-region width) - 1`, NOT
`amr_ref_ratio*width`.** The `amr_ref_ratio*width` form is correct only for the level-1 initial
block; nested boxes compound by `amr_ref_ratio` per level (`amr_ref_ratio**level`). Every
fine-extent computation uses `amr_ref_ratio**amr_block_level` — geometry
(`s_set_amr_fine_geometry`), the restart-reader extent check, load-weight, `fmul`.
Assuming `amr_ref_ratio*width` rejects level≥2 blocks as corrupt (the exact bug that bit the
multi-level restart reader).
- **"coarse" in the AMR coupling routines means the block's PARENT level (`l-1`), not the
base grid (level 0).** For a level-1 block the parent IS L0; for level≥2 the block folds
to/from its parent block's fine array. `s_amr_gather_coarse_patch`,
`s_interpolate_coarse_to_fine`, and the restrict/reflux path all operate in the
parent-fine frame — assuming L0 silently corrupts level≥2 coupling.
- **The fine advance SWAPS the coarse grid globals (`m/n/p`, `idwint/idwbuff`, coords,
`acoustic_source`, `ab_active`) to a fine block and restores them after — see the SWAP
CONTRACT block at the `sw_*` declarations in `m_amr.fpp`.** Any module-level variable
DERIVED from the grid that a kernel reads during the fine advance must be swapped there or
refreshed per fine call at its use site; if it is `GPU_DECLARE`'d, its DEVICE copy must be
refreshed too. A stale device copy of coarse bounds reads out of range on the fine grid
under **CCE OpenACC only** (NVHPC/CCE-omp evaluate bounds host-side) — this was the `ab_int`
regression, fixed by an unconditional `GPU_UPDATE` in `s_compute_rhs`. `amr_rvw` (cyl_coord
radius weights) is the next candidate, currently safe only via a `m_checker.fpp` gate.
A CPU-only or NVHPC-acc pass proves NOTHING here; this class is CCE-acc-specific.

## GPU

- NEVER put a `GPU_PARALLEL_LOOP` inside a Fortran `block` construct: amdflang compiles
it clean but silently DROPS the region from the device image — the first launch dies
with `HSA_STATUS_ERROR_INVALID_SYMBOL_NAME` naming an `__omp_offloading_*` symbol.
Hoist the kernel into its own module subroutine.
- WARNING: do NOT wrap `GPU_LOOP` in `GPU_PARALLEL` for spatial loops — `GPU_LOOP` emits
empty directives on Cray and AMD, causing silent serial execution. Spatial loops always
use `GPU_PARALLEL_LOOP`/`END_GPU_PARALLEL_LOOP`. Macro API:
Expand All @@ -36,6 +65,15 @@ covered in `docs/documentation/contributing.md`.
- `@:ACC_SETUP_VFs(...)`/`@:ACC_SETUP_SFs(...)` GPU pointer setup compiles only under
Cray. Around MPI: `GPU_UPDATE(host=...)` before send, `GPU_UPDATE(device=...)` after
receive.
- **Never `GPU_UPDATE` a NON-CONTIGUOUS array section.** `GPU_UPDATE(device='[q%sf(a:b,
c:d, e:f)]')` on a sub-box emits correct OpenMP, but AMD flang copies it as
`size(section)` CONTIGUOUS elements starting at the first: only the leading run lands
where it is named and the rest overwrites neighbouring cells with stale data — no error,
no warning. A leading section (`arr(1:n)`, or a fixed trailing index like
`freg(d)%lo(:,:,:,k)`) IS contiguous and safe; anything that strides is not. To move a
sub-box, pack/unpack it with a device kernel (`s_l0_pack_unpack_block`,
`s_amr_restrict_pack_device`) — that is why those exist. Measured: 10 of 60 covered
cells delivered in the AMR cross-rank restrict, mass off 1.4e-5 per regrid.
- An array whose bound is a device global (`dimension(num_fluids)`, `dimension(num_species)`) may be
passed to a device routine **from a parallel-loop body, but not from inside another
`GPU_ROUTINE(parallelism='[seq]')`**. CCE OpenACC rejects the second form with
Expand Down
1 change: 1 addition & 0 deletions .lychee.toml
Original file line number Diff line number Diff line change
Expand Up @@ -33,4 +33,5 @@ exclude = [
"https://code\\.visualstudio\\.com/?$", # Root page returns 403 to automated requests
"https://stackoverflow\\.com", # Returns 403 to automated requests
"https://marketplace\\.visualstudio\\.com", # Returns 503 to automated requests
"_8md\\.html$", # Doxygen auto-links backticked *.md filenames in prose to per-file pages it never generates for markdown inputs; the real md_*.html page links are still checked
]
6 changes: 6 additions & 0 deletions .typos.toml
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,9 @@ extend-ignore-identifiers-re = [
AttributeIDSupressMenu = "AttributeIDSupressMenu"

[default.extend-words]
# Cray CCE spells it this way in the lib-4425 runtime error; quoted verbatim in the AMR ledger so the
# message stays greppable against what the machine actually prints.
Unitialized = "Unitialized"
INOUT = "INOUT"
WRONLY = "WRONLY"
nd = "nd"
Expand All @@ -22,6 +25,9 @@ TKE = "TKE"
HSA = "HSA"
infp = "infp"
Sur = "Sur"
thi = "thi" # AMR clustering local: tagged-box hi index (tlo/thi)
alo = "alo" # AMR clustering local: accepted-box lo array (alo/ahi)
thr = "thr" # AMR clustering local: min-separation merge threshold
equil = "equil" # abbreviation for "equilibrium" (flamelet chemistry)
chioces = "chioces" # typo for "choices" - tests constraint key validation
reqires = "reqires" # typo for "requires" - tests dependency key validation
Expand Down
6 changes: 6 additions & 0 deletions cmake/GPU.cmake
Original file line number Diff line number Diff line change
Expand Up @@ -90,9 +90,15 @@ elseif (CMAKE_Fortran_COMPILER_ID STREQUAL "Cray")
add_link_options("SHELL:-hkeepfiles")

if (CMAKE_BUILD_TYPE STREQUAL "Debug")
# -h bounds: array-bounds and pointer checking, the Cray equivalent of gfortran's
# -fcheck=bounds,pointer / Intel's -check bounds / NVHPC's -Mbounds, all of which the
# debug branches above already set. Cray was the ONLY compiler whose debug build had no
# bounds checking, so an out-of-bounds write showed up here only as a later, unrelated
# allocation failing with an uninitialised descriptor.
add_compile_options(
"SHELL:-h acc_model=auto_async_none"
"SHELL: -h acc_model=no_fast_addr"
"SHELL: -h bounds"
"SHELL: -K trap=fp" "SHELL: -g" "SHELL: -O0"
)
add_link_options("SHELL: -K trap=fp" "SHELL: -g" "SHELL: -O0")
Expand Down
Loading
Loading