Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
52 commits
Select commit Hold shift + click to select a range
da35285
Level 3 first batch: audit, build strategy, LAMMPS/SPARTA/WarpX/SPECF…
Sep 5, 2026
366b72f
Level 3 first batch: correctness/reproducibility fixes and nekRS coar…
Sep 5, 2026
9327c26
Level 3 tools: per-profile paths and static-cudart backend check for …
Sep 6, 2026
769482f
Level 3 second batch: Nyx 26.09 (CUDA sm_100) built, validated at 1/2…
Sep 6, 2026
0b9daaf
Level 3 tools: clear conda build variables for system-toolchain build…
Sep 6, 2026
1f7e14f
Nyx: heating/cooling variant (SUNDIALS 7.2.1 CUDA) validated at 1/2/4…
Sep 6, 2026
b04db93
Level 3 second batch: CP2K v2026.2 (CUDA sm_100, DBCSR 2.10.0) built,…
Sep 6, 2026
352cbb3
Level 3 tools: negative/positive tests for the second-batch applicati…
Sep 6, 2026
166657e
Level 3 second batch: QMCPACK 4.4.0 (OpenMP offload + CUDA, sm_100) w…
Sep 6, 2026
32e25e9
Level 3 tools: clear the login shell's foreign include/library search…
Sep 6, 2026
62f48e5
Level 3 second batch: DFT-FE 1.2.0 (CUDA sm_100) with a private deal.…
Sep 6, 2026
7408614
CP2K: link and run against the toolchain OpenBLAS (not the conda pthr…
Sep 6, 2026
c7b90f9
Nyx: heat/cool strong-scaling timings and 8/40/80 dry-run record
Sep 6, 2026
1e5d00d
QMCPACK: walker-population guard, 256-walker strong/weak series, results
Sep 6, 2026
e354870
GEOS: develop b7a0f133 + TPL 361-1070 on CUDA 13.2/sm_100, beam-bendi…
Sep 6, 2026
e91119c
DFT-FE: ELPA GPU probe parses ELPA's analytic-test errors; validation…
Sep 6, 2026
fbaa962
GEOS: record the unit-test probe (254/261; five flow/well physics tes…
Sep 6, 2026
feaf8a3
Level 3 second batch: status, audit and build-strategy documentation
Sep 6, 2026
0bf3eb2
docs: add CLAUDE.md (working notes for Claude Code sessions in this r…
Sep 6, 2026
6711980
Level 3 tools: allow-listed environment wrapper and HPCPERF_L3_RUN_SU…
Sep 7, 2026
f231217
Nyx: strict fcompare wrapper, I_R_CHECK_PENDING verdict, tolerance pr…
Sep 7, 2026
fc4d2a1
Nyx run.sh: run directory through $L3_RUN_SUBDIR (left out of 6711980)
Sep 7, 2026
74c1d4d
CLAUDE.md: regression run trees (HPCPERF_L3_RUN_SUBDIR) and the allow…
Sep 7, 2026
1a6a60e
Level 3 tools/Nyx: box-layout classification, verdict classes, creden…
Sep 8, 2026
0dd99a1
Level 3 status: joint-HEAD regression recorded, Nyx heat/cool PENDING…
Sep 8, 2026
0128eab
Level 3 source distribution tools: freeze, materialize, workspace con…
Sep 8, 2026
4bca9dc
Level 3: freeze all ten applications; build/run/validate read only th…
Sep 8, 2026
142aea0
Level 3 docs: source distribution layout, SOURCE_ARCHIVES.md, per-app…
Sep 8, 2026
3a24883
Level 3 source distribution scheme 3: external source artifacts, cont…
bowencui123 Sep 10, 2026
8d6ce73
Level 3: migrate the ten frozen sources to external artifacts (schema…
bowencui123 Sep 10, 2026
d083797
Level 3: ExaCA 2.1.0 as the replacement candidate for the tenth slot …
bowencui123 Sep 10, 2026
73e50b4
Level 3 docs: external-artifact design, LFS-to-artifact migration rec…
bowencui123 Sep 10, 2026
14ff87a
Level 3 workspace integrity: layered verdicts (REFUSED 6 / BUILD_FAIL…
bowencui123 Sep 11, 2026
fb12061
Level 3 ExaCA: validation protocol v2 (calibration/holdout separated)…
bowencui123 Sep 11, 2026
40cb958
Level 3 docs: per-application evidence matrix, LAMMPS 2-GPU workspace…
bowencui123 Sep 11, 2026
5f79c64
Level 3 release preparation: plan generator, gated GitHub Release ada…
bowencui123 Sep 11, 2026
51c6ffa
Level 3: release plan for level3-source-hpcperf-l3-v1-rc1 (11 assets,…
bowencui123 Sep 11, 2026
d0a4947
Level 3 workspace validation: register this iteration's runs, keep bu…
bowencui123 Sep 11, 2026
4541373
Level 3 release publication: strict HTTP/JSON/identity checking, and …
bowencui123 Sep 11, 2026
f463c93
Level 3 ExaCA: run the upstream GoogleTest unit tests and record the …
bowencui123 Sep 11, 2026
bdae7ce
Level 3 README: rewrite for first-time users; move the project policy…
bowencui123 Sep 11, 2026
bfd8dd7
Level 3: regenerate the release plan at the reviewed code commit (11 …
bowencui123 Sep 11, 2026
daff71f
Level 3 ExaCA: attribute all 22 upstream unit-test failures (0 applic…
bowencui123 Sep 11, 2026
065795a
Level 3 release publisher: verify the real Git tag and the remote ass…
bowencui123 Sep 11, 2026
c72d669
Level 3 docs: separate this project's test totals from an application…
bowencui123 Sep 11, 2026
bd85b90
Level 3: regenerate the release plan and manifest after the ExaCA acc…
bowencui123 Sep 11, 2026
16dcf18
Level 3 release prerequisites: strict asset-state check and ExaCA sta…
bowencui123 Sep 12, 2026
89ee004
Level 3 release plan regenerated at the publication target commit
bowencui123 Sep 12, 2026
8fdfcba
Level 3 source locks: real published URLs after the anonymous downloa…
bowencui123 Sep 12, 2026
ba82b1c
Level 3: publication evidence and documentation for the published pre…
bowencui123 Sep 12, 2026
4a854e8
Level 3: remove optimization_scope.yaml; integrity = protected by def…
bowencui123 Sep 14, 2026
b46f27b
Level 3: PyYAML in the project environment; release tag is not a prep…
bowencui123 Sep 14, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
15 changes: 15 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -23,3 +23,18 @@ __pycache__/
*.sqlite
rocprof*.csv
rocprof*.json

# Level 3 source distribution (scheme 3: external source artifacts + automatic materialization, 2026-09-10)
# materialized source trees, the content-addressed artifact cache, materialization staging, per-run agent
# workspaces and trusted workspace baselines are local state, never committed; source artifacts live in
# the maintainer's staging / the external artifact storage (never in git, never in Git LFS)
level3/*/src/
level3/*/deps/
level3/*/.hpcperf-materialized.yaml
level3/*/.materialize.tmp.*
level3/*/.discarded-*
level3/.materialize-staging/
.artifacts/
.hpcperf/
workspaces/
*.tar.zst
287 changes: 287 additions & 0 deletions CLAUDE.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ An End-to-End AI Framework for Performance Prediction and Optimization in HPC Ap
|-------|---------|--------|
| [level1/](level1/) | 50 standalone GPU benchmarks (independently buildable / runnable / validatable) | 50/50 working, CUDA-validated |
| [level2/](level2/) | 20 proxy applications / mini-apps (upstream build systems kept, wrapped by `build.sh` / `run.sh` / `validate.sh`) | 19/20 working, CUDA-validated; MiniEM pending (Trilinos) |
| [level3/](level3/) | production / end-to-end HPC applications | planned |
| [level3/](level3/) | 10 production / end-to-end HPC applications, multi-GPU, with frozen source artifacts restored by `tools/prepare_benchmark.sh` (see [level3/README.md](level3/README.md)) | 10 applications brought up and CUDA-validated at 1/2/4 GPUs on one node (criteria and evidence levels per application); source artifacts **published** on 2026-09-12 as the prerelease `level3-source-hpcperf-l3-v1-rc1` and verified by an anonymous download; the code itself is **not merged into `main`** yet, so check out the PR branch `level3/source-freeze` (the release tag is the artifacts' source identity, not an entry point) |

Level 1 benchmarks are extracted (or faithfully ported) from six upstream
suites -- HeCBench, RAJAPerf, NPB-GPU, Hetero-Mark, Rodinia, Kokkos Kernels --
Expand Down
5 changes: 5 additions & 0 deletions check_env.sh
Original file line number Diff line number Diff line change
Expand Up @@ -87,6 +87,11 @@ if python -c "import numpy, scipy" 2>/dev/null; then
else
fail "numpy/scipy not importable in the active python"
fi
if python -c "import yaml" 2>/dev/null; then
ok "PyYAML importable ($(python -c 'import yaml; print(yaml.__version__)'); needed by the Level 3 harness tools)"
else
fail "PyYAML not importable in the active python (tools/prepare_benchmark.sh, check_workspace.py and the freeze tools need it; rerun ./setup_env.sh)"
fi

# --- CUDA (system-provided; NOT installed by setup_env.sh)
if [ -n "${CUDA_HOME:-}" ] && [ -x "$CUDA_HOME/bin/nvcc" ]; then
Expand Down
1 change: 1 addition & 0 deletions environment.yml
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@ dependencies:
- numpy=2.5.2 # used by the verify.py correctness checks
- scipy=1.18.0
- perl=5.32.1 # runs .tools/bin/cloc (cloc is not on conda-forge)
- pyyaml=6.0.3 # Level 3 harness tools (source locks, benchmark.yaml, freeze specs, workspace checks)
# ---- Level 2 mini-app dependencies (added 2026-09-01) ----------------------
# CUDA-aware Open MPI (conda-forge "cuda" build variant, pinned by build
# string so the solver picks the same variant on every machine).
Expand Down
506 changes: 506 additions & 0 deletions level3/APPLICATION_AUDIT.md

Large diffs are not rendered by default.

274 changes: 274 additions & 0 deletions level3/BUILD_STRATEGY.md

Large diffs are not rendered by default.

101 changes: 101 additions & 0 deletions level3/CORRECTNESS_FIXES.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,101 @@
# Level 3 first batch -- correctness / reproducibility fixes

Round after the da35285 review. Base for this work: HEAD was
`da352857561daa8f754161a856d0d53875d6f3ad` (verified; working tree clean at
start). No application source, patch, input, or tolerance was changed to obtain
a PASS; the five applications, their patches, inputs and prior results are kept.

Shared mechanism lives in `level3/tools/l3_common.sh` (+ `l3_check.py`); the
CPU-only tests are `level3/tools/tests/` (`run_all.sh` -> `test_l3_infra.sh`,
13/13 passing).

## 1. False PASS / exit codes

| Review point | Fix | Where |
|---|---|---|
| A failed run must not PASS on a stale log | validators delete/rewrite only this-run output; run.sh removes its target log before launching; validators require the run's real exit code == 0 | all `validate.sh`; `lammps/sparta run.sh` `rm -f "$LOG"` |
| SPECFEM solver failure swallowed by `\| tee \| grep \|\| true` | solver now runs to a file and its exit code is captured directly (`rc=0; cmd > log \|\| rc=$?`); the grep is display-only afterwards | `specfem3d/run.sh` |
| Separate execution from log filtering; capture launcher/app/validator real exit codes | validators run the app into a stdout file and gate on `rc`; `l3_capture` returns the command's status, not tee's; LAMMPS/SPARTA run.sh use `\|\| rc=$?` (not `; rc=$?` which `set -e` would abort) so the code and manifest are always recorded | `l3_common.sh` `l3_capture`; every `run.sh`/`validate.sh` |
| timeout / missing file / analysis exception / nonzero -> FAIL; validate only new output | `timeout` wraps each run (rc 124 -> FAIL); missing output -> FAIL; python raises `ValidationError` -> FAIL | every `validate.sh` |

Negative test: `test_l3_infra.sh` case 3 shows a failed run (rc=1) FAILs the gate
even when a PASS-looking stale log is present; case 2 shows `l3_capture` returns
the real code (7), not tee's 0.

## 2. Numerical finiteness / completeness

| Review point | Fix |
|---|---|
| Reject NaN/Inf in data/reference/error | `l3_check.require_finite` names and rejects any non-finite compared quantity; used by LAMMPS/SPARTA/WarpX validators and the SPECFEM sample scan |
| LAMMPS: expected thermo fields + final step | requires Step 0 and Step 100 rows and the fields Temp/E_pair/TotEng/Press present, else FAIL |
| SPARTA: final step, complete stats interval, fields | requires the benchmark block to span steps 30..130 (equilibration boundary to final) and the fields Np/temp/Natt present |
| SPECFEM: all required reference traces, sampling range, comparison | requires the compared-trace count == number of REF_SEIS traces (12), and scans every produced `.semd` for non-finite samples before trusting the correlation |
| nekRS: complete cimode check set (not passed>0) | requires `passed+failed == EXPECT_CHECKS` (9) AND failed==0 AND rc==0 AND coarse-location matches the cimode |
| WarpX: final step, expected particle count, field completeness, reader robustness | reader FAILs on missing field, box/fab mismatch, a truncated FAB, or boxes not covering 100% of the domain; requires the langmuir plotfile at step 40 and the uniform-plasma NP series to reach step 10 with a constant finite count |
| Adapted-subset labelling | LAMMPS/SPARTA validators print "adapted subset" and name exactly what upstream check they re-implement |
| CPU-only negative tests | `test_l3_infra.sh` (NaN/Inf rejection, rc-gate, dry-run sentinel, patch/cache) |

Thresholds unchanged (LAMMPS 1e-8/1e-5; SPARTA Np-exact/2%/15%; WarpX 5e-2/1e-11;
SPECFEM 0.8/1%/0.01s; nekRS EPS 0.3).

## 3. Result management

| Review point | Fix |
|---|---|
| Unique run_id; full stdout/stderr, command, exit code, source/binary/input hash | `l3_run_id` + `l3_manifest` write `run_manifest.txt` per real run with run_id, exit_code, binary+input sha256, backend, ranks, sizes, timer; the launcher already logs the command and per-rank GPU audit into the captured stdout |
| backend, GPU/rank/node, CPU/GPU binding, transport, timer, validation | recorded in the manifest and in the launcher lines of the captured stdout |
| dry-run must not delete/overwrite/rewrite real results; sentinel test | `l3_rundir` routes a dry-run to a `.dryrun/` scratch dir and refuses paths outside `build/level3/`; WarpX/SPECFEM/nekRS run.sh (which `rm -rf`'d the real dir before the dry-run check) now go through it; LAMMPS/SPARTA redirect their log into `.dryrun/`. Verified live: an 8-GPU dry-run left a real `log.smoke.np1` untouched and used `.dryrun/` (`test_l3_infra` case 4 + the live sentinel run) |
| Historical UNKNOWN exit codes stay UNKNOWN | not back-filled; the review bundle already labels them UNKNOWN |

## 4. Fingerprint / cache

| Review point | Fix |
|---|---|
| Patch full path + ordered content hash, not basename | `l3_fingerprint_text` records `patch[i]=<name> sha256=<hash>` in order and a `patch_series_sha256` over the ordered contents (schema bumped l3-1 -> l3-2) |
| Missing / unhashable patch -> error | `l3_fingerprint_text` returns non-zero on a missing patch; `nekrs/build.sh` also checks each patch exists before building |
| Source-cache key includes upstream SHA + patch content hash; renamed-but-changed invalidates | nekRS src stamp is now `SHA <ordered-patch-content-hash>` (was basenames) |
| build/install/cache isolated by backend/toolchain/dependency/config | per-app `.deps/level3/<app>/{src,build,install,logs}`; nekRS is further split per variant (`hypregpu` legacy paths, others under `.deps/level3/nekrs/<variant>/` with their own build dir and JIT cache) |
| Verify binary backend before running (no HIP request on a CUDA install) | `l3_binary_backend_check` (libcudart vs libamdhip64) available in `l3_common.sh` |
| Post-hoc fingerprints marked | `l3_fingerprint_write` stamps `built=<UTC> (build-time record)`; nothing back-dates |

Negative test: `test_l3_infra` case 5 (missing patch -> error; same-name changed
content -> different series hash; empty series -> `none`).

## 5. Dependency isolation

| Review point | Fix / finding |
|---|---|
| Check LAMMPS/SPARTA actual Kokkos helper source | The recorded (polluted-env) `CMakeCache.txt` had `Kokkos_NVCC_WRAPPER`/`Kokkos_COMPILE_LAUNCHER` pointing at **Level 2's** `.deps/install/kokkos`; the actual compile/link commands had **0** Level 2 references and used the bundled `nvcc_wrapper`+includes (so the binaries were clean, the cache entry was an inert stale detection). |
| Reconfigure without Level 2 prefixes; rebuild only affected apps | `l3_isolate_build_env` strips `$R/.deps/install/` from CMAKE_PREFIX_PATH/LD_LIBRARY_PATH in every build.sh. LAMMPS + SPARTA rebuilt in the isolated env (221 s / 570 s): `CMakeCache.txt` now has 0 Level 2 refs and `Kokkos_NVCC_WRAPPER` points at the bundled Kokkos. WarpX/nekRS/SPECFEM already had 0 refs (CXX = conda/mpicxx, autotools); the isolation call was added to their build.sh too but they were not rebuilt. Re-validated LAMMPS/SPARTA 1/2/4 -> identical numbers, PASS. |
| env_profiles base/head 9/11 tracked separately, not green | unchanged Level 2 issue (conda cmake activation drops a user CMAKE_PREFIX_PATH); documented in the review bundle `50_issues/env_profiles`; NOT a Level 3 regression and NOT marked passing here. |

## 6. Status / build strategy

| Review point | Fix |
|---|---|
| Separate PASS (smoke/analytic) from COMPLETED (strong/weak) | status table distinguishes validated-correctness runs from run-completed runs; strong/weak remain COMPLETED unless a numerical criterion applies |
| printed-values-equal != full bitwise | wording corrected: LAMMPS reports the four state variables agree to printed precision, not full-state bitwise identity |
| Don't widen to untested algorithm paths | LAMMPS stays LJ; WarpX stays FFT=OFF Yee PIC; no new solver paths added |
| Spack doc vs local facts; concretization NOT_RUN | recorded in the review bundle `50_issues/build_strategy` (local Spack 1.0.0.dev0, recipe versions) and `COMPATIBILITY.md`; no concretization run this round |
| Missing container runtime is an environment limit, not a feasibility verdict | stated as such; only podman present, not evaluated |

## nekRS-specific (see COMPATIBILITY.md)

- Confirmed the six logs' `COARSE SOLVER LOCATION: CPU` against upstream source
defaults and `ci.inc`: cimode 2 = CPU coarse (GPU main app), cimode 3 = DEVICE
(GPU) coarse.
- The cimode-2 "9/9" validates the CUDA main app + CPU coarse only. GPU HYPRE is
now separately verified with cimode 3 (`hypregpu` variant): 9/9, coarse=DEVICE,
at 1 and 4 GPUs.
- The `pair`/`reverse_iterator` build errors are missing-include (visibility)
issues; only `thrust::not1` is a genuinely removed API. Patches are labelled
project-local (no upstream backport SHA located).
- Minimal candidate `cpucoarse` (`ENABLE_HYPRE_GPU=OFF`, **0 patches**, isolated
variant tree, 113 s build): cimode 2 (CPU coarse) PASS 1/2/4 GPU (9/9,
coarse=CPU, main app on distinct GPUs); cimode 3 (DEVICE requested) is
**explicitly rejected** by nekRS (`HYPRE+DEVICE not enabled! Recompile with
-DENABLE_HYPRE_GPU=ON`, exit 1) -- no silent CPU fallback. So the three patches
are needed only for GPU HYPRE; the current CPU-coarse workload needs none.
- `hypregpu` (patched) cimode 3 (DEVICE/GPU coarse) PASS 1/4 GPU (9/9,
coarse=DEVICE): the GPU HYPRE coarse solve the patches enable is verified
correct, not merely compiled.
Loading