ExaCA -- "An exascale-capable cellular automaton for nucleation and grain growth" (LLNL / ExaAM, Exascale Computing Project): a complete additive-manufacturing microstructure application (multilayer melt-pool solidification from time-temperature histories, spot/directional/single-grain analytic problems, optional Finch heat-transport coupling, grain-analysis post-processing), not a mini-app or proxy. Motifs: cellular automata, grain nucleation, decentred-octahedron grain growth, neighbourhood updates over a steering vector of active cells, MPI 1-D domain decomposition with halo exchange, Kokkos GPU execution.
Status (2026-09-11, node dgx003, 4x B200, CUDA 13.2.78): BUILD_PASS; project-defined statistical validation
of the dirsolid smoke case (128^3, seed 0) at 1/2/4 GPUs -- calibration 2026-09-10 (8 runs), frozen protocol
v2, independent holdout 2026-09-11 9/9 PASS; strong/weak COMPLETED at 1/2/4 GPUs (no acceptance claim);
40/80 GPUs DRY-RUN only; HIP UNTESTED (no ROCm); multi-node UNVERIFIED. Admitted 2026-09-11 as the tenth
application of the default suite (replacement for GEOS) on this basis; only the smoke case is validated by that
protocol. Release acceptance was put ON_HOLD the same day when 22 of 52 upstream unit tests were found to
fail, and restored after every one of those failures was attributed (no failure fell into the APPLICATION
class, a statement about these 22 failures in this environment and not a general claim about the application)
and upstream's own small full-application cases were shown to pass at 1/2/4 real GPUs -- see the two sections
below.
| item | value |
|---|---|
| official repository | https://github.com/LLNL/ExaCA (README: README.md, license: LICENSE) |
| selected release | tag 2.1.0, commit d26e59cd51e241a327c5267d43fd70537e5425f7 (2025-10-22) -- latest release |
| license | MIT (LICENSE + NOTICE, LLNL); example material files and the 10,000-orientation grain files are repository content |
Kokkos (benchmark-specific dependency, deps/kokkos) |
tag 4.7.04, commit 82799e4577568f9666bde36265ac15d78da3e6c8 (2026-04-09), Apache-2.0 WITH LLVM-exception; ExaCA requires Kokkos >= 4.0 |
nlohmann_json (deps/json/json-3.12.0.tar.xz) |
release tarball v3.12.0, sha256 42f6e95cad6ec532fd372391373363b62a14af6d771056dbfc86160e6dfff7aa, MIT -- the exact file ExaCA's CMake would FetchContent-download; bundled so that the build needs no network |
| not bundled | Finch (optional coupled heat transport; not used, no coupled workflow is claimed), GoogleTest (unit tests not built), ExaCA-Data (external temperature-history data for FromFile problems; the benchmark uses the analytic Directional problem) |
| patches | none (modification class A: build flags only) |
| source artifact | exaca-hpcperf-l3-v1.tar.zst (see the source-distribution section at the end; identity in provenance/source.lock.yaml) |
LOC (cloc code lines, provenance/LOC.md) |
application-owned 6,512 (29 files: 5,976 header lines, 473 C++, 63 CMake), tests 2,760, benchmark-specific dependency source (Kokkos) 278,336, total materialized 287,697 |
Kokkos version choice: the first freeze used Kokkos 4.6.02; its bin/nvcc_wrapper still defaults to
-arch=sm_70, which CUDA 13.2 rejects in CMake's compiler test ("Unsupported gpu architecture 'sm_70'").
Kokkos 4.7.04 (default sm_80, "Support CUDA 13", "Fix compiling with Cuda 13.1" in its changelog) builds
unchanged; the artifact was re-frozen with it before any publication (same source_version,
hpcperf-l3-v1, because nothing had been released).
Three out-of-source stages from the materialized artifact only ($HERE/src, $HERE/deps; no fetch, no
patch): (1) Kokkos 4.7.04, CUDA backend, Kokkos_ARCH_BLACKWELL100 (sm_100, detected), Serial host
backend, nvcc_wrapper with the conda GCC 13.3.0 host compiler -> .deps/level3/exaca/<profile>/install/kokkos;
(2) nlohmann_json 3.12.0 from the bundled tarball (header-only, JSON_BuildTests=OFF) ->
install/json; (3) ExaCA with ExaCA_REQUIRE_EXTERNAL_JSON=ON, ExaCA_ENABLE_TESTING=OFF, no Finch ->
build/level3/exaca/<profile> and install/exaca (bin/ExaCA, bin/ExaCA-GrainAnalysis, material and
orientation data under share/ExaCA, which is where ExaCA resolves the file names of the deck). MPI: conda
Open MPI 5.0.10 (CUDA-aware, not required: ExaCA stages its halos through host buffers). Verified
build time 109 s at -j32 (all three stages). HIP: the same script selects hipcc + Kokkos_ENABLE_HIP
(ExaCA documents Serial/OpenMP/Threads/CUDA/HIP), untested here (no ROCm). Fingerprint
(.hpcperf-l3-fingerprint): ExaCA commit, source tree sha256, Kokkos version/commit, json version, arch,
compiler, CUDA, MPI, CMake options.
One MPI rank per GPU (Kokkos CUDA, device 0 under the launcher's per-rank CUDA_VISIBLE_DEVICES
wrapper; the launcher audits the rank->GPU mapping: "N verified, 0 mismatch"). ExaCA decomposes the domain
in Y only (1-D; every rank count is legal as long as each rank gets >= 2 cells in Y); halos are exchanged
every time step. The run is ONE global simulation (the log records NumberMPIRanks and the subdomain
sizes/offsets of the single decomposition; validate.sh checks they cover the box exactly once).
Deck: inputs/dirsolid.template.json -- the physics of upstream's flagship example
examples/Inp_DirSolidification.json (Inconel 625 interfacial response, cell size 1 um, time step
0.0666667 us, thermal gradient G = 5e5 K/m, cooling rate R = 3e5 K/s, nucleation density 10 x 10^12 m^-3,
mean nucleation undercooling 5 K, sigma 0.5 K, RandomSeed 0), with ONE deliberate difference: the bottom
substrate uses SurfaceSiteDensity: 0.25 (grains placed by a global RNG at physical locations, identical
for every decomposition) instead of SurfaceSiteFraction (upstream's mode (i), whose grain placement is
rank-local and therefore decomposition-dependent). Nuclei are drawn from one global RNG stream and
distributed to the owning ranks, so the initial condition is decomposition-independent. run.sh sets only
Domain.Nx/Ny/Nz, RandomSeed, the output path/name and the printing policy (the GrainID field is
written in smoke mode; in strong/weak mode only ExaCA's JSON log is written unless HPCPERF_EXACA_PRINT=1,
because the 4 B/cell field write is not the CA work being measured).
mode (HPCPERF_SCALE_MODE) |
box (cells) | per rank at N GPUs | steps | purpose |
|---|---|---|---|---|
| smoke (default) | 128 x 128 x 128 = 2.10 M | 2.10 M / N | 7,000 | correctness (validate.sh) |
| strong | 512 x 256 x 512 = 67.1 M (fixed) | 67.1 M / N | 17,000 | strong scaling; a single rank cannot hold 512^3: ExaCA indexes its 26-neighbour octahedron arrays with int (134 M x 26 overflows: "failed to allocate 1.678e+07 TiB"), an upstream limitation |
| weak | 512 x (128 N) x 512 | 33.5 M | 17,000 | weak scaling (Y extent grows with N) |
| 40 / 80 GPUs | strong box | 1.68 M / 0.84 M | -- | `HPCPERF_DRY_RUN=1 HPCPERF_NODES=10 |
Overrides: HPCPERF_EXACA_NX/NY/NZ (Ny per rank in weak mode), HPCPERF_EXACA_SEED. A second positional
deck on the command line is rejected. Results land in build/level3/exaca/<profile>/<run>/dirsolid.<mode>.np<N>/
with run_manifest.txt (run id, binary sha256, deck sha256, sizes, exit code).
Layout since 2026-09-15 (backend/profile isolation): profile cuda (or hip; override HPCPERF_EXACA_PROFILE,
must name the backend); build tree build/level3/exaca/<profile>/, install/logs
.deps/level3/exaca/<profile>/{install,logs} with the fingerprint in the profile's install; results under
build/level3/exaca/<profile>/run*/. The results recorded above were produced with the pre-profile layout
(.deps/level3/exaca/install, build/level3/exaca/cuda), which is kept as historical state and is never read by
the current scripts. Until 2026-09-15 the install used per-backend suffixes inside a shared root (install/kokkos-cuda, install/exaca-cuda); the profile root (.deps/level3/exaca/<profile>/install/{kokkos,json,exaca}) replaces that scheme.
ExaCA ships no reference output for this problem, and its final GrainID field is not bitwise
reproducible: identical 1-GPU runs differ in 2.2 % of the cells (46,499 of 2,097,152), 1- vs 4-rank runs in
6.1 %; the differing cells carry thousands of distinct (id, id) pairs and the grain-ID sets differ, so these are
real spatial differences of competitive growth, not a renumbering. Even a single growing grain (upstream
Inp_SmallEquiaxedGrain, 64^3) differs by 3 edge cells between identical 1-GPU runs and its termination cycle
by 1 between rank counts (references/calibration.json: singlegrain_probe). The code path consistent with this
(not an upstream statement): active cells are collected with Kokkos::atomic_fetch_add (src/CAupdate.hpp
lines 41/97/108/132) and a contested liquid cell is claimed with Kokkos::atomic_compare_exchange on
cell_type (line 208); the winner sets the new octahedron's centre. Upstream's own deterministic checks are its GoogleTest unit tests, and unit_test/tstUpdate.hpp additionally
runs two complete official simulations with numerical references (see the next section). The unit tests were
built and run on 2026-09-11 with the fused GoogleTest 1.11.0 that Kokkos vendors inside the artifact:
30 of 52 passed under CTest, and all 22 failures have since been attributed with before/after evidence
(references/upstream_unit_test_matrix.json): 7 test-fixture bugs, 13 host-space test variants that are invalid
inside a CUDA-enabled build, 2 caused by CTest launching mpiexec without per-rank GPU binding, 0 in the
APPLICATION class and 0 unresolved -- a classification of those 22 failures here, not a claim that the
application is free of defects in other environments or in code paths no test covers. For the project's own dirsolid case the validation remains project-defined and
statistical; the full protocol, the definition
of the spread, and the calibration/holdout separation are in references/validation_protocol.md.
Criteria (each must hold; references/validation_protocol.yaml names the frozen files):
- completeness / one global decomposition: run exits 0; field + log exist;
DIMENSIONS= deck; every cell solidified; logNumberMPIRanks= N; the Y subdomains tile the box exactly once (first offset 0,offset[i+1] = offset[i] + size[i] - 2, last end = Ny, every size >= 2, N entries -- the halo-sum formula alone would accept a missing + duplicated pair); - self-consistency: ExaCA's own
VolFractionNucleated(log) equals the value recomputed from the field (1e-3); - vs the frozen reference
references/dirsolid_smoke.reference.json(statistics of calibration run cal-08, 1 GPU, run id20260910T212610Z-216184-29676, binary sha256f2a8875c...); - with N > 1: the N-GPU statistics vs this build's 1-GPU run (rank-count independence).
Tolerances (frozen 2026-09-11 BEFORE the holdout; applied to a single run vs the reference; spread = range over the 8 calibration runs, rule = max(3 x range, floor); n = 8, an empirical range rule, no 3-sigma claim):
| statistic | meaning | calibration range (8 runs) | tolerance (v2) | holdout max deviation from the reference (9 runs) |
|---|---|---|---|---|
unsolidified_cells |
cells never captured | 0 | exactly 0 | 0 |
n_grains |
distinct grain IDs in the final field | 10 (3612-3622) | 1 % rel | 9 (0.25 %) |
n_nucleated |
distinct nucleated grains | 0 (20) | 5 % rel (= +-1) | 0 |
vol_fraction_nucleated |
volume fraction of nucleated grains | 0.0012 | 0.01 abs (floor) | 0.0015 |
top_layer_grains |
grains reaching the top surface | 2 (28-30) | 25 % rel | 2 (31 vs 29: 6.9 %) |
mean_misorientation_z_deg |
cell-weighted mean angle, closest <001> axis to +z | 0.033 deg | 0.25 deg (floor) | 0.039 deg |
mean_misorientation_z_top_deg |
the same over the top layer | 0.226 deg | 0.7 deg (3 x) | 0.155 deg |
mean_grain_volume_cells |
cells / grains | 1.6 | 1 % rel | 1.4 (0.25 %) |
Protocol v1 (2026-09-10, top_layer_grains 15 %, mean_misorientation_z_top_deg 0.5 deg = only 2.2 x the
range) validated 1/2/4 GPUs against the same data that set its thresholds; those runs are re-labelled
CALIBRATION (references/calibration.json, 8 runs incl. the reference). The holdout
(references/holdout.json): 3 independent sets x fresh 1/2/4-GPU validate.sh runs (9 runs, new run ids,
separate run directories run.holdout-001-{1,2,3}), evaluated with the frozen v2 rule: 9/9 PASS, launcher
audit N verified / 0 mismatch for every run. No tolerance was changed after the holdout; a failure would have
been kept and analysed, and a revised rule would need a new holdout set. Negative tests of the validator
(NaN/Inf statistics, unsolidified cells, gap/duplicate/short/wrong-rank decompositions, log/field inconsistency,
out-of-tolerance statistics, non-finite reference, a stale result standing in for a failed run):
level3/tools/tests/test_exaca_validator.sh (19 checks).
Two independent things live in upstream's test suite:
- kernel-level unit tests (
unit_test/tst*.hpp): 52 CTest entries here (26 per backend, SERIAL and CUDA, MPI suites at 1/2/4 ranks). - two complete official simulations inside
tstUpdate.hpp(full_simulations), each with an upstream numerical reference:Inp_SmallDirSolidification.json->VolFractionNucleated= 0.1882 ± 0.0100, andInp_SmallEquiaxedGrain.json->TimeStepOfOutput= 4820 ± 1 (upstream notes a possible race in a FIXME). These references belong to those two small official cases; they are not transferable to this benchmark's 128³dirsolidcase and are not a reason to relax its frozen protocol.
Results (2026-09-11, private test build, Kokkos 4.7.04, artifact source unchanged):
| run | result |
|---|---|
CTest as upstream defines it (mpiexec -n N, no per-rank GPU binding), after cmake --install |
30/52 pass, 22 fail, identical pattern in two runs |
| the same CUDA suites with test-side fixture patches and one GPU per rank | 23/23 pass (np1/2/4 where MPI-enabled) |
| the same SERIAL suites in a Serial-only Kokkos build | 23/23 pass |
| the two official full simulations through the project launcher | PASS at 1, 2 and 4 GPUs: VolFractionNucleated 0.1881 / 0.188 / 0.191 (expect 0.1882 ± 0.0100); TimeStepOfOutput 4820 / 4820 / 4820 (expect 4820 ± 1); launcher audit N verified, 0 mismatch |
the same official cases under plain mpiexec (no GPU binding) |
np1 PASS, np2 and np4 segfault inside runExaCA (all ranks on GPU 0) |
Attribution of the 22 failures (references/upstream_unit_test_matrix.json, per test: name, backend, ranks,
assertion or exit code, expected vs actual, command, class, evidence):
| class | count | what it is |
|---|---|---|
| TEST_FIXTURE | 7 | tstNucleation rebinds a local grain_id handle with create_mirror_view_and_copy instead of writing celldata's own subview (instrumented proof: the subview held 0,0,0 while the handle held 1,2,3); tstOrientation reads a create_mirror_view that was never copied; tstInterface calls two KOKKOS_INLINE_FUNCTIONs from host loops and indexes the device view octahedron_center_test ("DOCenter") on the host. Fixture-only patches in patches/upstream-tests/ make all of them pass; no expectation was changed. |
| UNSUPPORTED_CONFIG | 13 | the host-space ("SERIAL") test variants inside a CUDA-enabled build: Kokkos' default execution space is Cuda, so application kernels touch HostSpace views ("attempt to access inaccessible memory space"). They pass in a Serial-only build. |
| TEST_INFRA | 2 | ExaCA_Update_test_CUDA at 2 and 4 ranks: CTest launches mpiexec without per-rank GPU assignment. Through the validated launcher the same binary passes and meets both upstream metrics. (6 further tests had failed before cmake --install, for the same class of reason: the tests resolve data files through the install prefix.) |
| APPLICATION | 0 | none of these 22 failures was traced to the application's own kernels. The scope is exactly this matrix, on this machine, for the tests upstream ships at commit d26e59cd: it is evidence that these failures are not application defects, not evidence that the application has no defects. |
| UNRESOLVED | 0 |
The patches are test code only, labelled patched-upstream-tests, recorded with the upstream commit, the
target file hashes and their own hashes, and they are not part of the frozen artifact or the release payload
(patches/upstream-tests/README.json). Adopting them would require a new source_version and a new artifact.
Therefore: unmodified upstream tests do not all pass in this environment, and no unit-test result is counted
as a PASS in this repository's own totals.
RESULTS (2026-09-10, run/ tree; reference run id 20260910T212610Z-216184-29676, binary sha256 f2a8875c6565a619...)
| test | result |
|---|---|
| build (CUDA sm_100, Kokkos 4.7.04) | BUILD_PASS, 109 s |
| smoke 1/2/4 GPU, protocol v1 (2026-09-10) | CALIBRATION (same data as the thresholds): 3612 grains, 20 nucleated, vf 0.5496, top 29, misorientation 26.99 / 35.6 deg; CA 0.80 s (1 GPU), 1.5 s (4 GPUs); subdomains [65, 65] and [33, 34, 34, 33] |
| smoke holdout, protocol v2 (2026-09-11) | PASS 9/9 (3 sets x 1/2/4 GPUs; run ids in references/holdout.json; 1-GPU runs 3612 grains, 2-GPU 3614, 4-GPU 3619-3621; top-layer grains 28-31; vf 0.5481-0.5499) |
| strong 512x256x512, ExaCA "Time spent performing CA calculations" | 1 GPU 8.02 s, 2 GPU 6.26 s, 4 GPU 5.17 s (1.55x on 4 GPUs: the work per step is the active interface layer; per-step launch/halo overhead dominates at this size; COMPLETED, no correctness claim beyond the log statistics vf = 0.8833 / 0.8833 / 0.8833) |
| weak 512x128x512 per rank | 1 GPU 4.88 s, 2 GPU 6.16 s, 4 GPU 6.80 s (COMPLETED) |
| 40 / 80 GPUs | DRY-RUN (HYPOTHETICAL plan, run/.dryrun/), not executed |
| HIP | UNTESTED (no ROCm) |
| multi-node | UNVERIFIED (site transport limitation, see level3/README.md) |
Launcher audit of every executed run: "N verified, 0 mismatch, 0 unverified".
| # | criterion | status |
|---|---|---|
| 1 | complete production application, not a mini-app | yes (ExaAM application; feature set above); small code base (6.5 k application-owned LOC) -- stated plainly |
| 2 | exact upstream commit pinned | d26e59cd51e2... (tag 2.1.0) |
| 3 | license | MIT + NOTICE |
| 4 | dependency licenses (Kokkos, MPI, json, Finch) | Apache-2.0-LLVM, environment MPI, MIT; Finch not used |
| 5 | scientific input/data redistributable | material/orientation files are repository content (MIT); no external data |
| 6 | CUDA official path | Kokkos CUDA backend, documented; built and validated |
| 7 | HIP official path | Kokkos HIP backend, documented; untested here |
| 8 | MPI distributed path | 1-D Y decomposition, validated 1/2/4 ranks |
| 9 | real 1/2/4 GPU | holdout runs at 1/2/4 GPUs, launcher audit "N verified, 0 mismatch" for all 9 |
| 10 | one global run, not N copies | log decomposition check (criterion 1) |
| 11 | correctness criterion | project-defined statistical validation (protocol v2: invariants + frozen reference + rank-count independence; calibration/holdout separated; 9/9 holdout PASS); no upstream oracle exists |
| 12 | strong / weak definition | fixed 67.1 M-cell box / 33.5 M cells per rank |
| 13 | 40/80 GPU dry-run only | yes |
| 14 | multi-node not claimed | UNVERIFIED |
| 15 | scheme-3 source artifact | exaca-hpcperf-l3-v1.tar.zst, LOCAL_ARTIFACT_VERIFIED, REMOTE_FETCH_VERIFIED (published 2026-09-12) |
| 16 | benchmark.yaml | yes |
| 17 | source ownership recorded | source_scope in provenance/source.lock.yaml: application_owned src/src/*, src/bin/*, src/analysis/src/*, src/analysis/bin/*; benchmark deps under deps/ (Kokkos, nlohmann_json); tests separate. Descriptive metadata only: the benchmark does not prescribe what an optimization agent may modify |
| 18 | provenance/source.lock.yaml | schema 2, redistribution cleared |
| 19 | application-owned LOC | 6,512 |
| 20 | total materialized code LOC | 287,697 (includes the 278,336 Kokkos + nlohmann_json lines under deps/) |
Not claimed: Finch-ExaCA coupled workflow, FromFile/multilayer problem types, ExaCA unit tests (not
built), any performance optimization.
The application source is not in git and not read from _upstream/: tools/prepare_benchmark.sh level3 exaca materializes the frozen source artifact (<app>[-<variant>]-<source_version>.tar.zst, found in the local content-addressed cache .artifacts/sha256/ or downloaded from the immutable URL recorded in provenance/source.lock*.yaml once published; --artifact FILE for a local copy) into src/ (+ deps/), the only source build.sh/run.sh/validate.sh use. Archive size + sha256 and source_tree_sha256 are verified before anything is placed. Identity, patch series, licenses, redistribution status and the equivalence proof against the tree the results above were validated from are under provenance/ (source.lock*.yaml, patch_series*.txt, original_vs_baseline*.diff, LICENSES*.md, equivalence*.md, LOC*.md); benchmark.yaml is the machine-readable contract (entries, inputs, references, identity). The benchmark does not prescribe which part of the source an optimization agent may modify; the integrity layer only protects the harness and the validation assets. Remote status: see level3/SOURCE_ARTIFACTS.md.
| variant | artifact | source version | compressed / uncompressed | entries | source_tree_sha256 | archive sha256 | upstream | patches (pre-applied) | redistribution | equivalence | remote | LOC app-owned / bundled deps / benchmark deps / tests / total |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| - | exaca-hpcperf-l3-v1.tar.zst |
hpcperf-l3-v1 | 2.6 MB / 15.3 MB | 1579 | b88e84b74db34b8adb58bc9924ec1e0f3c7d416c58577c5b16bcfbda11f81f44 |
3de419c92222c0a76e6afdb621bb46f6d7b530dc1d064fcee34393f88c9fe316 |
2.1.0 d26e59cd51e2 |
none | cleared | src: EQUIVALENT, deps/kokkos: EQUIVALENT | REMOTE_FETCH_VERIFIED | 6512 / 0 / 278336 / 2760 / 287697 |
LOC = cloc 2.06 code lines of the materialized tree (no blank/comment lines, documentation and data excluded); source-ownership categories from provenance/source.lock*.yaml (source_scope, descriptive metadata written at freeze time). Dependencies are counted per benchmark, so totals overlap across benchmarks that ship the same dependency. The validated results recorded above were produced from trees proven content-equivalent to this artifact (provenance/equivalence*.md); they are not re-run by the migration.