Community Code
Open-source HPC codes
Releases, changelogs and project news from the community codes and libraries HPC actually runs — simulation engines, solvers, I/O layers, MPI stacks, schedulers and visualization tools — rewritten in plain language with a straight link to the original.
Code registry
147 projects · baseline version and last release date, read from upstreamMolecular & Materials
Classical molecular dynamics engine for materials — atoms, polymers, metals and granular systems at very large scale.
General-purpose particle simulation engine for MD and coarse-grained models, scriptable in Python and fast on GPUs.
High-throughput molecular dynamics tuned for biomolecules; heavily GPU-optimized and common in drug and protein work.
Quantum Monte Carlo code for high-accuracy electronic structure of molecules and solids.
Atomistic simulation package for solids, liquids and interfaces, best known for fast ab-initio molecular dynamics.
Parallel molecular dynamics for large biomolecular systems, built on Charm++ and used for million-atom simulations.
Density-functional theory suite for electronic structure and materials modeling with plane waves and pseudopotentials.
Plugin library that adds enhanced sampling and free-energy calculations to most molecular dynamics engines.
Biomolecular simulation suite with GPU-accelerated MD and its own widely used force fields.
Computational chemistry suite covering quantum chemistry and molecular dynamics on parallel machines.
Python library for reading, analyzing and post-processing molecular dynamics trajectories in most common formats.
Widely used (commercially licensed) plane-wave DFT package for electronic structure and materials properties.
Long-standing biomolecular simulation package and force field family for proteins, lipids and nucleic acids.
Classical molecular dynamics package from Daresbury Laboratory, tuned for large systems on massively parallel machines.
Massively parallel MD for classical and polarizable force fields, neural-network potentials and QM/MM on thousands of GPUs.
Interactive molecular visualization and analysis for biomolecular systems — the standard front end to NAMD.
CFD & Engineering
Open-source computational fluid dynamics toolbox — meshing, solvers and post-processing for industrial flow problems.
Compressible CFD and adjoint-based shape optimization suite, popular in aerospace design studies.
Wind-energy CFD solver for turbine blade and wind-farm wake simulations, part of the ExaWind stack.
Spectral-element CFD solver for high-fidelity turbulence and thermal-hydraulics simulations.
Climate & Earth
Community Earth System Model — coupled atmosphere, ocean, land and ice climate simulation.
Model for Prediction Across Scales — unstructured-mesh atmosphere/ocean cores that can move resolution locally without remeshing.
MIT General Circulation Model — flexible ocean/atmosphere engine famous for adjoint-based state estimation.
Weather Research and Forecasting model used for regional forecasting and atmospheric research.
DOE's Energy Exascale Earth System Model, built for high-resolution climate runs on leadership systems.
Regional Ocean Modeling System — free-surface, terrain-following ocean model used for coastal and regional studies.
Spectral-element code simulating seismic wave propagation through the Earth — the standard for earthquake and full-waveform studies.
Finite-Volume Cubed-Sphere dynamical core — the atmosphere engine of NOAA's operational models and UFS.
Icosahedral Nonhydrostatic model from DWD/MPI-M — the operational forecast model for Germany and a growing climate tool.
Modular Ocean Model 6 — the ocean engine behind NOAA operational forecasts and GFDL climate models.
Math Libraries & Solvers
Toolkit of parallel linear and nonlinear solvers and time steppers underpinning many PDE applications.
Suite of ODE, DAE and nonlinear solvers used for time integration inside simulation codes.
Finite-element library for high-order discretizations on unstructured meshes, GPU-capable.
Block-structured AMR framework powering many exascale applications on CPUs and GPUs.
Collection of solver, preconditioner and discretization packages for large-scale multiphysics simulation.
Scalable linear solvers and multigrid preconditioners for very large sparse systems.
Parallel unstructured mesh infrastructure with adaptation and load balancing for FEM codes.
Optimized open-source BLAS implementation tuned per CPU architecture — the default numeric backbone on most clusters.
AMD-backed BLAS framework with pluggable kernels, exposing both BLAS and CBLAS interfaces.
AMD's open-source ROCm BLAS implementation for Instinct GPUs, with hipBLASLt underneath for tuned GEMMs.
The standard portable fast Fourier transform library for C and Fortran codes.
Distributed-memory version of LAPACK for dense linear algebra across many nodes.
Dense linear algebra library for GPU-accelerated nodes, a heterogeneous counterpart to LAPACK.
High-performance sparse linear algebra library built around portable abstractions for CPU and GPU back ends.
DOE Exascale Computing Project library bringing dense linear algebra (ScaLAPACK-class) to heterogeneous GPU clusters.
Direct sparse LU solver for general unsymmetric systems, with sequential and distributed versions.
Reference dense linear algebra library — the numerical foundation under most scientific software.
Preconditioned iterative eigensolver for large Hermitian and non-Hermitian problems, tuned for both CPUs and GPUs.
Sparse direct solver and preconditioner using low-rank compression for large structured systems.
Framework for block-structured adaptive mesh refinement PDE solvers from Berkeley Lab.
Implicitly restarted Arnoldi/Lanczos eigensolver for large sparse matrices — the classic library behind many eigenshells.
Serial graph and mesh partitioning library — the classic tool for splitting unstructured meshes across ranks.
Intel's optimized math library (commercially licensed) — BLAS, LAPACK, FFTs, sparse solvers and RNGs for Intel CPUs and GPUs.
NVIDIA's GPU-accelerated BLAS/LAPACK library (closed source), the workhorse under most CUDA numerical software.
NVIDIA's CUDA FFT library (closed source) — the fast Fourier transform standard on NVIDIA GPUs.
Massively parallel eigenvalue solvers for symmetric dense matrices — the standard inside DFT codes like VASP and Quantum ESPRESSO.
Multifrontal sparse direct solver for symmetric and general systems, shared by many FEM and multiphysics codes.
MPI-parallel extension of METIS for partitioning very large graphs and meshes across the whole machine.
Graph and mesh partitioning framework with its MPI-parallel PT-Scotch variant, widely used as a MUMPS/Trilinos ordering backend.
Frameworks & Portability
LLNL's C++ abstraction layer for portable loop execution across GPU and CPU back ends.
C++ performance-portability layer so one source tree runs on NVIDIA, AMD, Intel GPUs and CPUs.
Programming framework from LANL for building multiphysics apps on task-based runtimes.
Managed array abstraction that moves data between CPU and GPU automatically alongside RAJA.
Portable memory-management library for allocating and pooling across host and device memory.
Task-based runtime that schedules work and data movement automatically across heterogeneous nodes.
Kokkos-based particle simulation toolkit for MD, PIC and related methods.
I/O & Data
Self-describing file format and parallel I/O library used for most large scientific datasets.
Lightweight I/O characterization tool that shows how an application actually hits the filesystem.
Array-oriented data format and library, the standard in climate, weather and ocean science.
High-performance I/O and streaming framework for coupling simulations with analysis at scale.
Error-bounded lossy compressor for scientific floating-point data with tunable accuracy.
Compressed floating-point array library offering fast lossy or lossless compression for big fields.
File Systems & Storage
The dominant parallel file system on leadership-class systems — the backbone of most HPC center storage.
Distributed storage platform (object, block and file) that scales to exabytes on commodity hardware.
High-performance parallel file system from ThinkParQ, popular on mid-size clusters for its easy setup.
Distributed Asynchronous Object Storage — Intel's NVMe-native storage tier built for exascale I/O.
Policy engine for Lustre and POSIX storage — automated tiering, purging and metadata management.
The standard benchmark for measuring parallel file system throughput across ranks.
Metadata benchmark that stresses file and directory operations — the other half of storage performance testing. Ships in the same repository as IOR.
IBM's commercially licensed parallel file system — common on enterprise and leadership storage, formerly GPFS.
Communication & Runtime
Reference MPI implementation that many vendor MPI stacks are derived from.
NVIDIA's OpenSHMEM implementation with GPU-initiated communication, letting kernels communicate across nodes directly from device code.
NVIDIA Collective Communications Library — GPU-optimized collectives (all-reduce, all-gather) underpinning training frameworks like PyTorch and JAX.
Open Fabrics interfaces library exposing low-level provider fabrics (RDMA, sockets, shared memory) to MPI, OSSS and other middleware.
Open-source MPI implementation used as the default message-passing layer on many clusters.
Unified Communication X — the point-to-point communication framework that Open MPI and MPICH route over for InfiniBand, RoCE and shared memory.
Process Management Interface — the standard wiring between schedulers like Slurm and MPI/GPU runtimes for startup, coordination and fault handling.
AMD's ROCm Communication Collectives Library — the HIP port of NCCL for multi-GPU collectives on AMD accelerators.
Partitioned global address space (PGAS) API for one-sided communication — an alternative programming model to message passing.
Systems & Packaging
Workload management system for high-throughput computing — long-running opportunistic jobs across pools of machines.
Secure single-file container system built for HPC — the successor to Singularity, standard on shared clusters.
Community-built software stack of packaged HPC components for provisioning Linux clusters.
Lua-based module system from TACC for managing software environments on shared systems — the modern successor to Environment Modules.
Next-generation hierarchical resource manager from LLNL, aiming to run heterogeneous workflows across the full system.
The dominant open-source workload manager and job scheduler for HPC clusters.
Build and installation framework for scientific software on HPC systems, with a large community-maintained collection of easyconfigs.
Package manager for HPC that builds many versions and compiler combinations side by side.
Open-source PBS workload manager — the community version of the scheduler used on many government and vendor systems.
Bio & Workflows
Workflow manager that runs containerized pipelines portably across clusters and clouds.
Python-based workflow engine for reproducible, cluster-aware analysis pipelines.
Genome Analysis Toolkit — the standard variant-calling pipeline for sequencing data.
Fast short-read aligner for mapping sequencing reads to reference genomes.
Bioinformatics
High-performance toolkit for molecular simulation with first-class GPU support and a Python-friendly API.
The standard toolkit for reading, writing and manipulating alignments in SAM/BAM/CRAM format.
The versatile aligner for long-read sequencing — noisy reads, splice-aware RNA mapping and chromosome-scale alignments.
The faster, SIMD-accelerated rewrite of the BWA short-read aligner, ~3x quicker at identical output.
Ultrafast RNA-seq splice-aware aligner, the most widely used in transcriptomics pipelines.
DeepMind's protein structure prediction system — the open AlphaFold 2 and 3 implementations that reshaped structural biology.
Visualization
Visualization Toolkit — the underlying C++ library behind ParaView, VisIt and many custom tools.
Parallel visualization and analysis tool for large mesh and field data from simulations.
Parallel visualization application for exploring very large simulation datasets interactively or in batch.
In-situ visualization and analysis library that renders while the simulation is still running.
Performance Tools
Source-code annotation library for collecting application performance data with low overhead.
Portable interface to hardware performance counters on CPUs, GPUs and other components.
Sampling-based performance analysis suite for CPU and GPU codes with call-path attribution.
Profiling and tracing toolkit for parallel programs across MPI, OpenMP and GPU runtimes.
Physics & Multiphysics
GPU-capable particle-in-cell code for plasma and accelerator physics, born from the DOE Exascale effort.
Automated finite-element system with a Python interface, sharing the FEniCS equation interface but its own compiler stack.
C++ finite-element library with mature adaptive meshes and GPU support, used in physics and engineering PDE codes.
Monte Carlo neutron and photon transport code for reactor and shielding analysis, Python-driven and GPU-ready.
Next-generation FEniCS — Python/C++ finite elements with Just-in-time compiling kernels and MPI-parallel assembly.
GPU-offloaded spectral-element CFD solver built for exascale thermal-hydraulics and turbulent flows.
Idaho Lab's Multiphysics Object-Oriented Simulation Environment — the framework behind many nuclear and materials codes.
Multiphysics simulation framework from the University of Chicago — astrophysical flames, supernovae and laser-plasma experiments.
CERN's toolkit for simulating particles through matter — the reference in particle, nuclear and medical physics.
Proxy Apps & Benchmarks
Particle-in-cell kernel library and proxy used by the WarpX plasma code.
Adaptive mesh refinement mini-app that stresses dynamic load balancing.
Deterministic transport proxy app used to study sweep algorithms and data layouts.
Material point method proxy app for particle-mesh solid and fluid mechanics.
Particle-in-cell proxy app built on Cabana for plasma simulation patterns.
Monte Carlo neutron transport proxy app dominated by random memory lookups.
High-order Lagrangian hydrodynamics miniapp from the CEED project, built on MFEM.
Graph community-detection mini-app representing irregular graph workloads.
Monte Carlo transport proxy app with irregular, branch-heavy work patterns.
LANL discrete-ordinates neutral particle transport proxy app.
Sparse solver benchmark that ranks systems on memory and communication rather than peak flops.
Seismic wave propagation proxy app derived from SW4 for earthquake simulation kernels.
The dense LINPACK benchmark behind the TOP500 ranking.
Algebraic multigrid solver proxy app derived from hypre.
Multi-purpose I/O proxy app for benchmarking parallel filesystem behavior.
Compact version of QMCPACK's kernels for performance porting studies.
Shock hydrodynamics proxy app widely used for evaluating machines and programming models.
Communication pattern mini-apps for evaluating interconnect and MPI behavior.
Kokkos-based MD proxy app for testing portability across accelerators.
Spectral-element solver kernel extracted from Nek5000 for machine evaluation.
Molecular dynamics proxy app capturing the core cost of classical MD kernels.
Distributed 3D FFT proxy app taken from the HACC cosmology code.
