Skip to content
29syliuPublic

About

A High-Performance Code for Gyrokinetic-MHD Hybrid Simulation on GPUs with CUDA C++

Topics

Resources

Stars

17 stars

Watchers

0 watching

Forks

Latest commit

 

History

76 Commits

Folders and files

Repository files navigation

cuGMEC

cuGMEC: A High-Performance Code for Gyrokinetic-MHD Hybrid Simulation on GPUs with CUDA C++

CPC DOI CUDA C++ GPLv3 Version 1.27.2609

NCSX visualization


📌 Overview

cuGMEC is the fully reconstructed GPU version of the GMEC (Gyrokinetic-MHD Energetic-particle Code) family:

  • P. Y. Jiang, Z. Y. Liu, S. Y. Liu, J. Bao, and G. Y. Fu
    Development of a gyrokinetic-MHD energetic particle simulation code. I. MHD version
    Physics of Plasmas 31, 073904 (2024)
    DOI: 10.1063/5.0203252

  • Z. Y. Liu, P. Y. Jiang, S. Y. Liu, L. L. Zhang, and G. Y. Fu
    Development of a gyrokinetic-MHD energetic particle simulation code. II. Linear simulations of Alfvén eigenmodes driven by energetic particles
    Physics of Plasmas 31, 073905 (2024)
    DOI: 10.1063/5.0206762

  • S. Y. Liu, P. Y. Jiang, and G. Y. Fu
    cuGMEC: A High-Performance Code for Gyrokinetic-MHD Hybrid Simulation on GPUs with CUDA C++
    Computer Physics Communications 327, 110249 (2026)
    DOI: 10.1016/j.cpc.2026.110249


🧲 Physical Model

cuGMEC solves a nonlinear gyrokinetic-MHD hybrid model with fluid electrons and gyrokinetic ions, including both thermal ions and energetic particles.

The code can be used to study energetic-particle-driven Alfvén eigenmodes, such as TAE, RSAE, and BAE. It can also be applied to drift-wave and electromagnetic microinstabilities, such as ITG and KBM.

For details, please refer to our CPC paper.

MHD Component

The MHD component includes:

  • Gyrokinetic vorticity equation
  • Gyrokinetic Poisson equation
  • Parallel Ampère’s law
  • Parallel Ohm’s law
  • Electron continuity equation
  • Electron isothermal condition

PIC Component

Thermal ions and energetic particles are modeled gyrokinetically and advanced using the delta-f particle-in-cell method. PIC-computed pressures enter the gyrokinetic vorticity equation through the pressure-curvature term.


🧮 Numerical Methods

cuGMEC uses:

  • Shifted metric coordinates
  • Delta-f method for particle simulation
  • Five-point central finite-difference scheme for spatial discretization
  • Fourth-order Runge-Kutta method for time integration

For details, please refer to our CPC paper.


⚡ Performance

The performance reported in our CPC paper no longer represents cuGMEC's best performance, as both the MHD and PIC components have since been optimized. The updated single-GPU speed tests below are based on the ITPA TAE benchmark case, using a 256x64x16 grid and dt = 0.025 Alfvén time (major radius divided by Alfvén velocity), and we believe the results are quite fast among codes of the same type. The MHD component uses double precision; the PIC component uses 400 keV fast ions with a maximum velocity of about 1.2 times the Alfvén velocity and is tested in both double and float precision. Note that particle deposition is performed in double precision in both double- and float-precision PIC runs. GYRO denotes the number of gyro-average points, and P/G denotes the average number of particles per grid. For convenience, all GPUs are tested using the same gridDim and blockDim, so the timings may differ from the absolute optimum.

Click the triangle below for results.

NVIDIA GeForce RTX 4090 D

The MHD per-step time is 14.6ms. The PIC per-step times are shown in the table.

double
GYRO P/G 0 4 8 16
32 30.5ms 102ms 165ms 296ms
64 60.1ms 203ms 329ms 584ms
128 119ms 403ms 656ms 1.16s
256 237ms 806ms 1.33s 2.31s
float
GYRO P/G 0 4 8 16
32 2.82ms 8.96ms 15.0ms 28.5ms
64 4.27ms 15.5ms 26.1ms 49.7ms
128 7.50ms 28.5ms 48.5ms 93.1ms
256 14.0ms 54.9ms 94.0ms 178ms
NVIDIA A800-SXM4-80GB

The MHD per-step time is 13.8ms. The PIC per-step times are shown in the table.

double
GYRO P/G 0 4 8 16
32 7.26ms 26.3ms 44.4ms 78.4ms
64 13.5ms 48.9ms 83.8ms 148ms
128 26.2ms 94.2ms 162ms 288ms
256 51.2ms 184ms 319ms 566ms
float
GYRO P/G 0 4 8 16
32 3.72ms 13.9ms 23.1ms 42.0ms
64 6.65ms 25.4ms 41.7ms 75.8ms
128 12.5ms 48.6ms 79.8ms 143ms
256 24.4ms 95.7ms 161ms 281ms
NVIDIA RTX PRO 6000 Blackwell Server Edition

The MHD per-step time is 9.76ms. The PIC per-step times are shown in the table.

double
GYRO P/G 0 4 8 16
32 22.4ms 76.5ms 125ms 222ms
64 41.0ms 140ms 229ms 407ms
128 81.6ms 280ms 457ms 809ms
256 162ms 556ms 906ms 1.61s
float
GYRO P/G 0 4 8 16
32 1.68ms 6.24ms 9.71ms 25.4ms
64 2.65ms 10.4ms 17.0ms 43.6ms
128 4.09ms 18.7ms 31.0ms 66.3ms
256 7.90ms 36.6ms 60.3ms 127ms
NVIDIA B200

The MHD per-step time is 9.89ms. The PIC per-step times are shown in the table.

double
GYRO P/G 0 4 8 16
32 0.906ms 6.60ms 15.5ms 32.6ms
64 2.03ms 16.9ms 33.8ms 66.6ms
128 7.20ms 37.5ms 70.4ms 135ms
256 18.1ms 81.2ms 151ms 294ms
float
GYRO P/G 0 4 8 16
32 0.805ms 3.27ms 7.80ms 17.0ms
64 1.01ms 7.08ms 16.3ms 33.0ms
128 3.33ms 18.2ms 39.2ms 73.1ms
256 5.82ms 40.5ms 89.2ms 162ms

The total runtime of a simulation can be estimated as (MHD per-step time + PIC per-step time) × the total number of time steps. For other numerical parameters or multi-GPU runs, the runtime can be estimated by simple scaling.

For example, an ITER steady-state full-torus case up to toroidal mode number 36 uses a 384x32x576 grid (about 7.1 million points), P/G = 240 (about 5.1 billion particles across three ion species), GYRO = 4, dt = 0.025 Alfvén time, 20,000 steps, double-precision MHD, and float-precision PIC. Using only linear scaling from the tables above gives 5.31 hours on 8 NVIDIA A800-SXM4-80GB GPUs, compared with a measured 4.97 hours, a 6% difference; on 8 NVIDIA B200 GPUs, it gives 2.32 hours, compared with a measured 2.23 hours, a 4% difference. This agreement indicates that the tables above capture performance from the simplest benchmarks to more demanding cases.


💻 How to Use

A typical cuGMEC workflow is:

  1. Compute the tokamak equilibrium with scripts/equilibrium/compute2D_0170.ipynb, and export it with scripts/equilibrium/output2D_0170.ipynb.
  2. Generate the cuGMEC input files with scripts/preprocess/generateInput2D.m. If phase-space diagnostics are needed, also run scripts/preprocess/generatePhaseSpaceMapping2D.m.
  3. Configure src/cuGMEC_param.h according to docs/en-US/parameters.md.
  4. Configure the repository-root Makefile for the target environment according to docs/en-US/environment.md, and then compile cuGMEC.
  5. Run the simulation locally or submit it to a computing cluster.
  6. Visualize the simulation output with scripts/postprocess/visualizeMHD.m and scripts/postprocess/visualizePIC.m.

🚀 Future Work

Part of our effort is going into extending cuGMEC to support non-axisymmetric devices, such as stellarators.
Other future directions remain open.


📄 License

cuGMEC is released under the GNU General Public License v3.0.


✉️ Contact

Using cuGMEC generally requires nontrivial preprocessing and postprocessing, including equilibrium preparation, metric conversion, input-file generation, and output analysis. If you are interested in using cuGMEC, it is strongly recommended to contact the developer first.

For questions, suggestions, or collaboration, please contact:


About

A High-Performance Code for Gyrokinetic-MHD Hybrid Simulation on GPUs with CUDA C++

Topics

Resources

Stars

17 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages