Skip to content

Architecture

The Duke Compute Cluster (DCC) is a heterogeneous high-performance computing (HPC) system that supports traditional HPC simulations, high-throughput computing, data analysis, and GPU-accelerated machine learning and AI workloads. The cluster combines multiple generations of CPU and GPU hardware, owned by either Duke or individual research groups.

System overview

Component Specification
Login nodes 5 × Intel Xeon Platinum 8462Y+ with 8 CPUs
Compute nodes 505
CPUs 44,114
GPUs 1,084
Aggregate RAM ~270 TB
Storage ~7 PB
CPU architectures Intel Xeon (Skylake through Emerald Rapids), AMD EPYC (Zen 3)
GPU architectures NVIDIA Turing through Hopper (see GPUs by type)
Interconnect Ethernet, 10/40 Gb/s on older node groups, up to 100 Gb/s bonded on newer node groups; 200 Gb/s InfiniBand on H200 nodes
Operating system Alma Linux 9

Because the cluster is heterogeneous, nodes differ in CPU microarchitecture, core count, memory capacity, GPU configuration, and network bandwidth. Performance-sensitive and tightly coupled multi-node jobs benefit from running on a homogeneous group of nodes.

GPUs by type

GPU Architecture Slurm GRES type Count
GeForce RTX 2080 Ti Turing 2080 609
RTX A5000 Ampere a5000 162
RTX 5000 Ada Ada Lovelace 5000_ada, 5000_ada_generation 88
H200 Hopper h200 72
RTX 6000 Ada Ada Lovelace 6000_ada, 6000_ada_generation 68
RTX A6000 Ampere a6000 30
A100 Ampere a100, nvidia_a100-sxm4-80gb 28
Quadro RTX 6000 Turing 6000 23
H100 Hopper h100 4

GPUs are requested by their GRES type,

#SBATCH --gres=gpu:<gpu_type>:<num_gpus>

E.g.,
#SBATCH --gres=gpu:a5000:2

CPU architectures

DCC includes Intel Xeon and AMD EPYC nodes spanning several microarchitecture generations. Each generation is exposed as a Slurm feature that can be requested with --constraint as,

#SBATCH --constraint=<feature>
Feature CPUs
amd EPYC 7513, 7763
intel All Xeons
zen3 EPYC 7513, 7763 (Milan)
skylake 6142, 6148, 6152, 6154
cascadelake 5218R, 5220R, 6226, 6248R, 6252, 6254
icelake 5317, 5320, 6336Y
sapphirerapids 5418Y, 5420+, 6444Y, 8462Y+
emeraldrapids 5520+, 6544Y, 8562Y+, 8568Y+

For example, to request a Slurm job with only Sapphire Rapids nodes,

#SBATCH --constraint=sapphirerapids

This will provide a set of nodes with multiple variants of Sapphire Rapids CPU architectures such as 5418Y, 5420+, 6444Y, 8462Y+. If users want more fine grained control over the node selection, they can request specific sets of nodes with the --nodelist=<node-list> or -w <node-list> option. For example, the model UCSX-210C-M7 cluster of nodes networked through a 100Gb/s bonded Ethernet connection can be requested with,

#SBATCH -w dcc-core-ferc-u-ab39-5-[1-8],dcc-core-ferc-u-ab39-6-[1-6]

The tables in Node configurations list the node groups available in each partition.
See Scheduling Jobs for general job submission documentation and Partitions for partition limits and access.

Node configurations

The tables below list the node groups available in the general-use partitions, grouped by CPU architecture. Homogeneous node groups are recommended for multi-node MPI jobs and reproducible benchmarking.

As detailed in Partitions, the general-use partitions on the DCC are common, gpu-common, scavenger, scavenger-gpu, interactive, and scavenger-h200. scavenger and scavenger-gpu partitions consist of lab-owned nodes and provide opportunistic low-priority access to all users and are subject to preemption by higher-priority jobs.

common

CPU architecture CPU model Total CPUs Physical CPUs RAM (GB) nodes nodelist
sapphirerapids 8462Y+ 128 64 1007 14 dcc-core-ferc-u-ab39-5-[1-8],dcc-core-ferc-u-ab39-6-[1-6]
emeraldrapids 8562Y+ 128 64 1007 8 dcc-core-ferc-u-ab25-3-[1-8]
cascadelake 6252 96 48 754 48 dcc-dhvimdcore-gpu-ferc-s-o15-5,dcc-dhvimdcore-gpu-ferc-s-i11-[4-7,9-16,18-20],dcc-dhvimdcore-gpu-ferc-s-o15-[1-4,6-9],dcc-dhvimdcore-gpu-ferc-s-p15-[1-19,21-24],dcc-dhvimdcore-gpu-ferc-s-i11-8
cascadelake 6248R 96 48 754 7 dcc-core-ferc-u-ac39-4-[1-3,5-8]
cascadelake 6252 96 48 629 1 dcc-dhvimdcore-gpu-ferc-s-i11-17
icelake 6336Y 96 48 503 5 dcc-core-ferc-u-ab39-1-[7-8],dcc-core-ferc-u-ab39-2-[1-2,8]
cascadelake 5220R 96 48 503 2 dcc-core-ferc-u-ac39-2-[5-6]
skylake 6154 72 36 754 1 dcc-core-ferc-u-y32-1-5

gpu-common

CPU architecture CPU model Total CPUs Physical CPUs RAM (GB) GPU nodes nodelist
sapphirerapids 5420+ 112 56 1007 8 x RTX5000ADA 3 dcc-core-gpu-ferc-s-aa32-[1,5,9]
cascadelake 6252 96 48 754 3 x RTX2080TI 1 dcc-core-gpu-ferc-s-p15-20
cascadelake 6252 96 48 376 4 x RTX2080TI 2 dcc-core-ferc-s-z25-20,dcc-core-ferc-s-z25-21
skylake 6152 88 44 376 4 x RTX2080TI 3 dcc-core-gpu-ferc-s-h36-9,dcc-core-gpu-ferc-s-h36-[5-6]
icelake 5320 12 12 119 1 x A5000 9 dcc-rental-gpu-[07-15]

scavenger

CPU architecture CPU model Total CPUs Physical CPUs RAM (GB) nodes nodelist
emeraldrapids 8562Y+ 128 64 1511 2 dcc-mism-ferc-u-ab25-4-[2-3]
cascadelake 6248R 96 48 1007 1 dcc-pbenfeylab-ferc-u-ac39-2-1
cascadelake 6248R 96 48 754 2 dcc-allenlab-ferc-u-ac39-1-5,dcc-zjhuanglab-ferc-u-ab39-1-3
cascadelake 5220R 96 48 503 3 dcc-caperlab-ferc-u-ab39-1-5,dcc-valdivialab-ferc-u-ab39-1-4,dcc-velmeshevlab-ferc-u-ac39-1-6
cascadelake 6248R 96 48 376 3 dcc-adrc-ferc-u-q18-5-2,dcc-delairelab-ferc-u-q18-5-3,dcc-katzlab-ferc-u-q18-5-6
sapphirerapids 5418Y 96 48 251 3 dcc-schmidler-ferc-u-ab25-2-1,dcc-schmidler-ferc-u-ab25-1-8,dcc-barthellab-ferc-u-ab25-4-1
cascadelake 5218R 80 40 376 2 dcc-cagpm-ferc-u-ac39-5-5,dcc-kirschlab-ferc-u-ac39-5-4
skylake 6148 80 40 376 5 dcc-barthellab-ferc-u-y32-2-[2-3],dcc-fergusonlab-ferc-u-q18-4-6,dcc-pcharbon-ferc-u-y32-3-1,dcc-physics-ferc-u-y32-5-1
skylake 6154 72 36 754 9 dcc-dolbowlab-ferc-u-y32-1-1,dcc-dunsonlab-ferc-u-q18-1-1,dcc-volfovskylab-ferc-u-y32-1-2,dcc-dunsonlab-ferc-u-q18-2-2,dcc-fergusonlab-ferc-u-q18-4-8,dcc-fergusonlab-ferc-u-q18-5-1,dcc-fergusonlab-ferc-u-y32-3-[2-4]
skylake 6154 72 36 376 1 dcc-pcharbon-ferc-u-y32-1-4
sapphirerapids 6444Y 64 32 503 2 dcc-qcd-ferc-u-ab39-3-1,dcc-dmc-ferc-u-ab39-3-2
cascadelake 6248R 48 48 688 1 dcc-katzlab-02

scavenger-gpu

CPU architecture CPU model Total CPUs Physical CPUs RAM (GB) GPU nodes nodelist
sapphirerapids 5420+ 112 56 1007 8 x RTX6000ADA 1 dcc-allenlab-gpu-ferc-s-ab32-15
icelake 5320 104 52 503 4 x RTX6000ADA 2 dcc-majoroslab-gpu-ferc-a-r32-34,dcc-majoroslab-gpu-ferc-g-r32-29
icelake 5317 48 24 503 4 x A6000 1 dcc-carlsonlab-gpu-ferc-s-h36-15
icelake 5317 48 24 503 4 x A5000 1 dcc-vossenlab-gpu-ferc-s-g36-19
cascadelake 6252 96 48 754 4 x RTX2080TI 1 dcc-chsi-gpu-ferc-s-i11-1
cascadelake 6252 96 48 376 4 x A5000 1 dcc-viplab-gpu-ferc-s-aa25-4
cascadelake 6252 96 48 376 4 x RTX2080TI 6 dcc-carlsonlab-gpu-ferc-s-o15-10,dcc-pearsonlab-gpu-ferc-s-o15-17,dcc-plusds-gpu-ferc-s-j11-[17-18],dcc-plusds-gpu-ferc-s-z25-23,dcc-carlsonlab-gpu-ferc-s-h36-23
cascadelake 6252 96 48 376 3 x RTX2080TI 1 dcc-carlsonlab-gpu-ferc-s-h36-24
skylake 6152 88 44 376 4 x RTX2080TI 10 dcc-gehmlab-gpu-ferc-s-n32-[10-13,23-24],dcc-gehmlab-gpu-ferc-s-z25-[15,17-19]
icelake 5320 10 12 117 1 x RTX5000ADA 3 dcc-plusds-gpu-[01-03]
sapphirerapids 5420+ 12 12 119 1 x RTX5000ADA 1 dcc-rental-gpu-01
icelake 5320 12 12 119 1 x A5000 3 dcc-rental-gpu-[02-04]

interactive

CPU architecture CPU model Total CPUs Physical CPUs RAM (GB) nodes nodelist
sapphirerapids 8462Y+ 128 64 1007 14 dcc-core-ferc-u-ab39-5-[1-8],dcc-core-ferc-u-ab39-6-[1-6]

H200 nodes

Access to the H200 nodes is provided through the scavenger-h200 and h200-hp partitions that share the same nodes.
DCC includes 9 nodes with a total of 72 NVIDIA H200 Tensor Core GPUs (8 GPUs per node). Each H200 GPU provides,

  • 141 GB of HBM3e memory per GPU for handling the largest deep learning and HPC workloads.
  • Up to 4.8 TB/s of memory bandwidth, enabling extremely fast data movement for AI model training and large-scale simulation.
  • NVLink high-speed GPU interconnect supporting 900 GB/s bidirectional bandwidth per GPU, allowing multiple H200s in the same node to communicate as a single large memory space with near-zero CPU intervention.
  • PCIe Gen5 connectivity for rapid communication with CPUs and storage subsystems.
  • Full support for NVIDIA's CUDA, cuDNN, and AI/ML acceleration libraries, enabling seamless scaling from single-GPU to multi-GPU workloads.

See H200 for access and job submission details.

Node information for the H200 nodes is summarized in the table below.

CPU architecture CPU model Total CPUs Physical CPUs RAM (GB) nodes nodelist
emeraldrapids 8568Y+ 192 96 2013 9 dcc-h200-gpu-ferc-l-ac46-[1,9,17,25],dcc-h200-gpu-ferc-l-e36-[1,9,17,25,33]

Most DCC nodes have simultaneous multithreading (hyper-threading) enabled, so the total CPU count is twice the physical core count in the tables above. If your application does not benefit from hyper-threading, it can be disabled for a job with,

#SBATCH --hint=nomultithread

Querying node information

To obtain node information for lab-owned partitions not listed above, users may run the following command,

sinfo -p <partition_name> -o "%n %c %m"

This will give a list of nodes in the specified partition, along with the number of CPUs and memory (in MB) available on each node. To obtain detailed information about a specific node, run,

ssh <node_name> "lscpu; free -mh; nvidia-smi"

Interconnect

DCC compute nodes are connected primarily via Ethernet, with bandwidth varying by node generation. Older node groups use 10 Gb/s or 40 Gb/s Ethernet, while newer node generations use high-bandwidth bonded Ethernet. For example, some Sapphire Rapids node groups in the common partition are networked through 100 Gb/s bonded Ethernet. The H200 nodes are connected through 200 Gb/s InfiniBand.

Network characteristics vary by node generation and node group. Tightly coupled multi-node MPI jobs are sensitive to interconnect bandwidth and latency and should be placed on a single homogeneous node group using --constraint or --nodelist.

Compiling for a specific CPU architecture

Applications compiled with architecture-specific optimizations can run significantly faster than generic builds. All Intel Xeon generations on DCC from Skylake onward support the AVX-512 instruction set. Target a specific microarchitecture with the -march flag (GCC) or the -x flag (Intel compilers) during compilation for best performance.

Generation GCC flag Intel compiler flag Notable instruction sets
skylake -march=skylake-avx512 -xSKYLAKE-AVX512 AVX-512
cascadelake -march=cascadelake -xCASCADELAKE AVX-512, VNNI
icelake -march=icelake-server -xICELAKE-SERVER AVX-512, VNNI
sapphirerapids -march=sapphirerapids -xSAPPHIRERAPIDS AVX-512, VNNI, AMX
emeraldrapids -march=sapphirerapids -xSAPPHIRERAPIDS AVX-512, VNNI, AMX
zen3 -march=znver3 -march=core-avx2 AVX2

With Intel compilers, -xCORE-AVX512 can be used as a common AVX-512 baseline that runs on all Skylake and newer Intel nodes, and the newer LLVM-based Intel compilers (icx/ifx) also accept the GCC-style -march values. Note that Intel -x flags generate code that only runs on genuine Intel CPUs. For the AMD zen3 nodes, use -march=core-avx2 (or build with GCC and -march=znver3) instead.

For example, a typical optimized build for Sapphire Rapids nodes with automake,

CFLAGS="-O3 -march=sapphirerapids"
CPPFLAGS="-O3 -march=sapphirerapids"
FCFLAGS="-O3 -march=sapphirerapids"

and its CMake equivalent,

set(CMAKE_C_FLAGS "-O3 -march=sapphirerapids")
set(CMAKE_CXX_FLAGS "-O3 -march=sapphirerapids")
set(CMAKE_Fortran_FLAGS "-O3 -march=sapphirerapids")

A binary built with -march=<arch> will fail with an illegal-instruction error on nodes older than the targeted architecture. Always pair architecture-specific builds with the matching Slurm constraint, #SBATCH --constraint=sapphirerapids.

Binaries compiled for an older architecture (e.g., skylake-avx512) will run on newer Intel nodes, at the cost of not using newer instruction set extensions. Instruction sets are cumulative, so any CPU that supports AVX-512 also supports AVX2; a binary built for AVX2 (e.g., -march=core-avx2 or -mavx2) will therefore run on every general-use DCC compute node, Intel and AMD alike, making it the most portable, though least optimized (beyond a generic build), choice. Alternatively, applications can provide multiple binaries, each built for one architecture, and select the appropriate binary at job submission time based on the requested constraint.

Compilers, MPI libraries, and math libraries are available through environment modules; see Modules.

Compiling for a specific GPU architecture

CUDA code is similarly compiled for specific GPU architectures, identified by their compute capability. The GPU generations on DCC correspond to the following nvcc targets.

GPU Architecture Compute capability nvcc flag
RTX 2080 Ti, Quadro RTX 6000 Turing 7.5 -arch=sm_75
A100 Ampere 8.0 -arch=sm_80
A5000, A6000 Ampere 8.6 -arch=sm_86
RTX 5000 Ada, RTX 6000 Ada Ada Lovelace 8.9 -arch=sm_89
H100, H200 Hopper 9.0 -arch=sm_90

To build a single binary that runs on all DCC GPU generations, compile for multiple architectures. For example, with CMake,

set(CMAKE_CUDA_ARCHITECTURES "75;80;86;89;90")

As with CPU builds, code compiled only for a newer compute capability will not run on older GPUs, so pair GPU-specific builds with the matching --gres=gpu:<gpu_type>:<num_gpus> request.

Example Slurm job script

To summarize the details in this guide, the following example job submission script requests 4 Intel Sapphire Rapids nodes with 64 tasks per node (disabling hyper-threading), and 500 GB of memory per node. The job loads the FHIaims module and runs the aims-spr.x executable optimized for Sapphire Rapids CPUs with 256 MPI tasks.

#!/bin/bash
#SBATCH -J jobname       # Job name
#SBATCH -p common        # Partition name
#SBATCH -N 4             # Total no. of nodes
#SBATCH --ntasks-per-node=64  # Tasks per node
#SBATCH --mem=500G       # Memory per node
#SBATCH -t 02:00:00      # Walltime limit (hh:mm:ss)
#SBATCH --hint=nomultithread # Disable hyper-threading
#SBATCH --constraint=sapphirerapids # Request Sapphire Rapids nodes

# Load modules
module load FHIaims

# Execution
cd $SLURM_SUBMIT_DIR
mpirun -n $SLURM_NTASKS aims-spr.x > aims.out 2> aims.err