AMD Instinct MI100 vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Instinct MI100

CORE STATE Arcturus
VRAM 32 GB
CLOCK SPEED 1502 MHz
TDP 300 W
BUS WIDTH 4096 bit
ARCHITECTURE CDNA 1.0
nm
PROCESS 7 nm
LAUNCH DATE 2020
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
139,035
93,395
geekbench_vulkan
N/A
77,879

Analysis: AMD Instinct MI100 vs NVIDIA CMP 40HX

Head-to-Head Benchmarks

The recorded data includes a single head-to-head comparison: Geekbench OpenCL. In this test, the AMD Instinct MI100 scores 139,035, while the NVIDIA CMP 40HX scores 93,395. The result is a 48.9% advantage for the AMD part. This is a decisive margin, placing the MI100 well ahead of the CMP 40HX in raw compute workloads measured through OpenCL.

Looking at the broader database context, the MI100 sits at the 96th percentile among all GPUs, while the CMP 40HX sits at the 93rd. The MI100’s average benchmark score is 139,035, and its nearest rivals are all within a narrow band: the NVIDIA Tesla V100 PCIe 16 GB (138,063, 0.7% slower), the Tesla V100 SXM2 32 GB (137,731, 0.9% slower), the AMD Radeon PRO V620 (136,472, 1.9% slower), and the AMD Radeon Pro W6800X Duo (135,774, 2.4% slower). The MI100 is effectively at the top of its immediate competitive cluster, with no rival in the list ahead of it.

The CMP 40HX, by contrast, has an average benchmark score of 85,637, which is pulled down by its second recorded result: a Geekbench Vulkan score of 77,879. Its OpenCL score of 93,395 is the stronger of the two. The nearest rivals for the CMP 40HX include the AMD Radeon PRO W7600 (87,108, 1.7% faster), the NVIDIA Quadro GP100 (87,445, 2.1% faster), the AMD Radeon PRO W6600 (81,995, 4.4% slower), and the AMD Radeon Pro Vega 64X (80,959, 5.8% slower). So in its own tier, the CMP 40HX is competitive but not dominant, sitting behind two rivals and ahead of two others.

The head-to-head delta of 48.9% is far larger than any gap observed within either card’s nearest rival group. This indicates that the two products are not direct competitors in the same performance class. The MI100 is operating in a higher tier, while the CMP 40HX is a mid-range part by comparison.

Architecture Differences

The two GPUs come from fundamentally different design philosophies. The AMD Instinct MI100 is built on the CDNA 1.0 architecture, using the Arcturus chip, and is fabricated on a 7 nm process at TSMC. It packs 25,600 million transistors onto a die size of 750 mm², resulting in a transistor density of 34.1 million per mm². The NVIDIA CMP 40HX uses the Turing architecture with the TU106 chip, fabricated on a 12 nm process, also at TSMC. It contains 10,800 million transistors on a 445 mm² die, giving a density of 24.3 million per mm². The MI100 has more than double the transistor count and a substantially larger die.

The compute resources differ dramatically. The MI100 has 7,680 shading units, 480 texture mapping units, and 64 ROPs. The CMP 40HX has 2,304 shading units, 144 TMUs, and 64 ROPs. While the ROP count is identical, the shading unit count is 3.3 times higher on the MI100. The CMP 40HX does include dedicated hardware that the MI100 lacks entirely: 36 RT cores and 288 tensor cores. The MI100 has no RT cores and no tensor cores listed in the database.

Clock speeds tell a different story. The CMP 40HX runs at a base clock of 1470 MHz and a boost clock of 1650 MHz, while the MI100 runs at 1000 MHz base and 1502 MHz boost. The NVIDIA part has higher raw clock rates, but the AMD part compensates with far more compute units. The pixel rate favors the CMP 40HX slightly: 105.6 GPixel/s versus 96.13 GPixel/s. The texture rate heavily favors the MI100: 721.0 GTexel/s versus 237.6 GTexel/s. FP32 throughput is 23.07 TFLOPS on the MI100 versus 7.603 TFLOPS on the CMP 40HX. FP16 throughput is 46.14 TFLOPS (2:1) versus 15.21 TFLOPS (2:1).

Memory architecture is another major divergence. The MI100 uses 32 GB of HBM2 on a 4096-bit bus, delivering 1.23 TB/s of bandwidth. The CMP 40HX uses 8 GB of GDDR6 on a 256-bit bus, delivering 448.0 GB/s. The MI100 has 2.7 times the bandwidth and 4 times the capacity.

Where Each One Wins

The AMD Instinct MI100 is the clear winner in OpenCL compute performance. Its 48.9% lead over the CMP 40HX in the head-to-head test is substantial, and its near-96th percentile standing among all GPUs places it among the top tier of accelerators. The MI100 also dominates in memory bandwidth, texture rate, and FP32/FP16 throughput, making it suitable for large-scale data processing, scientific simulation, and other memory-bound workloads.

The NVIDIA CMP 40HX, despite losing the OpenCL comparison, has specific strengths. It includes RT cores and tensor cores, which the MI100 does not have at all. This makes the CMP 40HX capable of ray tracing and tensor-accelerated workloads, areas where the MI100 has no listed capability. The CMP 40HX also has a higher pixel rate (105.6 GPixel/s versus 96.13 GPixel/s), which could benefit certain rasterization tasks, though the card has no display outputs and is not designed for graphics output.

The CMP 40HX also has a Vulkan benchmark result (77,879), while the MI100 has no Vulkan score recorded. This suggests the NVIDIA card at least has functional Vulkan support, while the MI100’s API support is listed as N/A for DirectX, OpenGL, and Vulkan. So for any workload relying on those APIs, the CMP 40HX is the only option between the two.

In terms of power efficiency, the CMP 40HX has a TDP of 185 W, while the MI100 has a TDP of 300 W. The NVIDIA card also requires only a single 8-pin power connector and a 450 W suggested PSU, while the MI100 needs two 8-pin connectors and a 700 W suggested PSU. The CMP 40HX is physically smaller as well: 229 mm in length versus 267 mm for the MI100.

Specification Differences

The database records the following differences between the two cards:

  • Process node: MI100 is 7 nm, CMP 40HX is 12 nm.
  • Transistor count: MI100 has 25,600 million, CMP 40HX has 10,800 million.
  • Die size: MI100 is 750 mm², CMP 40HX is 445 mm².
  • Transistor density: MI100 is 34.1M / mm², CMP 40HX is 24.3M / mm².
  • Base clock: MI100 is 1000 MHz, CMP 40HX is 1470 MHz.
  • Boost clock: MI100 is 1502 MHz, CMP 40HX is 1650 MHz.
  • Memory clock: MI100 is 1200 MHz (2.4 Gbps effective), CMP 40HX is 1750 MHz (14 Gbps effective).
  • Memory size: MI100 is 32 GB, CMP 40HX is 8 GB.
  • Memory type: MI100 is HBM2, CMP 40HX is GDDR6.
  • Memory bus width: MI100 is 4096 bit, CMP 40HX is 256 bit.
  • Memory bandwidth: MI100 is 1.23 TB/s, CMP 40HX is 448.0 GB/s.
  • Shading units: MI100 has 7,680, CMP 40HX has 2,304.
  • TMUs: MI100 has 480, CMP 40HX has 144.
  • RT cores: MI100 has none, CMP 40HX has 36.
  • Tensor cores: MI100 has none, CMP 40HX has 288.
  • Pixel rate: MI100 is 96.13 GPixel/s, CMP 40HX is 105.6 GPixel/s.
  • Texture rate: MI100 is 721.0 GTexel/s, CMP 40HX is 237.6 GTexel/s.
  • FP32: MI100 is 23.07 TFLOPS, CMP 40HX is 7.603 TFLOPS.
  • FP16: MI100 is 46.14 TFLOPS (2:1), CMP 40HX is 15.21 TFLOPS (2:1).
  • TDP: MI100 is 300 W, CMP 40HX is 185 W.
  • Power connectors: MI100 is 2x 8-pin, CMP 40HX is 1x 8-pin.
  • Suggested PSU: MI100 is 700 W, CMP 40HX is 450 W.
  • Bus interface: MI100 is PCIe 4.0 x16, CMP 40HX is PCIe 1.0 x4.
  • API support: MI100 lists N/A for DirectX, OpenGL, and Vulkan; CMP 40HX lists DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.
  • Dimensions: MI100 is 267 mm (10.5 inches) long and 111 mm (4.4 inches) high; CMP 40HX is 229 mm (9 inches) long, 111 mm (4.4 inches) high, and 35 mm (1.4 inches) wide.
  • Release date: MI100 was released on 2020-11-15, CMP 40HX on 2021-02-24.

The two cards share the same ROP count (64), the same dual-slot form factor, and both have no display outputs.

FAQ

Q: Which card has higher OpenCL performance?

A: The AMD Instinct MI100 scores 139,035 in Geekbench OpenCL, which is 48.9% higher than the NVIDIA CMP 40HX’s score of 93,395.

Q: Does the NVIDIA CMP 40HX have any compute features the MI100 lacks?

A: Yes, the CMP 40HX includes 36 RT cores and 288 tensor cores. The MI100 lists no RT cores and no tensor cores.

Q: What is the memory bandwidth difference?

A: The MI100 delivers 1.23 TB/s from 32 GB of HBM2 on a 4096-bit bus. The CMP 40HX delivers 448.0 GB/s from 8 GB of GDDR6 on a 256-bit bus.

Q: Which card supports modern graphics APIs?

A: The CMP 40HX lists DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The MI100 lists N/A for all three APIs.

Q: How do the power requirements compare?

A: The MI100 has a TDP of 300 W and requires two 8-pin power connectors with a 700 W suggested PSU. The CMP 40HX has a TDP of 185 W, one 8-pin connector, and a 450 W suggested PSU.

Q: What is the launch MSRP of the CMP 40HX?

A: The CMP 40HX has a launch MSRP of 699 USD. The MI100 has no launch MSRP recorded.

The Verdict

The data points to a clear split. The AMD Instinct MI100 is the superior choice for raw compute performance. Its OpenCL score is 48.9% higher, its FP32 throughput is 23.07 TFLOPS versus 7.603 TFLOPS, and its memory bandwidth is 1.23 TB/s versus 448.0 GB/s. The MI100 also offers 32 GB of HBM2 memory, which is critical for large datasets. It sits at the 96th percentile among all GPUs and leads its nearest rivals by margins of 0.7% to 2.4%. For anyone running OpenCL compute workloads, scientific simulations, or memory-intensive processing, the MI100 is the stronger accelerator.

The NVIDIA CMP 40HX is the better fit for workloads that require dedicated RT cores or tensor cores, as the MI100 has none. It also supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the MI100 lists no API support. The CMP 40HX has a higher pixel rate (105.6 GPixel/s versus 96.13 GPixel/s) and a lower TDP (185 W versus 300 W). It is also smaller physically, making it easier to install in constrained systems. Its launch MSRP is 699 USD, though the MI100 has no recorded launch price.

The CMP 40HX sits at the 93rd percentile, and its nearest rivals show it is competitive in its own tier, with deltas of -1.7%, -2.1%, +4.4%, and +5.8% against four other cards. But it is not in the same performance class as the MI100. The head-to-head delta of 48.9% is the deciding metric. If the task is pure compute, the MI100 wins decisively. If the task involves tensor or RT workloads, or requires modern graphics API support, the CMP 40HX is the only one of the two that can handle it.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI100
CMP 40HX
Core Specs
Shading Units
7,680
2,304 -70.0%
Shaders
7,680
2,304 -70.0%
TMUs
480
144 -70.0%
ROPs
64
64 0.0%
Compute Units
120
—
SM Count
—
36
Clocks
Base Clock
1000 MHz
1470 MHz
Boost Clock
1502 MHz
1650 MHz
Memory Clock
1200 MHz 2.4 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
32 GB
8 GB
VRAM (MB)
32,768
8,192 -75.0%
Memory Type
HBM2
GDDR6
Memory Bus
4096 bit
256 bit
Bandwidth
1.23 TB/s
448.0 GB/s
Cache
L1 Cache
16 KB (per CU)
64 KB (per SM)
L2 Cache
8 MB
4 MB
Performance
Pixel Rate
96.13 GPixel/s
105.6 GPixel/s
Texture Rate
721.0 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
23.07 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
11.54 TFLOPS (1:2)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
46.14 TFLOPS (2:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
—
36
Tensor Cores
—
288
Power
TDP
300 W
185 W
TDP (W)
300
185 -38.3%
Suggested PSU
700 W
450 W
Power Connectors
2x 8-pin
1x 8-pin
Architecture
Architecture
CDNA 1.0
Turing
GPU Name
Arcturus
TU106
Generation
Instinct (MIx)
Mining GPUs
Process Size
7 nm
12 nm
Transistors
25,600 million
10,800 million
Die Size
750 mm²
445 mm²
Foundry
TSMC
TSMC
Density
34.1M / mm²
24.3M / mm²
API Support
DirectX
—
12 Ultimate (12_2)
OpenGL
—
4.6
Vulkan
—
1.4
OpenCL
2.1
3.0
CUDA
—
7.5
Shader Model
—
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
229 mm 9 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 1.0 x4
Other
Launch Price
—
699 USD
Production
End-of-life
End-of-life
Predecessor
Radeon Instinct
—
View Instinct MI100 Details View CMP 40HX Details