GPU Comparison

AMD
RADEON

AMD Radeon VII

CORE STATE Vega 20
VRAM 16 GB
CLOCK SPEED 1750 MHz
TDP 295 W
BUS WIDTH 4096 bit
ARCHITECTURE GCN 5.1
nm
PROCESS 7 nm
LAUNCH DATE 2019
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
2,304
N/A
geekbench_metal
77,975
N/A
geekbench_opencl
91,947
93,395
geekbench_vulkan
91,788
77,879

Analysis: AMD Radeon VII vs NVIDIA CMP 40HX

NVIDIA CMP 40HX and AMD Radeon VII are both end-of-life products aimed at different audiences, with the CMP 40HX being a mining-focused GPU with no display outputs and the Radeon VII a consumer flagship with full display connectivity. Benchmark data shows the CMP 40HX edges out the Radeon VII in OpenCL (93395 vs 91947, a 1.6% lead) but falls significantly behind in Vulkan (77879 vs 91788, a 15.2% deficit). The Radeon VII holds a higher average benchmark score in its own rival context, but the CMP 40HX sits at the 93rd percentile versus the Radeon VII’s 90th percentile across all GPUs. These two cards occupy different niches, yet their compute-oriented specifications make direct comparisons relevant for workloads that leverage raw throughput.

The Verdict

From the data, the NVIDIA CMP 40HX is the pick for users who prioritize OpenCL compute performance and do not require display outputs. Its 93395 OpenCL score tops the Radeon VII’s 91947, and its 93rd percentile ranking places it above the Radeon VII’s 90th percentile. The CMP 40HX also has a lower TDP (185 W vs 295 W) and a smaller physical footprint (229 mm length vs 280 mm), making it easier to integrate into systems with tighter power and space constraints. However, its PCIe 1.0 x4 interface is a severe limitation for data transfer, and the lack of display outputs means it cannot serve as a primary graphics solution.

The AMD Radeon VII is the better choice for Vulkan-based workloads or for users who need a functional display output alongside compute capability. Its Vulkan score of 91788 is 15.2% higher than the CMP 40HX’s 77879, and it offers 16 GB of HBM2 memory with 1.02 TB/s bandwidth, more than double the CMP 40HX’s 8 GB GDDR6 at 448.0 GB/s. The Radeon VII’s 13.44 TFLOPS FP32 and 420.0 GTexel/s texture rate dwarf the CMP 40HX’s 7.603 TFLOPS and 237.6 GTexel/s, indicating superior raw compute throughput in most scenarios. Its PCIe 3.0 x16 interface and display outputs (1x HDMI 2.0b, 3x DisplayPort 1.4a) make it a more versatile card, despite a higher power draw of 295 W and a larger 280 mm length.

For mining-specific applications, the CMP 40HX’s lower power consumption and compact design are advantageous, but its PCIe 1.0 x4 bus could bottleneck performance in memory-intensive tasks. The Radeon VII, while older (released 2019-02-06 vs 2021-02-24), offers higher memory capacity and bandwidth, which benefits large dataset workloads. Both cards launched at the same MSRP (699 USD), but the data does not support a clear overall winner, each leads in different benchmark categories.

Architecture Differences

The NVIDIA CMP 40HX uses the TU106 chip on a 12 nm TSMC process, with 10,800 million transistors on a 445 mm² die (24.3M / mm² density). It is built on the Turing architecture and features 2304 shading units, 144 TMUs, 64 ROPs, 36 RT cores, and 288 tensor cores. Its FP32 throughput is 7.603 TFLOPS, with FP16 at 15.21 TFLOPS (2:1). The memory subsystem consists of 8 GB GDDR6 on a 256-bit bus, delivering 448.0 GB/s bandwidth. It supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, but has no display outputs.

The AMD Radeon VII employs the Vega 20 chip on a 7 nm TSMC process, with 13,230 million transistors on a 331 mm² die (40.0M / mm² density). This GCN 5.1 architecture has 3840 shading units, 240 TMUs, and 64 ROPs, with no dedicated RT or tensor cores. Its FP32 performance is 13.44 TFLOPS, and FP16 reaches 26.88 TFLOPS (2:1). Memory is 16 GB of HBM2 on a 4096-bit bus, providing 1.02 TB/s bandwidth, a massive advantage over the CMP 40HX. The Radeon VII supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.3, and includes 1x HDMI 2.0b and 3x DisplayPort 1.4a outputs.

The process node difference (12 nm vs 7 nm) contributes to the Radeon VII’s higher transistor density (40.0M / mm² vs 24.3M / mm²) and lower die size (331 mm² vs 445 mm²) despite having more transistors. The CMP 40HX’s Turing architecture brings RT and tensor cores, which the Radeon VII lacks entirely. However, the Radeon VII’s raw compute resources (3840 shading units vs 2304) and memory bandwidth (1.02 TB/s vs 448.0 GB/s) are substantially higher, explaining its lead in Vulkan and its higher theoretical peak rates.

FAQ

Q: Which card has higher FP32 performance?

A: The AMD Radeon VII delivers 13.44 TFLOPS FP32, which is 76.8% higher than the NVIDIA CMP 40HX’s 7.603 TFLOPS.

Q: How do their memory bandwidths compare?

A: The Radeon VII’s HBM2 memory provides 1.02 TB/s bandwidth, while the CMP 40HX’s GDDR6 offers 448.0 GB/s, the Radeon VII has more than double the bandwidth.

Q: What are the power requirements for each card?

A: The CMP 40HX has a TDP of 185 W with a suggested PSU of 450 W and one 8-pin connector. The Radeon VII has a TDP of 295 W, a suggested PSU of 600 W, and requires two 8-pin connectors.

Q: Can either card be used for display output?

A: No. The CMP 40HX has no display outputs. The Radeon VII has 1x HDMI 2.0b and 3x DisplayPort 1.4a outputs.

Q: Which card performs better in Vulkan benchmarks?

A: The Radeon VII scores 91788 in Geekbench Vulkan, which is 15.2% higher than the CMP 40HX’s 77879.

Q: What is the memory capacity difference?

A: The Radeon VII has 16 GB of HBM2, while the CMP 40HX has 8 GB of GDDR6, the Radeon VII offers twice the memory capacity.

Specification Differences

| Specification | NVIDIA CMP 40HX | AMD Radeon VII |

|----------------|-----------------|----------------|

| Process Node | 12 nm | 7 nm |

| Transistors | 10,800 million | 13,230 million |

| Die Size | 445 mm² | 331 mm² |

| Transistor Density | 24.3M / mm² | 40.0M / mm² |

| Base Clock | 1470 MHz | 1400 MHz |

| Boost Clock | 1650 MHz | 1750 MHz |

| Memory Clock | 1750 MHz (14 Gbps effective) | 1000 MHz (2 Gbps effective) |

| Memory Size | 8 GB GDDR6 | 16 GB HBM2 |

| Memory Bus Width | 256 bit | 4096 bit |

| Memory Bandwidth | 448.0 GB/s | 1.02 TB/s |

| Shading Units | 2304 | 3840 |

| TMUs | 144 | 240 |

| RT Cores | 36 | null |

| Tensor Cores | 288 | null |

| Pixel Rate | 105.6 GPixel/s | 112.0 GPixel/s |

| Texture Rate | 237.6 GTexel/s | 420.0 GTexel/s |

| FP32 | 7.603 TFLOPS | 13.44 TFLOPS |

| FP16 | 15.21 TFLOPS (2:1) | 26.88 TFLOPS (2:1) |

| TDP | 185 W | 295 W |

| Power Connectors | 1x 8-pin | 2x 8-pin |

| Suggested PSU | 450 W | 600 W |

| Bus Interface | PCIe 1.0 x4 | PCIe 3.0 x16 |

| Display Outputs | No outputs | 1x HDMI 2.0b, 3x DisplayPort 1.4a |

| DirectX | 12 Ultimate (12_2) | 12 (12_1) |

| Vulkan | 1.4 | 1.3 |

| Length | 229 mm | 280 mm |

| Height | 111 mm | 125 mm |

| Width | 35 mm | 40 mm |

| Release Date | 2021-02-24 | 2019-02-06 |

| Launch MSRP | 699 USD | 699 USD |

Head-to-Head Benchmarks

The head-to-head benchmark data covers two tests: Geekbench OpenCL and Geekbench Vulkan. In OpenCL, the NVIDIA CMP 40HX wins with a score of 93395 against the AMD Radeon VII’s 91947, a delta of 1.6%. This narrow margin suggests that despite the Radeon VII’s higher FP32 throughput (13.44 TFLOPS vs 7.603 TFLOPS), the CMP 40HX’s Turing architecture with its tensor cores and RT cores may be better optimized for OpenCL workloads in this specific benchmark. The CMP 40HX’s higher base clock (1470 MHz vs 1400 MHz) and boost clock (1650 MHz vs 1750 MHz) partially offset its lower core count, though the Radeon VII’s boost clock is 100 MHz higher.

In Vulkan, the results reverse dramatically. The AMD Radeon VII scores 91788, while the NVIDIA CMP 40HX manages only 77879, a 15.2% deficit for the CMP 40HX. This is a substantial gap that aligns with the Radeon VII’s superior texture rate (420.0 GTexel/s vs 237.6 GTexel/s) and pixel rate (112.0 GPixel/s vs 105.6 GPixel/s). The Radeon VII’s 16 GB of HBM2 with 1.02 TB/s bandwidth likely provides a significant advantage in memory-bound Vulkan scenes, whereas the CMP 40HX’s 8 GB GDDR6 at 448.0 GB/s may bottleneck under high-resolution textures.

The wins are split evenly: one benchmark victory for each card. However, the magnitude of the Vulkan loss for the CMP 40HX (15.2%) far exceeds its OpenCL gain (1.6%). This asymmetry indicates that the Radeon VII is the more consistently performant card in compute-heavy scenarios, despite its lower percentile ranking (90th vs 93rd). The CMP 40HX’s average benchmark score of 85637 is higher than the Radeon VII’s 66004, but this average includes different benchmark suites, the Radeon VII has additional tests (3DMark Steel Nomad DX12 at 2304 and Geekbench Metal at 77975) that pull its average down. The CMP 40HX’s nearest rivals include the AMD Radeon PRO W7600 (87108, -1.7% delta) and NVIDIA Quadro GP100 (87445, -2.1% delta), while the Radeon VII’s nearest rivals are NVIDIA Tesla T4 (66733, -1.1% delta) and Tesla P40 (65095, 1.4% delta). These rival comparisons confirm that the CMP 40HX sits in a higher performance tier by average score, but the Radeon VII’s Vulkan strength makes it a formidable competitor in specific APIs.

For users migrating from older GCN cards, the Radeon VII’s successor is Navi, while the CMP 40HX has no listed predecessor or successor, reflecting its specialized mining niche. The Radeon VII’s predecessor is Vega, and its release in 2019 makes it two years older than the CMP 40HX’s 2021 launch. Despite the age difference, the Radeon VII’s memory subsystem and compute resources remain competitive, as evidenced by its Vulkan lead. The CMP 40HX counters with lower power consumption (185 W vs 295 W) and a more compact design (229 mm vs 280 mm), which are meaningful for dense mining rigs but less so for general compute applications.

DETAILED SPECIFICATIONS

SPECIFICATION
VII
CMP 40HX
Core Specs
Shading Units
3,840
2,304 -40.0%
Shaders
3,840
2,304 -40.0%
TMUs
240
144 -40.0%
ROPs
64
64 0.0%
Compute Units
60
SM Count
36
Clocks
Base Clock
1400 MHz
1470 MHz
Boost Clock
1750 MHz
1650 MHz
Memory Clock
1000 MHz 2 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
16 GB
8 GB
VRAM (MB)
16,384
8,192 -50.0%
Memory Type
HBM2
GDDR6
Memory Bus
4096 bit
256 bit
Bandwidth
1.02 TB/s
448.0 GB/s
Cache
L1 Cache
16 KB (per CU)
64 KB (per SM)
L2 Cache
4 MB
4 MB
Performance
Pixel Rate
112.0 GPixel/s
105.6 GPixel/s
Texture Rate
420.0 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
13.44 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
3.360 TFLOPS (1:4)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
26.88 TFLOPS (2:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
36
Tensor Cores
288
Power
TDP
295 W
185 W
TDP (W)
295
185 -37.3%
Suggested PSU
600 W
450 W
Power Connectors
2x 8-pin
1x 8-pin
Architecture
Architecture
GCN 5.1
Turing
GPU Name
Vega 20
TU106
Generation
Vega II (Radeon VII)
Mining GPUs
Process Size
7 nm
12 nm
Transistors
13,230 million
10,800 million
Die Size
331 mm²
445 mm²
Foundry
TSMC
TSMC
Density
40.0M / mm²
24.3M / mm²
API Support
DirectX
12 (12_1)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
7.5
Shader Model
6.7
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
280 mm 11 inches
229 mm 9 inches
Height
125 mm 4.9 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.0b3x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 1.0 x4
Other
Launch Price
699 USD
699 USD
Production
End-of-life
End-of-life
Predecessor
Vega
Successor
Navi
View Radeon VII Details View CMP 40HX Details