AMD Radeon Pro Vega 48 vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Radeon Pro Vega 48

CORE STATE Vega 10
VRAM 8 GB
CLOCK SPEED
TDP
BUS WIDTH 2048 bit
ARCHITECTURE GCN 5.0
nm
PROCESS 14 nm
LAUNCH DATE 2019
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_metal
69,010
N/A
geekbench_opencl
53,757
93,395
geekbench_vulkan
57,653
77,879

Analysis: AMD Radeon Pro Vega 48 vs NVIDIA CMP 40HX

Where Each One Wins

The recorded data splits cleanly along API boundaries. The NVIDIA CMP 40HX wins both shared benchmark categories, and it wins them by substantial margins. In OpenCL, the CMP 40HX scores 93,395 against the Radeon Pro Vega 48's 53,757, a 73.7% advantage. In Vulkan, the gap narrows but remains decisive: 77,879 versus 57,653, a 35.1% lead. There is no benchmark category in the database where the AMD card comes out ahead. The Radeon Pro Vega 48 does have one exclusive result, a Metal score of 69,010, but that test has no corresponding entry for the NVIDIA part, so it cannot be used as a direct comparison.

The use-case split is therefore not about which card wins workloads, but about which environments are even accessible. The CMP 40HX is a mining-oriented part with no display outputs, so it exists purely for compute tasks that do not require a frame buffer output. The Radeon Pro Vega 48 is an integrated GPU for portable Mac systems, with display output described as "Portable Device Dependent". Its Metal benchmark score suggests it is intended for Apple ecosystem workloads, while the NVIDIA card never appears in Metal testing at all. The data indicates that the NVIDIA part dominates raw compute throughput in the two APIs where both are measured, while the AMD part serves a different platform niche entirely.

Architecture Differences

The two GPUs come from fundamentally different design philosophies. The NVIDIA CMP 40HX uses the TU106 chip on TSMC's 12 nm process, built on the Turing architecture. It packs 10,800 million transistors into a 445 mm² die, yielding a transistor density of 24.3 million per square millimeter. The AMD Radeon Pro Vega 48 uses the Vega 10 chip on GlobalFoundries' 14 nm process, using the older GCN 5.0 architecture. It carries 12,500 million transistors across a larger 495 mm² die, with a density of 25.3 million per square millimeter. The AMD chip is physically larger and has more transistors, yet it delivers lower benchmark scores, which points to architectural efficiency differences rather than raw resource counts.

Shader configuration tells a similar story. The Radeon Pro Vega 48 fields 3,072 shading units, 192 texture mapping units, and 64 ROPs. The CMP 40HX has 2,304 shading units, 144 TMUs, and 64 ROPs. Despite having 768 fewer shaders and 48 fewer TMUs, the NVIDIA part produces higher texture and pixel rates in the measured data: 237.6 GTexel/s versus 230.4 GTexel/s, and 105.6 GPixel/s versus 76.80 GPixel/s. The CMP 40HX also reaches 7.603 TFLOPS FP32 against the Vega 48's 7.373 TFLOPS, a modest edge. The NVIDIA card additionally includes 36 RT cores and 288 tensor cores, features entirely absent from the AMD part. Memory subsystems diverge sharply: the CMP 40HX uses 8 GB of GDDR6 on a 256 bit bus with 448.0 GB/s bandwidth, while the Vega 48 uses 8 GB of HBM2 on a 2048 bit bus with 402.4 GB/s bandwidth. The wide HBM2 bus cannot overcome its lower effective clock. Both cards support similar API levels, with NVIDIA reaching DirectX 12 Ultimate (12_2) and Vulkan 1.4, while AMD tops out at DirectX 12 (12_1) and Vulkan 1.3.

Head-to-Head Benchmarks

The OpenCL result is the single largest gap between the two cards. The CMP 40HX posts 93,395 versus 53,757 for the Radeon Pro Vega 48, a 73.7% difference. That is not a small margin; it is a dominant result across every workload that OpenCL covers. The CMP 40HX lands at the 93rd percentile among all GPUs in the database, with an average benchmark score of 85,637. The Vega 48 sits at the 88th percentile with an average of 60,140. The percentile gap of five points understates the score gap, because the field is crowded near the lower end.

The Vulkan test tells a more nuanced story. The CMP 40HX scores 77,879, and the Vega 48 scores 57,653. The 35.1% advantage is still large, but it is roughly half the OpenCL margin. This suggests the AMD architecture holds up relatively better under Vulkan's lower-level API overhead, though it still loses decisively. The CMP 40HX's nearest rivals in the database include the AMD Radeon PRO W7600 at 87,108 (1.7% lower), the NVIDIA Quadro GP100 at 87,445 (2.1% lower), the AMD Radeon PRO W6600 at 81,995 (4.4% higher), and the AMD Radeon Pro Vega 64X at 80,959 (5.8% higher). The Vega 48's nearest rivals are much closer in absolute terms: the Intel Arc Pro A60 at 60,326 (0.3% lower), the NVIDIA GeForce RTX 4090 at 60,347 (0.3% lower), the AMD Radeon PRO V710 at 58,657 (2.5% higher), and the NVIDIA P102-100 at 58,528 (2.8% higher). The RTX 4090 appearing at nearly the same average score as the Vega 48 is a reminder that average benchmark scores can compress very different performance profiles into a single number.

FAQ

Q: Which card has the higher average benchmark score?

A: The NVIDIA CMP 40HX has an average benchmark score of 85,637, while the AMD Radeon Pro Vega 48 averages 60,140, a difference of roughly 42%.

Q: Does the AMD card win any shared benchmark?

A: No. The CMP 40HX wins both shared tests, Geekbench OpenCL by 73.7% and Geekbench Vulkan by 35.1%. The Vega 48 has a Metal score of 69,010, but the NVIDIA card has no Metal result to compare against.

Q: How do their memory configurations differ?

A: Both cards have 8 GB of memory, but the CMP 40HX uses GDDR6 on a 256 bit bus with 448.0 GB/s bandwidth, while the Vega 48 uses HBM2 on a 2048 bit bus with 402.4 GB/s bandwidth.

Q: What is the transistor situation on each chip?

A: The CMP 40HX has 10,800 million transistors on a 445 mm² die, and the Vega 48 has 12,500 million transistors on a 495 mm² die. The AMD chip is larger in both measures.

Q: Do both cards support the same DirectX version?

A: No. The CMP 40HX supports DirectX 12 Ultimate (12_2), while the Vega 48 supports DirectX 12 (12_1).

Q: What is the CMP 40HX's percentile ranking?

A: The CMP 40HX falls in the 93rd percentile among all GPUs in the database, compared to the Vega 48's 88th percentile.

Specification Differences

The two cards differ across nearly every specification category. The CMP 40HX uses a 12 nm process from TSMC, while the Vega 48 uses a 14 nm process from GlobalFoundries. Transistor counts are 10,800 million versus 12,500 million, and die sizes are 445 mm² versus 495 mm². Transistor density is 24.3M per mm² for NVIDIA and 25.3M per mm² for AMD. Base and boost clocks are listed for the CMP 40HX at 1470 MHz and 1650 MHz, while the Vega 48 has no base or boost clock data recorded. Memory clocks also diverge: the CMP 40HX runs at 1750 MHz with 14 Gbps effective, and the Vega 48 runs at 786 MHz with 1572 Mbps effective.

Shader counts put AMD ahead numerically: 3,072 shading units and 192 TMUs versus 2,304 shading units and 144 TMUs. Both have 64 ROPs. The CMP 40HX includes 36 RT cores and 288 tensor cores, which the Vega 48 lacks entirely. Pixel rate is 105.6 GPixel/s versus 76.80 GPixel/s, and texture rate is 237.6 GTexel/s versus 230.4 GTexel/s. FP32 throughput is 7.603 TFLOPS versus 7.373 TFLOPS, and FP16 is 15.21 TFLOPS versus 14.75 TFLOPS. The CMP 40HX has a TDP of 185 W with a suggested 450 W PSU, while the Vega 48 has no TDP or PSU recommendation recorded. The NVIDIA card requires a dual-slot cooler and a single 8-pin power connector; the Vega 48 is an IGP with no power connectors. Bus interfaces are PCIe 1.0 x4 for the CMP 40HX and PCIe 3.0 x16 for the Vega 48. Display outputs are "No outputs" for NVIDIA and "Portable Device Dependent" for AMD. The CMP 40HX measures 229 mm by 111 mm by 35 mm, while the Vega 48 has no physical dimensions recorded. Release dates are February 2021 for the CMP 40HX and March 2019 for the Vega 48. The CMP 40HX has a launch MSRP of 699 USD; the Vega 48 has none recorded.

The Verdict

The data points to a clear performance hierarchy. The NVIDIA CMP 40HX outperforms the AMD Radeon Pro Vega 48 in every shared benchmark by margins ranging from 35.1% to 73.7%. It also achieves a higher average benchmark score, a higher percentile ranking, higher pixel and texture rates, higher FP32 and FP16 throughput, and higher memory bandwidth. The AMD card has more shading units, more TMUs, more transistors, and a wider memory bus, but none of those advantages translate into benchmark wins. The Vega 48's only exclusive strength is its Metal score of 69,010, which matters only in Apple platform contexts where the NVIDIA card cannot participate at all.

For compute workloads in OpenCL or Vulkan, the CMP 40HX is the superior choice by every recorded metric. Its lack of display outputs is irrelevant for mining or headless compute tasks, and its 93rd percentile ranking places it well above the Vega 48's 88th percentile. The Vega 48, by contrast, appears designed for portable Mac integration, with its IGP form factor, portable-device-dependent display outputs, and no power connectors. If the workload requires Metal support or lives inside a Mac portable chassis, the Vega 48 is the only option between the two. If the workload runs on a standard PCIe platform and uses OpenCL or Vulkan, the CMP 40HX wins outright. The recorded data does not support any scenario where the AMD card is faster in a directly comparable test, but it does support the conclusion that the two cards target different platforms altogether. Choose the CMP 40HX for raw compute throughput on a desktop system; choose the Vega 48 only when the platform demands it.

DETAILED SPECIFICATIONS

SPECIFICATION
Pro Vega 48
CMP 40HX
Core Specs
Shading Units
3,072
2,304 -25.0%
Shaders
3,072
2,304 -25.0%
TMUs
192
144 -25.0%
ROPs
64
64 0.0%
Compute Units
48
SM Count
36
Clocks
Base Clock
1470 MHz
Boost Clock
1650 MHz
GPU Clock
1200 MHz
Memory Clock
786 MHz 1572 Mbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
8 GB
8 GB
VRAM (MB)
8,192
8,192 0.0%
Memory Type
HBM2
GDDR6
Memory Bus
2048 bit
256 bit
Bandwidth
402.4 GB/s
448.0 GB/s
Cache
L1 Cache
16 KB (per CU)
64 KB (per SM)
L2 Cache
4 MB
4 MB
Performance
Pixel Rate
76.80 GPixel/s
105.6 GPixel/s
Texture Rate
230.4 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
7.373 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
460.8 GFLOPS (1:16)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
14.75 TFLOPS (2:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
36
Tensor Cores
288
Power
TDP
185 W
TDP (W)
185
Suggested PSU
450 W
Power Connectors
None
1x 8-pin
Architecture
Architecture
GCN 5.0
Turing
GPU Name
Vega 10
TU106
Generation
Radeon Pro Mac (Vega Series)
Mining GPUs
Process Size
14 nm
12 nm
Transistors
12,500 million
10,800 million
Die Size
495 mm²
445 mm²
Foundry
GlobalFoundries
TSMC
Density
25.3M / mm²
24.3M / mm²
API Support
DirectX
12 (12_1)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
7.5
Shader Model
6.7
6.8
Physical
Slot Width
IGP
Dual-slot
Length
229 mm 9 inches
Height
111 mm 4.4 inches
Outputs
Portable Device Dependent
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 1.0 x4
Other
Launch Price
699 USD
Production
End-of-life
End-of-life
View Radeon Pro Vega 48 Details View CMP 40HX Details