GPU Comparison

AMD
RADEON

AMD Radeon PRO W7600

CORE STATE Navi 33
VRAM 8 GB
CLOCK SPEED 2440 MHz
TDP 130 W
BUS WIDTH 128 bit
ARCHITECTURE RDNA 3.0
nm
PROCESS 6 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
81,528
93,395
geekbench_vulkan
92,688
77,879

Analysis: AMD Radeon PRO W7600 vs NVIDIA CMP 40HX

The AMD Radeon PRO W7600 and NVIDIA CMP 40HX are two very different products that happen to land near each other in aggregate benchmark scores, yet the data reveals a split decision depending on the workload. The AMD card averages 85,851 across two Geekbench tests, while the NVIDIA card averages 85,637, a margin of only 0.2% in favor of the AMD part. Both sit at the 94th percentile of all GPUs, placing them in the upper tier of performance. However, that near-identical average masks a fundamental divergence: the NVIDIA CMP 40HX dominates OpenCL compute, while the AMD Radeon PRO W7600 wins decisively in Vulkan graphics workloads. The two cards share the same 8 GB memory capacity and 64 ROPs, but nearly every other specification tells a story of two distinct design philosophies — one aimed at professional graphics, the other at mining operations.

Head-to-Head Benchmarks

The most striking result in the head-to-head data is the NVIDIA CMP 40HX’s performance in Geekbench OpenCL, where it scores 93,395 against the AMD Radeon PRO W7600’s 81,528. That is a 12.7% advantage for the NVIDIA card, a substantial lead in a compute-oriented API. The CMP 40HX’s advantage here is not marginal; it is a clear, double-digit gap that suggests its architecture is better suited to the type of parallel floating-point work that OpenCL often represents. For any workload that leans on OpenCL, the data indicates the NVIDIA card is the stronger choice by a meaningful margin.

The reverse is true in Geekbench Vulkan, where the AMD Radeon PRO W7600 scores 90,174 compared to the CMP 40HX’s 77,879. That gives AMD a 15.8% lead, which is even larger than NVIDIA’s OpenCL advantage. The Vulkan test typically stresses graphics pipeline throughput, and the AMD card’s victory here is decisive. The delta between the two cards in Vulkan is 12,295 points, a wider gap than the 11,867-point difference in OpenCL, making the AMD win the more pronounced of the two. Netting out the two tests, the AMD card’s average of 85,851 edges out the CMP 40HX’s 85,637 by 0.2%, but that aggregate number is less informative than the individual test results, which show each card winning one discipline outright.

Looking at the nearest rivals for context, the AMD Radeon PRO W7600 leads the NVIDIA GeForce RTX 5090 by 1.8% and the RTX 5090 D by 1.9%, while the NVIDIA CMP 40HX holds a 1.6% advantage over the RTX 5090 and a 1.7% edge over the RTX 5090 D. Both cards also outperform the RTX 5050 Mobile by roughly 2%. These deltas are small, but they show that both cards are competitive with far more modern flagship hardware, at least in these aggregate metrics.

Architecture Differences

The architectural divide is stark. The AMD Radeon PRO W7600 uses the Navi 33 chip built on RDNA 3.0 architecture, with a 6 nm process from TSMC. It packs 13,300 million transistors into a 204 mm² die, yielding a transistor density of 65.2 million per mm². The NVIDIA CMP 40HX, by contrast, uses the TU106 chip on the older Turing architecture, fabricated on a 12 nm process also from TSMC. It contains 10,800 million transistors spread across a much larger 445 mm² die, resulting in a density of just 24.3 million per mm². The AMD chip is more than twice as dense, reflecting the newer process node, and it runs at significantly higher clocks — 1720 MHz base and 2440 MHz boost versus the NVIDIA’s 1470 MHz base and 1650 MHz boost.

Memory subsystems diverge as well. Both cards have 8 GB of GDDR6, but the AMD card uses a 128-bit bus with 288.0 GB/s of bandwidth, while the NVIDIA card uses a 256-bit bus delivering 448.0 GB/s. That 160 GB/s bandwidth advantage for NVIDIA is substantial and likely contributes to its OpenCL win. The AMD card’s memory runs at 2250 MHz (18 Gbps effective), while the NVIDIA card’s memory runs at 1750 MHz (14 Gbps effective); despite the lower clock, the wider bus gives NVIDIA the bandwidth edge.

Compute resources tell a different story. The AMD card has 2048 shading units, 128 TMUs, 32 RT cores, and no tensor cores. The NVIDIA card has 2304 shading units, 144 TMUs, 36 RT cores, and 288 tensor cores. NVIDIA leads in raw shader and texture counts, and its tensor cores are absent from the AMD part entirely. Yet the AMD card’s FP32 throughput is 19.99 TFLOPS, more than 2.6 times the CMP 40HX’s 7.603 TFLOPS. The AMD card also leads in FP16 at 39.98 TFLOPS versus 15.21 TFLOPS. Pixel rates favor AMD at 156.2 GPixel/s versus 105.6 GPixel/s, and texture rates favor AMD at 312.3 GTexel/s versus 237.6 GTexel/s. Despite fewer shading units, AMD’s higher clocks and architecture efficiency produce far higher raw throughput numbers.

The CMP 40HX has no display outputs at all, a clear sign of its mining purpose, while the AMD card offers 4x DisplayPort 2.1 outputs. The NVIDIA card uses a PCIe 1.0 x4 bus interface, a curious and limiting choice, versus the AMD card’s PCIe 4.0 x8. Power requirements also differ: the AMD card draws 130 W TDP with a 300 W suggested PSU and a single 6-pin connector, while the NVIDIA card draws 185 W TDP with a 450 W suggested PSU and a single 8-pin connector. The AMD card is single-slot, the NVIDIA card dual-slot. The AMD card is 241 mm long, 115 mm tall; the NVIDIA card is 229 mm long, 111 mm tall, and 35 mm wide.

Where Each One Wins

The NVIDIA CMP 40HX wins in OpenCL compute by a 12.7% margin, which is its only benchmark victory but a significant one. That result, combined with its 448.0 GB/s of memory bandwidth and 288 tensor cores, suggests it is better suited to workloads that rely heavily on memory throughput and tensor operations. Its 2304 shading units and 144 TMUs also provide more raw geometry and texture processing elements than the AMD card. For compute tasks that scale with bandwidth and CUDA-style parallelism, the data points to the CMP 40HX.

The AMD Radeon PRO W7600 wins in Vulkan by 15.8%, a larger margin than NVIDIA’s OpenCL lead. Its 19.99 TFLOPS of FP32 performance, 156.2 GPixel/s pixel rate, and 312.3 GTexel/s texture rate are all far ahead of the NVIDIA card, which explains its Vulkan dominance. The AMD card also has display outputs, making it usable for actual graphics output and professional visualization, whereas the CMP 40HX cannot drive a monitor. The AMD card’s 130 W TDP and single-slot design make it far easier to integrate into a workstation, and its PCIe 4.0 x8 interface provides much higher host bandwidth than the CMP 40HX’s PCIe 1.0 x4.

Specification Differences

The two cards differ on nearly every specification except memory size (both 8 GB), memory type (both GDDR6), ROP count (both 64), and API support (both DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4). The AMD card is built on a 6 nm process, the NVIDIA on 12 nm. AMD has 13,300 million transistors on 204 mm²; NVIDIA has 10,800 million on 445 mm². AMD’s base clock is 1720 MHz versus 1470 MHz, and its boost clock is 2440 MHz versus 1650 MHz. AMD’s memory clock is 2250 MHz (18 Gbps effective) versus 1750 MHz (14 Gbps effective). AMD’s bus width is 128 bit versus 256 bit, and its bandwidth is 288.0 GB/s versus 448.0 GB/s. AMD has 2048 shading units, 128 TMUs, and 32 RT cores; NVIDIA has 2304 shading units, 144 TMUs, 36 RT cores, and 288 tensor cores. AMD’s FP32 is 19.99 TFLOPS versus 7.603 TFLOPS, and its FP16 is 39.98 TFLOPS versus 15.21 TFLOPS. Pixel rate is 156.2 GPixel/s versus 105.6 GPixel/s, and texture rate is 312.3 GTexel/s versus 237.6 GTexel/s. TDP is 130 W versus 185 W. AMD is single-slot with a 6-pin connector; NVIDIA is dual-slot with an 8-pin connector. Suggested PSU is 300 W versus 450 W. Bus interface is PCIe 4.0 x8 versus PCIe 1.0 x4. AMD has 4x DisplayPort 2.1 outputs; NVIDIA has none. AMD is 241 mm long and 115 mm tall; NVIDIA is 229 mm long, 111 mm tall, and 35 mm wide.

FAQ

Q: Which card has the higher average benchmark score?

A: The AMD Radeon PRO W7600 averages 85,851, which is 0.2% higher than the NVIDIA CMP 40HX’s 85,637.

Q: How much faster is the NVIDIA CMP 40HX in OpenCL?

A: The CMP 40HX scores 93,395 in Geekbench OpenCL versus the AMD card’s 81,528, a 12.7% advantage.

Q: How much faster is the AMD Radeon PRO W7600 in Vulkan?

A: The AMD card scores 90,174 in Geekbench Vulkan versus the CMP 40HX’s 77,879, a 15.8% lead.

Q: Which card has more memory bandwidth?

A: The NVIDIA CMP 40HX has 448.0 GB/s of bandwidth due to its 256-bit bus, compared to the AMD card’s 288.0 GB/s from a 128-bit bus.

Q: Which card has higher FP32 performance?

A: The AMD Radeon PRO W7600 delivers 19.99 TFLOPS of FP32, while the NVIDIA CMP 40HX delivers 7.603 TFLOPS.

Q: Can the NVIDIA CMP 40HX output video to a display?

A: No, the CMP 40HX has no display outputs, while the AMD Radeon PRO W7600 has 4x DisplayPort 2.1 outputs.

The Verdict

The data presents a clear choice based on use case. For OpenCL-heavy compute workloads, the NVIDIA CMP 40HX is the stronger performer, with a 12.7% lead that is backed by its wider 256-bit memory bus and 448.0 GB/s of bandwidth. Its 288 tensor cores and higher shading unit count also give it an edge in certain parallel compute tasks, even though its overall FP32 throughput is far lower than the AMD card’s. However, the CMP 40HX’s lack of display outputs, end-of-life production status, and older 12 nm process make it a niche product for specialized compute or mining tasks, not a general-purpose GPU.

For Vulkan-based graphics and any workload requiring display output, the AMD Radeon PRO W7600 is the definitive winner, holding a 15.8% advantage in that test. Its 19.99 TFLOPS of FP32, 156.2 GPixel/s pixel rate, and 312.3 GTexel/s texture rate are all dramatically higher than the NVIDIA card’s, and its 6 nm process, single-slot design, and 130 W TDP make it a far more practical workstation component. The AMD card is also actively in production, with a 94th percentile ranking matching the NVIDIA card. If the workload is primarily graphics or Vulkan compute, the AMD card wins outright. If the workload is specifically OpenCL and does not require a display, the NVIDIA card’s 12.7% lead in that single test is the deciding factor. The aggregate scores differ by only 0.2%, but the distribution of wins could not be more lopsided — each card wins its own domain by double digits.

DETAILED SPECIFICATIONS

SPECIFICATION
PRO W7600
CMP 40HX
Core Specs
Shading Units
2,048
2,304 +12.5%
Shaders
2,048
2,304 +12.5%
TMUs
128
144 +12.5%
ROPs
64
64 0.0%
Compute Units
32
SM Count
36
Clocks
Base Clock
1720 MHz
1470 MHz
Boost Clock
2440 MHz
1650 MHz
Memory Clock
2250 MHz 18 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
8 GB
8 GB
VRAM (MB)
8,192
8,192 0.0%
Memory Type
GDDR6
GDDR6
Memory Bus
128 bit
256 bit
Bandwidth
288.0 GB/s
448.0 GB/s
Cache
L1 Cache
128 KB per Array
64 KB (per SM)
L2 Cache
2 MB
4 MB
L3 Cache
32 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
156.2 GPixel/s
105.6 GPixel/s
Texture Rate
312.3 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
19.99 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
624.6 GFLOPS (1:32)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
39.98 TFLOPS (2:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
32
36 +12.5%
Tensor Cores
288
Matrix Cores
64
Power
TDP
130 W
185 W
TDP (W)
130
185 +42.3%
Suggested PSU
300 W
450 W
Power Connectors
1x 6-pin
1x 8-pin
Architecture
Architecture
RDNA 3.0
Turing
GPU Name
Navi 33
TU106
Codename
Hotpink Bonefish
Generation
Radeon Pro Navi (Navi III Series)
Mining GPUs
Process Size
6 nm
12 nm
Transistors
13,300 million
10,800 million
Die Size
204 mm²
445 mm²
Foundry
TSMC
TSMC
Density
65.2M / mm²
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.2
3.0
CUDA
7.5
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Dual-slot
Length
241 mm 9.5 inches
229 mm 9 inches
Height
115 mm 4.5 inches
111 mm 4.4 inches
Outputs
4x DisplayPort 2.1
No outputs
Bus Interface
PCIe 4.0 x8
PCIe 1.0 x4
Other
Launch Price
599 USD
699 USD
Production
Active
End-of-life
Predecessor
Radeon Pro Vega
View Radeon PRO W7600 Details View CMP 40HX Details