AMD Radeon PRO W7900 vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Radeon PRO W7900

CORE STATE Navi 31
VRAM 48 GB
CLOCK SPEED 2495 MHz
TDP 295 W
BUS WIDTH 384 bit
ARCHITECTURE RDNA 3.0
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
84,379
93,395
geekbench_vulkan
137,070
77,879

Analysis: AMD Radeon PRO W7900 vs NVIDIA CMP 40HX

Head-to-Head Benchmarks

The benchmark data splits the two contenders almost perfectly down the middle, with each card claiming one decisive victory. In the Geekbench OpenCL test, the NVIDIA CMP 40HX posts a score of 93,395, beating the AMD Radeon PRO W7900’s 84,379 by 9.7%. That is a meaningful margin in raw compute throughput, and it flips the expected hierarchy — the smaller, mining-focused NVIDIA part outperforms AMD’s workstation flagship in this particular API. However, the Geekbench Vulkan results tell a completely different story. The AMD Radeon PRO W7900 scores 137,070, which is a staggering 76% higher than the CMP 40HX’s 77,879. That is not a small gap; it is a generational chasm in graphics-oriented workloads.

Looking at the broader average benchmark scores, the AMD card lands at 110,725, while the NVIDIA card sits at 85,637. That puts the W7900 roughly 29% ahead on aggregate performance, even though the CMP 40HX wins the OpenCL head-to-head. The percentile rankings are close — the W7900 sits at the 94th percentile of all GPUs, and the CMP 40HX is at the 93rd — but the absolute scores reveal where each card’s strengths lie. The W7900’s nearest rivals include the NVIDIA Tesla V100 SXM2 16 GB at 114,395 (-3.2%) and the NVIDIA RTX A5500 Mobile at 113,944 (-2.8%), both of which it trails slightly. The CMP 40HX, by contrast, competes with the AMD Radeon PRO W7600 at 87,108 (-1.7%) and the NVIDIA Quadro GP100 at 87,445 (-2.1%), which it edges out.

The Vulkan result is the single most telling datapoint. A 76% advantage in Vulkan means the W7900 is not just faster — it is in a different performance class for modern graphics APIs. The OpenCL win for the CMP 40HX, while real, is narrower in relative terms and may reflect the card’s compute-oriented design. For anyone weighing these two, the choice hinges on which API dominates their workload. Raw OpenCL compute favors the NVIDIA card; anything graphics-heavy or Vulkan-based heavily favors the AMD card. The data does not support a single “winner” — it supports two different tools for two different jobs.

Architecture Differences

The architectural gulf between these two cards is vast, and it explains nearly every benchmark result. The AMD Radeon PRO W7900 is built on TSMC’s 5 nm process with the Navi 31 chip, codenamed Plum Bonito, using the RDNA 3.0 architecture. The NVIDIA CMP 40HX uses the TU106 chip on a 12 nm process with the older Turing architecture. That process gap alone — 5 nm versus 12 nm — accounts for major efficiency and density differences. The W7900 packs 57,700 million transistors into a 529 mm² die, yielding a transistor density of 109.1 million per square millimeter. The CMP 40HX has just 10,800 million transistors on a 445 mm² die, for a density of 24.3 million per square millimeter. The AMD chip crams more than four times as many transistors into a comparable physical area.

Memory configurations are equally divergent. The W7900 comes with 48 GB of GDDR6 on a 384-bit bus, delivering 864.0 GB/s of bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus, with 448.0 GB/s — roughly half the bandwidth and one-sixth the capacity. Clock speeds also favor the AMD card: the W7900 runs at a base of 1760 MHz and boosts to 2495 MHz, while the CMP 40HX operates at 1470 MHz base and 1650 MHz boost. The memory clock on the W7900 is 2250 MHz (18 Gbps effective), versus 1750 MHz (14 Gbps effective) on the CMP 40HX.

Compute resources tell the same story. The W7900 has 6,144 shading units, 384 TMUs, and 192 ROPs, with 96 ray tracing cores. The CMP 40HX has 2,304 shading units, 144 TMUs, and 64 ROPs, with 36 ray tracing cores and 288 tensor cores — the latter being a Turing-era feature the AMD card lacks entirely. The AMD card’s FP32 throughput is 61.32 TFLOPS, with FP16 at the same 61.32 TFLOPS (1:1 ratio). The NVIDIA card manages 7.603 FP32 TFLOPS and 15.21 FP16 TFLOPS (2:1 ratio). Pixel rate is 479.0 GPixel/s on the W7900 versus 105.6 GPixel/s on the CMP 40HX; texture rate is 958.1 GTexel/s versus 237.6 GTexel/s. Every architectural metric points the same way: the W7900 is a modern, high-throughput design, while the CMP 40HX is a scaled-down Turing part repurposed for mining.

Power and physical specs also diverge. The W7900 draws 295 W TDP and needs a 600 W suggested PSU with two 8-pin connectors. The CMP 40HX consumes 185 W, requires a 450 W PSU, and uses a single 8-pin connector. The W7900 is a triple-slot card measuring 280 mm long, 110 mm tall, and 51 mm wide; the CMP 40HX is a dual-slot card at 229 mm long, 111 mm tall, and 35 mm wide. The bus interface differs too: PCIe 4.0 x16 on the AMD card versus PCIe 1.0 x4 on the NVIDIA card — a massive bottleneck for the CMP 40HX in any data-transfer-heavy scenario. Display outputs are another differentiator: the W7900 offers 3x DisplayPort 2.1 and 1x mini-DisplayPort 2.1, while the CMP 40HX has no display outputs at all, reflecting its mining-only purpose.

The Verdict

The data points to a clear split decision. If your work is dominated by Vulkan or graphics-heavy rendering, the AMD Radeon PRO W7900 is the only rational choice — its 137,070 Vulkan score is 76% higher than the CMP 40HX’s 77,879, and that gap is too large to ignore. The W7900 also offers 48 GB of VRAM, 864.0 GB/s of bandwidth, and a modern 5 nm architecture with PCIe 4.0 x16 support. It is a genuine workstation card with display outputs, active production status, and a 94th-percentile overall ranking.

If your workload is purely OpenCL compute — and you do not need display output, modern PCIe bandwidth, or large memory capacity — the NVIDIA CMP 40HX wins the specific head-to-head with a 93,395 score, 9.7% ahead of the W7900’s 84,379. It also draws 110 W less power, fits in a smaller dual-slot form factor, and costs far less at launch MSRP of 699 USD versus the W7900’s 3,999 USD. However, the CMP 40HX is end-of-life, uses PCIe 1.0 x4, has no display outputs, and its 8 GB VRAM is a hard ceiling for large datasets. The 93rd-percentile ranking is respectable, but the aggregate score deficit (85,637 versus 110,725) means it loses on overall performance. For most professional use cases, the W7900 is the safer, more capable card. The CMP 40HX is only preferable if you have a narrow OpenCL-only pipeline where its specific compute advantage matters more than everything else.

FAQ

Q: Which card wins in Geekbench Vulkan?

A: The AMD Radeon PRO W7900 scores 137,070 in Vulkan, which is 76% higher than the NVIDIA CMP 40HX’s 77,879.

Q: Does the NVIDIA CMP 40HX beat the AMD card in any benchmark?

A: Yes, the CMP 40HX wins the Geekbench OpenCL test with a score of 93,395 versus the W7900’s 84,379, a 9.7% difference.

Q: What is the memory capacity difference?

A: The AMD Radeon PRO W7900 has 48 GB of GDDR6 memory, while the NVIDIA CMP 40HX has 8 GB of GDDR6.

Q: Can the NVIDIA CMP 40HX connect to a display?

A: No, the CMP 40HX has no display outputs, while the W7900 offers 3x DisplayPort 2.1 and 1x mini-DisplayPort 2.1.

Q: Which card has a higher transistor count?

A: The AMD Radeon PRO W7900 has 57,700 million transistors on a 529 mm² die, versus 10,800 million on the CMP 40HX’s 445 mm² die.

Q: What are the power requirements for each card?

A: The W7900 has a 295 W TDP and suggests a 600 W PSU, while the CMP 40HX has a 185 W TDP and suggests a 450 W PSU.

Where Each One Wins

The AMD Radeon PRO W7900 wins in every scenario that involves modern graphics APIs, large memory footprints, or display output. Its 48 GB VRAM and 864.0 GB/s bandwidth make it suitable for massive datasets, high-resolution textures, or GPU-accelerated rendering that exceeds 8 GB. The 76% Vulkan advantage means any DirectX 12 Ultimate, Vulkan, or OpenGL 4.6 workload will run dramatically better on the W7900. The card’s 96 ray tracing cores and 61.32 TFLOPS FP32 throughput provide hardware-accelerated ray tracing and compute headroom that the CMP 40HX cannot match. It also supports PCIe 4.0 x16, ensuring fast host-device transfers, and its triple-slot cooler and 295 W TDP reflect a design built for sustained professional workloads. The W7900 is active in production, released in 2023, and sits at the 94th percentile of all GPUs.

The NVIDIA CMP 40HX wins specifically in OpenCL compute performance, where its 93,395 score edges out the W7900 by 9.7%. That makes it a potential choice for OpenCL-only pipelines that do not need Vulkan, display output, or more than 8 GB of memory. It also draws significantly less power — 185 W versus 295 W — and uses a dual-slot cooler, making it easier to fit in dense systems. The 288 tensor cores are a Turing feature that could accelerate certain AI or inference workloads, though the card’s PCIe 1.0 x4 interface will bottleneck data transfers. The CMP 40HX is end-of-life and was released in 2021, so it is a legacy part with limited future support. For raw OpenCL number-crunching at lower power, it holds a narrow but real advantage; for virtually anything else, the W7900 dominates.

Specification Differences

| Specification | AMD Radeon PRO W7900 | NVIDIA CMP 40HX |

|---|---|---|

| Architecture | RDNA 3.0 | Turing |

| Process Node | 5 nm | 12 nm |

| Transistors | 57,700 million | 10,800 million |

| Die Size | 529 mm² | 445 mm² |

| Transistor Density | 109.1M / mm² | 24.3M / mm² |

| Base Clock | 1760 MHz | 1470 MHz |

| Boost Clock | 2495 MHz | 1650 MHz |

| Memory Clock | 2250 MHz (18 Gbps effective) | 1750 MHz (14 Gbps effective) |

| Memory Size | 48 GB | 8 GB |

| Memory Bus Width | 384 bit | 256 bit |

| Memory Bandwidth | 864.0 GB/s | 448.0 GB/s |

| Shading Units | 6,144 | 2,304 |

| TMUs | 384 | 144 |

| ROPs | 192 | 64 |

| Ray Tracing Cores | 96 | 36 |

| Tensor Cores | N/A | 288 |

| Pixel Rate | 479.0 GPixel/s | 105.6 GPixel/s |

| Texture Rate | 958.1 GTexel/s | 237.6 GTexel/s |

| FP32 | 61.32 TFLOPS | 7.603 TFLOPS |

| FP16 | 61.32 TFLOPS (1:1) | 15.21 TFLOPS (2:1) |

| TDP | 295 W | 185 W |

| Slot Width | Triple-slot | Dual-slot |

| Power Connectors | 2x 8-pin | 1x 8-pin |

| Suggested PSU | 600 W | 450 W |

| Bus Interface | PCIe 4.0 x16 | PCIe 1.0 x4 |

| Display Outputs | 3x DisplayPort 2.1, 1x mini-DisplayPort 2.1 | No outputs |

| Production Status | Active | End-of-life |

| Release Date | 2023-05-25 | 2021-02-24 |

| Launch MSRP | 3,999 USD | 699 USD |

| Avg Benchmark Score | 110,725 | 85,637 |

| Percentile | 94 | 93 |

DETAILED SPECIFICATIONS

SPECIFICATION
PRO W7900
CMP 40HX
Core Specs
Shading Units
6,144
2,304 -62.5%
Shaders
6,144
2,304 -62.5%
TMUs
384
144 -62.5%
ROPs
192
64 -66.7%
Compute Units
96
SM Count
36
Clocks
Base Clock
1760 MHz
1470 MHz
Boost Clock
2495 MHz
1650 MHz
Memory Clock
2250 MHz 18 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
48 GB
8 GB
VRAM (MB)
49,152
8,192 -83.3%
Memory Type
GDDR6
GDDR6
Memory Bus
384 bit
256 bit
Bandwidth
864.0 GB/s
448.0 GB/s
Cache
L1 Cache
256 KB per Array
64 KB (per SM)
L2 Cache
6 MB
4 MB
L3 Cache
96 MB
L0 Cache
64 KB per WGP
Performance
Pixel Rate
479.0 GPixel/s
105.6 GPixel/s
Texture Rate
958.1 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
61.32 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
1.916 TFLOPS (1:32)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
61.32 TFLOPS (1:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
96
36 -62.5%
Tensor Cores
288
Matrix Cores
192
Power
TDP
295 W
185 W
TDP (W)
295
185 -37.3%
Suggested PSU
600 W
450 W
Power Connectors
2x 8-pin
1x 8-pin
Architecture
Architecture
RDNA 3.0
Turing
GPU Name
Navi 31
TU106
Codename
Plum Bonito
Generation
Radeon Pro Navi (Navi III Series)
Mining GPUs
Process Size
5 nm
12 nm
Transistors
57,700 million
10,800 million
Die Size
529 mm²
445 mm²
Foundry
TSMC
TSMC
Density
109.1M / mm²
24.3M / mm²
AMD MCM
GCD Transistors
45,400 million
GCD Die Size
304.35 mm²
MCD Transistors
2,050 million x6
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.2
3.0
CUDA
7.5
Shader Model
6.9
6.8
Physical
Slot Width
Triple-slot
Dual-slot
Length
280 mm 11 inches
229 mm 9 inches
Height
110 mm 4.3 inches
111 mm 4.4 inches
Outputs
3x DisplayPort 2.11x mini-DisplayPort 2.1
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 1.0 x4
Other
Launch Price
3,999 USD
699 USD
Production
Active
End-of-life
Predecessor
Radeon Pro Vega
View Radeon PRO W7900 Details View CMP 40HX Details