AMD Radeon RX 9070 GRE vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Radeon RX 9070 GRE

CORE STATE Navi 48
VRAM 12 GB
CLOCK SPEED 2790 MHz
TDP 220 W
BUS WIDTH 192 bit
ARCHITECTURE RDNA 4.0
nm
PROCESS 4 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
5,424
N/A
geekbench_opencl
109,309
93,395
geekbench_vulkan
N/A
77,879

Analysis: AMD Radeon RX 9070 GRE vs NVIDIA CMP 40HX

The NVIDIA CMP 40HX and AMD Radeon RX 9070 GRE make for an unusual pairing: one is a Turing-era mining card stripped of display outputs, the other an active RDNA 4 gaming product. The recorded data shows the RX 9070 GRE winning the single head-to-head benchmark available, yet the CMP 40HX sits in the higher percentile position in the database. That tension is worth unpacking, because these two cards were built for entirely different jobs and the numbers reflect it.

Where Each One Wins

The direct comparison is short: there is exactly one shared benchmark in the database, Geekbench OpenCL, and the RX 9070 GRE wins it with a score of 109309 against 93395 for the CMP 40HX. That is a 14.6 percent margin in AMD's favor, or looked at the other way, the CMP 40HX trails by that same margin. The overall head-to-head tally stands at one win for the RX 9070 GRE and zero for the CMP 40HX.

Yet the percentile placements complicate the picture. The CMP 40HX sits at the 93rd percentile versus all GPUs in the database, while the RX 9070 GRE sits at the 87th. How can the card that loses the direct benchmark rank higher overall? The answer lies in how the average scores are computed. The CMP 40HX carries an average benchmark score of 85637 built from two Geekbench submissions (OpenCL at 93395 and Vulkan at 77879), while the RX 9070 GRE averages 57367 across a different mix of tests that includes a Steel Nomad DirectX 12 score of 5424 alongside its OpenCL result. Different test weightings produce different averages, so percentile and head-to-head results answer different questions.

Context from each card's nearest rivals reinforces the split. The CMP 40HX clusters tightly around professional cards like the AMD Radeon PRO W7600 (average score 87108, just 1.7 percent ahead), the NVIDIA Quadro GP100 (87445, 2.1 percent ahead), the AMD Radeon PRO W6600 (81995, 4.4 percent behind), and the AMD Radeon Pro Vega 64X (80959, 5.8 percent behind). The RX 9070 GRE, meanwhile, keeps company with the Intel Arc A580 (57756), the AMD Radeon RX 5600 OEM (58085), the Intel Arc A570M (58239), and the AMD Radeon RX 6950 XT (58392), all fractionally ahead of it by under 2 percent. In raw compute throughput, though, the RX 9070 GRE is in another league: 34.28 TFLOPS FP32 against 7.603 TFLOPS for the CMP 40HX.

Architecture Differences

These two chips come from different eras and philosophies. The CMP 40HX uses NVIDIA's TU106 die on the Turing architecture, manufactured by TSMC on a 12 nm process. It packs 10,800 million transistors into a 445 mm² die, yielding a density of 24.3 million transistors per square millimeter. The RX 9070 GRE uses the Navi 48 die on RDNA 4.0, built on a 4 nm TSMC process, with 53,900 million transistors in a smaller 357 mm² die. That works out to 151.0 million transistors per square millimeter, roughly six times the density, and it explains how AMD fits five times the transistor count into less silicon area.

The internal resource allocation diverges sharply. The CMP 40HX has 2304 shading units, 144 TMUs, 64 ROPs, 36 RT cores, and 288 tensor cores. The RX 9070 GRE has 3072 shading units, 192 TMUs, 96 ROPs, and 48 RT cores, but no tensor cores listed at all. That last point is telling: the mining card was built around integer and tensor-heavy parallel workloads, while the gaming card invests elsewhere. Both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Clock behavior differs enormously. The CMP 40HX runs a base clock of 1470 MHz and boosts to 1650 MHz, with memory at 1750 MHz (14 Gbps effective). The RX 9070 GRE starts lower at a 1420 MHz base but boosts to 2790 MHz, with a game clock of 2220 MHz and memory at 2250 MHz (18 Gbps effective). Memory configuration also splits them: 8 GB of GDDR6 on a 256-bit bus for the CMP 40HX, delivering 448.0 GB/s of bandwidth, versus 12 GB of GDDR6 on a 192-bit bus for the RX 9070 GRE, delivering 432.0 GB/s. The older card actually holds a slight bandwidth edge despite the narrower capacity, but the newer card offers half again the memory and faster per-pin speeds.

Theoretical throughput numbers make the generation gap vivid. Pixel rate is 105.6 GPixel/s versus 267.8 GPixel/s, texture rate is 237.6 GTexel/s versus 535.7 GTexel/s, and FP16 behavior is qualitatively different: the CMP 40HX manages 15.21 TFLOPS at a 2:1 ratio, while the RX 9070 GRE matches its FP32 figure of 34.28 TFLOPS at 1:1.

Connectivity and usage differences round out the story. The CMP 40HX connects over PCIe 1.0 x4 and has no display outputs at all, a deliberate choice for a mining product. The RX 9070 GRE uses PCIe 5.0 x16 and offers one HDMI 2.1b plus three DisplayPort 2.1a outputs. Power delivery differs too: the CMP 40HX draws 185 W through a single 8-pin connector with a suggested 450 W supply, while the RX 9070 GRE draws 220 W through two 8-pin connectors with a suggested 550 W supply. Both are dual-slot cards.

FAQ

Q: Which card is faster in the shared benchmark?

A: The RX 9070 GRE wins Geekbench OpenCL with 109309 versus 93395, a 14.6 percent advantage over the CMP 40HX.

Q: Can the CMP 40HX drive a monitor?

A: No. It has no display outputs whatsoever. The RX 9070 GRE has one HDMI 2.1b port and three DisplayPort 2.1a ports.

Q: Are either of these cards still in production?

A: The CMP 40HX is end-of-life; it launched in February 2021 with a launch MSRP of 699 USD. The RX 9070 GRE is active, released in May 2025 with a launch MSRP of 549 USD.

Q: Which has more memory?

A: The RX 9070 GRE has 12 GB of GDDR6 against 8 GB for the CMP 40HX, though the CMP 40HX has higher total bandwidth (448.0 GB/s versus 432.0 GB/s) thanks to its wider 256-bit bus.

Q: How do they compare in raw compute?

A: The RX 9070 GRE delivers 34.28 TFLOPS FP32, roughly 4.5 times the 7.603 TFLOPS of the CMP 40HX. The CMP 40HX is the only one of the two with tensor cores, at 288.

Q: What rivals does each card compete with in the database?

A: The CMP 40HX trades blows with workstation cards such as the Radeon PRO W7600 and Quadro GP100; the RX 9070 GRE sits within about 2 percent of cards like the Arc A580 and Radeon RX 6950 XT.

Specification Differences

  • Architecture and chip: TU106 on Turing (CMP 40HX) versus Navi 48 on RDNA 4.0 (RX 9070 GRE)
  • Process node: 12 nm versus 4 nm, both TSMC
  • Transistors and die: 10,800 million on 445 mm² (24.3M/mm²) versus 53,900 million on 357 mm² (151.0M/mm²)
  • Clocks: 1470 MHz base, 1650 MHz boost, no game clock listed versus 1420 MHz base, 2790 MHz boost, 2220 MHz game clock
  • Memory: 8 GB GDDR6, 256-bit, 448.0 GB/s, 14 Gbps effective versus 12 GB GDDR6, 192-bit, 432.0 GB/s, 18 Gbps effective
  • Shader resources: 2304 shading units, 144 TMUs, 64 ROPs, 36 RT cores, 288 tensor cores versus 3072 shading units, 192 TMUs, 96 ROPs, 48 RT cores, no tensor cores listed
  • Throughput: 7.603 TFLOPS FP32 and 15.21 TFLOPS FP16 (2:1) versus 34.28 TFLOPS FP32 and 34.28 TFLOPS FP16 (1:1); 105.6 versus 267.8 GPixel/s; 237.6 versus 535.7 GTexel/s
  • Power: 185 W TDP, one 8-pin, 450 W suggested supply versus 220 W TDP, two 8-pin, 550 W suggested supply
  • Interface and outputs: PCIe 1.0 x4 with no display outputs versus PCIe 5.0 x16 with HDMI 2.1b and three DisplayPort 2.1a
  • Dimensions: the CMP 40HX measures 229 mm long, 111 mm tall, and 35 mm wide; the RX 9070 GRE has no dimensions recorded
  • Status: end-of-life versus active

Head-to-Head Benchmarks

Geekbench OpenCL is the only test where both cards appear, and the RX 9070 GRE takes it decisively: 109309 against 93395, a 14.6 percent gap. For a card that boosts to nearly 2790 MHz with triple the FP32 throughput on paper, one might ask why the margin is not larger, and the CMP 40HX's 256-bit bus with 448.0 GB/s of bandwidth is a plausible partial explanation: OpenCL workloads often reward bandwidth, and here the older card actually leads on that specification.

The database also holds results the two cards do not share. The CMP 40HX scored 77879 in Geekbench Vulkan, a result within a few points per day of its professional-card rivals. The RX 9070 GRE posted 5424 in 3DMark Steel Nomad DirectX 12, the kind of modern gaming workload the mining card was never intended to run and has no recorded result for. In the OpenCL comparison alone, the CMP 40HX's 93395 would place it ahead of every nearest rival listed for the RX 9070 GRE on average score, which hints at how differently the two averages are composed.

The Verdict

The data points to a clear split. Anyone needing a working display output, current production availability, PCIe 5.0 x16 bandwidth, more memory, and the only recorded win in the direct comparison should land on the RX 9070 GRE. Its 34.28 TFLOPS FP32, 12 GB of GDDR6, and 267.8 GPixel/s pixel rate make it the stronger all-around performer, and its 87th percentile placement comes with a benchmark mix that includes modern gaming tests.

The case for the CMP 40HX is narrower and stranger, but the data supports it. Its 93rd percentile position, higher average score of 85637, and cluster of professional rivals (the PRO W7600, Quadro GP100, PRO W6600, and Pro Vega 64X) suggest a compute card whose OpenCL and Vulkan results still hold up in workstation-adjacent territory. Its tensor cores, higher memory bandwidth, and lower 185 W draw through a single 8-pin connector suit it to headless parallel work. But with no display outputs, a PCIe 1.0 x4 interface, and end-of-life status, its usefulness outside that niche is limited. For everything the shared data measures head-to-head, the RX 9070 GRE wins; for the specialized footprint the CMP 40HX occupies, the database still ranks it among the top few percent of GPUs recorded.

DETAILED SPECIFICATIONS

SPECIFICATION
RX 9070 GRE
CMP 40HX
Core Specs
Shading Units
3,072
2,304 -25.0%
Shaders
3,072
2,304 -25.0%
TMUs
192
144 -25.0%
ROPs
96
64 -33.3%
Compute Units
48
SM Count
36
Clocks
Base Clock
1420 MHz
1470 MHz
Boost Clock
2790 MHz
1650 MHz
Game Clock
2220 MHz
Memory Clock
2250 MHz 18 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
12 GB
8 GB
VRAM (MB)
12,288
8,192 -33.3%
Memory Type
GDDR6
GDDR6
Memory Bus
192 bit
256 bit
Bandwidth
432.0 GB/s
448.0 GB/s
Cache
L1 Cache
64 KB (per SM)
L2 Cache
8 MB
4 MB
L3 Cache
48 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
267.8 GPixel/s
105.6 GPixel/s
Texture Rate
535.7 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
34.28 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
1,071.4 GFLOPS (1:32)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
34.28 TFLOPS (1:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
48
36 -25.0%
Tensor Cores
288
Matrix Cores
96
Power
TDP
220 W
185 W
TDP (W)
220
185 -15.9%
Suggested PSU
550 W
450 W
Power Connectors
2x 8-pin
1x 8-pin
Architecture
Architecture
RDNA 4.0
Turing
GPU Name
Navi 48
TU106
Generation
Navi IV (RX 9000)
Mining GPUs
Process Size
4 nm
12 nm
Transistors
53,900 million
10,800 million
Die Size
357 mm²
445 mm²
Foundry
TSMC
TSMC
Density
151.0M / mm²
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.2
3.0
CUDA
7.5
Shader Model
6.9
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
229 mm 9 inches
Height
111 mm 4.4 inches
Outputs
1x HDMI 2.1b3x DisplayPort 2.1a
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 1.0 x4
Other
Launch Price
549 USD
699 USD
Production
Active
End-of-life
Predecessor
Navi III
View Radeon RX 9070 GRE Details View CMP 40HX Details