AMD Radeon RX 9060 XT LP vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Radeon RX 9060 XT LP

CORE STATE Navi 44
VRAM 16 GB
CLOCK SPEED 3050 MHz
TDP 140 W
BUS WIDTH 128 bit
ARCHITECTURE RDNA 4.0
nm
PROCESS 4 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
88,183
93,395
geekbench_vulkan
39,476
77,879

Analysis: AMD Radeon RX 9060 XT LP vs NVIDIA CMP 40HX

Where Each One Wins

The recorded benchmark data splits this comparison clearly: the NVIDIA CMP 40HX wins both tracked tests, while the AMD Radeon RX 9060 XT LP does not claim a single victory in the database. However, the magnitude of those wins matters more than the win count. In Geekbench OpenCL, the CMP 40HX scores 93,395 against 88,183 for the RX 9060 XT LP, a 5.9% advantage. That is a modest margin, the kind that suggests the two parts trade blows depending on workload characteristics. In Geekbench Vulkan, the story changes dramatically: the CMP 40HX scores 77,879 versus 39,476, a 97.3% lead, meaning it nearly doubles the AMD card in that specific API test.

The use-case split is therefore about API preference rather than raw compute class. The CMP 40HX is the stronger option for Vulkan-centric applications, where its recorded score places it far ahead. The RX 9060 XT LP is competitive in OpenCL, trailing by only 5.9%, which is close enough that driver optimizations or specific workload tuning could flip the result. For users prioritizing Vulkan performance, the NVIDIA card is the clear choice based on the database. For OpenCL-heavy tasks, the AMD card is not far behind, and its architectural advantages in other areas may compensate in real-world scenarios that the two tracked benchmarks do not fully capture.

Looking at the average benchmark scores, the CMP 40HX averages 85,637 across both tests, while the RX 9060 XT LP averages 63,830. That 21,807-point gap is driven almost entirely by the Vulkan disparity, since the OpenCL scores are relatively close. The percentile rankings reinforce this: the CMP 40HX sits in the 93rd percentile of all GPUs, the RX 9060 XT LP in the 89th. Both are high performers, but the NVIDIA part occupies a higher overall standing in the database.

Architecture Differences

The two cards come from fundamentally different design philosophies. The NVIDIA CMP 40HX uses the TU106 chip, built on Turing architecture, fabricated on a 12 nm process at TSMC. It packs 10,800 million transistors into a 445 mm² die, yielding a transistor density of 24.3 million per square millimeter. The AMD Radeon RX 9060 XT LP uses the Navi 44 chip, built on RDNA 4.0 architecture, also fabricated at TSMC but on a much smaller 4 nm process. It contains 29,700 million transistors on a 199 mm² die, achieving a density of 149.2 million per square millimeter. This is a generational leap in integration: the AMD chip crams nearly three times the transistors into less than half the silicon area.

The compute resources differ in configuration. The CMP 40HX has 2,304 shading units, 144 texture mapping units, and 64 render output units. It also includes 36 ray tracing cores and 288 tensor cores, features inherited from the Turing architecture. The RX 9060 XT LP has 2,048 shading units, 128 TMUs, and 64 ROPs. It includes 32 ray tracing cores but no tensor cores, reflecting RDNA 4.0’s different approach to acceleration. The AMD card’s raw FP32 throughput is 24.99 TFLOPS, more than triple the CMP 40HX’s 7.603 TFLOPS. Its FP16 performance is identical at 24.99 TFLOPS with a 1:1 ratio, whereas the CMP 40HX reaches 15.21 TFLOPS at a 2:1 ratio. The AMD part is clearly the higher-throughput compute engine on paper.

Memory configurations also diverge sharply. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus, delivering 448.0 GB/s of bandwidth. The RX 9060 XT LP has 16 GB of GDDR6 on a 128-bit bus, delivering 322.3 GB/s. The NVIDIA card has higher bandwidth but half the capacity; the AMD card has double the memory but lower bandwidth. Clock behavior tells a similar story: the CMP 40HX runs at 1470 MHz base and 1650 MHz boost, while the RX 9060 XT LP runs at 1380 MHz base, 2450 MHz game clock, and 3050 MHz boost. The AMD part’s boost clock is nearly double the NVIDIA card’s, which explains its enormous FP32 advantage despite fewer shading units.

The bus interface is another major difference. The CMP 40HX uses PCIe 1.0 x4, a severely limited connection that likely constrains data transfer in certain workloads. The RX 9060 XT LP uses PCIe 5.0 x16, a modern high-bandwidth interface. This matters for any task that streams data to and from the GPU. The CMP 40HX also has no display outputs at all, as it was designed for mining, while the RX 9060 XT LP provides 1x HDMI 2.1b and 2x DisplayPort 2.1a.

Head-to-Head Benchmarks

The most striking result in the database is the Vulkan test. The CMP 40HX scores 77,879, while the RX 9060 XT LP scores 39,476. That is a delta of 97.3%, meaning the NVIDIA card is nearly twice as fast in this API. This is not a marginal difference; it is a dominant showing. The RX 9060 XT LP’s Vulkan score is actually lower than its own OpenCL score by a wide margin, 39,476 versus 88,183, which suggests the AMD driver stack or hardware implementation does not handle Vulkan as efficiently in this particular benchmark. The CMP 40HX, by contrast, scores higher in OpenCL than in Vulkan but remains strong in both.

The OpenCL test is much closer. The CMP 40HX scores 93,395, the RX 9060 XT LP scores 88,183, a 5.9% gap. Given that the RX 9060 XT LP has vastly higher theoretical FP32 throughput, 24.99 TFLOPS versus 7.603 TFLOPS, one might expect it to win this test outright. The data shows otherwise, indicating that raw compute throughput does not translate directly to benchmark performance in this case. Memory bandwidth, driver efficiency, and workload characteristics all play a role. The CMP 40HX’s 448.0 GB/s bandwidth versus 322.3 GB/s for the AMD card may be a factor, as OpenCL workloads often stress memory access patterns.

The average benchmark scores tell a similar story. The CMP 40HX averages 85,637, placing it in the 93rd percentile. The RX 9060 XT LP averages 63,830, placing it in the 89th percentile. The nearest rivals for the CMP 40HX include the AMD Radeon PRO W7600 at 87,108 (-1.7%), the NVIDIA Quadro GP100 at 87,445 (-2.1%), the AMD Radeon PRO W6600 at 81,995 (+4.4%), and the AMD Radeon Pro Vega 64X at 80,959 (+5.8%). The RX 9060 XT LP’s nearest rivals are much closer in score: the NVIDIA CMP 30HX at 63,842 (0%), the AMD Radeon RX 7600M at 63,775 (+0.1%), the AMD Radeon Pro Vega 56 at 63,693 (+0.2%), and the AMD Radeon Pro WX 9100 at 64,212 (-0.6%). This clustering shows that the RX 9060 XT LP sits in a tightly contested performance band, while the CMP 40HX occupies a higher tier with more separation from its closest competitors.

Specification Differences

The two cards differ across nearly every specification category in the database. The process node is 12 nm for the CMP 40HX versus 4 nm for the RX 9060 XT LP. Transistor count is 10,800 million versus 29,700 million. Die size is 445 mm² versus 199 mm². Transistor density is 24.3 million per square millimeter versus 149.2 million per square millimeter. Base clock is 1470 MHz versus 1380 MHz. Boost clock is 1650 MHz versus 3050 MHz. The RX 9060 XT LP also lists a game clock of 2450 MHz, which the CMP 40HX does not have. Memory speed is 14 Gbps effective for the NVIDIA card versus 20.1 Gbps effective for the AMD card. Memory size is 8 GB versus 16 GB. Bus width is 256-bit versus 128-bit. Bandwidth is 448.0 GB/s versus 322.3 GB/s. Shading units are 2,304 versus 2,048. TMUs are 144 versus 128. ROPs are equal at 64. Ray tracing cores are 36 versus 32. Tensor cores are 288 for the CMP 40HX and null for the RX 9060 XT LP. Pixel rate is 105.6 GPixel/s versus 195.2 GPixel/s. Texture rate is 237.6 GTexel/s versus 390.4 GTexel/s. FP32 is 7.603 TFLOPS versus 24.99 TFLOPS. FP16 is 15.21 TFLOPS versus 24.99 TFLOPS. TDP is 185 W versus 140 W. Suggested PSU is 450 W versus 300 W. The bus interface is PCIe 1.0 x4 versus PCIe 5.0 x16. Display outputs are none versus 1x HDMI 2.1b and 2x DisplayPort 2.1a. Dimensions are available for the CMP 40HX at 229 mm length, 111 mm height, and 35 mm width, while the RX 9060 XT LP has no recorded dimensions. Production status is end-of-life for the NVIDIA card and active for the AMD card. Release dates are February 2021 for the CMP 40HX and December 2025 for the RX 9060 XT LP. The CMP 40HX has a launch MSRP of 699 USD, while the RX 9060 XT LP has no recorded launch MSRP.

FAQ

Q: Which card has a higher average benchmark score?

A: The NVIDIA CMP 40HX averages 85,637 across both tracked tests, while the AMD Radeon RX 9060 XT LP averages 63,830. The CMP 40HX also sits in the 93rd percentile of all GPUs, compared to the 89th percentile for the RX 9060 XT LP.

Q: How large is the Vulkan performance gap?

A: In the Geekbench Vulkan test, the CMP 40HX scores 77,879 versus 39,476 for the RX 9060 XT LP, a delta of 97.3%. This is the largest single-test margin in the comparison.

Q: Does the RX 9060 XT LP win any benchmark?

A: No. The database records two head-to-head tests, and the CMP 40HX wins both. The RX 9060 XT LP has zero wins in this comparison.

Q: Which card has more memory?

A: The RX 9060 XT LP has 16 GB of GDDR6 memory, while the CMP 40HX has 8 GB. However, the CMP 40HX has a wider 256-bit bus and higher bandwidth at 448.0 GB/s, versus 322.3 GB/s for the AMD card.

Q: What are the power requirements?

A: The CMP 40HX has a TDP of 185 W and a suggested PSU of 450 W. The RX 9060 XT LP has a TDP of 140 W and a suggested PSU of 300 W. Both use a single 8-pin power connector.

Q: Are the compute capabilities similar?

A: No. The RX 9060 XT LP has 24.99 TFLOPS of FP32 throughput, more than triple the CMP 40HX’s 7.603 TFLOPS. The AMD card also has 32 ray tracing cores, while the NVIDIA card has 36 ray tracing cores plus 288 tensor cores that the AMD card lacks.

DETAILED SPECIFICATIONS

SPECIFICATION
RX 9060 XT LP
CMP 40HX
Core Specs
Shading Units
2,048
2,304 +12.5%
Shaders
2,048
2,304 +12.5%
TMUs
128
144 +12.5%
ROPs
64
64 0.0%
Compute Units
32
—
SM Count
—
36
Clocks
Base Clock
1380 MHz
1470 MHz
Boost Clock
3050 MHz
1650 MHz
Game Clock
2450 MHz
—
Memory Clock
2518 MHz 20.1 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
16 GB
8 GB
VRAM (MB)
16,384
8,192 -50.0%
Memory Type
GDDR6
GDDR6
Memory Bus
128 bit
256 bit
Bandwidth
322.3 GB/s
448.0 GB/s
Cache
L1 Cache
—
64 KB (per SM)
L2 Cache
4 MB
4 MB
L3 Cache
32 MB
—
L0 Cache
32 KB per WGP
—
Performance
Pixel Rate
195.2 GPixel/s
105.6 GPixel/s
Texture Rate
390.4 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
24.99 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
780.8 GFLOPS (1:32)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
24.99 TFLOPS (1:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
32
36 +12.5%
Tensor Cores
—
288
Matrix Cores
64
—
Power
TDP
140 W
185 W
TDP (W)
140
185 +32.1%
Suggested PSU
300 W
450 W
Power Connectors
1x 8-pin
1x 8-pin
Architecture
Architecture
RDNA 4.0
Turing
GPU Name
Navi 44
TU106
Generation
Navi IV (RX 9000)
Mining GPUs
Process Size
4 nm
12 nm
Transistors
29,700 million
10,800 million
Die Size
199 mm²
445 mm²
Foundry
TSMC
TSMC
Density
149.2M / mm²
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.2
3.0
CUDA
—
7.5
Shader Model
6.9
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
—
229 mm 9 inches
Height
—
111 mm 4.4 inches
Outputs
1x HDMI 2.1b2x DisplayPort 2.1a
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 1.0 x4
Other
Launch Price
—
699 USD
Production
Active
End-of-life
Predecessor
Navi III
—
View Radeon RX 9060 XT LP Details View CMP 40HX Details