NVIDIA CMP 90HX vs NVIDIA Tesla P40 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 90HX

CORE STATE GA102
VRAM 10 GB
CLOCK SPEED 1710 MHz
TDP 320 W
BUS WIDTH 320 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Tesla P40

CORE STATE GP102
VRAM 24 GB
CLOCK SPEED 1531 MHz
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_opencl
69,000
62,017
geekbench_vulkan
N/A
68,172

Analysis: NVIDIA CMP 90HX vs NVIDIA Tesla P40

The data presents a clear performance hierarchy between two very different NVIDIA compute accelerators. The NVIDIA CMP 90HX, built on the modern Ampere architecture, decisively wins the only available head-to-head benchmark, while the NVIDIA Tesla P40 offers a substantially larger memory pool from an older generation. This analysis examines where each card excels, the architectural gulf between them, and which workloads each is suited for based strictly on the provided benchmark results and specifications.

Where Each One Wins

The benchmark results are unambiguous in raw compute: the NVIDIA CMP 90HX wins the single head-to-head test, a Geekbench OpenCL run, with a score of 69000 against the Tesla P40's 62017. That is a delta of 11.3%, placing the CMP 90HX firmly ahead in general-purpose compute throughput. The CMP 90HX also posts a higher average benchmark score of 69000 versus the Tesla P40's 65095, reinforcing its status as the faster card in this comparison.

However, the Tesla P40 wins in a category not captured by a single benchmark score: memory capacity. The P40 carries 24 GB of GDDR5 memory, more than double the CMP 90HX's 10 GB of GDDR6X. For workloads that require large datasets to reside on the GPU itself—such as certain inference models or in-memory databases—the P40's capacity advantage is the deciding factor, even if its raw throughput is lower. The P40 also has a wider 384-bit memory bus compared to the CMP 90HX's 320-bit bus, although the CMP 90HX's faster GDDR6X memory gives it a massive bandwidth lead of 760.3 GB/s versus 347.1 GB/s.

In terms of relative standing among peers, the CMP 90HX sits at the 90th percentile of all GPUs, while the Tesla P40 is at the 89th. This near-identical percentile ranking suggests that despite the 11.3% gap between them, both are high-performing accelerators in the broader GPU landscape. The CMP 90HX's closest rival is the Intel Arc A770, which trails by just 0.3%, indicating that the CMP 90HX is competitive with modern consumer-grade cards. The Tesla P40, meanwhile, is within 1.4% of the AMD Radeon Pro WX 9100 and the AMD Radeon VII, showing it still holds its own against newer professional and enthusiast cards.

Architecture Differences

The architectural divide between these two cards is generational and profound. The CMP 90HX is built on the GA102 chip using the Ampere architecture on an 8 nm Samsung process. It packs 28,300 million transistors onto a 628 mm² die, yielding a transistor density of 45.1 million per square millimeter. The Tesla P40, in contrast, uses the GP102 chip from the older Pascal architecture on a 16 nm TSMC process, with 11,800 million transistors on a 471 mm² die and a density of just 25.1 million per square millimeter. This means the CMP 90HX has over twice the transistor count in a larger but far denser package.

The compute resources differ sharply. The CMP 90HX has 6400 shading units, 200 TMUs, and 80 ROPs, while the Tesla P40 has 3840 shading units, 240 TMUs, and 96 ROPs. The CMP 90HX compensates for fewer ROPs with a higher pixel rate of 136.8 GPixel/s versus the P40's 147.0 GPixel/s, but the P40 has a higher texture rate of 367.4 GTexel/s versus the CMP 90HX's 342.0 GTexel/s. In raw FP32 throughput, the CMP 90HX delivers 21.89 TFLOPS, nearly double the P40's 11.76 TFLOPS.

The most significant feature gap is in specialized cores. The CMP 90HX includes 50 RT cores for ray tracing and 200 tensor cores for AI acceleration, features the Tesla P40 lacks entirely. The FP16 performance tells this story: the CMP 90HX achieves 21.89 TFLOPS at a 1:1 ratio with FP32, while the P40 manages only 183.7 GFLOPS at a 1:64 ratio. This makes the CMP 90HX vastly superior for mixed-precision and AI workload, while the P40 is effectively limited to FP32 compute. The CMP 90HX also supports DirectX 12 Ultimate (12_2), while the P40 is limited to DirectX 12 (12_1).

Other differences matter for deployment. The CMP 90HX uses a PCIe 1.0 x4 interface, a legacy bus that severely limits data transfer speeds despite the card's compute power. The Tesla P40 uses PCIe 3.0 x16, a far more standard and faster connection. The CMP 90HX draws 320 W and requires two 8-pin power connectors, while the P40 draws 250 W with a single 8-pin EPS connector. Both are dual-slot cards with no display outputs, indicating their intended role as compute-only accelerators. The CMP 90HX is slightly longer at 285 mm versus the P40's 267 mm.

The Verdict

For raw compute performance, the NVIDIA CMP 90HX is the clear winner. Its 11.3% lead in Geekbench OpenCL, combined with roughly double the FP32 throughput and a massive advantage in FP16 and tensor operations, makes it the superior choice for compute-heavy tasks like machine learning inference, scientific simulation, or any workload that can leverage its RT and tensor cores. The data shows no scenario in the provided benchmarks where the Tesla P40 outperforms the CMP 90HX in speed.

However, the Tesla P40 is not without merit. Its 24 GB of memory is the largest capacity in this comparison, and for workloads that are memory-bound rather than compute-bound, such as loading large models or datasets that exceed 10 GB, the P40 is the only option. The P40's lower power draw of 250 W and standard PCIe 3.0 x16 interface also make it easier to integrate into existing systems without the need for a newer motherboard or power supply.

The choice ultimately depends on workload priorities. If the task requires maximum compute throughput and can fit within 10 GB of memory, the CMP 90HX is the obvious pick. If the task requires more than 10 GB of memory or needs to run on older infrastructure with standard PCIe, the Tesla P40 is the practical choice. The CMP 90HX's PCIe 1.0 x4 interface is a significant bottleneck that could negate its compute advantage in real-world data transfer scenarios.

FAQ

Q: Which card is faster in the head-to-head benchmark?

A: The NVIDIA CMP 90HX wins the only available head-to-head test, Geekbench OpenCL, with a score of 69000 against the Tesla P40's 62017, a delta of 11.3%.

Q: What is the memory capacity difference?

A: The Tesla P40 has 24 GB of GDDR5 memory, while the CMP 90HX has 10 GB of GDDR6X. The P40 offers more capacity, but the CMP 90HX has higher bandwidth at 760.3 GB/s versus 347.1 GB/s.

Q: Does the Tesla P40 support ray tracing or tensor cores?

A: No. The Tesla P40 has no RT cores or tensor cores, while the CMP 90HX includes 50 RT cores and 200 tensor cores.

Q: What are the FP16 performance differences?

A: The CMP 90HX achieves 21.89 TFLOPS FP16 at a 1:1 ratio with FP32, while the Tesla P40 manages only 183.7 GFLOPS at a 1:64 ratio, making the CMP 90HX vastly superior for mixed-precision work.

Q: Which card has a better PCIe interface?

A: The Tesla P40 uses PCIe 3.0 x16, while the CMP 90HX uses PCIe 1.0 x4, which is a much older and slower interface.

Q: What are the power requirements?

A: The CMP 90HX has a TDP of 320 W and requires two 8-pin power connectors, while the Tesla P40 has a TDP of 250 W and uses a single 8-pin EPS connector.

Head-to-Head Benchmarks

The single benchmark result available paints a definitive picture. In Geekbench OpenCL, the NVIDIA CMP 90HX scores 69000, while the NVIDIA Tesla P40 scores 62017. The CMP 90HX wins this test by a margin of 11.3%, a substantial gap that reflects its architectural advantages. This result is consistent with the raw specifications: the CMP 90HX delivers 21.89 TFLOPS FP32 versus the P40's 11.76 TFLOPS, a near-doubling of compute throughput that translates directly into the benchmark score.

Beyond the raw score, the CMP 90HX's advantage is even more pronounced in specialized workloads. The presence of 200 tensor cores and full-rate FP16 (21.89 TFLOPS at 1:1) means that any AI or machine learning task will see a significantly larger performance gap than the 11.3% shown in the OpenCL test, which likely exercises general FP32 compute. The Tesla P40's FP16 performance of 183.7 GFLOPS is a fraction of its FP32 throughput, making it unsuitable for modern mixed-precision training or inference workloads.

The memory bandwidth difference further amplifies the CMP 90HX's lead. With 760.3 GB/s of bandwidth, the CMP 90HX can feed its compute units far faster than the P40's 347.1 GB/s. This is critical for memory-intensive algorithms, where the P40 may stall waiting for data. However, the P40's 24 GB capacity means it can hold larger working sets without spilling to system memory, a factor that could mitigate its bandwidth disadvantage in specific scenarios.

The nearest rivals data provides context for both cards. The CMP 90HX is nearly tied with the Intel Arc A770 (delta 0.3%) and the AMD Radeon Instinct MI25 (delta 0.6%), suggesting it is on par with modern mid-range accelerators. The Tesla P40, despite its age, is within 1.4% of the AMD Radeon Pro WX 9100 and the AMD Radeon VII, and just 2% ahead of the NVIDIA CMP 30HX. This indicates that while the P40 trails the CMP 90HX, it remains a competitive option against other older or lower-tier cards.

In the only direct comparison, the verdict is clear: the CMP 90HX is the faster card by a meaningful margin. The Tesla P40's only winning argument is its larger memory pool, which the benchmark data does not evaluate. For pure compute, the CMP 90HX wins decisively.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 90HX
Tesla P40
Core Specs
Shading Units
6,400
3,840 -40.0%
Shaders
6,400
3,840 -40.0%
TMUs
200
240 +20.0%
ROPs
80
96 +20.0%
SM Count
50
30 -40.0%
Clocks
Base Clock
1500 MHz
1303 MHz
Boost Clock
1710 MHz
1531 MHz
Memory Clock
1188 MHz 19 Gbps effective
1808 MHz 7.2 Gbps effective
Memory
Memory Size
10 GB
24 GB
VRAM (MB)
10,240
24,576 +140.0%
Memory Type
GDDR6X
GDDR5
Memory Bus
320 bit
384 bit
Bandwidth
760.3 GB/s
347.1 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SM)
L2 Cache
5 MB
3 MB
Performance
Pixel Rate
136.8 GPixel/s
147.0 GPixel/s
Texture Rate
342.0 GTexel/s
367.4 GTexel/s
FP32 (TFLOPS)
21.89 TFLOPS
11.76 TFLOPS
FP64 (TFLOPS)
342.0 GFLOPS (1:64)
367.4 GFLOPS (1:32)
FP16 (TFLOPS)
21.89 TFLOPS (1:1)
183.7 GFLOPS (1:64)
AI/RT
RT Cores
50
Tensor Cores
200
Power
TDP
320 W
250 W
TDP (W)
320
250 -21.9%
Suggested PSU
700 W
600 W
Power Connectors
2x 8-pin
8-pin EPS
Architecture
Architecture
Ampere
Pascal
GPU Name
GA102
GP102
Generation
Mining GPUs
Tesla Pascal (Pxx)
Process Size
8 nm
16 nm
Transistors
28,300 million
11,800 million
Die Size
628 mm²
471 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
25.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
285 mm 11.2 inches
267 mm 10.5 inches
Height
112 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 1.0 x4
PCIe 3.0 x16
Other
Launch Price
5,699 USD
Production
End-of-life
End-of-life
Predecessor
Tesla Maxwell
Successor
Tesla Volta
View CMP 90HX Details View Tesla P40 Details