NVIDIA CMP 90HX vs NVIDIA Tesla P40 Comparison
NVIDIA CMP 90HX
Tesla P40
PERFORMANCE BENCHMARKS
Analysis: NVIDIA CMP 90HX vs NVIDIA Tesla P40
The data presents a clear performance hierarchy between two very different NVIDIA compute accelerators. The NVIDIA CMP 90HX, built on the modern Ampere architecture, decisively wins the only available head-to-head benchmark, while the NVIDIA Tesla P40 offers a substantially larger memory pool from an older generation. This analysis examines where each card excels, the architectural gulf between them, and which workloads each is suited for based strictly on the provided benchmark results and specifications.
Where Each One Wins
The benchmark results are unambiguous in raw compute: the NVIDIA CMP 90HX wins the single head-to-head test, a Geekbench OpenCL run, with a score of 69000 against the Tesla P40's 62017. That is a delta of 11.3%, placing the CMP 90HX firmly ahead in general-purpose compute throughput. The CMP 90HX also posts a higher average benchmark score of 69000 versus the Tesla P40's 65095, reinforcing its status as the faster card in this comparison.
However, the Tesla P40 wins in a category not captured by a single benchmark score: memory capacity. The P40 carries 24 GB of GDDR5 memory, more than double the CMP 90HX's 10 GB of GDDR6X. For workloads that require large datasets to reside on the GPU itself—such as certain inference models or in-memory databases—the P40's capacity advantage is the deciding factor, even if its raw throughput is lower. The P40 also has a wider 384-bit memory bus compared to the CMP 90HX's 320-bit bus, although the CMP 90HX's faster GDDR6X memory gives it a massive bandwidth lead of 760.3 GB/s versus 347.1 GB/s.
In terms of relative standing among peers, the CMP 90HX sits at the 90th percentile of all GPUs, while the Tesla P40 is at the 89th. This near-identical percentile ranking suggests that despite the 11.3% gap between them, both are high-performing accelerators in the broader GPU landscape. The CMP 90HX's closest rival is the Intel Arc A770, which trails by just 0.3%, indicating that the CMP 90HX is competitive with modern consumer-grade cards. The Tesla P40, meanwhile, is within 1.4% of the AMD Radeon Pro WX 9100 and the AMD Radeon VII, showing it still holds its own against newer professional and enthusiast cards.
Architecture Differences
The architectural divide between these two cards is generational and profound. The CMP 90HX is built on the GA102 chip using the Ampere architecture on an 8 nm Samsung process. It packs 28,300 million transistors onto a 628 mm² die, yielding a transistor density of 45.1 million per square millimeter. The Tesla P40, in contrast, uses the GP102 chip from the older Pascal architecture on a 16 nm TSMC process, with 11,800 million transistors on a 471 mm² die and a density of just 25.1 million per square millimeter. This means the CMP 90HX has over twice the transistor count in a larger but far denser package.
The compute resources differ sharply. The CMP 90HX has 6400 shading units, 200 TMUs, and 80 ROPs, while the Tesla P40 has 3840 shading units, 240 TMUs, and 96 ROPs. The CMP 90HX compensates for fewer ROPs with a higher pixel rate of 136.8 GPixel/s versus the P40's 147.0 GPixel/s, but the P40 has a higher texture rate of 367.4 GTexel/s versus the CMP 90HX's 342.0 GTexel/s. In raw FP32 throughput, the CMP 90HX delivers 21.89 TFLOPS, nearly double the P40's 11.76 TFLOPS.
The most significant feature gap is in specialized cores. The CMP 90HX includes 50 RT cores for ray tracing and 200 tensor cores for AI acceleration, features the Tesla P40 lacks entirely. The FP16 performance tells this story: the CMP 90HX achieves 21.89 TFLOPS at a 1:1 ratio with FP32, while the P40 manages only 183.7 GFLOPS at a 1:64 ratio. This makes the CMP 90HX vastly superior for mixed-precision and AI workload, while the P40 is effectively limited to FP32 compute. The CMP 90HX also supports DirectX 12 Ultimate (12_2), while the P40 is limited to DirectX 12 (12_1).
Other differences matter for deployment. The CMP 90HX uses a PCIe 1.0 x4 interface, a legacy bus that severely limits data transfer speeds despite the card's compute power. The Tesla P40 uses PCIe 3.0 x16, a far more standard and faster connection. The CMP 90HX draws 320 W and requires two 8-pin power connectors, while the P40 draws 250 W with a single 8-pin EPS connector. Both are dual-slot cards with no display outputs, indicating their intended role as compute-only accelerators. The CMP 90HX is slightly longer at 285 mm versus the P40's 267 mm.
The Verdict
For raw compute performance, the NVIDIA CMP 90HX is the clear winner. Its 11.3% lead in Geekbench OpenCL, combined with roughly double the FP32 throughput and a massive advantage in FP16 and tensor operations, makes it the superior choice for compute-heavy tasks like machine learning inference, scientific simulation, or any workload that can leverage its RT and tensor cores. The data shows no scenario in the provided benchmarks where the Tesla P40 outperforms the CMP 90HX in speed.
However, the Tesla P40 is not without merit. Its 24 GB of memory is the largest capacity in this comparison, and for workloads that are memory-bound rather than compute-bound, such as loading large models or datasets that exceed 10 GB, the P40 is the only option. The P40's lower power draw of 250 W and standard PCIe 3.0 x16 interface also make it easier to integrate into existing systems without the need for a newer motherboard or power supply.
The choice ultimately depends on workload priorities. If the task requires maximum compute throughput and can fit within 10 GB of memory, the CMP 90HX is the obvious pick. If the task requires more than 10 GB of memory or needs to run on older infrastructure with standard PCIe, the Tesla P40 is the practical choice. The CMP 90HX's PCIe 1.0 x4 interface is a significant bottleneck that could negate its compute advantage in real-world data transfer scenarios.
FAQ
Q: Which card is faster in the head-to-head benchmark?
A: The NVIDIA CMP 90HX wins the only available head-to-head test, Geekbench OpenCL, with a score of 69000 against the Tesla P40's 62017, a delta of 11.3%.
Q: What is the memory capacity difference?
A: The Tesla P40 has 24 GB of GDDR5 memory, while the CMP 90HX has 10 GB of GDDR6X. The P40 offers more capacity, but the CMP 90HX has higher bandwidth at 760.3 GB/s versus 347.1 GB/s.
Q: Does the Tesla P40 support ray tracing or tensor cores?
A: No. The Tesla P40 has no RT cores or tensor cores, while the CMP 90HX includes 50 RT cores and 200 tensor cores.
Q: What are the FP16 performance differences?
A: The CMP 90HX achieves 21.89 TFLOPS FP16 at a 1:1 ratio with FP32, while the Tesla P40 manages only 183.7 GFLOPS at a 1:64 ratio, making the CMP 90HX vastly superior for mixed-precision work.
Q: Which card has a better PCIe interface?
A: The Tesla P40 uses PCIe 3.0 x16, while the CMP 90HX uses PCIe 1.0 x4, which is a much older and slower interface.
Q: What are the power requirements?
A: The CMP 90HX has a TDP of 320 W and requires two 8-pin power connectors, while the Tesla P40 has a TDP of 250 W and uses a single 8-pin EPS connector.
Head-to-Head Benchmarks
The single benchmark result available paints a definitive picture. In Geekbench OpenCL, the NVIDIA CMP 90HX scores 69000, while the NVIDIA Tesla P40 scores 62017. The CMP 90HX wins this test by a margin of 11.3%, a substantial gap that reflects its architectural advantages. This result is consistent with the raw specifications: the CMP 90HX delivers 21.89 TFLOPS FP32 versus the P40's 11.76 TFLOPS, a near-doubling of compute throughput that translates directly into the benchmark score.
Beyond the raw score, the CMP 90HX's advantage is even more pronounced in specialized workloads. The presence of 200 tensor cores and full-rate FP16 (21.89 TFLOPS at 1:1) means that any AI or machine learning task will see a significantly larger performance gap than the 11.3% shown in the OpenCL test, which likely exercises general FP32 compute. The Tesla P40's FP16 performance of 183.7 GFLOPS is a fraction of its FP32 throughput, making it unsuitable for modern mixed-precision training or inference workloads.
The memory bandwidth difference further amplifies the CMP 90HX's lead. With 760.3 GB/s of bandwidth, the CMP 90HX can feed its compute units far faster than the P40's 347.1 GB/s. This is critical for memory-intensive algorithms, where the P40 may stall waiting for data. However, the P40's 24 GB capacity means it can hold larger working sets without spilling to system memory, a factor that could mitigate its bandwidth disadvantage in specific scenarios.
The nearest rivals data provides context for both cards. The CMP 90HX is nearly tied with the Intel Arc A770 (delta 0.3%) and the AMD Radeon Instinct MI25 (delta 0.6%), suggesting it is on par with modern mid-range accelerators. The Tesla P40, despite its age, is within 1.4% of the AMD Radeon Pro WX 9100 and the AMD Radeon VII, and just 2% ahead of the NVIDIA CMP 30HX. This indicates that while the P40 trails the CMP 90HX, it remains a competitive option against other older or lower-tier cards.
In the only direct comparison, the verdict is clear: the CMP 90HX is the faster card by a meaningful margin. The Tesla P40's only winning argument is its larger memory pool, which the benchmark data does not evaluate. For pure compute, the CMP 90HX wins decisively.