NVIDIA CMP 90HX vs NVIDIA Tesla T4 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 90HX

CORE STATE GA102
VRAM 10 GB
CLOCK SPEED 1710 MHz
TDP 320 W
BUS WIDTH 320 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Tesla T4

CORE STATE TU104
VRAM 16 GB
CLOCK SPEED 1590 MHz
TDP 70 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2018

PERFORMANCE BENCHMARKS

geekbench_opencl
69,000
61,276
geekbench_vulkan
N/A
72,190

Analysis: NVIDIA CMP 90HX vs NVIDIA Tesla T4

The NVIDIA CMP 90HX and NVIDIA Tesla T4 represent two fundamentally different approaches to GPU design, despite sharing the same manufacturer. The data shows a clear split between raw compute throughput and specialized efficiency. In the sole head-to-head benchmark available, the Geekbench OpenCL test, the CMP 90HX delivers a 12.6% higher score, but the Tesla T4 counters with superior memory capacity, dramatically lower power draw, and a different architectural focus. This analysis breaks down where each card excels based strictly on the provided benchmark data and specifications.

Where Each One Wins

The NVIDIA CMP 90HX is the undisputed winner in raw compute performance. Its Geekbench OpenCL score of 69000 places it in the 90th percentile of all GPUs, matching the Tesla T4's percentile ranking but with a substantial absolute lead. The CMP 90HX's nearest rivals include the Intel Arc A770 (68809, +0.3%) and AMD Radeon Instinct MI25 (68562, +0.6%), while it trails the AMD Radeon Pro WX 8200 (69870, -1.2%) and NVIDIA Quadro P6000 (69986, -1.4%) by narrow margins. This positions the CMP 90HX as a high-end performer that edges out most competitors in its immediate class.

The Tesla T4 wins in memory capacity and power efficiency. Its 16 GB GDDR6 memory is 60% larger than the CMP 90HX's 10 GB, which is critical for models or datasets that exceed the smaller card's capacity. The T4's 70 W TDP is a fraction of the CMP 90HX's 320 W, making it suitable for dense server deployments where thermal and power budgets are constrained. The T4 also offers a Geekbench Vulkan score of 72190, a benchmark the CMP 90HX was not tested on, suggesting the T4 handles graphics-level compute tasks effectively. Its nearest rivals include the AMD Radeon VII (66004, +1.1%) and NVIDIA Tesla P40 (65095, +2.5%), showing the T4 outperforms these alternatives despite its lower raw FP32 throughput.

Architecture Differences

The architectural gap between these two GPUs is generational and fundamental. The CMP 90HX uses the GA102 chip on the Ampere architecture, built on Samsung's 8 nm process. It packs 28,300 million transistors into a 628 mm² die, achieving a transistor density of 45.1M per mm². This chip is designed for maximum compute throughput, featuring 6400 shading units, 200 texture mapping units, 80 ROPs, 50 ray tracing cores, and 200 tensor cores. The FP32 performance is 21.89 TFLOPS, and FP16 performance is identical at 21.89 TFLOPS with a 1:1 ratio, indicating no specialized half-precision acceleration.

The Tesla T4 utilizes the TU104 chip on the older Turing architecture, built on TSMC's 12 nm process. It contains 13,600 million transistors on a 545 mm² die, with a lower transistor density of 25.0M per mm². The T4 has 2560 shading units, 160 TMUs, 64 ROPs, 40 ray tracing cores, and 320 tensor cores. Its FP32 performance is 8.141 TFLOPS, but FP16 performance doubles to 16.28 TFLOPS with a 2:1 ratio, showing a deliberate design choice to accelerate half-precision workloads. The T4's higher tensor core count relative to shading units (320 vs 2560, or 1:8) compared to the CMP 90HX (200 vs 6400, or 1:32) indicates the T4 is more focused on AI inference tasks that rely heavily on tensor operations.

Memory subsystems differ significantly. The CMP 90HX uses 10 GB of GDDR6X on a 320-bit bus, delivering 760.3 GB/s bandwidth. The T4 uses 16 GB of GDDR6 on a 256-bit bus, delivering 320.0 GB/s bandwidth. The CMP 90HX's bandwidth advantage is clear, but the T4's larger capacity with lower speed favors workloads that need to hold large working sets rather than stream data rapidly. The CMP 90HX also has a peculiar PCIe 1.0 x4 interface, which is severely limited compared to the T4's PCIe 3.0 x16 interface, potentially bottlenecking data transfer in some scenarios. Both cards have no display outputs, confirming their compute-only purpose.

Head-to-Head Benchmarks

The only direct benchmark comparison available is the Geekbench OpenCL test, where the NVIDIA CMP 90HX scores 69000 against the Tesla T4's 61276. This represents a 12.6% advantage for the CMP 90HX. In raw OpenCL compute, the CMP 90HX's higher shading unit count and faster clock speeds (1500 MHz base, 1710 MHz boost vs 585 MHz base, 1590 MHz boost) translate directly into superior performance. The CMP 90HX's FP32 throughput of 21.89 TFLOPS is 2.7 times the T4's 8.141 TFLOPS, which explains the substantial lead in this workload.

However, the benchmark data also reveals the T4's strengths. Its Geekbench Vulkan score of 72190 exceeds its own OpenCL score and would likely beat the CMP 90HX's OpenCL result of 69000 if directly compared. Vulkan is a lower-level API that can better expose the T4's tensor core capabilities, and the T4's higher tensor core count (320 vs 200) suggests it excels in workloads that leverage these units. The T4's FP16 performance of 16.28 TFLOPS is 75% of the CMP 90HX's FP16 throughput, a much closer margin than the FP32 comparison, indicating the T4 narrows the gap when half-precision is used.

The deltaPct values in the nearest rivals data further contextualize these scores. The CMP 90HX's closest competitor, the Intel Arc A770, is only 0.3% behind, showing the CMP 90HX sits at the top of a tight cluster. The T4's closest rival, the AMD Radeon VII, is 1.1% behind, while the T4 trails the AMD Radeon Instinct MI25 by 2.7% and the Intel Arc A770 by 3%. This suggests the T4 is competitive but not dominant in its performance tier, unlike the CMP 90HX which leads its immediate rivals by small but consistent margins.

The Verdict

The data presents a clear choice: the NVIDIA CMP 90HX is the superior choice for raw compute performance, as evidenced by its 12.6% OpenCL lead and 21.89 TFLOPS FP32 throughput. Its 760.3 GB/s memory bandwidth and 10 GB GDDR6X memory make it well-suited for bandwidth-intensive tasks like real-time ray tracing or large matrix operations. The 90th percentile ranking with a 69000 average score confirms its position as a high-end performer, narrowly beating the Intel Arc A770 and AMD Radeon Instinct MI25 while staying within 1.4% of the NVIDIA Quadro P6000.

The Tesla T4 is the better option for memory-constrained and power-sensitive environments. Its 16 GB memory capacity is 60% larger, accommodating larger models or datasets without swapping. The 70 W TDP versus 320 W means the T4 can be deployed in far denser configurations, and its single-slot design with no power connectors simplifies installation. The T4's FP16 performance of 16.28 TFLOPS and higher tensor core ratio (1:8 vs 1:32) make it more suitable for AI inference workloads that rely on half-precision tensor operations. Its Vulkan score of 72190 also suggests strong performance in Vulkan-based compute, which the CMP 90HX was not tested on.

The CMP 90HX is end-of-life and uses a PCIe 1.0 x4 interface, which is a significant limitation for data transfer in modern servers. The T4, while also end-of-life, uses PCIe 3.0 x16 and was designed for the Tesla Turing server generation, with a predecessor in Tesla Volta and successor in Server Ampere. For users prioritizing raw compute in a single-GPU scenario, the CMP 90HX is the data-backed choice. For users prioritizing memory capacity, power efficiency, and tensor-heavy AI workloads, the Tesla T4 is clearly preferable. The benchmark results show a 12.6% performance gap in favor of the CMP 90HX, but the T4's architectural advantages in other dimensions make it the more versatile option for varied server workloads.

FAQ

Q: Which GPU has higher raw compute performance in OpenCL?

A: The NVIDIA CMP 90HX scores 69000 in Geekbench OpenCL, which is 12.6% higher than the Tesla T4's 61276 score. The CMP 90HX also delivers 21.89 TFLOPS FP32 versus the T4's 8.141 TFLOPS.

Q: How do memory capacities compare between the two cards?

A: The Tesla T4 has 16 GB of GDDR6 memory, which is 60% larger than the CMP 90HX's 10 GB of GDDR6X. However, the CMP 90HX has much higher bandwidth at 760.3 GB/s compared to the T4's 320.0 GB/s.

Q: What is the power consumption difference?

A: The Tesla T4 has a 70 W TDP, while the CMP 90HX has a 320 W TDP. The T4 also requires no power connectors and suggests a 250 W PSU, whereas the CMP 90HX needs 2x 8-pin connectors and a 700 W PSU.

Q: Which card is better for AI inference workloads?

A: The Tesla T4 has 320 tensor cores versus the CMP 90HX's 200, and its FP16 performance of 16.28 TFLOPS is 75% of the CMP 90HX's 21.89 TFLOPS, despite having far lower FP32 throughput. The T4's 2:1 FP16 ratio indicates deliberate half-precision acceleration.

Q: How do the two cards rank relative to their nearest rivals?

A: The CMP 90HX leads the Intel Arc A770 by 0.3% and the AMD Radeon Instinct MI25 by 0.6%, but trails the AMD Radeon Pro WX 8200 by 1.2% and NVIDIA Quadro P6000 by 1.4%. The Tesla T4 leads the AMD Radeon VII by 1.1% and NVIDIA Tesla P40 by 2.5%, but trails the AMD Radeon Instinct MI25 by 2.7% and Intel Arc A770 by 3%.

Q: What are the interface and form factor differences?

A: The CMP 90HX uses a PCIe 1.0 x4 interface, is dual-slot, and measures 285 mm in length. The Tesla T4 uses PCIe 3.0 x16, is single-slot, and measures 168 mm in length. Both cards have no display outputs.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 90HX
Tesla T4
Core Specs
Shading Units
6,400
2,560 -60.0%
Shaders
6,400
2,560 -60.0%
TMUs
200
160 -20.0%
ROPs
80
64 -20.0%
SM Count
50
40 -20.0%
Clocks
Base Clock
1500 MHz
585 MHz
Boost Clock
1710 MHz
1590 MHz
Memory Clock
1188 MHz 19 Gbps effective
1250 MHz 10 Gbps effective
Memory
Memory Size
10 GB
16 GB
VRAM (MB)
10,240
16,384 +60.0%
Memory Type
GDDR6X
GDDR6
Memory Bus
320 bit
256 bit
Bandwidth
760.3 GB/s
320.0 GB/s
Cache
L1 Cache
128 KB (per SM)
64 KB (per SM)
L2 Cache
5 MB
4 MB
Performance
Pixel Rate
136.8 GPixel/s
101.8 GPixel/s
Texture Rate
342.0 GTexel/s
254.4 GTexel/s
FP32 (TFLOPS)
21.89 TFLOPS
8.141 TFLOPS
FP64 (TFLOPS)
342.0 GFLOPS (1:64)
254.4 GFLOPS (1:32)
FP16 (TFLOPS)
21.89 TFLOPS (1:1)
16.28 TFLOPS (2:1)
AI/RT
RT Cores
50
40 -20.0%
Tensor Cores
200
320 +60.0%
Power
TDP
320 W
70 W
TDP (W)
320
70 -78.1%
Suggested PSU
700 W
250 W
Power Connectors
2x 8-pin
None
Architecture
Architecture
Ampere
Turing
GPU Name
GA102
TU104
Generation
Mining GPUs
Tesla Turing (Txx)
Process Size
8 nm
12 nm
Transistors
28,300 million
13,600 million
Die Size
628 mm²
545 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
25.0M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
7.5
Shader Model
6.8
6.9
Physical
Slot Width
Dual-slot
Single-slot
Length
285 mm 11.2 inches
168 mm 6.6 inches
Height
112 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 1.0 x4
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Volta
Successor
Server Ampere
View CMP 90HX Details View Tesla T4 Details