NVIDIA GeForce RTX 4090 vs NVIDIA Tesla P40 Comparison
NVIDIA GeForce RTX 4090
Tesla P40
PERFORMANCE BENCHMARKS
Analysis: NVIDIA GeForce RTX 4090 vs NVIDIA Tesla P40
The NVIDIA Tesla P40 and the NVIDIA GeForce RTX 4090 represent two distinct eras of GPU design, separated by six years of architectural evolution. The data in the FACT PACK shows a stark contrast: the RTX 4090 dominates every shared benchmark, yet the Tesla P40 holds its own in the broader performance distribution. This comparison is not about a close contest; it is about understanding where a specialized compute card from 2016 stands against a modern flagship, and what that means for different workloads.
The Verdict
Based strictly on the benchmark data, the NVIDIA GeForce RTX 4090 is the clear winner in raw compute performance. In the two benchmarks shared between the cards, the RTX 4090 achieves a score of 255,416 in Geekbench OpenCL, which is 75.7% higher than the Tesla P40's 62,017. Similarly, in Geekbench Vulkan, the RTX 4090 scores 271,631 versus the P40's 68,172, a 74.9% lead. These deltaPct values are massive and unambiguous.
However, the percentile rankings tell a more nuanced story. The Tesla P40 sits at the 89th percentile of all GPUs, slightly above the RTX 4090's 88th percentile. This is counterintuitive given the RTX 4090's raw score advantage, but it reflects the average benchmark scores: the P40 averages 65,095, while the RTX 4090 averages 60,347. The difference arises because the RTX 4090's average is dragged down by its Passmark DirectX scores (which are low, ranging from 150 to 397), while the P40 only has two high Geekbench scores. For buyers, the verdict is simple: if your workload is generic compute or gaming-like APIs, the RTX 4090 is overwhelmingly superior. If your application relies solely on OpenCL or Vulkan compute and you need a card that slots into a server without display outputs, the Tesla P40 remains a viable option.
Where Each One Wins
The RTX 4090 wins in every head-to-head benchmark recorded in the FACT PACK, so its advantage is across the board. It excels in Geekbench OpenCL and Vulkan, which are general-purpose compute tests. Its Passmark scores, while low relative to its other scores, still cover DirectX 9 through 12, G2D, G3D, and GPU compute, indicating a wide range of capability. The RTX 4090 is also the only card with display outputs (1x HDMI 2.1 and 3x DisplayPort 1.4a), making it suitable for any workload that requires visual output.
The Tesla P40 wins in the context of its niche. It has no display outputs, which means it is designed exclusively for headless compute. Its wins are not in raw performance but in efficiency of purpose: it is a dual-slot card with a 250W TDP, compared to the RTX 4090's triple-slot design and 450W TDP. The P40 also uses an 8-pin EPS power connector, which is common in server environments, whereas the RTX 4090 requires a 16-pin connector. For a server rack with strict power and cooling budgets, the P40's lower power draw and dual-slot form factor could be a practical win, even if its compute scores are far behind.
Architecture Differences
The architectural gap between these two cards is generational. The Tesla P40 uses the GP102 chip on the Pascal architecture, built on a 16 nm process at TSMC. The RTX 4090 uses the AD102 chip on the Ada Lovelace architecture, built on a 5 nm process, also at TSMC. This process shrink allows for a dramatic increase in transistor count: the P40 has 11,800 million transistors on a 471 mm² die, while the RTX 4090 packs 76,300 million transistors into a 609 mm² die. The transistor density jumps from 25.1M per mm² on the P40 to 125.3M per mm² on the RTX 4090.
The RTX 4090 also introduces hardware that the P40 lacks entirely: 128 RT cores for ray tracing and 512 tensor cores for AI acceleration. The P40 has no such dedicated units. In terms of raw compute units, the RTX 4090 has 16,384 shading units, 512 TMUs, and 176 ROPs, compared to the P40's 3,840 shading units, 240 TMUs, and 96 ROPs. Clock speeds are also higher on the RTX 4090, with a base of 2235 MHz and boost of 2520 MHz versus the P40's 1303 MHz base and 1531 MHz boost. The FP32 throughput tells the story: 82.58 TFLOPS on the RTX 4090 versus 11.76 TFLOPS on the P40. Notably, the RTX 4090 has 1:1 FP16 to FP32 ratio (82.58 TFLOPS), while the P40's FP16 is a paltry 183.7 GFLOPS at a 1:64 ratio.
FAQ
Q: Which card has a higher average benchmark score?
A: The Tesla P40 has a higher average benchmark score of 65,095, compared to the RTX 4090's 60,347. This is despite the RTX 4090 winning both head-to-head benchmarks.
Q: Does the Tesla P40 support any modern APIs?
A: Yes, the Tesla P40 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The RTX 4090 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.
Q: What is the memory configuration difference?
A: Both cards have 24 GB of memory on a 384-bit bus. The Tesla P40 uses GDDR5 with 347.1 GB/s bandwidth and 7.2 Gbps effective speed. The RTX 4090 uses GDDR6X with 1.01 TB/s bandwidth and 21 Gbps effective speed.
Q: Which card has a higher pixel rate?
A: The RTX 4090 has a pixel rate of 443.5 GPixel/s, which is significantly higher than the Tesla P40's 147.0 GPixel/s.
Q: Are both cards still in production?
A: No. Both are marked as end-of-life in the FACT PACK. The Tesla P40 was released on 2016-09-12, and the RTX 4090 was released on 2022-09-19.
Q: What is the launch MSRP of each card?
A: The Tesla P40 has a launch MSRP of 5,699 USD, while the RTX 4090 has a launch MSRP of 1,599 USD.
Head-to-Head Benchmarks
The only two benchmarks where both cards have scores are Geekbench OpenCL and Geekbench Vulkan. In Geekbench OpenCL, the RTX 4090 scores 255,416 against the Tesla P40's 62,017. The deltaPct is -75.7%, meaning the RTX 4090 is 75.7% faster. In Geekbench Vulkan, the RTX 4090 scores 271,631 against the P40's 68,172, a deltaPct of -74.9%. These are the largest wins in the data.
Beyond these, the RTX 4090 has additional benchmarks that the P40 lacks. In 3DMark Steel Nomad DX12, it scores 9,223. Its Passmark scores include 38,194 in G3D, 26,613 in GPU Compute, and 1,299 in G2D. The DirectX-specific Passmark scores are 397 (DX9), 326 (DX11), 224 (DX10), and 150 (DX12). These results indicate that the RTX 4090's performance is not uniform across APIs; it performs best in G3D and compute, but its DirectX scores are relatively low, which contributes to its lower average benchmark score compared to the P40.
Specification Differences
The table below highlights only the fields where the two cards differ, based on the FACT PACK data.
| Specification | NVIDIA Tesla P40 | NVIDIA GeForce RTX 4090 |
|---|---|---|
| Architecture | Pascal | Ada Lovelace |
| Process Node | 16 nm | 5 nm |
| Transistors | 11,800 million | 76,300 million |
| Die Size | 471 mm² | 609 mm² |
| Transistor Density | 25.1M / mm² | 125.3M / mm² |
| Base Clock | 1303 MHz | 2235 MHz |
| Boost Clock | 1531 MHz | 2520 MHz |
| Memory Type | GDDR5 | GDDR6X |
| Memory Speed | 7.2 Gbps effective | 21 Gbps effective |
| Memory Bandwidth | 347.1 GB/s | 1.01 TB/s |
| Shading Units | 3840 | 16384 |
| TMUs | 240 | 512 |
| ROPs | 96 | 176 |
| RT Cores | None | 128 |
| Tensor Cores | None | 512 |
| Pixel Rate | 147.0 GPixel/s | 443.5 GPixel/s |
| Texture Rate | 367.4 GTexel/s | 1,290.2 GTexel/s |
| FP32 Performance | 11.76 TFLOPS | 82.58 TFLOPS |
| FP16 Performance | 183.7 GFLOPS (1:64) | 82.58 TFLOPS (1:1) |
| TDP | 250 W | 450 W |
| Slot Width | Dual-slot | Triple-slot |
| Power Connectors | 8-pin EPS | 1x 16-pin |
| Suggested PSU | 600 W | 850 W |
| Bus Interface | PCIe 3.0 x16 | PCIe 4.0 x16 |
| Display Outputs | No outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a |
| DirectX Support | 12 (12_1) | 12 Ultimate (12_2) |
| Dimensions (L x H x W) | 267 mm x 111 mm | 304 mm x 137 mm x 61 mm |
| Release Date | 2016-09-12 | 2022-09-19 |
| Launch MSRP | 5,699 USD | 1,599 USD |
| Predecessor | Tesla Maxwell | GeForce 30 |
| Successor | Tesla Volta | GeForce 50 |