NVIDIA GeForce RTX 5090 Mobile vs NVIDIA Tesla P40 Comparison
NVIDIA GeForce RTX 5090 Mobile
Tesla P40
PERFORMANCE BENCHMARKS
Analysis: NVIDIA GeForce RTX 5090 Mobile vs NVIDIA Tesla P40
Head-to-Head Benchmarks
The recorded data contains two shared benchmark tests between the NVIDIA Tesla P40 and the NVIDIA GeForce RTX 5090 Mobile: Geekbench OpenCL and Geekbench Vulkan. The results are decisively one-sided. The RTX 5090 Mobile wins both tests, and by substantial margins.
In Geekbench OpenCL, the RTX 5090 Mobile scores 201,834, while the Tesla P40 scores 62,017. The delta is 69.3% in favor of the RTX 5090 Mobile. This means the mobile part delivers more than three times the raw compute throughput in this workload. The OpenCL test is a general-purpose compute benchmark that stresses shader arithmetic, memory bandwidth, and parallel execution. The RTX 5090 Mobile's score places it far above the P40, and the performance gap is large enough that the P40 cannot offset it with any architectural advantage.
In Geekbench Vulkan, the spread narrows slightly but remains overwhelming. The RTX 5090 Mobile scores 198,405 against the P40's 68,172, a delta of 65.6% in favor of the newer card. Vulkan is a low-level graphics API, and the result indicates that the RTX 5090 Mobile handles modern graphics workloads with far greater efficiency. The P40's Pascal architecture, while competent for its era, lacks the dedicated hardware features and the raw shader throughput of the Blackwell 2.0-based mobile part.
The head-to-head table shows the Tesla P40 wins zero tests, and the RTX 5090 Mobile wins two. There is no benchmark in the shared set where the P40 takes the lead. The magnitude of the losses, 69.3% and 65.6%, is consistent across both API types. That consistency suggests the gap is not a quirk of one workload but a fundamental difference in capability. The RTX 5090 Mobile also has a broader suite of recorded benchmarks, including 3DMark Steel Nomad DX12, several Passmark tests (DirectX 9/10/11/12, G2D, G3D, and GPU Compute), which show scores of 5,871 for Steel Nomad, 30034 for Passmark G3D, and 13,401 for Passmark GPU Compute. The P40 only has the two Geekbench entries.
Architecture Differences
The two cards come from different architectural eras. The Tesla P40 uses the GP102 chip, based on the Pascal architecture, fabricated on a 16 nm process at TSMC. The RTX 5090 Mobile uses the GB203 chip, based on Blackwell 2.0, fabricated on a 5 nm process, also at TSMC. The transistor count tells a major story: the P40 packs 11,800 million transistors on a 471 mm² die, for a transistor density of 25.1M per mm². The RTX 5090 Mobile has 45,600 million transistors on a smaller 378 mm² die, yielding a density of 120.6M per mm². That is roughly a 4.8x advantage in density, which directly translates into far more functional units available to the 5090 Mobile.
The core configuration mirrors that difference. The P40 has 3,840 shading units, 240 texture mapping units, and 96 render output units. The RTX 5090 Mobile has 10,496 shading units, 328 TMUs, and 112 ROPs. The shading unit count is more than 2.7x higher on the RTX part. Additionally, the RTX 5090 Mobile includes 82 ray tracing cores and 328 tensor cores, neither of which exist on the P40. Those dedicated hardware units enable hardware-accelerated ray tracing and AI-based tensor operations, features completely absent from the Pascal design.
Memory configurations are similar in capacity but different in nature. Both cards have 24 GB of memory. The P40 uses GDDR5 on a 384-bit bus, yielding 347.1 GB/s of bandwidth. The RTX 5090 Mobile uses GDDR7 on a narrower 256-bit bus, but its effective speed of 28 Gbps gives it 896.0 GB/s of bandwidth, which is 2.58x the P40's. The P40's memory clock is listed as 1808 MHz (7.2 Gbps effective), while the RTX part runs at 1750 MHz (28 Gbps effective). The higher data rate per pin compensates for the narrower bus.
Clock speeds are interesting: the P40 has a higher base clock of 1303 MHz and a boost of 1531 MHz, versus the RTX 5090 Mobile's base of 990 MHz and boost of 1515 MHz. Despite lower clocks, the RTX part's massive core count and memory bandwidth overwhelm the clock advantage. The FP32 throughputs confirm: the P40 produces 11.76 TFLOPS, while the RTX 5090 Mobile produces 31.80 TFLOPS, a 2.7x difference. The FP16 comparison is even starker. The P40 manages only 183.7 GFLOPS with a 1:64 ratio, while the RTX 5090 Mobile delivers 31.80 TFLOPS with a 1:1 ratio. That is a 173x difference in half-precision throughput, which matters for AI and compute workloads.
The RTX 5090 Mobile also supports DirectX 12 Ultimate (12_2), while the P40 is limited to DirectX 12 (12_1). Both support OpenGL 4.6 and Vulkan 1.4. The PCIe interface differs too: the P40 uses PCIe 3.0 x16, the RTX part uses PCIe 5.0 x16.
Where Each One Wins
Based solely on the benchmark data, the RTX 5090 Mobile wins every shared test. The OpenCL and Vulkan scores are both heavily in its favor, by 69.3% and 65.6% respectively. This indicates the RTX part is more suitable for any workload that runs through those APIs, including compute, rendering, and general-purpose GPU tasks.
The RTX 5090 Mobile's additional benchmark suite expands its win profile. It has a Passmark G3D score of 30,034, a Passmark GPU Compute score of 13,401, and 3DMark Steel Nomad DX12 score of 5,871. These numbers are not directly comparable to the P40, but they illustrate the RTX part's strength across multiple testing methodologies. The P40 has no recorded wins in any benchmark.
The P40's only potential advantage comes from its physical design and deployment context, not from performance scores. It is a dual-slot card with an 8-pin EPS power connector and a 250 W TDP. The RTX 5090 Mobile is an IGP (integrated graphics processor) for laptops, with no power connectors and a 95 W TDP. The P40 has no display outputs, while the RTX part's outputs are listed as "portable device dependent." The P40 is end-of-life, released in 2016, while the RTX 5090 Mobile is active, released in 2025.
The Verdict
The benchmark data is unambiguous. The RTX 5090 Mobile is the superior performer in every head-to-head measure available. It also has the advantage of a modern architecture with ray tracing and tensor cores, a higher transistor density, and over 2.5x the memory bandwidth. The P40, while having a higher base clock and a wider memory bus on paper, simply cannot overcome the generational gap in compute throughput.
The RTX 5090 Mobile should be the choice for any workload that uses OpenCL or Vulkan, as it delivers 65-69% more performance in those tests. It also has a full suite of additional Passmark and 3DMark scores that the P40 lacks, indicating a more well-rounded testing profile. The P40, by contrast, has no recorded wins and lacks the hardware features required for modern graphics effects like real-time ray tracing.
The P40 is an older product with a higher TDP, larger physical footprint, and no display outputs. Its 24 GB of GDDR5 memory is the same capacity as the RTX part, but the bandwidth is far lower. The data shows the P40's strengths are limited to the fact that it is a professional-grade card with a 5,699 USD launch MSRP, but that price point does not translate to any performance advantage in the recorded tests. For pure compute or graphics performance, the RTX 5090 Mobile is the clear winner.
FAQ
Q: How much faster is the RTX 5090 Mobile than the P40 in Geekbench OpenCL?
A: The RTX 5090 Mobile scores 201,834, while the P40 scores 62,017. That is a 69.3% difference in favor of the RTX part.
Q: Does the P40 have any benchmark where it beats the RTX 5090 Mobile?
A: No. In the two shared tests (Geekbench OpenCL and Vulkan), the RTX 5090 Mobile wins both. The P40 has zero head-to-head wins.
Q: What is the memory bandwidth difference?
A: The RTX 5090 Mobile has 896.0 GB/s of bandwidth, while the P40 has 347.1 GB/s. The RTX part provides roughly 2.58 times the bandwidth.
Q: Which card has more shading units?
A: The RTX 5090 Mobile has 10,496 shading units, whereas the P40 has 3,840. The RTX part has more than 2.7 times the shading units.
Q: Does the P40 support hardware ray tracing?
A: No. The P40 has no ray cores. The RTX 5090 Mobile has 82 ray cores.
Q: What is the transistor density difference?
A: The P40 has 25.1M transistors per mm², while the RTX 5090 Mobile has 120.6M per mm².
Specification Differences
| Field | NVIDIA Tesla P40 | NVIDIA GeForce RTX 5090 Mobile |
|-------|------------------|--------------------------------|
| Architecture | Pascal | Blackwell 2.0 |
| Process Node | 16 nm | 5 nm |
| Transistors | 11,800 million | 45,600 million |
| Die Size | 471 mm² | 378 mm² |
| Transistor Density | 25.1M / mm² | 120.6M / mm² |
| Base Clock | 1303 MHz | 990 MHz |
| Boost Clock | 1531 MHz | 1515 MHz |
| Memory Clock | 1807 MHz (7.2 Gbps effective) | 1750 MHz (28 Gbps effective) |
| Memory Type | GDDR5 | GDDR7 |
| Memory Bus Width | 384 bit | 256 bit |
| Memory Bandwidth | 347.1 GB/s | 896.0 GB/s |
| Shading Units | 3840 | 10496 |
| TMUs | 240 | 328 |
| ROPs | 96 | 112 |
| RT Cores | 0 | 82 |
| Tensor Cores | 0 | 328 |
| Pixel Rate | 147.0 GPixel/s | 169.7 GPixel/s |
| Texture Rate | 367.4 GTexel/s | 496.9 GTexel/s |
| FP32 | 11.76 TFLOPS | 31.80 TFLOPS |
| FP16 | 183.7 GFLOPS | 31.80 TFLOPS |
| TDP | 250 W | 95 W |
| Slot Width | Dual-slot | IGP |
| Power Connectors | 8-pin EPS | None |
| Suggested PSU | 600 W | None |
| Bus Interface | PCIe 3.0 x16 | PCIe 5.0 x16 |
| Display Outputs | No outputs | Portable Device Dependent |
| DirectX | 12 (12_1) | 12 Ultimate (12_2) |
| Release Date | 2016-09-12 | 2025-03-26 |
| Production Status | End-of-life | Active |
| Launch MSRP | 5,699 USD | None |
The data clearly shows that the RTX 5090 Mobile is the newer, more advanced part in nearly every specification that affects performance. The only areas where the P40 has a higher number are base clock, boost clock, and memory bus width, but those do not compensate for the RTX part's massive lead in core count, bandwidth, and feature set.