NVIDIA GeForce RTX 4090 vs NVIDIA Tesla P40 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4090

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 450 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022
VS
NVIDIA
GEFORCE

Tesla P40

CORE STATE GP102
VRAM 24 GB
CLOCK SPEED 1531 MHz
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
9,223
N/A
geekbench_opencl
255,416
62,017
geekbench_vulkan
271,631
68,172
passmark_directx_10
224
N/A
passmark_directx_11
326
N/A
passmark_directx_12
150
N/A
passmark_directx_9
397
N/A
passmark_g2d
1,299
N/A
passmark_g3d
38,194
N/A
passmark_gpu_compute
26,613
N/A

Analysis: NVIDIA GeForce RTX 4090 vs NVIDIA Tesla P40

The NVIDIA Tesla P40 and the NVIDIA GeForce RTX 4090 represent two distinct eras of GPU design, separated by six years of architectural evolution. The data in the FACT PACK shows a stark contrast: the RTX 4090 dominates every shared benchmark, yet the Tesla P40 holds its own in the broader performance distribution. This comparison is not about a close contest; it is about understanding where a specialized compute card from 2016 stands against a modern flagship, and what that means for different workloads.

The Verdict

Based strictly on the benchmark data, the NVIDIA GeForce RTX 4090 is the clear winner in raw compute performance. In the two benchmarks shared between the cards, the RTX 4090 achieves a score of 255,416 in Geekbench OpenCL, which is 75.7% higher than the Tesla P40's 62,017. Similarly, in Geekbench Vulkan, the RTX 4090 scores 271,631 versus the P40's 68,172, a 74.9% lead. These deltaPct values are massive and unambiguous.

However, the percentile rankings tell a more nuanced story. The Tesla P40 sits at the 89th percentile of all GPUs, slightly above the RTX 4090's 88th percentile. This is counterintuitive given the RTX 4090's raw score advantage, but it reflects the average benchmark scores: the P40 averages 65,095, while the RTX 4090 averages 60,347. The difference arises because the RTX 4090's average is dragged down by its Passmark DirectX scores (which are low, ranging from 150 to 397), while the P40 only has two high Geekbench scores. For buyers, the verdict is simple: if your workload is generic compute or gaming-like APIs, the RTX 4090 is overwhelmingly superior. If your application relies solely on OpenCL or Vulkan compute and you need a card that slots into a server without display outputs, the Tesla P40 remains a viable option.

Where Each One Wins

The RTX 4090 wins in every head-to-head benchmark recorded in the FACT PACK, so its advantage is across the board. It excels in Geekbench OpenCL and Vulkan, which are general-purpose compute tests. Its Passmark scores, while low relative to its other scores, still cover DirectX 9 through 12, G2D, G3D, and GPU compute, indicating a wide range of capability. The RTX 4090 is also the only card with display outputs (1x HDMI 2.1 and 3x DisplayPort 1.4a), making it suitable for any workload that requires visual output.

The Tesla P40 wins in the context of its niche. It has no display outputs, which means it is designed exclusively for headless compute. Its wins are not in raw performance but in efficiency of purpose: it is a dual-slot card with a 250W TDP, compared to the RTX 4090's triple-slot design and 450W TDP. The P40 also uses an 8-pin EPS power connector, which is common in server environments, whereas the RTX 4090 requires a 16-pin connector. For a server rack with strict power and cooling budgets, the P40's lower power draw and dual-slot form factor could be a practical win, even if its compute scores are far behind.

Architecture Differences

The architectural gap between these two cards is generational. The Tesla P40 uses the GP102 chip on the Pascal architecture, built on a 16 nm process at TSMC. The RTX 4090 uses the AD102 chip on the Ada Lovelace architecture, built on a 5 nm process, also at TSMC. This process shrink allows for a dramatic increase in transistor count: the P40 has 11,800 million transistors on a 471 mm² die, while the RTX 4090 packs 76,300 million transistors into a 609 mm² die. The transistor density jumps from 25.1M per mm² on the P40 to 125.3M per mm² on the RTX 4090.

The RTX 4090 also introduces hardware that the P40 lacks entirely: 128 RT cores for ray tracing and 512 tensor cores for AI acceleration. The P40 has no such dedicated units. In terms of raw compute units, the RTX 4090 has 16,384 shading units, 512 TMUs, and 176 ROPs, compared to the P40's 3,840 shading units, 240 TMUs, and 96 ROPs. Clock speeds are also higher on the RTX 4090, with a base of 2235 MHz and boost of 2520 MHz versus the P40's 1303 MHz base and 1531 MHz boost. The FP32 throughput tells the story: 82.58 TFLOPS on the RTX 4090 versus 11.76 TFLOPS on the P40. Notably, the RTX 4090 has 1:1 FP16 to FP32 ratio (82.58 TFLOPS), while the P40's FP16 is a paltry 183.7 GFLOPS at a 1:64 ratio.

FAQ

Q: Which card has a higher average benchmark score?

A: The Tesla P40 has a higher average benchmark score of 65,095, compared to the RTX 4090's 60,347. This is despite the RTX 4090 winning both head-to-head benchmarks.

Q: Does the Tesla P40 support any modern APIs?

A: Yes, the Tesla P40 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The RTX 4090 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Q: What is the memory configuration difference?

A: Both cards have 24 GB of memory on a 384-bit bus. The Tesla P40 uses GDDR5 with 347.1 GB/s bandwidth and 7.2 Gbps effective speed. The RTX 4090 uses GDDR6X with 1.01 TB/s bandwidth and 21 Gbps effective speed.

Q: Which card has a higher pixel rate?

A: The RTX 4090 has a pixel rate of 443.5 GPixel/s, which is significantly higher than the Tesla P40's 147.0 GPixel/s.

Q: Are both cards still in production?

A: No. Both are marked as end-of-life in the FACT PACK. The Tesla P40 was released on 2016-09-12, and the RTX 4090 was released on 2022-09-19.

Q: What is the launch MSRP of each card?

A: The Tesla P40 has a launch MSRP of 5,699 USD, while the RTX 4090 has a launch MSRP of 1,599 USD.

Head-to-Head Benchmarks

The only two benchmarks where both cards have scores are Geekbench OpenCL and Geekbench Vulkan. In Geekbench OpenCL, the RTX 4090 scores 255,416 against the Tesla P40's 62,017. The deltaPct is -75.7%, meaning the RTX 4090 is 75.7% faster. In Geekbench Vulkan, the RTX 4090 scores 271,631 against the P40's 68,172, a deltaPct of -74.9%. These are the largest wins in the data.

Beyond these, the RTX 4090 has additional benchmarks that the P40 lacks. In 3DMark Steel Nomad DX12, it scores 9,223. Its Passmark scores include 38,194 in G3D, 26,613 in GPU Compute, and 1,299 in G2D. The DirectX-specific Passmark scores are 397 (DX9), 326 (DX11), 224 (DX10), and 150 (DX12). These results indicate that the RTX 4090's performance is not uniform across APIs; it performs best in G3D and compute, but its DirectX scores are relatively low, which contributes to its lower average benchmark score compared to the P40.

Specification Differences

The table below highlights only the fields where the two cards differ, based on the FACT PACK data.

| Specification | NVIDIA Tesla P40 | NVIDIA GeForce RTX 4090 |

|---|---|---|

| Architecture | Pascal | Ada Lovelace |

| Process Node | 16 nm | 5 nm |

| Transistors | 11,800 million | 76,300 million |

| Die Size | 471 mm² | 609 mm² |

| Transistor Density | 25.1M / mm² | 125.3M / mm² |

| Base Clock | 1303 MHz | 2235 MHz |

| Boost Clock | 1531 MHz | 2520 MHz |

| Memory Type | GDDR5 | GDDR6X |

| Memory Speed | 7.2 Gbps effective | 21 Gbps effective |

| Memory Bandwidth | 347.1 GB/s | 1.01 TB/s |

| Shading Units | 3840 | 16384 |

| TMUs | 240 | 512 |

| ROPs | 96 | 176 |

| RT Cores | None | 128 |

| Tensor Cores | None | 512 |

| Pixel Rate | 147.0 GPixel/s | 443.5 GPixel/s |

| Texture Rate | 367.4 GTexel/s | 1,290.2 GTexel/s |

| FP32 Performance | 11.76 TFLOPS | 82.58 TFLOPS |

| FP16 Performance | 183.7 GFLOPS (1:64) | 82.58 TFLOPS (1:1) |

| TDP | 250 W | 450 W |

| Slot Width | Dual-slot | Triple-slot |

| Power Connectors | 8-pin EPS | 1x 16-pin |

| Suggested PSU | 600 W | 850 W |

| Bus Interface | PCIe 3.0 x16 | PCIe 4.0 x16 |

| Display Outputs | No outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a |

| DirectX Support | 12 (12_1) | 12 Ultimate (12_2) |

| Dimensions (L x H x W) | 267 mm x 111 mm | 304 mm x 137 mm x 61 mm |

| Release Date | 2016-09-12 | 2022-09-19 |

| Launch MSRP | 5,699 USD | 1,599 USD |

| Predecessor | Tesla Maxwell | GeForce 30 |

| Successor | Tesla Volta | GeForce 50 |

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4090
Tesla P40
Core Specs
Shading Units
16,384
3,840 -76.6%
Shaders
16,384
3,840 -76.6%
TMUs
512
240 -53.1%
ROPs
176
96 -45.5%
SM Count
128
30 -76.6%
Clocks
Base Clock
2235 MHz
1303 MHz
Boost Clock
2520 MHz
1531 MHz
Memory Clock
1313 MHz 21 Gbps effective
1808 MHz 7.2 Gbps effective
Memory
Memory Size
24 GB
24 GB
VRAM (MB)
24,576
24,576 0.0%
Memory Type
GDDR6X
GDDR5
Memory Bus
384 bit
384 bit
Bandwidth
1.01 TB/s
347.1 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SM)
L2 Cache
72 MB
3 MB
Performance
Pixel Rate
443.5 GPixel/s
147.0 GPixel/s
Texture Rate
1,290.2 GTexel/s
367.4 GTexel/s
FP32 (TFLOPS)
82.58 TFLOPS
11.76 TFLOPS
FP64 (TFLOPS)
1,290.2 GFLOPS (1:64)
367.4 GFLOPS (1:32)
FP16 (TFLOPS)
82.58 TFLOPS (1:1)
183.7 GFLOPS (1:64)
AI/RT
RT Cores
128
Tensor Cores
512
Power
TDP
450 W
250 W
TDP (W)
450
250 -44.4%
Suggested PSU
850 W
600 W
Power Connectors
1x 16-pin
8-pin EPS
Architecture
Architecture
Ada Lovelace
Pascal
GPU Name
AD102
GP102
Generation
GeForce 40
Tesla Pascal (Pxx)
Process Size
5 nm
16 nm
Transistors
76,300 million
11,800 million
Die Size
609 mm²
471 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
25.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Triple-slot
Dual-slot
Length
304 mm 12 inches
267 mm 10.5 inches
Height
137 mm 5.4 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 3.0 x16
Other
Launch Price
1,599 USD
5,699 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 30
Tesla Maxwell
Successor
GeForce 50
Tesla Volta
View GeForce RTX 4090 Details View Tesla P40 Details