GPU Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4090

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 450 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022
VS
NVIDIA
GEFORCE

Tesla P40

CORE STATE GP102
VRAM 24 GB
CLOCK SPEED 1531 MHz
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
9,223
N/A
geekbench_opencl
254,291
62,017
geekbench_vulkan
259,107
68,172
passmark_directx_10
224
N/A
passmark_directx_11
326
N/A
passmark_directx_12
150
N/A
passmark_directx_9
397
N/A
passmark_g2d
1,299
N/A
passmark_g3d
38,194
N/A
passmark_gpu_compute
26,613
N/A

Analysis: NVIDIA GeForce RTX 4090 vs NVIDIA Tesla P40

The benchmark data places the NVIDIA GeForce RTX 4090 and the NVIDIA Tesla P40 in the same overall performance percentile, both ranking at the 91st percentile among all GPUs. Their average benchmark scores are remarkably close, with the RTX 4090 posting 66,473 against the Tesla P40’s 66,127, a difference of just 0.5%. However, this near-parity in aggregate scores masks a profound divergence in individual workloads, where the RTX 4090 demonstrates overwhelming superiority in compute-centric tests. The Tesla P40, a Pascal-era server card, holds its own only in the context of its narrow benchmark profile, while the Ada Lovelace-based RTX 4090 delivers generational leaps in raw throughput.

Head-to-Head Benchmarks

The most dramatic separation occurs in the Geekbench OpenCL test, where the RTX 4090 scores 317,684 against the Tesla P40’s 62,017. This represents a 412.3% advantage for the newer card, a delta that underscores the architectural gulf between the two. The RTX 4090’s FP32 compute rate of 82.58 TFLOPS versus the P40’s 11.76 TFLOPS explains the magnitude of this gap; the Ada Lovelace card is processing over seven times the floating-point operations per second. In Vulkan workloads, the margin narrows but remains decisive: the RTX 4090 scores 270,615 against the P40’s 70,237, a 285.3% lead. This test reflects graphics and async compute efficiency, where the RTX 4090’s 128 RT cores and 512 tensor cores provide hardware acceleration that the P40 lacks entirely, as it has no RT or tensor core equivalents listed.

Despite these lopsided wins, the aggregate average scores tell a different story. The RTX 4090’s average of 66,473 is only 0.5% higher than the P40’s 66,127. This occurs because the P40’s benchmark results are limited to just two Geekbench entries, while the RTX 4090’s average includes nine tests spanning DirectX 9 through 12, Passmark G2D/G3D, and compute workloads. The RTX 4090’s Passmark scores show a mixed profile: it achieves 38,194 in G3D and 26,613 in GPU compute, but its DirectX 12 score is a modest 150, and its DirectX 9 score is 397. The P40’s absence from these DirectX and Passmark tests means its average is calculated solely from its strong OpenCL and Vulkan numbers, which, while far lower in absolute terms, do not drag its average down with DX9-era legacy tests. Thus, the head-to-head data indicates that the RTX 4090 wins both direct comparisons decisively, but the aggregate percentile ranking places both cards at 91%, a statistical artifact of the P40’s limited test footprint.

Examining the nearest rivals for each card further contextualizes the scores. The RTX 4090 sits 0.4% below the Tesla T4’s average of 66,733 and 0.9% below the AMD Radeon Pro Vega 56’s 67,097. The Tesla P40, meanwhile, is 0.9% below the T4 and 1.4% below the Radeon Pro Vega 56. Both cards trail the Quadro P6000, which averages 67,320, with the RTX 4090 1.3% behind and the P40 1.8% behind. This clustering suggests that in the narrow set of benchmarks where all these cards are measured, the architectural differences between Ada Lovelace and Pascal are less relevant than the specific test composition. The RTX 4090’s raw compute advantages manifest only in the Geekbench OpenCL and Vulkan tests, which are the same tests where the P40 is competitive enough to maintain its percentile standing.

The Verdict

From the data, the choice between these two cards depends entirely on workload profile. For any task that leverages OpenCL or Vulkan compute, the RTX 4090 is the unequivocal pick, delivering 412.3% higher OpenCL performance and 285.3% higher Vulkan performance. Its 82.58 TFLOPS FP32 throughput, 1.01 TB/s memory bandwidth, and 24 GB of GDDR6X memory make it a formidable compute accelerator for modern workloads. The Tesla P40’s 11.76 TFLOPS FP32 and 347.1 GB/s bandwidth are a fraction of those figures, and its 24 GB of GDDR5 memory, while equal in capacity, operates at less than a third of the bandwidth. For users running legacy DirectX 9 or 10 workloads, the RTX 4090’s Passmark scores of 397 and 224, respectively, suggest it handles those APIs, but the P40 has no comparable data, making a direct judgment impossible.

The Tesla P40’s case rests on its narrow benchmark profile and its server-oriented design. It offers 24 GB of memory in a dual-slot, 250 W package with an 8-pin EPS connector, making it suitable for dense server deployments where power and space are constrained. Its 91st percentile ranking, matching the RTX 4090, indicates that in the limited tests where it competes, it is not embarrassingly far behind. However, the absence of RT cores, tensor cores, and any DirectX benchmark results means its utility for modern graphics or AI inference is severely limited. The data shows no scenario where the P40 outperforms the RTX 4090; it wins zero head-to-head benchmarks. Thus, for any user prioritizing compute performance, the RTX 4090 is the rational choice. For a datacenter operator needing a low-power, 24 GB inference card for older CUDA workloads, the P40 may suffice, but the performance data offers no evidence that it is superior in any measurable way.

Architecture Differences

The two GPUs represent entirely different architectural generations. The RTX 4090 is built on the AD102 chip using the Ada Lovelace architecture, fabricated on TSMC’s 5 nm process. It integrates 76,300 million transistors on a 609 mm² die, yielding a transistor density of 125.3 million per square millimeter. The Tesla P40 uses the GP102 chip on the Pascal architecture, produced on TSMC’s 16 nm process with 11,800 million transistors on a 471 mm² die, for a density of 25.1 million per square millimeter. This density difference — a 5x improvement — is the foundational enabler of the RTX 4090’s performance advantage.

The RTX 4090 features 16,384 shading units, 512 TMUs, 176 ROPs, 128 RT cores, and 512 tensor cores. The P40 has 3,840 shading units, 240 TMUs, and 96 ROPs, with no RT cores or tensor cores recorded. The RTX 4090 also supports DirectX 12 Ultimate (12_2), while the P40 is limited to DirectX 12 (12_1). Both cards support OpenGL 4.6 and Vulkan 1.4, but the RTX 4090’s hardware ray tracing and tensor acceleration are absent from the P40. The FP16 compute rates highlight this divide: the RTX 4090 achieves 82.58 TFLOPS at a 1:1 ratio with FP32, while the P40 delivers 183.7 GFLOPS at a 1:64 ratio, meaning its FP16 performance is a tiny fraction of its FP32 throughput. This makes the P40 wholly unsuitable for workloads requiring half-precision arithmetic, a common requirement in AI training and inference.

Specification Differences

The specification sheets reveal stark contrasts in nearly every measurable field. The RTX 4090 has a base clock of 2235 MHz and a boost clock of 2520 MHz, compared to the P40’s 1303 MHz base and 1531 MHz boost. Memory subsystems differ fundamentally: the RTX 4090 uses 24 GB of GDDR6X at 1313 MHz (21 Gbps effective) delivering 1.01 TB/s bandwidth, while the P40 uses 24 GB of GDDR5 at 1808 MHz (7.2 Gbps effective) for 347.1 GB/s. The RTX 4090’s pixel rate is 443.5 GPixel/s versus the P40’s 147.0 GPixel/s, and its texture rate is 1,290.2 GTexel/s versus 367.4 GTexel/s. Power requirements diverge as well: the RTX 4090 has a 450 W TDP with a 1x 16-pin connector and an 850 W suggested PSU, while the P40 has a 250 W TDP with an 8-pin EPS connector and a 600 W suggested PSU. The RTX 4090 is triple-slot and measures 304 mm by 137 mm by 61 mm, while the P40 is dual-slot and measures 267 mm by 111 mm with no width specified. The bus interfaces differ, with the RTX 4090 on PCIe 4.0 x16 and the P40 on PCIe 3.0 x16. Display outputs are another differentiator: the RTX 4090 offers 1x HDMI 2.1 and 3x DisplayPort 1.4a, while the P40 has no display outputs, reflecting its server-only design. The RTX 4090’s launch MSRP was 1,599 USD, while the P40’s was 5,699 USD, though no current pricing is provided.

FAQ

Q: Which GPU has higher overall average benchmark scores?

A: The NVIDIA GeForce RTX 4090 has an average benchmark score of 66,473, which is 0.5% higher than the Tesla P40’s 66,127.

Q: How much faster is the RTX 4090 in OpenCL compute?

A: The RTX 4090 scores 317,684 in Geekbench OpenCL versus the P40’s 62,017, a 412.3% advantage.

Q: Does the Tesla P40 support hardware ray tracing?

A: No, the P40’s specification lists no RT cores, whereas the RTX 4090 includes 128 RT cores.

Q: What is the memory bandwidth difference between the two cards?

A: The RTX 4090 offers 1.01 TB/s bandwidth with GDDR6X memory, while the P40 provides 347.1 GB/s with GDDR5 memory.

Q: Are both cards in the same performance percentile?

A: Yes, both the RTX 4090 and the Tesla P40 rank at the 91st percentile among all GPUs.

Q: Which card has a higher FP16 compute rate?

A: The RTX 4090 achieves 82.58 TFLOPS FP16, while the P40 delivers 183.7 GFLOPS FP16, a 1:64 ratio versus the RTX 4090’s 1:1 ratio.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4090
Tesla P40
Core Specs
Shading Units
16,384
3,840 -76.6%
Shaders
16,384
3,840 -76.6%
TMUs
512
240 -53.1%
ROPs
176
96 -45.5%
SM Count
128
30 -76.6%
Clocks
Base Clock
2235 MHz
1303 MHz
Boost Clock
2520 MHz
1531 MHz
Memory Clock
1313 MHz 21 Gbps effective
1808 MHz 7.2 Gbps effective
Memory
Memory Size
24 GB
24 GB
VRAM (MB)
24,576
24,576 0.0%
Memory Type
GDDR6X
GDDR5
Memory Bus
384 bit
384 bit
Bandwidth
1.01 TB/s
347.1 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SM)
L2 Cache
72 MB
3 MB
Performance
Pixel Rate
443.5 GPixel/s
147.0 GPixel/s
Texture Rate
1,290.2 GTexel/s
367.4 GTexel/s
FP32 (TFLOPS)
82.58 TFLOPS
11.76 TFLOPS
FP64 (TFLOPS)
1,290.2 GFLOPS (1:64)
367.4 GFLOPS (1:32)
FP16 (TFLOPS)
82.58 TFLOPS (1:1)
183.7 GFLOPS (1:64)
AI/RT
RT Cores
128
Tensor Cores
512
Power
TDP
450 W
250 W
TDP (W)
450
250 -44.4%
Suggested PSU
850 W
600 W
Power Connectors
1x 16-pin
8-pin EPS
Architecture
Architecture
Ada Lovelace
Pascal
GPU Name
AD102
GP102
Generation
GeForce 40
Tesla Pascal (Pxx)
Process Size
5 nm
16 nm
Transistors
76,300 million
11,800 million
Die Size
609 mm²
471 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
25.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Triple-slot
Dual-slot
Length
304 mm 12 inches
267 mm 10.5 inches
Height
137 mm 5.4 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 3.0 x16
Other
Launch Price
1,599 USD
5,699 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 30
Tesla Maxwell
Successor
GeForce 50
Tesla Volta
View GeForce RTX 4090 Details View Tesla P40 Details