NVIDIA GB10 vs NVIDIA Tesla P40 Comparison

NVIDIA
GEFORCE

NVIDIA GB10

CORE STATE GB20B
VRAM 128 GB
CLOCK SPEED 2418 MHz
TDP 140 W
BUS WIDTH 256 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

Tesla P40

CORE STATE GP102
VRAM 24 GB
CLOCK SPEED 1531 MHz
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_opencl
120,137
62,017
geekbench_vulkan
114,648
68,172

Analysis: NVIDIA GB10 vs NVIDIA Tesla P40

Head-to-Head Benchmarks

The benchmark data presents a clear, unambiguous picture in the direct comparison between the NVIDIA GB10 and the NVIDIA Tesla P40. Across the two recorded tests, the GB10 wins both, and the margins are substantial. In the geekbench_opencl test, the GB10 scores 120137 against the Tesla P40's 62017, a delta of 93.7%. In geekbench_vulkan, the GB10 records 114648 while the Tesla P40 manages 68172, a delta of 68.2%. These are not close contests; the GB10 nearly doubles the P40's OpenCL performance and sits far ahead in Vulkan as well.

Looking at the broader database context, the GB10's average benchmark score of 117393 places it in the 95th percentile of all GPUs. Its nearest rivals include the NVIDIA RTX 4000 SFF Ada Generation at 117088 (a delta of 0.3%, essentially a statistical tie), the AMD Radeon PRO W7700 at 118976 (the GB10 trails by 1.3%), the NVIDIA Tesla V100 SXM2 16 GB at 114395 (the GB10 leads by 2.6%), and the NVIDIA RTX A5500 Mobile at 113944 (a 3% lead). The data suggests the GB10 sits in a competitive performance band, trading blows with professional workstation cards but consistently edging out older data center accelerators.

The Tesla P40, by contrast, has an average benchmark score of 65095, landing in the 89th percentile of all GPUs. Its nearest rivals are the AMD Radeon Pro WX 9100 at 64212 (the P40 leads by 1.4%), the AMD Radeon VII at 66004 (the P40 trails by 1.4%), and the NVIDIA CMP 30HX and AMD Radeon RX 9060 XT LP, both at 63842 and 63830 respectively (the P40 leads by 2% in both cases). The P40's scores place it in a lower performance tier, closer to consumer and mining-oriented hardware than to modern professional compute silicon.

What the head-to-head deltas reveal is not just a generational gap, but a fundamental shift in capability. The 93.7% OpenCL advantage for the GB10 suggests that in compute-heavy workloads, the newer architecture is extracting far more from its hardware. The 68.2% Vulkan lead, while still decisive, is comparatively smaller, hinting that the P40's older Pascal architecture retains some relative strength in graphics-oriented tasks, though it remains thoroughly outclassed.

The Verdict

The recorded data is unambiguous: the NVIDIA GB10 is the superior performer in every measured benchmark. If the choice is purely about raw compute scores, the GB10 wins outright. The 120137 OpenCL score and 114648 Vulkan score are both roughly double or better than the P40's respective 62017 and 68172. For any workload that relies on OpenCL or Vulkan, the GB10 is the correct selection based on these measurements.

However, the data also indicates that these are different classes of hardware with different intended lifecycles. The Tesla P40 is marked end-of-life, while the GB10 is active production. The P40's release date in the database is 2016, whereas the GB10's is 2025. In a deployment scenario where longevity and current support matter, the GB10 is the only rational choice from the data available.

The P40 does retain some niche appeal. Its average score of 65095 places it in the 89th percentile, which is respectable for hardware of its era. If a system requires a dual-slot card with an 8-pin EPS connector and no display outputs, the P40 fits that physical profile. But the performance gap is too wide to recommend the P40 on any compute metric. The verdict from the data is simple: the GB10 wins on performance, the GB10 is the active product, and the GB10 should be chosen unless a specific legacy interface or form factor requirement forces the P40 into consideration.

Architecture Differences

The architectural divide between these two NVIDIA parts is vast, spanning two distinct GPU generations. The GB10 is built on the Blackwell 2.0 architecture, using the GB20B chip, fabricated on a 5 nm process at TSMC. The Tesla P40 uses the Pascal architecture with the GP102 chip, fabricated on a 16 nm process, also at TSMC. The process node difference alone, 5 nm versus 16 nm, explains a significant portion of the performance and efficiency gap.

The die sizes are revealing. The GB10 has a die size of 382 mm², while the P40's die is larger at 471 mm². Despite the smaller die, the GB10 delivers vastly more performance, which speaks to the density improvements of the newer node. The P40's transistor count is listed at 11,800 million, with a density of 25.1M per mm². The GB10's transistor count is unknown in the database, but given its die size and node, the density is presumably far higher.

Memory architecture also diverges sharply. The GB10 uses 128 GB of LPDDR5X on a 256-bit bus, delivering 273.2 GB/s of bandwidth. The P40 uses 24 GB of GDDR5 on a 384-bit bus, delivering 347.1 GB/s. The P40 actually has higher memory bandwidth, but the GB10 has over five times the memory capacity. This suggests different workload profiles: the P40 was designed for bandwidth-sensitive tasks, while the GB10 prioritizes capacity for large models or datasets.

The compute resources are also different in kind. The GB10 has 6144 shading units, 384 TMUs, 48 ROPs, 48 RT cores, and 384 tensor cores. The P40 has 3840 shading units, 240 TMUs, 96 ROPs, and no RT or tensor cores. The GB10's inclusion of RT and tensor cores is a fundamental architectural addition that the Pascal-based P40 lacks entirely. This means the GB10 can handle ray tracing and tensor-driven workloads, while the P40 is purely a raster and compute engine.

Specification Differences

The specification tables show several key differences between the two cards. The GB10 has a base clock of 1665 MHz and a boost clock of 2418 MHz, while the P40 runs at 1303 MHz base and 1531 MHz boost. The GB10's clocks are substantially higher, contributing to its performance lead. Memory clocks differ as well: the GB10's memory runs at 1067 MHz with 8.5 Gbps effective, while the P40's memory runs at 1808 MHz with 7.2 Gbps effective.

Pixel and texture rates show mixed results. The GB10 achieves 116.1 GPixel/s and 928.5 GTexel/s, while the P40 achieves 147.0 GPixel/s and 367.4 GTexel/s. The P40 has a higher pixel rate despite its older architecture, but the GB10 dominates in texture rate by a wide margin. The FP32 compute figures are 29.71 TFLOPS for the GB10 against 11.76 TFLOPS for the P40. The FP16 comparison is stark: the GB10 delivers 29.71 TFLOPS (1:1 ratio), while the P40 delivers only 183.7 GFLOPS (1:64 ratio), a massive difference in half-precision capability.

Power and physical specifications differ too. The GB10 has a TDP of 140 W, is an IGP form factor, uses no power connectors, and suggests a 300 W PSU. The P40 has a TDP of 250 W, is dual-slot, uses an 8-pin EPS connector, and suggests a 600 W PSU. The bus interfaces are PCIe 5.0 x16 for the GB10 and PCIe 3.0 x16 for the P40. The GB10 has one HDMI output, while the P40 has no display outputs. Dimensions are also different: the GB10 is 150 mm by 51 mm by 150 mm, while the P40 is 267 mm by 111 mm. The API support differs as well, with the P40 listing DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4, while the GB10 lists N/A for all three, indicating it is not intended for traditional graphics API workloads in the same way.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA GB10, with an average benchmark score of 117393, compared to the NVIDIA Tesla P40's 65095.

Q: How large is the performance gap in OpenCL?

A: The GB10 scores 120137 in geekbench_opencl, while the P40 scores 62017, giving the GB10 a 93.7% advantage.

Q: Does the Tesla P40 have any architectural features the GB10 lacks?

A: The P40 has a higher pixel rate (147.0 GPixel/s versus 116.1 GPixel/s) and higher memory bandwidth (347.1 GB/s versus 273.2 GB/s), but it lacks RT cores and tensor cores entirely, which the GB10 includes.

Q: What is the memory capacity difference?

A: The GB10 has 128 GB of LPDDR5X memory, while the P40 has 24 GB of GDDR5 memory.

Q: Which card is still in production?

A: The GB10 is marked as active production, while the P40 is end-of-life.

Q: How do they compare in power consumption?

A: The GB10 has a TDP of 140 W, while the P40 has a TDP of 250 W.

Where Each One Wins

The GB10 wins in nearly every measurable category. It wins both head-to-head benchmarks, has a higher average score, higher FP32 and FP16 compute, a smaller process node, newer architecture, more memory capacity, tensor and RT cores, higher clocks, and a more modern PCIe interface. Its wins are decisive in compute-heavy workloads, particularly those that leverage tensor cores or FP16 precision, where its 29.71 TFLOPS dwarfs the P40's 183.7 GFLOPS. For AI inference, large model loading, or any modern compute task, the GB10 is the clear winner based on the recorded data.

The P40 does have specific wins in a few narrow categories. Its pixel rate of 147.0 GPixel/s is higher than the GB10's 116.1 GPixel/s, which could matter in fill-rate-limited rasterization tasks. Its memory bandwidth of 347.1 GB/s exceeds the GB10's 273.2 GB/s, potentially benefiting workloads that are highly bandwidth-sensitive but do not need large capacity. It also supports traditional graphics APIs like DirectX 12, OpenGL 4.6, and Vulkan 1.4, while the GB10 lists N/A for these, suggesting the P40 retains compatibility with legacy graphics software. The P40's dual-slot form factor and 8-pin EPS connector may suit systems that expect that physical configuration.

In practical terms, the data suggests the P40's wins are narrow and situational. A user with an existing PCIe 3.0 system that requires a dual-slot card with no display outputs and needs maximum pixel fill rate might find the P40 acceptable. But the performance gap is so wide in the benchmarks that the GB10 should be the default choice for any new deployment. The GB10 wins on compute, wins on capacity, wins on efficiency, and is the only one of the two still in active production. The P40's remaining strengths are legacy features and specific bandwidth or pixel-rate scenarios, which do not compensate for its massive compute deficit in the recorded tests.

DETAILED SPECIFICATIONS

SPECIFICATION
GB10
Tesla P40
Core Specs
Shading Units
6,144
3,840 -37.5%
Shaders
6,144
3,840 -37.5%
TMUs
384
240 -37.5%
ROPs
48
96 +100.0%
SM Count
48
30 -37.5%
Clocks
Base Clock
1665 MHz
1303 MHz
Boost Clock
2418 MHz
1531 MHz
Memory Clock
1067 MHz 8.5 Gbps effective
1808 MHz 7.2 Gbps effective
Memory
Memory Size
128 GB
24 GB
VRAM (MB)
131,072
24,576 -81.3%
Memory Type
LPDDR5X
GDDR5
Memory Bus
256 bit
384 bit
Bandwidth
273.2 GB/s
347.1 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SM)
L2 Cache
50 MB
3 MB
Performance
Pixel Rate
116.1 GPixel/s
147.0 GPixel/s
Texture Rate
928.5 GTexel/s
367.4 GTexel/s
FP32 (TFLOPS)
29.71 TFLOPS
11.76 TFLOPS
FP64 (TFLOPS)
464.3 GFLOPS (1:64)
367.4 GFLOPS (1:32)
FP16 (TFLOPS)
29.71 TFLOPS (1:1)
183.7 GFLOPS (1:64)
AI/RT
RT Cores
48
Tensor Cores
384
Power
TDP
140 W
250 W
TDP (W)
140
250 +78.6%
Suggested PSU
300 W
600 W
Power Connectors
None
8-pin EPS
Architecture
Architecture
Blackwell 2.0
Pascal
GPU Name
GB20B
GP102
Generation
Server Blackwell (Bxx)
Tesla Pascal (Pxx)
Process Size
5 nm
16 nm
Transistors
unknown
11,800 million
Die Size
382 mm²
471 mm²
Foundry
TSMC
TSMC
Density
25.1M / mm²
API Support
DirectX
12 (12_1)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
12.1
6.1
Shader Model
6.8
Physical
Slot Width
IGP
Dual-slot
Length
150 mm 5.9 inches
267 mm 10.5 inches
Height
51 mm 2 inches
111 mm 4.4 inches
Outputs
1x HDMI
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 3.0 x16
Other
Launch Price
3,999 USD
5,699 USD
Production
Active
End-of-life
Predecessor
Server Hopper
Tesla Maxwell
Successor
Server Rubin
Tesla Volta
View GB10 Details View Tesla P40 Details