NVIDIA Tesla P40 vs NVIDIA Tesla T4 Comparison
NVIDIA Tesla P40
Tesla T4
PERFORMANCE BENCHMARKS
Analysis: NVIDIA Tesla P40 vs NVIDIA Tesla T4
The NVIDIA Tesla T4 and NVIDIA Tesla P40 are both end-of-life server accelerators from NVIDIA, but they represent two very different architectural generations and design philosophies. This analysis compares the Turing-based T4 against the Pascal-based P40 using benchmark data and specification differences. The data shows a split decision: the T4 wins the Vulkan workload, while the P40 takes the OpenCL workload, with the overall average scores placing them within 2.5% of each other.
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA Tesla T4 has a higher average benchmark score of 66,733, compared to the NVIDIA Tesla P40’s 65,095. This places the T4 at the 90th percentile of all GPUs, while the P40 sits at the 89th percentile.
Q: How do the two cards compare in the Geekbench Vulkan test?
A: The Tesla T4 wins the Vulkan test with a score of 72,190, which is 5.9% higher than the Tesla P40’s score of 68,172.
Q: What about the Geekbench OpenCL test?
A: The Tesla P40 wins the OpenCL test with a score of 62,017, which is 1.2% higher than the Tesla T4’s score of 61,276.
Q: Which card has more memory and bandwidth?
A: The Tesla P40 has more memory at 24 GB of GDDR5, with a bus width of 384 bit and bandwidth of 347.1 GB/s. The Tesla T4 has 16 GB of GDDR6, a 256-bit bus, and 320.0 GB/s bandwidth.
Q: What are the power consumption differences?
A: The Tesla T4 has a 70 W TDP, while the Tesla P40 has a 250 W TDP. The T4 is single-slot and requires no power connectors, whereas the P40 is dual-slot and requires an 8-pin EPS connector.
Q: Which card has tensor and ray tracing cores?
A: The Tesla T4 has 320 tensor cores and 40 RT cores. The Tesla P40 has neither; its specifications list no RT cores and no tensor cores.
Architecture Differences
The architectural split between these two cards is fundamental. The Tesla T4 is built on the Turing architecture, using the TU104 chip, fabricated on a 12 nm process at TSMC. The Tesla P40 uses the older Pascal architecture, with the GP102 chip, on a 16 nm process, also from TSMC. This generational gap explains many of the performance and feature differences.
The transistor counts are close, with the T4 at 13,600 million and the P40 at 11,800 million, but the die sizes differ. The T4’s die is 545 mm², while the P40’s is 471 mm². This results in nearly identical transistor densities of 25.0M / mm² for the T4 and 25.1M / mm² for the P40. The T4 includes dedicated hardware for AI inference and ray tracing, with 320 tensor cores and 40 RT cores, features entirely absent from the Pascal-based P40.
Clock speeds tell a story of different design goals. The T4 has a low base clock of 585 MHz but boosts to 1590 MHz, while the P40 runs at a higher base of 1303 MHz and boosts to 1531 MHz. The memory subsystems also diverge: the T4 uses 10 Gbps effective GDDR6, while the P40 uses 7.2 Gbps effective GDDR5. The API support reflects the newer architecture, with the T4 supporting DirectX 12 Ultimate (12_2) and the P40 limited to DirectX 12 (12_1). Both support OpenGL 4.6 and Vulkan 1.4.
Head-to-Head Benchmarks
The benchmark results from the head-to-head testing show a clear split between the two workloads. In the Geekbench OpenCL test, the Tesla P40 edges out the T4 with a score of 62,017 against 61,276, a delta of 1.2% in favor of the P40. This is a narrow margin, suggesting that for raw compute tasks that scale with shading units and memory bandwidth, the older Pascal architecture remains competitive.
The Geekbench Vulkan test tells the opposite story. The Tesla T4 scores 72,190, which is 5.9% higher than the P40’s 68,172. This is a more substantial margin, indicating that the Turing architecture’s newer features, such as RT cores and improved scheduling, provide a measurable advantage in Vulkan workloads. The T4’s Vulkan score is also higher than its own OpenCL score by 17.8%, while the P40’s Vulkan score is only 9.9% higher than its OpenCL result.
When looking at the broader competitive landscape, the T4’s average score of 66,733 puts it 1.1% ahead of the AMD Radeon VII (66,004) and 2.5% ahead of the P40. The P40’s average score of 65,095 is 1.4% behind the Radeon VII and 1.4% ahead of the AMD Radeon Pro WX 9100 (64,212). Both cards are closely grouped, with the T4 also being 2.7% behind the AMD Radeon Instinct MI25 (68,562) and 3% behind the Intel Arc A770 (68,809). The P40 is 2% ahead of both the NVIDIA CMP 30HX (63,842) and the AMD Radeon RX 9060 XT LP (63,830).
Specification Differences
The most significant specification differences between the two cards are in memory, power, and compute features. The Tesla P40 offers more memory capacity at 24 GB compared to the T4’s 16 GB. The P40 also has a wider 384-bit memory bus versus the T4’s 256-bit bus, and higher peak bandwidth at 347.1 GB/s compared to 320.0 GB/s. However, the T4 uses newer GDDR6 memory, while the P40 uses GDDR5.
Compute resources differ considerably. The P40 has 3840 shading units, 240 TMUs, and 96 ROPs, while the T4 has 2560 shading units, 160 TMUs, and 64 ROPs. This gives the P40 a raw pixel rate of 147.0 GPixel/s and a texture rate of 367.4 GTexel/s, both higher than the T4’s 101.8 GPixel/s and 254.4 GTexel/s. The P40 also leads in FP32 throughput at 11.76 TFLOPS versus the T4’s 8.141 TFLOPS. The T4 counters with FP16 performance of 16.28 TFLOPS (2:1), while the P40’s FP16 is a meager 183.7 GFLOPS (1:64).
Power and physical dimensions are starkly different. The T4 has a 70 W TDP and is a single-slot card with no power connectors, measuring 168 mm in length. The P40 has a 250 W TDP, is dual-slot, requires an 8-pin EPS connector, and is 267 mm long with a height of 111 mm. The suggested PSU for the T4 is 250 W, while the P40 suggests a 600 W PSU. The release dates are four years apart, with the T4 launching on 2018-09-12 and the P40 on 2016-09-12. The P40 has a launch MSRP of 5,699 USD.
Where Each One Wins
The Tesla T4 wins in scenarios that benefit from its Turing architecture’s modern features. The 5.9% Vulkan advantage indicates better performance in APIs that leverage newer hardware capabilities, including ray tracing and tensor operations. The T4’s FP16 throughput of 16.28 TFLOPS is dramatically higher than the P40’s 183.7 GFLOPS, making the T4 the clear choice for workloads that use FP16 math, such as AI inference and certain machine learning tasks. Its 320 tensor cores and 40 RT cores provide dedicated hardware that the P40 completely lacks. The T4’s low 70 W TDP and single-slot design also make it suitable for dense server configurations where power and space are limited.
The Tesla P40 wins in scenarios that favor raw compute throughput and memory capacity. Its OpenCL score of 62,017 is 1.2% higher, and its FP32 performance of 11.76 TFLOPS is 44% higher than the T4’s 8.141 TFLOPS. The P40’s 24 GB memory capacity and 347.1 GB/s bandwidth exceed the T4’s 16 GB and 320.0 GB/s, making it better suited for large datasets that need to reside in GPU memory. Its higher pixel rate (147.0 GPixel/s) and texture rate (367.4 GTexel/s) suggest an advantage in tasks like traditional rasterization and image processing that rely on these fixed-function units.
The Verdict
The data does not declare a single overall winner, as the head-to-head benchmark results are split at one win each. The choice between these two cards depends entirely on the workload. For applications that are FP32-heavy, require large memory capacities, or rely on traditional graphics throughput, the Tesla P40 is the stronger option. Its 24 GB of memory and 11.76 TFLOPS FP32 performance are directly reflected in its OpenCL victory.
For applications that can leverage FP16 arithmetic, tensor cores, or ray tracing, the Tesla T4 is the only viable choice. The T4’s 16.28 TFLOPS FP16 performance and dedicated tensor and RT cores give it capabilities the P40 cannot match, and its 5.9% Vulkan win demonstrates the practical benefit of these features. The T4 also offers a massive efficiency advantage with its 70 W TDP versus the P40’s 250 W, making it the preferred option for power-constrained environments. The average benchmark scores favor the T4 by 2.5%, but the P40’s larger memory and FP32 compute make it a compelling option for specific professional workloads. Ultimately, the T4 is the forward-looking choice for modern compute tasks, while the P40 remains a strong contender for traditional, FP32-centric server applications.