NVIDIA Quadro GP100 vs NVIDIA Tesla P40 Comparison
NVIDIA Quadro GP100
Tesla P40
PERFORMANCE BENCHMARKS
Analysis: NVIDIA Quadro GP100 vs NVIDIA Tesla P40
The Verdict
The NVIDIA Quadro GP100 and Tesla P40 are both end-of-life Pascal workstation cards, but they serve fundamentally different purposes. The data shows the Quadro GP100 is the clear winner in raw compute performance, scoring 87,445 in Geekbench OpenCL compared to the Tesla P40's 62,017, a 41% advantage. The Quadro GP100 also holds a higher overall percentile rank, sitting at 93 versus the Tesla P40's 89.
Pick the Quadro GP100 if your priority is maximum compute throughput in professional applications, especially those leveraging OpenCL. Its 41% lead in the recorded benchmark is decisive. The Quadro GP100 also has display outputs, making it usable in a workstation with monitors attached.
Pick the Tesla P40 only if you specifically need 24 GB of memory capacity and can live with significantly lower compute performance. It is a compute-only card with no display outputs, designed for server environments. Its Vulkan score of 68,172 is notable, but its OpenCL score lags far behind the Quadro GP100. The Tesla P40 does have a launch MSRP of 5,699 USD, but that does not change its performance profile.
For the vast majority of professional workloads measured in this database, the Quadro GP100 is the superior choice. The Tesla P40's advantage is confined to its larger memory pool, not its speed.
Architecture Differences
Both cards are built on NVIDIA's Pascal architecture and use a 16 nm process at TSMC, but they use different chips. The Quadro GP100 is built on the GP100 chip, while the Tesla P40 uses the GP102 chip. This is the core architectural divergence that drives their performance characteristics.
The GP100 chip is physically larger, with a die size of 610 mm² and 15,300 million transistors. The GP102 is smaller at 471 mm² with 11,800 million transistors. Interestingly, both have the same transistor density of 25.1M per mm², reflecting the identical process node.
Memory architecture is where the two diverge most sharply. The Quadro GP100 uses 16 GB of HBM2 on a 4096-bit bus, delivering 732.2 GB/s of bandwidth. The Tesla P40 uses 24 GB of GDDR5 on a 384-bit bus, delivering 347.1 GB/s. That is a massive difference: the Quadro GP100 has more than double the memory bandwidth despite having less capacity.
The compute unit counts also differ. The Tesla P40 has more shading units (3,840 versus 3,584), more texture mapping units (240 versus 224), and the same number of ROPs (96). However, the Quadro GP100's HBM2 memory and wider bus more than compensate in the recorded benchmarks.
Clock speeds are close on paper. The Quadro GP100 runs at 1304 MHz base and 1443 MHz boost, while the Tesla P40 runs at 1303 MHz base and 1531 MHz boost. The Tesla P40 has a higher boost clock, but the Quadro GP100 still wins decisively in OpenCL.
Neither card has ray tracing cores or tensor cores. Both support DirectX 12 (12_1) and OpenGL 4.6. The Tesla P40 supports Vulkan 1.4, while the Quadro GP100 supports Vulkan 1.3.
Where Each One Wins
The recorded data gives the Quadro GP100 one win in the head-to-head benchmarks: Geekbench OpenCL. The Tesla P40 has zero wins in the same comparison. However, the database does include a Vulkan score for the Tesla P40 (68,172), which is not available for the Quadro GP100 in the records.
The Quadro GP100 wins in raw compute performance. Its OpenCL score of 87,445 places it in the 93rd percentile of all GPUs. It sits 0.4% ahead of the AMD Radeon PRO W7600 (87,108) and 2.1% ahead of the NVIDIA CMP 40HX (85,637). It trails the NVIDIA RTX A4500 Mobile (91,134) by 4% and the NVIDIA RTX A4500 (91,671) by 4.6%. This is a strong mid-to-high tier compute performer.
The Tesla P40 wins in memory capacity. With 24 GB of GDDR5, it offers 50% more memory than the Quadro GP100's 16 GB. This matters for workloads that need to hold large datasets or models in memory, even if the compute throughput is lower. Its Vulkan score of 68,172 also shows some API-specific strength. In the overall database, the Tesla P40 sits at the 89th percentile, close to the AMD Radeon Pro WX 9100 (64,212, 1.4% behind) and the AMD Radeon VII (66,004, 1.4% ahead). It is 2% ahead of both the NVIDIA CMP 30HX (63,842) and the AMD Radeon RX 9060 XT LP (63,830).
For practical use, the Quadro GP100 is the choice for compute-heavy tasks where memory bandwidth matters more than capacity. The Tesla P40 is the choice for tasks that are memory-capacity bound, where 24 GB is required and the slower bandwidth is acceptable.
FAQ
Q: Which card is faster in OpenCL?
A: The Quadro GP100 is 41% faster, scoring 87,445 versus the Tesla P40's 62,017 in Geekbench OpenCL.
Q: Does the Tesla P40 have any benchmark where it wins?
A: In the head-to-head comparison, the Tesla P40 has zero wins. It does have a recorded Geekbench Vulkan score of 68,172, but no comparable Vulkan score exists for the Quadro GP100 in the database.
Q: Which card has more memory?
A: The Tesla P40 has 24 GB of GDDR5 memory, while the Quadro GP100 has 16 GB of HBM2. However, the Quadro GP100 has much higher bandwidth at 732.2 GB/s versus 347.1 GB/s.
Q: Can I use either card as a display adapter?
A: No, only the Quadro GP100 has display outputs, which are 1x DVI and 4x DisplayPort 1.4a. The Tesla P40 has no display outputs and is compute-only.
Q: What is the power requirement difference?
A: The Quadro GP100 has a TDP of 235 W with a suggested PSU of 550 W and uses a single 8-pin power connector. The Tesla P40 has a TDP of 250 W with a suggested PSU of 600 W and uses an 8-pin EPS connector.
Q: Which card is better for the price?
A: The database does not include a launch MSRP for the Quadro GP100. The Tesla P40 has a launch MSRP of 5,699 USD. Performance data shows the Quadro GP100 is significantly faster, but no price comparison can be made without both values.
Head-to-Head Benchmarks
The single recorded head-to-head benchmark is Geekbench OpenCL, and it is a decisive victory for the Quadro GP100.
The Quadro GP100 scored 87,445, while the Tesla P40 scored 62,017. This is a 41% delta in favor of the Quadro GP100. To put that in perspective, the Quadro GP100 sits 0.4% above the AMD Radeon PRO W7600 and 2.1% above the NVIDIA CMP 40HX. The Tesla P40, by contrast, sits just 1.4% above the AMD Radeon Pro WX 9100 and 2% above the NVIDIA CMP 30HX. The gap between the two cards is roughly the same as the gap between the Tesla P40 and much lower-tier cards.
This massive OpenCL delta is likely explained by the memory architecture. The Quadro GP100's HBM2 memory with a 4096-bit bus delivers 732.2 GB/s, while the Tesla P40's GDDR5 on a 384-bit bus delivers only 347.1 GB/s. More than double the bandwidth means compute kernels that are memory-bound will finish much faster on the Quadro GP100.
The Tesla P40 does have a Vulkan score of 68,172, which is higher than its OpenCL score. This suggests the card performs better under the Vulkan API than under OpenCL. However, without a comparable Vulkan score for the Quadro GP100, it is impossible to determine which card wins in that API.
The overall average benchmark score tells the same story. The Quadro GP100's average is 87,445, placing it at the 93rd percentile. The Tesla P40's average is 65,095, placing it at the 89th percentile. The 22,350-point average gap is substantial.
Specification Differences
The two cards share the same architecture family and process node, but nearly every other specification differs.
Chip and Die: The Quadro GP100 uses the GP100 chip with a 610 mm² die and 15,300 million transistors. The Tesla P40 uses the GP102 chip with a 471 mm² die and 11,800 million transistors.
Clocks: The Quadro GP100 runs at 1304 MHz base and 1443 MHz boost. The Tesla P40 runs at 1303 MHz base and 1531 MHz boost. The Tesla P40 has a higher boost clock by 88 MHz.
Memory: The Quadro GP100 has 16 GB of HBM2 on a 4096-bit bus with 732.2 GB/s bandwidth. The Tesla P40 has 24 GB of GDDR5 on a 384-bit bus with 347.1 GB/s bandwidth.
Compute Units: The Quadro GP100 has 3,584 shading units, 224 TMUs, and 96 ROPs. The Tesla P40 has 3,840 shading units, 240 TMUs, and 96 ROPs.
Pixel and Texture Rates: The Quadro GP100 delivers 138.5 GPixel/s and 323.2 GTexel/s. The Tesla P40 delivers 147.0 GPixel/s and 367.4 GTexel/s. The Tesla P40 is faster here due to its higher clock and more units.
FP32 and FP16: The Quadro GP100 delivers 10.34 TFLOPS FP32 and 20.69 TFLOPS FP16 (2:1). The Tesla P40 delivers 11.76 TFLOPS FP32 and only 183.7 GFLOPS FP16 (1:64). The Tesla P40 is faster in FP32, but the Quadro GP100 crushes it in FP16.
Power: The Quadro GP100 has a 235 W TDP and a 550 W suggested PSU. The Tesla P40 has a 250 W TDP and a 600 W suggested PSU.
Power Connectors: The Quadro GP100 uses a single 8-pin connector. The Tesla P40 uses an 8-pin EPS connector.
Display Outputs: The Quadro GP100 has 1x DVI and 4x DisplayPort 1.4a. The Tesla P40 has no outputs.
Vulkan API: The Tesla P40 supports Vulkan 1.4, while the Quadro GP100 supports Vulkan 1.3.
Release Date: The Tesla P40 was released on 2016-09-12, and the Quadro GP100 was released on 2016-09-30. Both are end-of-life.
Dimensions: Both cards are 267 mm long and 111 mm tall, dual-slot designs.