NVIDIA Quadro P6000 vs NVIDIA Tesla P40 Comparison
NVIDIA Quadro P6000
Tesla P40
PERFORMANCE BENCHMARKS
Analysis: NVIDIA Quadro P6000 vs NVIDIA Tesla P40
NVIDIA’s Pascal-generation GP102 silicon powers two very different cards in the Quadro P6000 and Tesla P40. Both are end-of-life, dual-slot, 250 W boards with 24 GB of memory, but the benchmark data shows the Quadro P6000 is consistently faster in compute workloads. In the two head-to-head tests recorded, the Quadro P6000 wins both, with a 7% lead in Geekbench OpenCL (66,382 vs. 62,017) and a 7.9% lead in Geekbench Vulkan (73,590 vs. 68,172). The average benchmark score reinforces this gap: the Quadro P6000 sits at 69,986, which is 7.5% above the Tesla P40’s 65,095. This is not a marginal difference — it is a clear, repeatable performance tier separation between the two.
Head-to-Head Benchmarks
The most decisive victory for the Quadro P6000 comes in the Geekbench Vulkan test, where it scores 73,590 against the Tesla P40’s 68,172. That 7.9% delta is the largest single-test margin in this comparison. Vulkan is a low-overhead API that scales well with raw shader throughput, and the Quadro’s higher clock speeds translate directly into a measurable advantage here. The OpenCL result tells a similar story, albeit slightly narrower: the Quadro P6000 posts 66,382 versus 62,017 for the Tesla P40, a 7% edge. OpenCL workloads often stress memory bandwidth alongside compute, and while both cards share a 384-bit bus, the Quadro’s faster memory subsystem gives it an additional edge that shows up in the aggregate score.
Looking at the broader context, the Quadro P6000’s average score of 69,986 places it in the 90th percentile of all GPUs, while the Tesla P40’s 65,095 lands in the 89th percentile. That one-percentile gap may sound small, but the rival comparison tables reveal how tightly clustered this performance band is. The Quadro P6000’s nearest rival is the NVIDIA RTX A3000 Mobile at 70,140, which is just 0.2% faster, and the AMD Radeon Pro WX 8200 at 69,870, which is 0.2% slower. This means the Quadro P6000 is essentially trading blows with much newer mobile and workstation parts. The Tesla P40, by contrast, sits in a slightly lower tier: its closest competitor is the AMD Radeon VII at 66,004, which is 1.4% faster, and the AMD Radeon Pro WX 9100 at 64,212, which is 1.4% slower. The Tesla P40 is competitive within its class, but it is not matching the Quadro’s ability to hang with the next generation.
The per-test deltas are consistent, which suggests the performance gap is rooted in clock speeds rather than architectural differences. The Quadro P6000 runs a 1506 MHz base clock and 1645 MHz boost, whereas the Tesla P40 is clocked at 1303 MHz base and 1531 MHz boost. That is a 203 MHz difference at base and 114 MHz at boost — roughly 13% and 7% respectively, which aligns almost perfectly with the observed 7-8% benchmark deltas. The Quadro’s higher clocks are not a gimmick; they translate into real-world compute wins across both OpenCL and Vulkan, and they explain why the Quadro maintains a lead even though both cards share the same GP102 chip, 3840 shading units, 240 TMUs, and 96 ROPs.
The Verdict
The data is unambiguous: the NVIDIA Quadro P6000 is the faster card in every measured benchmark. If the decision is purely about compute performance, the Quadro P6000 wins on both Geekbench OpenCL and Geekbench Vulkan, with a 7% and 7.9% advantage respectively. Its average benchmark score of 69,986 is 7.5% higher than the Tesla P40’s 65,095, and it sits in a higher percentile (90th vs. 89th) of all GPUs. For users running OpenCL or Vulkan workloads, the Quadro P6000 delivers more throughput per frame or per kernel, and it does so without any trade-off in power draw — both cards are rated at 250 W TDP with a 600 W suggested PSU.
However, the Tesla P40 is not without its own rationale. It is a compute-focused card with no display outputs, which makes it suitable for headless server deployments where video output is irrelevant. The Quadro P6000, by contrast, offers 1x DVI and 4x DisplayPort 1.4a outputs, making it a workstation card that can drive displays while also accelerating compute. If the workload is purely server-side inference or batch processing, the Tesla P40’s lack of display outputs is not a disadvantage — it is a design choice that frees up resources for compute. But the benchmark data does not show any compensating performance benefit; the Tesla P40 is simply slower in both tests.
The launch MSRP reflects the positioning: the Quadro P6000 launched at 5,999 USD, while the Tesla P40 launched at 5,699 USD. That 300 USD difference at launch is the only price signal in the data, and it aligns with the performance gap — you pay more for the Quadro, and you get more compute per dollar in the benchmarks. For a buyer choosing between these two end-of-life cards, the Quadro P6000 is the pick if you need display outputs or maximum compute throughput. The Tesla P40 is the pick only if you require its specific 8-pin EPS power connector or have a strict no-display-output requirement for security or form-factor reasons. In pure performance terms, the Quadro P6000 wins 2-0, and the data suggests that lead will persist across most compute workloads.
FAQ
Q: Which card is faster in Geekbench OpenCL?
A: The NVIDIA Quadro P6000 wins with a score of 66,382, which is 7% higher than the Tesla P40’s 62,017.
Q: How big is the Vulkan performance gap?
A: The Quadro P6000 scores 73,590 in Geekbench Vulkan, while the Tesla P40 scores 68,172. The Quadro leads by 7.9%, which is the largest margin in any head-to-head test between the two.
Q: Do both cards have the same memory capacity?
A: Yes, both the Quadro P6000 and Tesla P40 feature 24 GB of memory. However, the Quadro uses GDDR5X with 432.8 GB/s bandwidth, while the Tesla uses GDDR5 with 347.1 GB/s bandwidth.
Q: Are there any benchmark tests where the Tesla P40 wins?
A: No. In the recorded head-to-head benchmarks, the Quadro P6000 wins both tests (OpenCL and Vulkan). The Tesla P40 has zero wins in the comparison.
Q: What is the average benchmark score for each card?
A: The Quadro P6000 has an average benchmark score of 69,986, placing it in the 90th percentile of all GPUs. The Tesla P40 averages 65,095, placing it in the 89th percentile.
Q: Do these cards support the same graphics APIs?
A: Yes, both support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. There is no difference in API support between the two.
Specification Differences
The most consequential difference is memory type and bandwidth. The Quadro P6000 uses GDDR5X memory with an effective speed of 9 Gbps, delivering 432.8 GB/s across a 384-bit bus. The Tesla P40 uses standard GDDR5 at 7.2 Gbps effective, yielding 347.1 GB/s on the same 384-bit bus. That is a 85.7 GB/s bandwidth deficit for the Tesla, which is roughly 20% lower — a significant gap that directly impacts memory-bound compute tasks.
Clock speeds also differ substantially. The Quadro P6000 runs at 1506 MHz base and 1645 MHz boost, while the Tesla P40 is clocked at 1303 MHz base and 1531 MHz boost. This translates into higher fill rates and compute throughput for the Quadro: its pixel rate is 157.9 GPixel/s versus 147.0 GPixel/s for the Tesla, and its texture rate is 394.8 GTexel/s versus 367.4 GTexel/s. FP32 performance follows the same pattern, with the Quadro hitting 12.63 TFLOPS and the Tesla at 11.76 TFLOPS. FP16 is also higher on the Quadro at 197.4 GFLOPS (1:64) versus 183.7 GFLOPS (1:64).
Power delivery is another differentiator. The Quadro P6000 uses a single 8-pin PCIe power connector, while the Tesla P40 uses an 8-pin EPS connector — a server-style plug that may require an adapter in standard workstation builds. Both cards share the same 250 W TDP and 600 W suggested PSU, and both are dual-slot with identical physical dimensions (267 mm length, 111 mm height). Display outputs are the final major split: the Quadro P6000 has 1x DVI and 4x DisplayPort 1.4a, while the Tesla P40 has no display outputs at all, making it a headless compute accelerator.
Architecture Differences
Both cards are built on the same GP102 chip using NVIDIA’s Pascal architecture, fabricated by TSMC on a 16 nm process. They share identical transistor counts (11,800 million), die size (471 mm²), and transistor density (25.1M / mm²). The core configuration is also identical: 3840 shading units, 240 texture mapping units, and 96 raster output pipelines. Neither card has dedicated ray tracing cores or tensor cores, which is consistent with the Pascal generation predating NVIDIA’s RTX hardware.
The architectural differences are limited to clock behavior and memory configuration, but the memory subsystem is where the cards diverge most meaningfully. The Quadro P6000’s GDDR5X memory runs at 1127 MHz (9 Gbps effective), while the Tesla P40’s GDDR5 runs at 1808 MHz (7.2 Gbps effective). Although the Tesla’s memory clock is numerically higher, GDDR5X carries twice the data rate per clock compared to GDDR5, which is why the Quadro achieves higher effective bandwidth despite a lower MHz figure. This is a fundamental architectural advantage for the Quadro in bandwidth-sensitive workloads.
The generation lineage also differs. The Quadro P6000 belongs to the Quadro Pascal (Px000) family, with its predecessor being Quadro Maxwell and its successor Quadro Volta. The Tesla P40 belongs to the Tesla Pascal (Pxx) family, with Tesla Maxwell as its predecessor and Tesla Volta as its successor. Both were released in September 2016, with the Tesla P40 arriving on September 12 and the Quadro P6000 on September 30 — an 18-day gap that likely reflects their different target markets. The Quadro is a workstation visualization and compute hybrid, while the Tesla is a pure compute accelerator for data centers. Despite those different missions, the silicon underneath is the same, and the Quadro’s higher clocks and faster memory make it the superior performer in every benchmark recorded.