NVIDIA P106-100 vs NVIDIA Tesla K40m Comparison
NVIDIA P106-100
Tesla K40m
PERFORMANCE BENCHMARKS
Analysis: NVIDIA P106-100 vs NVIDIA Tesla K40m
Head-to-Head Benchmarks
The recorded data includes only one direct comparison between the NVIDIA P106-100 and the NVIDIA Tesla K40m, and it is a decisive victory for the newer card. In the Geekbench OpenCL test, the P106-100 scores 35,951 points against the Tesla K40m's 19,885 points. That is a delta of 80.8%, meaning the P106-100 delivers roughly 81% higher performance in this compute-oriented workload. This is not a marginal win; it is a generational leap in raw throughput for general-purpose GPU compute.
Looking at the broader benchmark database, the P106-100 also has results in two additional tests that the Tesla K40m does not appear in. In 3DMark Steel Nomad (DX12), the P106-100 scores 899 points. In Geekbench Vulkan, it scores 32,897 points. The Tesla K40m has no recorded scores for either of these tests, so no direct comparison is possible there. However, the average benchmark score across all recorded tests tells a similar story: the P106-100 averages 23,249 points, while the Tesla K40m averages 19,885 points. That puts the P106-100 about 16.9% ahead in average aggregate performance.
The delta between the two cards is consistent with what the architecture shift would suggest. The P106-100 is built on a much newer design with higher clocks and a more efficient process. The Tesla K40m compensates somewhat with more shading units and more memory bandwidth, but in the one workload where both were measured, it simply cannot keep pace. The 80.8% lead in OpenCL is the headline number for this matchup.
FAQ
Q: Which card has the higher average benchmark score?
A: The NVIDIA P106-100 has an average benchmark score of 23,249 points, compared to the Tesla K40m's 19,885 points. The P106-100 sits at the 68th percentile of all GPUs in the database, while the Tesla K40m sits at the 65th percentile.
Q: How much faster is the P106-100 in OpenCL compute?
A: In the Geekbench OpenCL test, the P106-100 scores 35,951 against the Tesla K40m's 19,885, a lead of 80.8%.
Q: Does the Tesla K40m have any benchmark wins over the P106-100?
A: No. In the head-to-head benchmark data, the P106-100 wins the only shared test (Geekbench OpenCL). The wins tally is 1 for the P106-100 and 0 for the Tesla K40m.
Q: What is the memory configuration difference?
A: The Tesla K40m has 12 GB of GDDR5 memory on a 384-bit bus, yielding 288.4 GB/s of bandwidth. The P106-100 has 6 GB of GDDR5 on a 192-bit bus, yielding 192.2 GB/s of bandwidth.
Q: Which card has better API support?
A: The P106-100 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The Tesla K40m supports DirectX 12 (11_1), OpenGL 4.6, and Vulkan 1.2.175. The P106-100 has the higher DirectX feature level and a newer Vulkan version.
Q: What is the launch MSRP of the Tesla K40m?
A: The Tesla K40m had a launch MSRP of 7,699 USD. The P106-100 has no recorded launch MSRP in the database.
Architecture Differences
The two cards come from completely different NVIDIA architectures, and the data reflects that split. The P106-100 is built on the Pascal architecture, using the GP106 chip on a 16 nm process from TSMC. It packs 4,400 million transistors into a 200 mm² die, giving a transistor density of 22.0 million per mm². The Tesla K40m is a Kepler architecture part, using the GK110B chip on a 28 nm process, also from TSMC. It holds 7,080 million transistors on a much larger 561 mm² die, with a lower transistor density of 12.6 million per mm².
Clock speeds tell the rest of the story. The P106-100 runs at a base clock of 1506 MHz and boosts to 1709 MHz. The Tesla K40m runs at just 745 MHz base and 876 MHz boost. That is a massive clock advantage for the Pascal card, which is why it can outperform the Kepler card despite having fewer than half the shading units. The P106-100 has 1,280 shading units, 80 texture mapping units, and 48 ROPs. The Tesla K40m has 2,880 shading units, 240 TMUs, and 48 ROPs. So the Kepler card has more raw execution resources, but the Pascal card's much higher clocks and architectural efficiency win out in practice.
The compute feature sets also differ. The P106-100 achieves 4.375 TFLOPS of FP32 performance and has a token FP16 rate of 68.36 GFLOPS (1:64). The Tesla K40m achieves 5.046 TFLOPS of FP32 and has no recorded FP16 capability. Despite the Tesla K40m's higher theoretical FP32 peak, the P106-100 still wins the real-world OpenCL test by a wide margin.
Both cards are compute-focused in that they have no display outputs. The P106-100 is classified in the database as part of the "Mining GPUs" generation, while the Tesla K40m belongs to the "Tesla Kepler (Kxx)" generation. Neither card has ray tracing cores or tensor cores. The Tesla K40m lists a predecessor of Tesla Fermi and a successor of Tesla Maxwell, while the P106-100 has no predecessor or successor listed.
Specification Differences
The two cards differ across nearly every major specification category. The process node is 16 nm for the P106-100 versus 28 nm for the Tesla K40m. Transistor count is 4,400 million versus 7,080 million. Die size is 200 mm² versus 561 mm². Transistor density is 22.0 million per mm² versus 12.6 million per mm².
Clocks: the P106-100 runs at 1506 MHz base and 1709 MHz boost. The Tesla K40m runs at 745 MHz base and 876 MHz boost. Memory clocks are 2002 MHz (8 Gbps effective) for the P106-100 versus 1502 MHz (6 Gbps effective) for the Tesla K40m.
Memory capacity is 6 GB versus 12 GB. Bus width is 192 bit versus 384 bit. Memory bandwidth is 192.2 GB/s versus 288.4 GB/s. The Tesla K40m has the clear advantage in memory capacity and bandwidth, which matters for workloads that are memory-bound, but it does not translate into a win in the recorded OpenCL test.
Shading units are 1,280 versus 2,880. TMUs are 80 versus 240. ROPs are identical at 48 for both. Pixel rate is 82.03 GPixel/s for the P106-100 versus 52.56 GPixel/s for the Tesla K40m. Texture rate is 136.7 GTexel/s versus 210.2 GTexel/s. FP32 is 4.375 TFLOPS versus 5.046 TFLOPS.
Power and physical specs differ as well. The P106-100 has a TDP of 120 W and requires a 300 W suggested PSU with a single 6-pin power connector. The Tesla K40m has a TDP of 245 W and requires a 550 W suggested PSU, with no power connector details listed. Both are dual-slot cards. The P106-100 is 250 mm (9.8 inches) long, while the Tesla K40m is 267 mm (10.5 inches) long.
The bus interface differs: the P106-100 uses PCIe 1.0 x16, while the Tesla K40m uses PCIe 3.0 x16. The API support differs in DirectX version (12_1 versus 11_1) and Vulkan version (1.4 versus 1.2.175). The release dates are far apart: the P106-100 launched in June 2017, the Tesla K40m in November 2013. Both are marked end-of-life in the database.
Where Each One Wins
The P106-100 wins where it matters most in this comparison: raw compute performance in the recorded benchmark. Its 80.8% lead in Geekbench OpenCL makes it the clear choice for general compute workloads that rely on OpenCL. It also holds a higher average benchmark score (23,249 versus 19,885), a higher percentile ranking (68th versus 65th), and additional benchmark results in 3DMark Steel Nomad and Geekbench Vulkan that the Tesla K40m simply does not have recorded. The Pascal architecture brings much higher clocks and a modern feature set, including DirectX 12 (12_1) and Vulkan 1.4 support. For anyone running compute tasks that can use these APIs, the P106-100 is the stronger card.
The Tesla K40m's advantages are in memory and raw throughput specifications. It has 12 GB of VRAM versus 6 GB, a 384-bit bus versus 192-bit, and 288.4 GB/s of bandwidth versus 192.2 GB/s. It also has more shading units (2,880 versus 1,280), more TMUs (240 versus 80), and a higher theoretical FP32 peak (5.046 TFLOPS versus 4.375 TFLOPS). These specs suggest the Tesla K40m could be preferable for workloads that are heavily memory-bound or that can fully utilize its massive shading unit count, if such workloads do not map well to the P106-100's architecture. Additionally, the Tesla K40m uses PCIe 3.0 x16, which is a more modern bus interface than the P106-100's PCIe 1.0 x16, a potential advantage for data transfer in some systems.
For pixel rate, the P106-100 wins at 82.03 GPixel/s versus 52.56 GPixel/s. For texture rate, the Tesla K40m wins at 210.2 GTexel/s versus 136.7 GTexel/s. So the split is not uniform: the P106-100 dominates in pixel throughput and compute tests, while the Tesla K40m leads in texture throughput and memory bandwidth.
The Verdict
The data points to the NVIDIA P106-100 as the better card for most purposes. Its 80.8% lead in the only shared benchmark, its higher average score, and its higher percentile ranking all favor the Pascal part. The P106-100 also has a lower TDP at 120 W versus 245 W, a smaller physical footprint, and a much later release date. For compute workloads that use OpenCL, Vulkan, or DirectX 12, the P106-100 is the sensible pick.
The Tesla K40m is not without merit. It offers double the VRAM (12 GB versus 6 GB), significantly more memory bandwidth (288.4 GB/s versus 192.2 GB/s), and a higher theoretical FP32 peak. Those specs could matter for specific memory-heavy tasks or for users who need the larger frame buffer. It also has PCIe 3.0 x16 support, which the P106-100 lacks. But none of these advantages show up in the recorded benchmark results, where the Tesla K40m loses the only direct comparison by a large margin. Its launch MSRP of 7,699 USD also reflects its enterprise positioning, though price is not part of this analysis.
For a builder choosing between these two end-of-life cards, the P106-100 is the default recommendation. It wins the measured performance tests, runs cooler, and supports newer APIs. The Tesla K40m should only be considered if the workload explicitly requires more than 6 GB of VRAM or if the memory bandwidth advantage is critical, since the benchmark data does not show it converting those specs into a win. The verdict, strictly from the data, is clear: the P106-100 wins the head-to-head, and the Tesla K40m remains a niche option for memory capacity needs.