NVIDIA P106-090 vs NVIDIA Tesla K20Xm Comparison
NVIDIA P106-090
Tesla K20Xm
PERFORMANCE BENCHMARKS
Analysis: NVIDIA P106-090 vs NVIDIA Tesla K20Xm
The Verdict
The data presents a clear, if narrow, picture: the NVIDIA P106-090 is the faster card in the only benchmark where both were tested, but the choice between these two end-of-life accelerators depends entirely on workload. The P106-090 wins the head-to-head Geekbench OpenCL test with a score of 21,304 against the Tesla K20Xm’s 17,215, a substantial 23.8% advantage. Its average benchmark score of 13,470 across all tests also edges out the Tesla’s 12,625, and it sits at the 54th percentile of all GPUs versus the Tesla’s 52nd. For general compute tasks, the P106-090 is the better performer.
However, the Tesla K20Xm is not without merit. It offers double the memory (6 GB vs 3 GB) on a wider 384-bit bus, yielding higher memory bandwidth of 249.6 GB/s versus 192.2 GB/s. This makes it the pick for workloads that are memory-capacity bound or rely heavily on raw bandwidth, even if its compute throughput is lower. The Tesla’s FP32 performance is also higher at 3.935 TFLOPS versus 2.352 TFLOPS, suggesting it may excel in certain floating-point-heavy tasks despite losing the OpenCL test. The P106-090 is the more efficient and architecturally modern solution, while the Tesla K20Xm is a larger, older, and power-hungrier card with a specific niche in high-memory, high-bandwidth scenarios. Ultimately, the P106-090 is the better all-rounder for compute benchmarks, while the Tesla K20Xm is a specialist for memory-intensive jobs.
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA P106-090 holds the edge with an average benchmark score of 13,470, compared to the NVIDIA Tesla K20Xm’s 12,625. This places the P106-090 at the 54th percentile of all GPUs, while the Tesla K20Xm sits at the 52nd percentile.
Q: How much faster is the P106-090 in the shared OpenCL test?
A: The P106-090 scores 21,304 in Geekbench OpenCL, while the Tesla K20Xm scores 17,215. This gives the P106-090 a 23.8% lead in that specific test.
Q: Which card has more memory and bandwidth?
A: The NVIDIA Tesla K20Xm has more memory, with 6 GB of GDDR5, and a wider 384-bit bus that provides 249.6 GB/s of bandwidth. The NVIDIA P106-090 has 3 GB of GDDR5 on a 192-bit bus, delivering 192.2 GB/s.
Q: What are the architecture and process node differences?
A: The P106-090 is based on the Pascal architecture (GP106 chip) built on a 16 nm process, while the Tesla K20Xm uses the older Kepler architecture (GK110 chip) on a 28 nm process. Both are manufactured by TSMC.
Q: Which card has higher raw FP32 compute?
A: The Tesla K20Xm has a higher FP32 rating of 3.935 TFLOPS, compared to the P106-090’s 2.352 TFLOPS. This suggests the Tesla may be better suited for certain compute tasks despite its lower benchmark scores.
Q: Do either of these cards have display outputs?
A: No. Both the NVIDIA P106-090 and the NVIDIA Tesla K20Xm have no display outputs, making them strictly compute or mining-oriented products.
Architecture Differences
The two cards come from vastly different eras and design philosophies. The P106-090 is a Pascal-generation part from the "Mining GPUs" family, built on a modern 16 nm TSMC process. Its GP106 chip packs 4,400 million transistors into a 200 mm² die, achieving a transistor density of 22.0M per mm². It supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4.
In contrast, the Tesla K20Xm is a Kepler-generation compute card from the "Tesla Kepler (Kxx)" line. It uses the GK110 chip on an older 28 nm process, housing 7,080 million transistors on a much larger 561 mm² die, with a lower transistor density of 12.6M per mm². Its API support is more limited, with DirectX 12 (11_0) and Vulkan 1.2.175, though it also supports OpenGL 4.6. The Tesla also has designated predecessors and successors (Tesla Fermi and Tesla Maxwell, respectively), while the P106-090 lists none.
The Tesla K20Xm is a much larger and more complex chip, with 2,688 shading units and 224 texture mapping units, compared to the P106-090’s 768 shading units and 48 TMUs. However, both have the same number of ROPs at 48. The Tesla’s higher FP32 throughput (3.935 TFLOPS) comes from this massive shader count, even at lower clocks, while the P106-090’s higher pixel rate (73.49 GPixel/s vs 40.99 GPixel/s) reflects its higher clock speeds and more efficient design. The P106-090 also has a documented FP16 rate of 36.74 GFLOPS (1:64), while the Tesla lists no FP16 capability.
Specification Differences
The two cards diverge significantly across nearly every core specification. The P106-090 has a base clock of 1354 MHz and a boost clock of 1531 MHz, while the Tesla K20Xm has no listed base or boost clocks. Memory speeds also differ: the P106-090 runs at 2002 MHz (8 Gbps effective), while the Tesla runs at 1300 MHz (5.2 Gbps effective). Consequently, the P106-090 achieves a higher pixel rate (73.49 GPixel/s) and texture rate (73.49 GTexel/s) relative to its core size, but the Tesla’s texture rate is higher in absolute terms at 164.0 GTexel/s.
Power and physical specs are starkly different. The P106-090 has a TDP of 75 W and requires a 250 W suggested PSU, with a single 6-pin power connector. The Tesla K20Xm has a TDP of 235 W and a 550 W suggested PSU, with no listed power connector. The P106-090 is shorter at 250 mm (9.8 inches) versus the Tesla’s 267 mm (10.5 inches). Both are dual-slot cards and have no display outputs. The bus interface also differs: the P106-090 uses a limited PCIe 1.0 x1 interface, while the Tesla uses the more capable PCIe 3.0 x16. The Tesla K20Xm was released earlier (2012-11-11) than the P106-090 (2017-07-30), and both are end-of-life products. The Tesla has a launch MSRP of 7,699 USD.
Head-to-Head Benchmarks
There is only one direct benchmark comparison available in the data, and it decisively favors the P106-090. In Geekbench OpenCL, the P106-090 scores 21,304 against the Tesla K20Xm’s 17,215, a 23.8% win for the Pascal card. This is a significant margin, indicating that despite the Tesla’s higher FP32 theoretical peak and much larger shader count, the P106-090 is more effective at executing this particular compute workload.
Beyond the direct head-to-head, the average benchmark scores reinforce the P106-090’s overall advantage. The P106-090 averages 13,470 across all its tests, which includes additional results in 3DMark Steel Nomad DX12 (509) and Geekbench Vulkan (18,596). The Tesla’s average of 12,625 is based on its Geekbench OpenCL (17,215) and Geekbench Metal (8,035) results. The P106-090’s 54th percentile ranking versus the Tesla’s 52nd further confirms this, though the difference is small.
The nearest rivals for each card provide context for their performance tiers. The P106-090’s average score is within 0.5% of the AMD Radeon Pro 555 and within 0.5% of the AMD Radeon RX 9070 XT, and it is nearly identical to the NVIDIA GeForce GTX 570 (-0.3%) and AMD Radeon HD 7770M (-0.5%). The Tesla K20Xm, meanwhile, is 0.7% behind the AMD Radeon RX 7600M XT, 1.2% behind the NVIDIA GeForce GTX 670, and 1.6% behind both the NVIDIA GeForce GTX 590 and AMD Radeon Pro 455. These comparisons show that while the P106-090 sits at a slightly higher performance tier on average, both cards are clustered in a similar mid-range bracket. The single head-to-head win, however, is the most direct evidence: the P106-090 is the superior card for OpenCL compute, while the Tesla’s only advantages lie in its larger memory pool and higher theoretical FP32 rate, which did not translate to a benchmark win in this dataset.