NVIDIA P106-090 vs NVIDIA Tesla K20c Comparison
NVIDIA P106-090
Tesla K20c
PERFORMANCE BENCHMARKS
Analysis: NVIDIA P106-090 vs NVIDIA Tesla K20c
Head-to-Head Benchmarks
The only shared benchmark in the database is Geekbench OpenCL, and the result is decisive. The NVIDIA P106-090 scores 21,304, while the NVIDIA Tesla K20c scores 11,479. That gives the P106-090 an 85.6% advantage, a massive gap that reflects far more than clock speed differences.
Looking at the broader database context, the P106-090's average benchmark score is 13,470, placing it at the 54th percentile of all GPUs. Its nearest rivals include the GeForce GTX 570 (13,515, just 0.3% ahead), the Radeon Pro 555 (13,407, 0.5% behind), and the Radeon HD 7770M (13,536, 0.5% behind). This puts the P106-090 in a tight cluster of mid-range performers from several generations ago.
The Tesla K20c, by contrast, averages 11,479 and sits at the 51st percentile. Its closest comparables are the Radeon Pro 5500M (11,528, 0.4% ahead), the Radeon RX 7800 XT (11,627, 1.3% ahead), and the GeForce GTX 1660 (11,680, 1.7% ahead). Interestingly, the GTX 780M trails it by 1.9%. The K20c is firmly in lower-midrange territory despite its compute-oriented heritage.
The 85.6% OpenCL delta is stark. In raw compute workloads that stress the shading units, the P106-090 wins outright. There is no benchmark in the database where the K20c comes out ahead. The P106-090 also shows up in 3DMark Steel Nomad (DX12) with a score of 509 and in Geekbench Vulkan with 18,596, while the K20c has no recorded results for those tests, so the comparison remains limited to OpenCL.
Architecture Differences
The two cards come from different NVIDIA generations and are built for different purposes. The P106-090 uses the GP106 chip on the Pascal architecture, fabricated on a 16 nm TSMC process. The K20c uses the GK110 chip on the Kepler architecture, fabricated on a 28 nm TSMC process. That process gap is significant: the P106-090 packs 4,400 million transistors into a 200 mm² die, yielding a transistor density of 22.0 million per square millimeter. The K20c carries 7,080 million transistors across a 561 mm² die, with a density of just 12.6 million per square millimeter. Pascal is roughly 75% denser, which explains how the smaller, newer chip can outperform the larger one.
The K20c compensates with raw scale. It has 2,496 shading units, 208 texture mapping units, and 40 render output units. The P106-090 has 768 shading units, 48 TMUs, and 48 ROPs. That is a 3.25x difference in shading units and over 4x in TMUs, yet the P106-090 still wins in OpenCL. Clock speeds tell part of the story: the P106-090 runs at a 1354 MHz base and 1531 MHz boost, while the K20c has no base or boost clocks recorded in the database. The memory clock also favors the P106-090 at 2002 MHz (8 Gbps effective) versus 1300 MHz (5.2 Gbps effective) for the K20c.
Memory configurations differ as well. The P106-090 has 3 GB of GDDR5 on a 192-bit bus, delivering 192.2 GB/s of bandwidth. The K20c has 5 GB of GDDR5 on a 320-bit bus, delivering 208.0 GB/s. The K20c actually has more capacity and slightly more bandwidth, but its lower memory clock and older architecture hold it back.
Feature support also diverges. The P106-090 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The K20c supports DirectX 12 (11_0), OpenGL 4.6, and Vulkan 1.2.175. Both have no display outputs, which marks them as compute or mining products rather than gaming cards. The P106-090 is classified under "Mining GPUs," while the K20c belongs to the Tesla Kepler (Kxx) compute lineup.
Power and connectivity show clear generational differences. The P106-090 is rated at 75 W TDP with a single 6-pin power connector and a suggested 250 W PSU. The K20c draws 225 W, needs a 6-pin plus 8-pin connector, and suggests a 550 W PSU. The P106-090 uses PCIe 1.0 x1, which is an odd limitation for a mining card, while the K20c uses PCIe 2.0 x16. Physical dimensions are close: the P106-090 is 250 mm (9.8 inches) long, the K20c is 267 mm (10.5 inches). Both are dual-slot cards.
Where Each One Wins
The P106-090 wins in every measurable category from the recorded data. Its OpenCL score is 85.6% higher, which is a decisive margin for any compute workload that uses OpenCL. The 3DMark Steel Nomad DX12 score of 509 and the Vulkan score of 18,596 further indicate it handles modern graphics APIs, though the K20c has no comparable results to measure against.
For practical use, the P106-090 is the better choice for anyone running OpenCL-based compute tasks, including machine learning inference, physics simulation, or rendering pipelines that rely on that API. Its lower power draw (75 W vs 225 W) and simpler power connector requirements mean it can drop into systems with modest PSUs. The PCIe 1.0 x1 interface is a bottleneck, but for compute tasks that do not hammer the bus, the raw GPU throughput still carries the day.
The K20c has one advantage from the recorded specs: memory capacity. Its 5 GB frame buffer is larger than the P106-090's 3 GB, and its 208.0 GB/s bandwidth edges out the 192.2 GB/s of the P106-090. Workloads that need to hold larger datasets in VRAM, such as certain scientific computing kernels or oversized textures, could prefer the K20c despite its lower compute throughput. The K20c also has more shading units and TMUs, which theoretically helps in highly parallel, memory-bound operations, but the benchmark data does not show any test where that translates into a win.
The K20c's predecessor is Tesla Fermi and its successor is Tesla Maxwell, indicating it sits in a compute family that evolved quickly. The P106-090 has no predecessor or successor listed, which reflects its status as a niche mining product.
FAQ
Q: Which card is faster in OpenCL compute?
A: The NVIDIA P106-090 scores 21,304 in Geekbench OpenCL, which is 85.6% higher than the Tesla K20c's 11,479. The database records only one head-to-head benchmark, and the P106-090 wins it.
Q: Does the Tesla K20c have more compute units?
A: Yes, the K20c has 2,496 shading units and 208 TMUs, compared to 768 shading units and 48 TMUs on the P106-090. Despite having over three times the shading units, the K20c still loses in OpenCL due to its older Kepler architecture and lower clocks.
Q: Which card has more memory?
A: The Tesla K20c has 5 GB of GDDR5 on a 320-bit bus, while the P106-090 has 3 GB on a 192-bit bus. The K20c also has slightly higher bandwidth at 208.0 GB/s versus 192.2 GB/s.
Q: What are the power requirements?
A: The P106-090 is rated at 75 W TDP and needs a single 6-pin connector with a 250 W suggested PSU. The K20c is rated at 225 W, requires a 6-pin and 8-pin connector, and suggests a 550 W PSU.
Q: Do either of these cards support display outputs?
A: No. Both the P106-090 and the K20c have no display outputs, making them unsuitable for gaming or any application that requires a monitor connection.
Q: Which card has better API support?
A: The P106-090 supports DirectX 12 (12_1) and Vulkan 1.4, while the K20c supports DirectX 12 (11_0) and Vulkan 1.2.175. Both support OpenGL 4.6.
Specification Differences
| Specification | NVIDIA P106-090 | NVIDIA Tesla K20c |
| --- | --- | --- |
| Chip | GP106 | GK110 |
| Architecture | Pascal | Kepler |
| Generation | Mining GPUs | Tesla Kepler (Kxx) |
| Process node | 16 nm | 28 nm |
| Transistors | 4,400 million | 7,080 million |
| Die size | 200 mm² | 561 mm² |
| Transistor density | 22.0M / mm² | 12.6M / mm² |
| Base clock | 1354 MHz | Not recorded |
| Boost clock | 1531 MHz | Not recorded |
| Memory clock | 2002 MHz (8 Gbps effective) | 1300 MHz (5.2 Gbps effective) |
| Memory size | 3 GB | 5 GB |
| Memory bus width | 192 bit | 320 bit |
| Memory bandwidth | 192.2 GB/s | 208.0 GB/s |
| Shading units | 768 | 2496 |
| TMUs | 48 | 208 |
| ROPs | 48 | 40 |
| Pixel rate | 73.49 GPixel/s | 36.71 GPixel/s |
| Texture rate | 73.49 GTexel/s | 146.8 GTexel/s |
| FP32 | 2.352 TFLOPS | 3.524 TFLOPS |
| FP16 | 36.74 GFLOPS (1:64) | Not recorded |
| TDP | 75 W | 225 W |
| Power connectors | 1x 6-pin | 1x 6-pin + 1x 8-pin |
| Suggested PSU | 250 W | 550 W |
| Bus interface | PCIe 1.0 x1 | PCIe 2.0 x16 |
| DirectX | 12 (12_1) | 12 (11_0) |
| Vulkan | 1.4 | 1.2.175 |
| Length | 250 mm (9.8 inches) | 267 mm (10.5 inches) |
| Release date | 2017-07-30 | 2012-11-11 |
| Predecessor | Not recorded | Tesla Fermi |
| Successor | Not recorded | Tesla Maxwell |
| Launch MSRP | Not recorded | 3,199 USD |
The Verdict
The data points to the NVIDIA P106-090 as the better card for almost any workload. Its 85.6% OpenCL lead over the K20c is the only direct comparison available, and it is not close. The P106-090 also draws a third of the power (75 W versus 225 W), requires a simpler power connector, and supports newer API versions. It wins on pixel rate (73.49 GPixel/s versus 36.71 GPixel/s) and has a much higher transistor density, showing the efficiency of the 16 nm Pascal process.
The K20c does have redeeming qualities. Its 5 GB memory capacity and 208.0 GB/s bandwidth exceed the P106-090, and its FP32 throughput of 3.524 TFLOPS is higher than the P106-090's 2.352 TFLOPS. For workloads that are memory-capacity-bound or that scale with raw FP32 operations, the K20c might still be serviceable. But the benchmark record shows no test where the K20c wins. Its 51st percentile ranking versus the P106-090's 54th percentile confirms the overall performance gap, even if the percentile difference looks modest.
The launch MSRP of 3,199 USD for the K20c reflects its original enterprise positioning, but the recorded data does not support that premium today. The P106-090, despite its mining-oriented design and PCIe 1.0 x1 interface, is simply the faster card in the one test that matters. Choose the P106-090 for OpenCL compute, lower power draw, and modern API support. Choose the K20c only if you specifically need 5 GB of VRAM or its higher FP32 ceiling, and you are willing to accept the 225 W power draw and older architecture. For most builders, the P106-090 is the clear pick from the database evidence.