NVIDIA GeForce GTX 1070 vs NVIDIA Tesla K20c Comparison
NVIDIA GeForce GTX 1070
Tesla K20c
PERFORMANCE BENCHMARKS
Analysis: NVIDIA GeForce GTX 1070 vs NVIDIA Tesla K20c
FAQ
Q: How does the NVIDIA Tesla K20c compare to the GeForce GTX 1070 in OpenCL performance?
A: The GTX 1070 scores 44,700 in Geekbench OpenCL, while the Tesla K20c scores 11,479. This gives the GTX 1070 a 74.3% advantage in this specific workload.
Q: What are the average benchmark scores for these two GPUs?
A: The Tesla K20c has an average benchmark score of 11,479, placing it in the 51st percentile of all GPUs. The GTX 1070 has an average score of 9,780, placing it in the 47th percentile.
Q: Which GPU has more memory and higher bandwidth?
A: The GTX 1070 has 8 GB of GDDR5 memory on a 256-bit bus, delivering 256.3 GB/s of bandwidth. The Tesla K20c has 5 GB of GDDR5 on a 320-bit bus, delivering 208.0 GB/s.
Q: What are the core specifications of each GPU?
A: The Tesla K20c uses the GK110 chip with 2,496 shading units, 208 texture mapping units, and 40 ROPs. The GTX 1070 uses the GP104 chip with 1,920 shading units, 120 TMUs, and 64 ROPs.
Q: What is the power consumption difference?
A: The Tesla K20c has a TDP of 225 W and requires a 550 W power supply. The GTX 1070 has a TDP of 150 W and requires a 450 W power supply.
Q: What API levels do these GPUs support?
A: The Tesla K20c supports DirectX 12 (11_0), OpenGL 4.6, and Vulkan 1.2.175. The GTX 1070 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4.
The Verdict
The data presents a clear separation between these two NVIDIA products. The GeForce GTX 1070 is the stronger choice for anyone whose workload includes OpenCL compute, given its 74.3% lead in that specific test. Its higher memory capacity of 8 GB versus 5 GB and faster bandwidth of 256.3 GB/s versus 208.0 GB/s also favor it for data-heavy tasks.
The Tesla K20c, however, holds its own in the raw pixel pipeline. It delivers a texture rate of 146.8 GTexel/s and a pixel rate of 36.71 GPixel/s, though these are lower than the GTX 1070's figures of 202.0 GTexel/s and 107.7 GPixel/s respectively. The K20c's advantage lies in its compute-oriented design: it has more shading units (2,496 vs 1,920) and more TMUs (208 vs 120), which historically benefits certain scientific workloads.
For users evaluating based on the database's recorded average, the Tesla K20c holds a higher percentile ranking (51st vs 47th) and a higher average score (11,479 vs 9,780). This suggests that in the aggregate of all tests, the K20c performs better than its rival, despite the GTX 1070's dominance in the single head-to-head test.
The GTX 1070 is the logical pick for most users because it offers a fully featured consumer card with display outputs (1x DVI, 1x HDMI 2.0, 3x DisplayPort 1.4a) and a lower power draw of 150 W. The K20c has no display outputs and a 225 W TDP, marking it as a compute-only product. For pure machine learning or scientific compute tasks where memory capacity is less critical, the GTX 1070's OpenCL score and efficiency are decisive. For tasks that rely heavily on raw shading throughput, the K20c's larger count of shading units may be preferable, but the available benchmark data does not demonstrate that advantage.
The GTX 1070 is the better overall product for general and compute use. The K20c remains a niche offering for specific, non-display workloads that prioritize the Kepler architecture's legacy.
Head-to-Head Benchmarks
The database contains one directly comparable benchmark between these two GPUs: Geekbench OpenCL. In this test, the GeForce GTX 1070 scores 44,700, while the Tesla K20c scores 11,479. This represents a delta of 74.3% in favor of the GTX 1070. The K20c's score is 74.3% lower than the GTX's, making it a very large performance gap in compute performance.
This result aligns with the GTX 1070's significantly higher FP32 throughput of 6.463 TFLOPS, compared to the K20c's 3.524 TFLOPS. The GTX 1070 also has a higher memory bandwidth of 256.3 GB/s, which likely contributes to its superior OpenCL performance. The K20c's 320-bit bus is wider than the GTX's 256-bit bus, but the GTX uses faster memory (8 Gbps effective vs 5.2 Gbps effective), resulting in higher net bandwidth.
The GTX 1070 also shows a massive lead in pixel fill rate: 107.7 GPixel/s versus 36.71 GPixel/s for the K20c. This is a 65.9% difference. Its texture rate is also significantly higher at 202.0 GTexel/s versus 146.8 GTexel/s.
The K20c's only advantages in specification are its higher shading unit count (2,496 vs 1,920) and its larger die size (561 mm² vs 314 mm²). These are architectural choices, not performance wins in the recorded benchmark.
Overall, the head-to-head data shows the GTX 1070 as the clear winner in the only test that has a direct comparison. The K20c's higher average score and percentile suggest it may perform better in other workloads, but the database has no direct evidence for those.
Specification Differences
The two cards differ in several key specifications. The GTX 1070 has 6.463 TFLOPS of FP32 performance, while the K20c delivers 3.524 TFLOPS. The GTX 1070 also has a higher texture rate of 202.0 GTexel/s versus 146.8 GTexel/s. The pixel rate is also in favor of the GTX 1070: 107.7 GPixel/s versus 36.71 GPixel/s.
Memory configurations differ significantly. The GTX 1070 has 8 GB of GDDR5 memory, while the K20c has 5 GB. The GTX 1070's memory runs at 8 Gbps effective, providing 256.3 GB/s of bandwidth. The K20c's memory runs at 5.2 Gbps effective, providing 208.0 GB/s. The K20c uses a 320-bit bus, while the GTX 1070 uses a 256-bit bus.
The K20c has a higher TDP of 225 W compared to the GTX 1070's 150 W. Consequently, the K20c requires a 550 W power supply, while the GTX 1070 only needs a 450 W unit. The K20c uses both a 6-pin and an 8-pin power connector, while the GTX 1070 uses a single 8-pin connector.
The bus interface differs: the K20c uses PCIe 2.0 x16, while the GTX 1070 uses PCIe 3.0 x16. The GTX 1070 offers display outputs (1x DVI, 1x HDMI 2.0, 3x DisplayPort 1.4a), while the K20c has no outputs.
The GTX 1070 has a higher DirectX support (12_1) compared to the K20c's DirectX 12 (11_0). The GTX 1070 supports Vulkan 1.4, while the K20c supports Vulkan 1.2.175. Both support OpenGL 4.6.
The GTX 1070 has a full set of dimensions recorded: 267 mm length, 112 mm height, 40 mm width. The K20c only has a length of 267 mm.
Architecture Differences
The two GPUs come from different architectures and design generations. The Tesla K20c is based on the Kepler architecture, using the GK110 chip. It was manufactured on a 28 nm process from TSMC, with 7,080 million transistors on a 561 mm² die. This gives a transistor density of 12.6 million per mm². Its generation is labeled "Tesla Kepler (Kxx)" and it was released on November 11, 2012.
The GeForce GTX 1070 is based on the Pascal architecture, using the GP104 chip. It was manufactured on a 16 nm process from TSMC, with 7,200 million transistors on a 314 mm² die. This results in a much higher transistor density of 22.9 million per mm². Its generation is "GeForce 10", and it was released on June 9, 2016.
The K20c has 2,496 shading units, 208 texture mapping units, and 40 ROPs. The GTX 1070 has 1,920 shading units, 120 TMUs, and 64 ROPs. This means the K20c has more shading units but fewer ROPs, while the GTX 1070 has fewer shading units but more ROPs.
The architecture also affects API support. The GTX 1070 supports DirectX 12 (12_1) and Vulkan 1.4, which are more advanced versions than the K20c's DirectX 12 (11_0) and Vulkan 1.2.175. The GTX 1070 also has a separate FP16 performance rating of 101.0 GFLOPS (1:64), while the K20c has no recorded FP16 data.
The K20c's predecessor is the Tesla Fermi, and its successor is the Tesla Maxwell. The GTX 1070's predecessor is the GeForce 900 series, and its successor is the GeForce 20 series.
Both cards are dual-slot, and both have a 267 mm length. The GTX 1070 has a 112 mm height and 40 mm width, while the K20c's height and width are not recorded. The K20c is a compute-only product with no display outputs, reflecting its server and workstation design, whereas the GTX 1070 is a consumer graphics card with full display support.
The GTX 1070's modern Pascal architecture delivers higher clock speeds, with a base of 1506 MHz and a boost of 1683 MHz, while the K20c has no recorded base or boost clocks. The GTX 1070's memory is also faster, running at 2002 MHz (8 Gbps effective) versus the K20c's 1300 MHz (5.2 Gbps effective). These architectural improvements, combined with the smaller 16 nm process, account for the GTX 1070's superior performance in the recorded tests.