NVIDIA GeForce GTX 870M vs NVIDIA Tesla K20c Comparison
NVIDIA GeForce GTX 870M
Tesla K20c
PERFORMANCE BENCHMARKS
Analysis: NVIDIA GeForce GTX 870M vs NVIDIA Tesla K20c
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA GeForce GTX 870M has an average benchmark score of 9959, while the NVIDIA Tesla K20c has an average benchmark score of 11479. The Tesla K20c sits at the 51st percentile among all GPUs, whereas the GTX 870M sits at the 48th percentile.
Q: How do the two compare in the only shared benchmark?
A: In the Geekbench OpenCL test, the GTX 870M scores 12630 versus the Tesla K20c’s 11479. That gives the GTX 870M a 9.1% advantage, and it is the only head-to-head benchmark recorded for these two cards.
Q: What memory configurations do these cards use?
A: The Tesla K20c has 5 GB of GDDR5 memory on a 320-bit bus, delivering 208.0 GB/s of bandwidth. The GTX 870M has 3 GB of GDDR5 memory on a 192-bit bus, delivering 120.0 GB/s of bandwidth.
Q: Which card has more shading units?
A: The Tesla K20c has 2496 shading units, while the GTX 870M has 1344. The Tesla also has 208 texture mapping units versus 112, and 40 ROPs versus 24.
Q: What is the power draw difference?
A: The Tesla K20c is rated at 225 W TDP, requires a dual-slot cooler, and uses 1x 6-pin plus 1x 8-pin power connectors. The GTX 870M is rated at 100 W TDP, uses an MXM module form factor, and has no separate power connectors.
Q: Which card is newer?
A: The Tesla K20c was released on 2012-11-11, while the GTX 870M was released on 2014-03-11. Both are end-of-life products, and both use the Kepler architecture on a 28 nm TSMC process.
Architecture Differences
Both GPUs are built on NVIDIA’s Kepler architecture and manufactured by TSMC on a 28 nm process, but they are fundamentally different chips. The Tesla K20c uses the GK110 chip, a large compute-oriented die with 7,080 million transistors and a die size of 561 mm². The GTX 870M uses the GK104 chip, which is a smaller, more power-efficient design with 3,540 million transistors and a die size of 294 mm². Transistor density is nearly identical: 12.6M per mm² for the Tesla and 12.0M per mm² for the GeForce.
The GK110 in the Tesla is a much larger silicon investment, which shows in the raw compute resources. The Tesla has 2496 shading units, 208 TMUs, and 40 ROPs. The GTX 870M has 1344 shading units, 112 TMUs, and 24 ROPs. That is roughly 86% more shading units on the Tesla, which suggests the GK110 was designed for heavy parallel compute workloads rather than gaming efficiency.
Clock behavior also differs. The Tesla K20c has no listed base or boost clock in the database, while the GTX 870M runs at a base clock of 941 MHz and a boost clock of 967 MHz. Memory clocks are close: the Tesla runs at 1300 MHz (5.2 Gbps effective) and the GTX 870M at 1250 MHz (5 Gbps effective). The Tesla’s memory interface is wider at 320 bits versus 192 bits, which directly explains its much larger bandwidth figure of 208.0 GB/s versus 120.0 GB/s.
Feature support is largely identical: both support DirectX 12 (11_0), OpenGL 4.6, and Vulkan 1.2.175. Neither has ray tracing cores or tensor cores, as expected for Kepler-generation parts. The Tesla has no display outputs, while the GTX 870M’s outputs are listed as portable-device dependent, meaning the laptop vendor decides the actual connectors.
The physical form factors are starkly different. The Tesla K20c is a 267 mm dual-slot card, 10.5 inches long, requiring a 550 W suggested power supply. The GTX 870M is an MXM module, a compact board standard for notebooks, with no separate power connectors and a much lower 100 W TDP. These are not competing in the same chassis; they target entirely different system types.
Where Each One Wins
The recorded data shows a clear split. The GTX 870M wins the only shared benchmark, Geekbench OpenCL, by 9.1%. That is a meaningful margin for a single test, and it suggests that in OpenCL workloads, the mobile chip’s higher clock-driven efficiency holds up well against the larger desktop compute card.
However, the Tesla K20c’s strengths are structural rather than benchmark-derived. It has more than double the memory capacity (5 GB versus 3 GB), nearly double the memory bandwidth (208.0 GB/s versus 120.0 GB/s), and far more shading units. For workloads that scale with raw parallel throughput and memory bandwidth, such as large matrix operations or data-parallel compute, the Tesla’s hardware resources are substantially larger. The GTX 870M’s win in OpenCL may reflect driver optimization or the specific workload’s sensitivity to clock speed rather than raw resource counts.
The GTX 870M also wins on power efficiency in a qualitative sense: it delivers 2.599 TFLOPS of FP32 performance at 100 W, while the Tesla delivers 3.524 TFLOPS at 225 W. Per watt, the GTX 870M is the more efficient part, even though the Tesla has higher absolute throughput.
The Tesla K20c wins in absolute compute throughput: 3.524 TFLOPS versus 2.599 TFLOPS, a 35.5% advantage in FP32 capability. It also has higher pixel rate (36.71 GPixel/s versus 27.08 GPixel/s) and texture rate (146.8 GTexel/s versus 108.3 GTexel/s). For tasks that are not memory-bound, the Tesla should be faster. For tasks that are memory-latency or driver-bound, the GTX 870M’s OpenCL result suggests it can be competitive or faster in practice.
Specification Differences
The following fields differ between the two cards:
- Chip: GK110 (Tesla K20c) versus GK104 (GTX 870M)
- Transistors: 7,080 million versus 3,540 million
- Die size: 561 mm² versus 294 mm²
- Transistor density: 12.6M / mm² versus 12.0M / mm²
- Base clock: None listed versus 941 MHz
- Boost clock: None listed versus 967 MHz
- Memory clock: 1300 MHz (5.2 Gbps effective) versus 1250 MHz (5 Gbps effective)
- Memory size: 5 GB versus 3 GB
- Memory bus width: 320 bit versus 192 bit
- Memory bandwidth: 208.0 GB/s versus 120.0 GB/s
- Shading units: 2496 versus 1344
- TMUs: 208 versus 112
- ROPs: 40 versus 24
- Pixel rate: 36.71 GPixel/s versus 27.08 GPixel/s
- Texture rate: 146.8 GTexel/s versus 108.3 GTexel/s
- FP32: 3.524 TFLOPS versus 2.599 TFLOPS
- TDP: 225 W versus 100 W
- Slot width: Dual-slot versus MXM Module
- Power connectors: 1x 6-pin + 1x 8-pin versus None
- Suggested PSU: 550 W versus not listed
- Bus interface: PCIe 2.0 x16 versus MXM-B (3.0)
- Display outputs: No outputs versus Portable Device Dependent
- Dimensions: 267 mm (10.5 inches) versus not listed
- Release date: 2012-11-11 versus 2014-03-11
- Predecessor: Tesla Fermi versus GeForce 700M
- Successor: Tesla Maxwell versus GeForce 900M
- Generation: Tesla Kepler (Kxx) versus GeForce 800M
- Launch MSRP: 3,199 USD versus not listed
Head-to-Head Benchmarks
The database records one direct comparison: Geekbench OpenCL. In that test, the GTX 870M scores 12630, and the Tesla K20c scores 11479. The GTX 870M wins by 9.1%. This is the only shared benchmark, so it carries the full weight of the head-to-head comparison.
The margin is notable because the Tesla has far more hardware resources. The GK110 chip has 2496 shading units versus 1344, and the Tesla’s FP32 throughput is 3.524 TFLOPS versus 2.599 TFLOPS, a 35.5% advantage. Yet in OpenCL, the GTX 870M comes out ahead by nearly ten percent. This suggests that the GTX 870M’s higher clock speeds (941 MHz base, 967 MHz boost) and possibly better driver support for OpenCL on mobile parts overcome the Tesla’s raw compute advantage in this particular workload.
Looking at the nearest rivals helps contextualize both scores. The Tesla K20c’s average score of 11479 is within 1.9% of the NVIDIA GeForce GTX 780M (11261), and it trails the AMD Radeon Pro 5500M by 0.4%, the AMD Radeon RX 7800 XT by 1.3%, and the NVIDIA GeForce GTX 1660 by 1.7%. The GTX 870M’s average score of 9959 sits near the AMD Radeon Pro 5300M (10013, 0.5% higher), the NVIDIA Quadro K5100M (10043, 0.8% higher), and the AMD Radeon R9 M375 (10070, 1.1% higher), while beating the NVIDIA Quadro 6000 by 1.1%.
The averages are lower for the GTX 870M because its second benchmark, Geekbench Metal, scores only 7288, dragging the average down. The Tesla has only one benchmark recorded, so its average equals its OpenCL score. This means the GTX 870M’s OpenCL performance is actually its stronger result, while the Tesla’s single data point may not represent its typical compute performance across other APIs.
The Verdict
The data does not support a single winner across all use cases. For anyone running OpenCL workloads on a laptop, the GTX 870M is the clear choice: it leads the only direct benchmark by 9.1% and does so at 100 W TDP with a compact MXM form factor. Its average benchmark score is lower only because of the weaker Metal result, which is irrelevant to OpenCL-only comparisons.
For compute-heavy tasks that benefit from raw FP32 throughput, memory bandwidth, and large memory capacity, the Tesla K20c is the stronger hardware. It delivers 3.524 TFLOPS, 208.0 GB/s of bandwidth, and 5 GB of memory, all figures that dwarf the GTX 870M. The 35.5% FP32 advantage and 73.3% bandwidth advantage are structural, even if the OpenCL result does not reflect them.
The GTX 870M’s win in OpenCL is a reminder that benchmark scores do not always align with hardware specifications. The GK104 chip is smaller, older in design intent, but clocked higher and optimized for a consumer-facing product line. The Tesla K20c is a workstation compute card from 2012, with a launch MSRP of 3,199 USD, aimed at scientific and professional workloads that may not map directly onto Geekbench OpenCL.
Ultimately, the choice depends on the system: the Tesla K20c is a dual-slot desktop card requiring a 550 W power supply and external power connectors, while the GTX 870M is a notebook MXM module with no external power needs. If the question is which GPU to pick for a portable system, the GTX 870M is the only option that fits. If the question is which GPU has more raw compute headroom, the Tesla K20c wins on every architectural metric except the single benchmark result. The recorded data favors the GTX 870M in the only direct comparison, but the Tesla’s hardware profile suggests it was built for workloads that Geekbench OpenCL does not fully capture.