NVIDIA Tesla K20m vs NVIDIA Tesla K40c Comparison
NVIDIA Tesla K20m
Tesla K40c
PERFORMANCE BENCHMARKS
Analysis: NVIDIA Tesla K20m vs NVIDIA Tesla K40c
Head-to-Head Benchmarks
The recorded data includes a single direct comparison between these two accelerators: the Geekbench OpenCL test. In this measurement, the NVIDIA Tesla K40c scores 17468, while the NVIDIA Tesla K20m scores 16241. The K40c wins this head-to-head by 7 percentage points. This is a meaningful gap in compute workloads, as OpenCL exercises raw shader and memory throughput. The K20m trails by that same 7% margin, which aligns with the underlying hardware differences detailed below. Beyond the direct head-to-head, the database shows the K20m has an average benchmark score of 19089 across its recorded tests, which includes both OpenCL and Vulkan results. The K40c has only one recorded benchmark, the OpenCL score of 17468, making its average score identical to that single result at 17468.
Looking at the overall percentile ranking, the K20m sits at the 64th percentile among all GPUs in the database, while the K40c sits at the 61st percentile. This is an interesting inversion: despite losing the direct OpenCL comparison, the K20m holds a higher global percentile, driven by its additional Vulkan score of 21936, which is not recorded for the K40c. When comparing to nearby rivals, the K20m is within 0.2% of the NVIDIA GeForce RTX 4050 Mobile (average score 19049) and within 0.3% of both the AMD Radeon RX 6600 (19036) and the NVIDIA Quadro K6000 (19030). It is 0.4% behind the NVIDIA GeForce GTX 780 (19164). The K40c's closest rival is the AMD Radeon Pro 460 (17509), which is 0.2% ahead of it. The AMD Radeon Pro 560 (17551) is 0.5% ahead, the AMD Radeon 780M (17588) is 0.7% ahead, and the NVIDIA GeForce RTX 4060 (17639) is 1% ahead. The data indicates that in synthetic compute benchmarks, the K40c's single OpenCL result places it in a competitive band with modern integrated and entry-level discrete solutions, while the K20m's dual-test average gives it a broader performance profile.
Architecture Differences
Both cards are built on the Kepler architecture and use the same 28 nm process at TSMC. Both have 7,080 million transistors on a 561 mm² die, yielding a transistor density of 12.6M per mm². The K20m uses the GK110 chip, while the K40c uses the GK180 chip. The GK180 is a further refinement of the same basic design, but with substantial increases in execution resources. The K20m has 2,496 shading units, 208 texture mapping units, and 40 raster operation units. The K40c increases shading units to 2,880, texture units to 240, and raster operations to 48. This is a 15.4% increase in shading units, a 15.4% increase in texture units, and a 20% increase in ROPs, all relative to the K20m.
Clock behavior also differs. The K20m has no recorded base or boost clock in the database, while the K40c has a base clock of 745 MHz and a boost clock of 876 MHz. Memory clocks differ as well: the K20m runs its GDDR5 at 1300 MHz (5.2 Gbps effective), while the K40c runs at 1502 MHz (6 Gbps effective). The memory subsystem itself is larger on the K40c: 12 GB versus 5 GB, with a 384-bit bus versus 320-bit, and bandwidth of 288.4 GB/s versus 208.0 GB/s. The K40c's bandwidth advantage is 38.7% over the K20m. The K40c also uses PCIe 3.0 x16, while the K20m uses PCIe 2.0 x16. Both cards have no display outputs, dual-slot cooling, and require a 1x 6-pin plus 1x 8-pin power connector, with a suggested 550 W power supply. The K40c has a higher thermal design power at 245 W versus 225 W. Both support DirectX 12 (11_0), OpenGL 4.6, and Vulkan 1.2.175. Neither card has ray tracing cores or tensor cores, and neither lists FP16 performance. The FP32 output is 5.046 TFLOPS for the K40c versus 3.524 TFLOPS for the K20m, a 43.2% advantage.
The Verdict
The data points clearly to the NVIDIA Tesla K40c as the stronger compute accelerator in the only direct benchmark recorded. Its 7% OpenCL lead over the K20m is consistent with its larger shading unit count, higher clocks, wider memory bus, and greater memory capacity. For workloads that are sensitive to raw FP32 throughput and memory bandwidth, the K40c is the better choice based on the recorded measurements. The K40c also offers 12 GB of VRAM versus 5 GB, which matters for datasets that exceed the smaller frame buffer.
However, the K20m is not without merit in the database. Its average benchmark score of 19089 is higher than the K40c's 17468, primarily due to a strong Vulkan result of 21936. This suggests the K20m may have better driver optimization or architectural characteristics for Vulkan compute workloads, even though its OpenCL performance is lower. The K20m also holds a higher global percentile (64th versus 61st), indicating that across all GPUs in the database, its combined benchmark scores place it above the K40c in relative standing. The K20m also has a lower TDP at 225 W versus 245 W, which could be a factor in dense deployments.
For users prioritizing OpenCL compute with large memory footprints, the K40c is the clear pick. For users prioritizing Vulkan compute or overall average benchmark scores, the K20m appears to have an edge, though the absence of a Vulkan score for the K40c makes a direct comparison on that API impossible from the recorded data. The launch MSRP of the K40c was 7,699 USD, while the K20m was 3,199 USD. Both are end-of-life products, with the K20m released earlier and the K40c later in 2013. Neither card has display outputs, so both are strictly compute accelerators.
Specification Differences
The following table highlights only the fields where the two cards differ in the database.
| Specification | NVIDIA Tesla K20m | NVIDIA Tesla K40c |
|---|---|---|
| Chip | GK110 | GK180 |
| Base clock | Not recorded | 745 MHz |
| Boost clock | Not recorded | 876 MHz |
| Memory clock | 1300 MHz (5.2 Gbps effective) | 1502 MHz (6 Gbps effective) |
| Memory size | 5 GB | 12 GB |
| Memory bus width | 320 bit | 384 bit |
| Memory bandwidth | 208.0 GB/s | 288.4 GB/s |
| Shading units | 2496 | 2880 |
| Texture mapping units | 208 | 240 |
| Raster operation units | 40 | 48 |
| Pixel rate | 36.71 GPixel/s | 52.56 GPixel/s |
| Texture rate | 146.8 GTexel/s | 210.2 GTexel/s |
| FP32 performance | 3.524 TFLOPS | 5.046 TFLOPS |
| TDP | 225 W | 245 W |
| Bus interface | PCIe 2.0 x16 | PCIe 3.0 x16 |
| Release date | 2013-01-04 | 2013-10-07 |
| Launch MSRP | 3,199 USD | 7,699 USD |
| Geekbench OpenCL score | 16241 | 17468 |
| Geekbench Vulkan score | 21936 | Not recorded |
| Average benchmark score | 19089 | 17468 |
| Percentile vs all GPUs | 64 | 61 |
FAQ
Q: Which card has more memory?
A: The NVIDIA Tesla K40c has 12 GB of GDDR5 memory, while the NVIDIA Tesla K20m has 5 GB.
Q: What is the performance difference in OpenCL?
A: In the Geekbench OpenCL test, the K40c scores 17468, which is 7% higher than the K20m's score of 16241.
Q: Does the K20m have any benchmark advantage?
A: Yes, the K20m has a Geekbench Vulkan score of 21936, which is not recorded for the K40c. This contributes to the K20m's higher average benchmark score of 19089 versus 17468.
Q: Which card has higher memory bandwidth?
A: The K40c has 288.4 GB/s of bandwidth, while the K20m has 208.0 GB/s.
Q: Are both cards the same physical size?
A: Yes, both have a length of 267 mm (10.5 inches) and are dual-slot cards.
Q: What power connectors do they require?
A: Both cards require a 1x 6-pin plus 1x 8-pin power connector and a suggested 550 W power supply.
Where Each One Wins
The NVIDIA Tesla K40c wins in raw compute throughput. Its FP32 rating of 5.046 TFLOPS is 43.2% higher than the K20m's 3.524 TFLOPS. Its pixel rate of 52.56 GPixel/s beats the K20m's 36.71 GPixel/s, and its texture rate of 210.2 GTexel/s beats 146.8 GTexel/s. In the single direct benchmark, it wins OpenCL by 7%. The K40c also wins on memory capacity (12 GB versus 5 GB) and memory bandwidth (288.4 GB/s versus 208.0 GB/s), making it the better option for memory-bound compute tasks. The K40c also supports PCIe 3.0, doubling the bus bandwidth available on the K20m's PCIe 2.0 interface.
The NVIDIA Tesla K20m wins in the Vulkan compute domain, at least based on the recorded data. Its Vulkan score of 21936 is substantially higher than its own OpenCL score, and no Vulkan result exists for the K40c for comparison. The K20m also has a higher average benchmark score (19089 versus 17468) and a higher global percentile (64 versus 61). The K20m has a lower thermal design power at 225 W versus 245 W, which may be preferable in power-constrained environments. Its lower launch MSRP of 3,199 USD compared to 7,699 USD is a historical data point, though both products are now end-of-life. For workloads that use Vulkan compute APIs, or for users who value the broader benchmark profile, the K20m appears to hold an advantage. For OpenCL-heavy workloads with large memory requirements, the K40c is the stronger card according to the measurements.