NVIDIA Tesla K20c vs NVIDIA Tesla M2090 Comparison

NVIDIA
GEFORCE

NVIDIA Tesla K20c

CORE STATE GK110
VRAM 5 GB
CLOCK SPEED
TDP 225 W
BUS WIDTH 320 bit
ARCHITECTURE Kepler
nm
PROCESS 28 nm
LAUNCH DATE 2012
VS
NVIDIA
GEFORCE

Tesla M2090

CORE STATE GF110
VRAM 6 GB
CLOCK SPEED
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Fermi 2.0
nm
PROCESS 40 nm
LAUNCH DATE 2011

PERFORMANCE BENCHMARKS

geekbench_opencl
11,479
13,075

Analysis: NVIDIA Tesla K20c vs NVIDIA Tesla M2090

Head-to-Head Benchmarks

The only recorded head-to-head benchmark between these two compute accelerators is Geekbench OpenCL, and the result is decisive. The NVIDIA Tesla M2090 scores 13,075 points, while the NVIDIA Tesla K20c scores 11,479 points. That gives the older Fermi card a 13.9% advantage over the newer Kepler part in this specific workload.

Context matters here. The M2090 sits at the 53rd percentile among all GPUs in the database, while the K20c lands at the 51st percentile. Both are mid-pack performers by modern standards, but the M2090's position is slightly higher. The raw delta of 1,596 points is substantial enough that it is unlikely to be noise.

Looking at the nearest rivals for each card puts the scores in perspective. The M2090's 13,075 is within 0.7% of the GeForce GTX 1660 SUPER (12,986), within 1% of the RTX 3050 Ti Mobile (12,940), and within 1.1% of the AMD Radeon RX 580 (12,928). It trails the GTX 950 (13,189) by 0.9%. The K20c's 11,479 is 0.4% behind the AMD Radeon Pro 5500M (11,528), 1.3% behind the Radeon RX 7800 XT (11,627), and 1.7% behind the GTX 1660 (11,680). It beats the GTX 780M (11,261) by 1.9%.

The interesting implication is that the M2090, despite being from an earlier architecture generation, outperforms the K20c in this compute test. The gap between them (13.9%) is larger than the gap between either card and its nearest rivals in the database. This suggests the benchmark favors the M2090's specific configuration, or that the K20c is held back by factors visible in its specifications.

Where Each One Wins

The M2090 wins the only benchmark category recorded in the database, which is Geekbench OpenCL. Its 13,075 score versus the K20c's 11,479 gives it a clean sweep of the head-to-head comparison. The wins tally shows 1 for the M2090 and 0 for the K20c.

However, the specification sheet tells a more complex story. The K20c has clear theoretical advantages in raw compute throughput. Its FP32 performance is listed at 3.524 TFLOPS, which is roughly 2.6 times the M2090's 1,332.2 GFLOPS (about 1.33 TFLOPS). The K20c also has far higher texture throughput at 146.8 GTexel/s versus 41.66 GTexel/s, and higher pixel throughput at 36.71 GPixel/s versus 20.83 GPixel/s.

Memory bandwidth is another K20c advantage. The K20c delivers 208.0 GB/s over a 320-bit bus with 5 GB of GDDR5, while the M2090 provides 177.4 GB/s over a 384-bit bus with 6 GB. The K20c's memory clock is 1300 MHz (5.2 Gbps effective) compared to the M2090's 924 MHz (3.7 Gbps effective).

The M2090's win in OpenCL despite these theoretical disadvantages raises questions. The M2090 has more memory capacity (6 GB vs 5 GB) and a wider memory bus (384-bit vs 320-bit), which could help in certain memory-bound workloads. It also has more ROPs (48 vs 40), though fewer shading units (512 vs 2,496) and fewer TMUs (64 vs 208).

The data suggests a workload-dependent split. For tasks that scale with raw FP32 throughput or texture rate, the K20c should be the stronger part. For tasks sensitive to memory capacity or bus width, the M2090's configuration could be preferable. The OpenCL result indicates that the M2090's design wins in at least one important general-purpose compute scenario.

Architecture Differences

The two cards come from different NVIDIA architectures and different manufacturing nodes. The M2090 uses the GF110 chip based on Fermi 2.0, built on a 40 nm process at TSMC. The K20c uses the GK110 chip based on Kepler, built on a 28 nm process, also at TSMC.

The transistor counts reflect the node difference. The K20c packs 7,080 million transistors into a 561 mm² die, giving a transistor density of 12.6M per mm². The M2090 contains 3,000 million transistors on a 520 mm² die, for a density of 5.8M per mm². The Kepler chip more than doubles the transistor count of Fermi while using only a slightly larger die.

Shading unit counts differ dramatically. The M2090 has 512 shading units, 64 TMUs, and 48 ROPs. The K20c has 2,496 shading units, 208 TMUs, and 40 ROPs. That is nearly five times more shading units and more than three times more TMUs on the Kepler card, though it has fewer ROPs.

Memory subsystems are configured differently. The M2090 has 6 GB of GDDR5 on a 384-bit bus with 177.4 GB/s bandwidth. The K20c has 5 GB of GDDR5 on a 320-bit bus with 208.0 GB/s bandwidth. The K20c achieves higher bandwidth despite a narrower bus by using faster memory: 1300 MHz versus 924 MHz.

Both cards are dual-slot designs with no display outputs, consistent with their Tesla compute-focused positioning. Both use the same power connector arrangement (1x 6-pin + 1x 8-pin) and PCIe 2.0 x16 interfaces. The K20c has a lower TDP at 225 W versus the M2090's 250 W, and a lower suggested PSU rating at 550 W versus 600 W.

API support differs slightly. Both support DirectX 12 (11_0) and OpenGL 4.6. The K20c adds Vulkan 1.2.175 support, while the M2090 has no Vulkan support listed. This could matter for modern compute frameworks that rely on Vulkan.

The M2090 is physically shorter at 248 mm (9.8 inches) versus the K20c's 267 mm (10.5 inches). Production status for both is end-of-life. The M2090 was released in July 2011, and the K20c followed in November 2012. The M2090's predecessor is listed as Tesla and its successor as Tesla Kepler, while the K20c's predecessor is Tesla Fermi and its successor is Tesla Maxwell.

FAQ

Q: Which card has the higher OpenCL benchmark score?

A: The NVIDIA Tesla M2090 scores 13,075 in Geekbench OpenCL, which is 13.9% higher than the NVIDIA Tesla K20c's score of 11,479.

Q: Does the K20c have any advantages in compute throughput?

A: Yes. The K20c is rated at 3.524 TFLOPS FP32, compared to the M2090's 1,332.2 GFLOPS. It also has higher texture rate (146.8 GTexel/s vs 41.66 GTexel/s) and pixel rate (36.71 GPixel/s vs 20.83 GPixel/s).

Q: How do the memory configurations compare?

A: The M2090 has 6 GB of GDDR5 on a 384-bit bus with 177.4 GB/s bandwidth. The K20c has 5 GB of GDDR5 on a 320-bit bus with 208.0 GB/s bandwidth. The K20c's memory runs at 1300 MHz (5.2 Gbps effective) versus 924 MHz (3.7 Gbps effective) on the M2090.

Q: Which card is more power-efficient?

A: The K20c has a lower TDP at 225 W compared to the M2090's 250 W. It also has a lower suggested PSU rating of 550 W versus 600 W.

Q: Do both cards support the same APIs?

A: Both support DirectX 12 (11_0) and OpenGL 4.6. The K20c additionally supports Vulkan 1.2.175, while the M2090 has no Vulkan support listed.

Q: What are the production status and release dates?

A: Both cards are end-of-life. The M2090 was released in July 2011, and the K20c was released in November 2012.

Specification Differences

| Specification | NVIDIA Tesla M2090 | NVIDIA Tesla K20c |

|---|---|---|

| Chip | GF110 | GK110 |

| Architecture | Fermi 2.0 | Kepler |

| Process node | 40 nm | 28 nm |

| Transistors | 3,000 million | 7,080 million |

| Die size | 520 mm² | 561 mm² |

| Transistor density | 5.8M / mm² | 12.6M / mm² |

| Memory clock | 924 MHz (3.7 Gbps effective) | 1300 MHz (5.2 Gbps effective) |

| Memory size | 6 GB | 5 GB |

| Memory bus width | 384 bit | 320 bit |

| Memory bandwidth | 177.4 GB/s | 208.0 GB/s |

| Shading units | 512 | 2,496 |

| TMUs | 64 | 208 |

| ROPs | 48 | 40 |

| Pixel rate | 20.83 GPixel/s | 36.71 GPixel/s |

| Texture rate | 41.66 GTexel/s | 146.8 GTexel/s |

| FP32 | 1,332.2 GFLOPS | 3.524 TFLOPS |

| TDP | 250 W | 225 W |

| Suggested PSU | 600 W | 550 W |

| Length | 248 mm (9.8 inches) | 267 mm (10.5 inches) |

| Vulkan support | None | 1.2.175 |

| Launch MSRP | Not listed | 3,199 USD |

| Release date | July 2011 | November 2012 |

The Verdict

The benchmark data presents a paradox. The K20c is the newer architecture with dramatically higher theoretical compute specifications, yet the M2090 wins the recorded OpenCL test by 13.9%. Anyone choosing between these two for compute workloads should weigh the actual measured result against the theoretical specifications.

The M2090 is the pick if the workload resembles the Geekbench OpenCL test, since that is the only direct comparison available and it favors the Fermi card. Its larger memory capacity (6 GB vs 5 GB) and wider memory bus (384-bit vs 320-bit) provide additional headroom for memory-heavy tasks. The M2090 also has more ROPs (48 vs 40), which could help with certain rendering-related operations.

The K20c is the pick if the workload is known to scale with raw FP32 throughput, texture rate, or pixel rate. Its 3.524 TFLOPS FP32 is more than double the M2090's 1,332.2 GFLOPS. Its texture rate of 146.8 GTexel/s is more than three times the M2090's 41.66 GTexel/s. The K20c also draws less power (225 W vs 250 W) and has Vulkan support, which the M2090 lacks.

The percentile data suggests both cards are near the middle of the database distribution. The M2090 sits at the 53rd percentile and the K20c at the 51st, so neither is a top-tier performer by current standards. The K20c's nearest rival deltas are small (within 1.9% of its closest competitors), and the M2090's are similarly tight (within 1.1% of its closest competitors).

The recorded data shows one clear winner in the only head-to-head test, but the specification table shows why the K20c was positioned as a successor. Users who trust the benchmark should choose the M2090. Users who need maximum compute throughput on paper should choose the K20c. The data does not reconcile these two perspectives, so the decision ultimately depends on whether the workload resembles the OpenCL test or the theoretical peak rates.

DETAILED SPECIFICATIONS

SPECIFICATION
Tesla K20c
Tesla M2090
Core Specs
Shading Units
2,496
512 -79.5%
Shaders
2,496
512 -79.5%
TMUs
208
64 -69.2%
ROPs
40
48 +20.0%
SM Count
16
Clocks
GPU Clock
706 MHz
651 MHz
Shader Clock
1301 MHz
Memory Clock
1300 MHz 5.2 Gbps effective
924 MHz 3.7 Gbps effective
Memory
Memory Size
5 GB
6 GB
VRAM (MB)
5,120
6,144 +20.0%
Memory Type
GDDR5
GDDR5
Memory Bus
320 bit
384 bit
Bandwidth
208.0 GB/s
177.4 GB/s
Cache
L1 Cache
16 KB (per SMX)
64 KB (per SM)
L2 Cache
1280 KB
768 KB
Performance
Pixel Rate
36.71 GPixel/s
20.83 GPixel/s
Texture Rate
146.8 GTexel/s
41.66 GTexel/s
FP32 (TFLOPS)
3.524 TFLOPS
1,332.2 GFLOPS
FP64 (TFLOPS)
1,174.8 GFLOPS (1:3)
666.1 GFLOPS (1:2)
Power
TDP
225 W
250 W
TDP (W)
225
250 +11.1%
Suggested PSU
550 W
600 W
Power Connectors
1x 6-pin + 1x 8-pin
1x 6-pin + 1x 8-pin
Architecture
Architecture
Kepler
Fermi 2.0
GPU Name
GK110
GF110
Generation
Tesla Kepler (Kxx)
Tesla Fermi (x20xx)
Process Size
28 nm
40 nm
Transistors
7,080 million
3,000 million
Die Size
561 mm²
520 mm²
Foundry
TSMC
TSMC
Density
12.6M / mm²
5.8M / mm²
API Support
DirectX
12 (11_0)
12 (11_0)
OpenGL
4.6
4.6
Vulkan
1.2.175
OpenCL
3.0
1.1
CUDA
3.5
2.0
Shader Model
6.5 (5.1)
5.1
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
248 mm 9.8 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 2.0 x16
PCIe 2.0 x16
Other
Launch Price
3,199 USD
Production
End-of-life
End-of-life
Predecessor
Tesla Fermi
Tesla
Successor
Tesla Maxwell
Tesla Kepler
View Tesla K20c Details View Tesla M2090 Details