GPU Comparison

NVIDIA
GEFORCE

NVIDIA Tesla K10

CORE STATE GK104
VRAM 4 GB
CLOCK SPEED
TDP 225 W
BUS WIDTH 256 bit
ARCHITECTURE Kepler
nm
PROCESS 28 nm
LAUNCH DATE 2012
VS
NVIDIA
GEFORCE

Tesla M2090

CORE STATE GF110
VRAM 6 GB
CLOCK SPEED
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Fermi 2.0
nm
PROCESS 40 nm
LAUNCH DATE 2011

PERFORMANCE BENCHMARKS

geekbench_opencl
14,029
13,075

Analysis: NVIDIA Tesla K10 vs NVIDIA Tesla M2090

The NVIDIA Tesla K10 and NVIDIA Tesla M2090 represent two distinct generations of NVIDIA's compute-focused accelerator line, with the K10 built on the Kepler architecture and the M2090 on the earlier Fermi 2.0 design. Benchmark data from Geekbench OpenCL shows a clear but modest performance edge for the newer K10, which posts a score of 14029 compared to the M2090’s 13075, a delta of 7.3% in favor of the K10. This single head-to-head result positions the K10 as the stronger raw compute performer, though the margin is not overwhelming, suggesting that architectural efficiency rather than raw specification dominance drives the difference.

Head-to-Head Benchmarks

The only available benchmark comparison, Geekbench OpenCL, gives the Tesla K10 the win with a score of 14029 against the Tesla M2090’s 13075. The 7.3% delta is notable because it reflects a generational leap in compute architecture, yet the gap is narrower than the specification sheet might suggest. The K10’s score places it in the 55th percentile of all GPUs, while the M2090 sits at the 53rd percentile, meaning both cards are near the middle of the performance distribution, but the K10 edges out slightly more than half of the tested GPUs.

Looking at the nearest rivals for context, the K10’s 14029 score is just 0.9% behind the NVIDIA GeForce GTX 680 (14150) and 1.1% ahead of the AMD Radeon RX 570X (13871). This places the K10 in a tight cluster where a few points separate competitors. The M2090, with 13075, is 0.7% behind the NVIDIA GeForce GTX 1660 SUPER (12986) and 0.9% ahead of the NVIDIA GeForce GTX 950 (13189), showing it competes with a slightly different performance tier despite its older architecture.

The 7.3% advantage for the K10 is meaningful in a compute context, but it does not represent a dominant victory. For workloads that are heavily parallel and benefit from higher shading unit counts, the K10’s advantage could widen, but the data shows a modest overall lead. The single benchmark result limits the depth of analysis, but the consistency of the K10’s position relative to its rivals, all within 1.6% of its score, suggests it is a stable performer, while the M2090’s rivals show similar tight clustering, indicating both cards are well-matched within their respective peer groups.

Architecture Differences

The two accelerators diverge fundamentally in their underlying designs. The Tesla K10 uses the GK104 chip on a 28 nm process at TSMC, packing 3,540 million transistors into a 294 mm² die, yielding a transistor density of 12.0 million per mm². In contrast, the Tesla M2090 relies on the GF110 chip on a larger 40 nm process, with 3,000 million transistors spread across a 520 mm² die, resulting in a much lower density of 5.8 million per mm². The K10’s smaller, denser die is a direct result of the newer manufacturing node, which allows more transistors in less space.

The shading unit counts tell a dramatic story. The K10 features 1,536 shading units, 128 texture mapping units (TMUs), and 32 render output units (ROPs). The M2090, by comparison, has only 512 shading units, 64 TMUs, and 48 ROPs. This means the K10 has three times the shading units and double the TMUs, but fewer ROPs. The K10’s pixel rate is 23.84 GPixel/s versus the M2090’s 20.83 GPixel/s, a 14.4% advantage, while the texture rate shows a more extreme gap: 95.36 GTexel/s for the K10 against 41.66 GTexel/s for the M2090, a 128.9% difference.

Memory configurations also differ significantly. The K10 offers 4 GB of GDDR5 on a 256-bit bus with a bandwidth of 160.0 GB/s, running at 1250 MHz (5 Gbps effective). The M2090 provides 6 GB of GDDR5 on a wider 384-bit bus, achieving 177.4 GB/s bandwidth at a lower 924 MHz (3.7 Gbps effective). The M2090’s larger memory pool and wider bus give it a 10.9% bandwidth advantage, which could benefit workloads with large datasets that exceed the K10’s 4 GB capacity.

The compute capabilities reflect the architectural shift. The K10 delivers 2.289 TFLOPS of FP32 performance, while the M2090 produces 1,332.2 GFLOPS (approximately 1.33 TFLOPS). This is a 71.8% advantage for the K10 in raw single-precision throughput, driven by the massive increase in shading units. The power profiles differ as well: the K10 has a 225 W TDP with a suggested 550 W PSU, while the M2090 draws 250 W and recommends a 600 W PSU. Both use dual-slot designs and identical power connectors (1x 6-pin + 1x 8-pin). The K10 is longer at 272 mm (10.7 inches) versus the M2090’s 248 mm (9.8 inches).

The bus interface also differs, with the K10 supporting PCIe 3.0 x16 and the M2090 limited to PCIe 2.0 x16. API support shows minor differences: both support DirectX 12 (11_0) and OpenGL 4.6, but the K10 adds Vulkan 1.2.175 support, while the M2090 has no Vulkan capability. Neither card has display outputs, confirming their compute-only purpose. The K10 was released on 2012-04-30, while the M2090 predates it by roughly nine months, launching on 2011-07-24.

FAQ

Q: Which card has higher raw compute performance in FP32?

A: The Tesla K10 delivers 2.289 TFLOPS, which is 71.8% higher than the Tesla M2090’s 1,332.2 GFLOPS. This is reflected in the Geekbench OpenCL score, where the K10 posts 14029 versus 13075 for the M2090.

Q: How do the memory configurations compare?

A: The M2090 offers more memory (6 GB) and higher bandwidth (177.4 GB/s) on a 384-bit bus, while the K10 has 4 GB on a 256-bit bus with 160.0 GB/s bandwidth. The M2090’s bandwidth advantage is 10.9%, which may benefit large dataset workloads.

Q: What are the key architectural differences?

A: The K10 uses the Kepler architecture on 28 nm with 3,540 million transistors on a 294 mm² die, while the M2090 uses Fermi 2.0 on 40 nm with 3,000 million transistors on a 520 mm² die. The K10 has 1,536 shading units versus 512 for the M2090.

Q: Which card has better texture processing performance?

A: The K10 is substantially ahead, with a texture rate of 95.36 GTexel/s compared to the M2090’s 41.66 GTexel/s, a 128.9% advantage. This stems from the K10’s 128 TMUs versus 64 on the M2090.

Q: Are there differences in power requirements?

A: Yes, the M2090 has a higher TDP of 250 W and suggests a 600 W PSU, while the K10 has a 225 W TDP and suggests a 550 W PSU. Both use dual-slot coolers and the same power connector configuration.

Q: Do these cards support the same APIs?

A: Both support DirectX 12 (11_0) and OpenGL 4.6, but the K10 adds Vulkan 1.2.175 support, which the M2090 lacks. Neither card has display outputs, as they are compute-only accelerators.

Specification Differences

The table below highlights only the fields where the two cards differ, based on the data available:

| Field | NVIDIA Tesla K10 | NVIDIA Tesla M2090 |

|-------|------------------|--------------------|

| Architecture | Kepler | Fermi 2.0 |

| Generation | Tesla Kepler (Kxx) | Tesla Fermi (x20xx) |

| Process Node | 28 nm | 40 nm |

| Transistors | 3,540 million | 3,000 million |

| Die Size | 294 mm² | 520 mm² |

| Transistor Density | 12.0M / mm² | 5.8M / mm² |

| Memory Clock | 1250 MHz (5 Gbps effective) | 924 MHz (3.7 Gbps effective) |

| Memory Size | 4 GB | 6 GB |

| Memory Bus Width | 256 bit | 384 bit |

| Memory Bandwidth | 160.0 GB/s | 177.4 GB/s |

| Shading Units | 1536 | 512 |

| TMUs | 128 | 64 |

| ROPs | 32 | 48 |

| Pixel Rate | 23.84 GPixel/s | 20.83 GPixel/s |

| Texture Rate | 95.36 GTexel/s | 41.66 GTexel/s |

| FP32 Performance | 2.289 TFLOPS | 1,332.2 GFLOPS |

| TDP | 225 W | 250 W |

| Suggested PSU | 550 W | 600 W |

| Bus Interface | PCIe 3.0 x16 | PCIe 2.0 x16 |

| Vulkan Support | 1.2.175 | null |

| Dimensions (Length) | 272 mm (10.7 inches) | 248 mm (9.8 inches) |

| Release Date | 2012-04-30 | 2011-07-24 |

| Predecessor | Tesla Fermi | Tesla |

| Successor | Tesla Maxwell | Tesla Kepler |

| Launch MSRP | 5,099 USD | null |

| Geekbench OpenCL Score | 14029 | 13075 |

| Percentile vs All GPUs | 55 | 53 |

Where Each One Wins

The Tesla K10 is the clear winner in raw compute throughput. Its 2.289 TFLOPS FP32 performance is 71.8% higher than the M2090’s 1,332.2 GFLOPS, making it the better choice for floating-point-intensive workloads such as scientific simulation, machine learning inference, and general-purpose GPU computing that relies on massive parallelism. The K10 also dominates in texture-heavy operations, with a 128.9% higher texture rate, which could benefit image processing or any workload that heavily utilizes texture fetches. Its 7.3% Geekbench OpenCL advantage confirms this edge in practical compute benchmarks.

The Tesla M2090 wins in memory capacity and bandwidth. With 6 GB versus the K10’s 4 GB, it can handle larger datasets without spilling to system memory, which is critical for workloads like large matrix operations or datasets that exceed the K10’s capacity. The M2090 also has 10.9% higher memory bandwidth (177.4 GB/s versus 160.0 GB/s), which can reduce memory-bound bottlenecks. Additionally, the M2090 has a higher ROP count (48 versus 32), giving it a 50% advantage in raster operations, though this is less relevant for compute-focused cards without display outputs.

The K10 also offers better power efficiency per unit of compute, delivering higher performance at a lower TDP (225 W versus 250 W) and a lower suggested PSU requirement (550 W versus 600 W). Its smaller physical footprint (272 mm versus 248 mm, though the K10 is longer) and newer PCIe 3.0 interface provide additional modern connectivity advantages. The K10’s Vulkan support is an extra feature that the M2090 lacks, though for compute-only cards, this may be of limited practical value.

In summary, the Tesla K10 is the superior choice for pure compute performance and efficiency, while the Tesla M2090 retains advantages in memory capacity and bandwidth that could make it preferable for specific memory-heavy workloads. The benchmark data shows the K10 ahead by 7.3%, but the M2090’s larger memory pool and wider bus mean it is not obsolete, particularly for tasks where data size matters more than raw throughput. Each card has a distinct role, and the choice between them should be guided by whether the workload is compute-bound or memory-bound.

DETAILED SPECIFICATIONS

SPECIFICATION
Tesla K10
Tesla M2090
Core Specs
Shading Units
1,536
512 -66.7%
Shaders
1,536
512 -66.7%
TMUs
128
64 -50.0%
ROPs
32
48 +50.0%
SM Count
16
Clocks
GPU Clock
745 MHz
651 MHz
Shader Clock
1301 MHz
Memory Clock
1250 MHz 5 Gbps effective
924 MHz 3.7 Gbps effective
Memory
Memory Size
4 GB
6 GB
VRAM (MB)
4,096
6,144 +50.0%
Memory Type
GDDR5
GDDR5
Memory Bus
256 bit
384 bit
Bandwidth
160.0 GB/s
177.4 GB/s
Cache
L1 Cache
16 KB (per SMX)
64 KB (per SM)
L2 Cache
512 KB
768 KB
Performance
Pixel Rate
23.84 GPixel/s
20.83 GPixel/s
Texture Rate
95.36 GTexel/s
41.66 GTexel/s
FP32 (TFLOPS)
2.289 TFLOPS
1,332.2 GFLOPS
FP64 (TFLOPS)
95.36 GFLOPS (1:24)
666.1 GFLOPS (1:2)
Power
TDP
225 W
250 W
TDP (W)
225
250 +11.1%
Suggested PSU
550 W
600 W
Power Connectors
1x 6-pin + 1x 8-pin
1x 6-pin + 1x 8-pin
Architecture
Architecture
Kepler
Fermi 2.0
GPU Name
GK104
GF110
Generation
Tesla Kepler (Kxx)
Tesla Fermi (x20xx)
Process Size
28 nm
40 nm
Transistors
3,540 million
3,000 million
Die Size
294 mm²
520 mm²
Foundry
TSMC
TSMC
Density
12.0M / mm²
5.8M / mm²
API Support
DirectX
12 (11_0)
12 (11_0)
OpenGL
4.6
4.6
Vulkan
1.2.175
OpenCL
3.0
1.1
CUDA
3.0
2.0
Shader Model
6.5 (5.1)
5.1
Physical
Slot Width
Dual-slot
Dual-slot
Length
272 mm 10.7 inches
248 mm 9.8 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 2.0 x16
Other
Launch Price
5,099 USD
Production
End-of-life
End-of-life
Predecessor
Tesla Fermi
Tesla
Successor
Tesla Maxwell
Tesla Kepler
View Tesla K10 Details View Tesla M2090 Details