GPU Comparison
NVIDIA Tesla K10
Tesla M2090
PERFORMANCE BENCHMARKS
Analysis: NVIDIA Tesla K10 vs NVIDIA Tesla M2090
The NVIDIA Tesla K10 and NVIDIA Tesla M2090 represent two distinct generations of NVIDIA's compute-focused accelerator line, with the K10 built on the Kepler architecture and the M2090 on the earlier Fermi 2.0 design. Benchmark data from Geekbench OpenCL shows a clear but modest performance edge for the newer K10, which posts a score of 14029 compared to the M2090’s 13075, a delta of 7.3% in favor of the K10. This single head-to-head result positions the K10 as the stronger raw compute performer, though the margin is not overwhelming, suggesting that architectural efficiency rather than raw specification dominance drives the difference.
Head-to-Head Benchmarks
The only available benchmark comparison, Geekbench OpenCL, gives the Tesla K10 the win with a score of 14029 against the Tesla M2090’s 13075. The 7.3% delta is notable because it reflects a generational leap in compute architecture, yet the gap is narrower than the specification sheet might suggest. The K10’s score places it in the 55th percentile of all GPUs, while the M2090 sits at the 53rd percentile, meaning both cards are near the middle of the performance distribution, but the K10 edges out slightly more than half of the tested GPUs.
Looking at the nearest rivals for context, the K10’s 14029 score is just 0.9% behind the NVIDIA GeForce GTX 680 (14150) and 1.1% ahead of the AMD Radeon RX 570X (13871). This places the K10 in a tight cluster where a few points separate competitors. The M2090, with 13075, is 0.7% behind the NVIDIA GeForce GTX 1660 SUPER (12986) and 0.9% ahead of the NVIDIA GeForce GTX 950 (13189), showing it competes with a slightly different performance tier despite its older architecture.
The 7.3% advantage for the K10 is meaningful in a compute context, but it does not represent a dominant victory. For workloads that are heavily parallel and benefit from higher shading unit counts, the K10’s advantage could widen, but the data shows a modest overall lead. The single benchmark result limits the depth of analysis, but the consistency of the K10’s position relative to its rivals, all within 1.6% of its score, suggests it is a stable performer, while the M2090’s rivals show similar tight clustering, indicating both cards are well-matched within their respective peer groups.
Architecture Differences
The two accelerators diverge fundamentally in their underlying designs. The Tesla K10 uses the GK104 chip on a 28 nm process at TSMC, packing 3,540 million transistors into a 294 mm² die, yielding a transistor density of 12.0 million per mm². In contrast, the Tesla M2090 relies on the GF110 chip on a larger 40 nm process, with 3,000 million transistors spread across a 520 mm² die, resulting in a much lower density of 5.8 million per mm². The K10’s smaller, denser die is a direct result of the newer manufacturing node, which allows more transistors in less space.
The shading unit counts tell a dramatic story. The K10 features 1,536 shading units, 128 texture mapping units (TMUs), and 32 render output units (ROPs). The M2090, by comparison, has only 512 shading units, 64 TMUs, and 48 ROPs. This means the K10 has three times the shading units and double the TMUs, but fewer ROPs. The K10’s pixel rate is 23.84 GPixel/s versus the M2090’s 20.83 GPixel/s, a 14.4% advantage, while the texture rate shows a more extreme gap: 95.36 GTexel/s for the K10 against 41.66 GTexel/s for the M2090, a 128.9% difference.
Memory configurations also differ significantly. The K10 offers 4 GB of GDDR5 on a 256-bit bus with a bandwidth of 160.0 GB/s, running at 1250 MHz (5 Gbps effective). The M2090 provides 6 GB of GDDR5 on a wider 384-bit bus, achieving 177.4 GB/s bandwidth at a lower 924 MHz (3.7 Gbps effective). The M2090’s larger memory pool and wider bus give it a 10.9% bandwidth advantage, which could benefit workloads with large datasets that exceed the K10’s 4 GB capacity.
The compute capabilities reflect the architectural shift. The K10 delivers 2.289 TFLOPS of FP32 performance, while the M2090 produces 1,332.2 GFLOPS (approximately 1.33 TFLOPS). This is a 71.8% advantage for the K10 in raw single-precision throughput, driven by the massive increase in shading units. The power profiles differ as well: the K10 has a 225 W TDP with a suggested 550 W PSU, while the M2090 draws 250 W and recommends a 600 W PSU. Both use dual-slot designs and identical power connectors (1x 6-pin + 1x 8-pin). The K10 is longer at 272 mm (10.7 inches) versus the M2090’s 248 mm (9.8 inches).
The bus interface also differs, with the K10 supporting PCIe 3.0 x16 and the M2090 limited to PCIe 2.0 x16. API support shows minor differences: both support DirectX 12 (11_0) and OpenGL 4.6, but the K10 adds Vulkan 1.2.175 support, while the M2090 has no Vulkan capability. Neither card has display outputs, confirming their compute-only purpose. The K10 was released on 2012-04-30, while the M2090 predates it by roughly nine months, launching on 2011-07-24.
FAQ
Q: Which card has higher raw compute performance in FP32?
A: The Tesla K10 delivers 2.289 TFLOPS, which is 71.8% higher than the Tesla M2090’s 1,332.2 GFLOPS. This is reflected in the Geekbench OpenCL score, where the K10 posts 14029 versus 13075 for the M2090.
Q: How do the memory configurations compare?
A: The M2090 offers more memory (6 GB) and higher bandwidth (177.4 GB/s) on a 384-bit bus, while the K10 has 4 GB on a 256-bit bus with 160.0 GB/s bandwidth. The M2090’s bandwidth advantage is 10.9%, which may benefit large dataset workloads.
Q: What are the key architectural differences?
A: The K10 uses the Kepler architecture on 28 nm with 3,540 million transistors on a 294 mm² die, while the M2090 uses Fermi 2.0 on 40 nm with 3,000 million transistors on a 520 mm² die. The K10 has 1,536 shading units versus 512 for the M2090.
Q: Which card has better texture processing performance?
A: The K10 is substantially ahead, with a texture rate of 95.36 GTexel/s compared to the M2090’s 41.66 GTexel/s, a 128.9% advantage. This stems from the K10’s 128 TMUs versus 64 on the M2090.
Q: Are there differences in power requirements?
A: Yes, the M2090 has a higher TDP of 250 W and suggests a 600 W PSU, while the K10 has a 225 W TDP and suggests a 550 W PSU. Both use dual-slot coolers and the same power connector configuration.
Q: Do these cards support the same APIs?
A: Both support DirectX 12 (11_0) and OpenGL 4.6, but the K10 adds Vulkan 1.2.175 support, which the M2090 lacks. Neither card has display outputs, as they are compute-only accelerators.
Specification Differences
The table below highlights only the fields where the two cards differ, based on the data available:
| Field | NVIDIA Tesla K10 | NVIDIA Tesla M2090 |
|-------|------------------|--------------------|
| Architecture | Kepler | Fermi 2.0 |
| Generation | Tesla Kepler (Kxx) | Tesla Fermi (x20xx) |
| Process Node | 28 nm | 40 nm |
| Transistors | 3,540 million | 3,000 million |
| Die Size | 294 mm² | 520 mm² |
| Transistor Density | 12.0M / mm² | 5.8M / mm² |
| Memory Clock | 1250 MHz (5 Gbps effective) | 924 MHz (3.7 Gbps effective) |
| Memory Size | 4 GB | 6 GB |
| Memory Bus Width | 256 bit | 384 bit |
| Memory Bandwidth | 160.0 GB/s | 177.4 GB/s |
| Shading Units | 1536 | 512 |
| TMUs | 128 | 64 |
| ROPs | 32 | 48 |
| Pixel Rate | 23.84 GPixel/s | 20.83 GPixel/s |
| Texture Rate | 95.36 GTexel/s | 41.66 GTexel/s |
| FP32 Performance | 2.289 TFLOPS | 1,332.2 GFLOPS |
| TDP | 225 W | 250 W |
| Suggested PSU | 550 W | 600 W |
| Bus Interface | PCIe 3.0 x16 | PCIe 2.0 x16 |
| Vulkan Support | 1.2.175 | null |
| Dimensions (Length) | 272 mm (10.7 inches) | 248 mm (9.8 inches) |
| Release Date | 2012-04-30 | 2011-07-24 |
| Predecessor | Tesla Fermi | Tesla |
| Successor | Tesla Maxwell | Tesla Kepler |
| Launch MSRP | 5,099 USD | null |
| Geekbench OpenCL Score | 14029 | 13075 |
| Percentile vs All GPUs | 55 | 53 |
Where Each One Wins
The Tesla K10 is the clear winner in raw compute throughput. Its 2.289 TFLOPS FP32 performance is 71.8% higher than the M2090’s 1,332.2 GFLOPS, making it the better choice for floating-point-intensive workloads such as scientific simulation, machine learning inference, and general-purpose GPU computing that relies on massive parallelism. The K10 also dominates in texture-heavy operations, with a 128.9% higher texture rate, which could benefit image processing or any workload that heavily utilizes texture fetches. Its 7.3% Geekbench OpenCL advantage confirms this edge in practical compute benchmarks.
The Tesla M2090 wins in memory capacity and bandwidth. With 6 GB versus the K10’s 4 GB, it can handle larger datasets without spilling to system memory, which is critical for workloads like large matrix operations or datasets that exceed the K10’s capacity. The M2090 also has 10.9% higher memory bandwidth (177.4 GB/s versus 160.0 GB/s), which can reduce memory-bound bottlenecks. Additionally, the M2090 has a higher ROP count (48 versus 32), giving it a 50% advantage in raster operations, though this is less relevant for compute-focused cards without display outputs.
The K10 also offers better power efficiency per unit of compute, delivering higher performance at a lower TDP (225 W versus 250 W) and a lower suggested PSU requirement (550 W versus 600 W). Its smaller physical footprint (272 mm versus 248 mm, though the K10 is longer) and newer PCIe 3.0 interface provide additional modern connectivity advantages. The K10’s Vulkan support is an extra feature that the M2090 lacks, though for compute-only cards, this may be of limited practical value.
In summary, the Tesla K10 is the superior choice for pure compute performance and efficiency, while the Tesla M2090 retains advantages in memory capacity and bandwidth that could make it preferable for specific memory-heavy workloads. The benchmark data shows the K10 ahead by 7.3%, but the M2090’s larger memory pool and wider bus mean it is not obsolete, particularly for tasks where data size matters more than raw throughput. Each card has a distinct role, and the choice between them should be guided by whether the workload is compute-bound or memory-bound.