NVIDIA Tesla K20Xm vs NVIDIA Tesla M2090 Comparison
NVIDIA Tesla K20Xm
Tesla M2090
PERFORMANCE BENCHMARKS
Analysis: NVIDIA Tesla K20Xm vs NVIDIA Tesla M2090
The NVIDIA Tesla M2090 and NVIDIA Tesla K20Xm are both end-of-life, dual-slot compute accelerators with no display outputs, but they represent two distinct generations of NVIDIA's Tesla line. The M2090 is a Fermi 2.0 part built on a 40 nm process, while the K20Xm is a Kepler part built on a 28 nm process. The benchmark data available for direct comparison is limited to a single OpenCL test, where the K20Xm shows a decisive advantage, but the architectural differences and specification sheets reveal a more nuanced picture for different compute workloads.
Where Each One Wins
The data shows a single head-to-head benchmark result, and it is a clear victory for the K20Xm. In the Geekbench OpenCL test, the K20Xm scores 17,215, while the M2090 scores 13,075. This represents a 24% lead for the K20Xm, as indicated by the deltaPct of -24 for the M2090 (meaning the M2090 trails by that margin). This suggests that for general-purpose compute tasks that scale well with the number of streaming multiprocessors and raw FP32 throughput, the K20Xm is the vastly superior card.
However, the M2090 is not without its own merits. While it loses the available benchmark, its specification sheet shows strengths that could translate to advantages in specific, non-benchmarked workloads. The M2090 has a higher memory clock at 924 MHz (3.7 Gbps effective) compared to the K20Xm's 1300 MHz (5.2 Gbps effective), but the K20Xm's memory bandwidth is ultimately higher at 249.6 GB/s versus 177.4 GB/s. The M2090's lower memory bandwidth is a significant drawback for memory-bound tasks. The M2090 also has a higher TDP of 250 W compared to the K20Xm's 235 W, which means it draws more power while delivering less compute performance. In terms of the available data, the K20Xm wins the only direct benchmark and also wins on the key metric of memory bandwidth, making it the stronger choice for most compute scenarios.
Where the M2090 might be considered is in its compatibility with older software ecosystems or its lower transistor density; however, these are not benchmark wins. The M2090's 40 nm process and Fermi architecture are older, but the card retains a 53rd percentile ranking among all GPUs, which is nearly identical to the K20Xm's 52nd percentile. This indicates that while the M2090 loses the direct OpenCL test, its overall standing in the broader GPU landscape is comparable, suggesting that its performance in other, untested applications may be more competitive than the single benchmark suggests.
Architecture Differences
The two cards are built on fundamentally different architectures. The M2090 uses the GF110 chip based on the Fermi 2.0 architecture, while the K20Xm uses the GK110 chip based on the Kepler architecture. This generational shift is reflected in the manufacturing process: the M2090 is fabricated on TSMC's 40 nm node, while the K20Xm uses the more advanced 28 nm node. This allows the K20Xm to pack significantly more transistors into a similar die size. The M2090 has 3,000 million transistors on a 520 mm² die, resulting in a transistor density of 5.8 million per mm². The K20Xm packs 7,080 million transistors onto a 561 mm² die, achieving a density of 12.6 million per mm².
The core configuration differences are stark. The M2090 has 512 shading units, 64 texture mapping units (TMUs), and 48 ROPs. The K20Xm massively expands this with 2,688 shading units, 224 TMUs, and the same 48 ROPs. This 5.25x increase in shading units and 3.5x increase in TMUs is the primary driver behind the K20Xm's compute advantage. The pixel rate of the K20Xm is 40.99 GPixel/s, nearly double the M2090's 20.83 GPixel/s, and the texture rate is 164.0 GTexel/s versus the M2090's 41.66 GTexel/s, a nearly 4x difference.
Memory subsystems are similar in capacity but not in speed. Both cards feature 6 GB of GDDR5 memory on a 384-bit bus. The K20Xm's memory operates at 1300 MHz (5.2 Gbps effective) yielding a bandwidth of 249.6 GB/s, while the M2090's memory runs at 924 MHz (3.7 Gbps effective) for a bandwidth of 177.4 GB/s. The K20Xm also supports PCIe 3.0 x16, whereas the M2090 is limited to PCIe 2.0 x16, doubling the potential host-to-device transfer bandwidth. Both cards support DirectX 12 (11_0) and OpenGL 4.6, but only the K20Xm lists Vulkan support (version 1.2.175).
Head-to-Head Benchmarks
The only available head-to-head benchmark is the Geekbench OpenCL test. In this test, the K20Xm scores 17,215, and the M2090 scores 13,075. The K20Xm wins this test by a margin of 24%. This is a substantial gap that aligns with the massive difference in shading units (2,688 vs 512) and FP32 compute power (3.935 TFLOPS vs 1,332.2 GFLOPS). The K20Xm's raw compute throughput is nearly three times that of the M2090, so a 24% lead in a real-world OpenCL workload, while significant, actually shows that the M2090 is not completely outclassed in all aspects of the test.
The M2090's nearest rivals in the database include the NVIDIA GeForce GTX 1660 SUPER (avg score 12,986) and the AMD Radeon RX 580 (avg score 12,928), with the M2090 being 0.7% and 1.1% ahead of them, respectively. The K20Xm's nearest rivals include the NVIDIA GeForce GTX 670 (avg score 12,773) and the NVIDIA GeForce GTX 590 (avg score 12,830), with the K20Xm being 1.2% and 1.6% behind them, respectively. This places the K20Xm's average score (12,625) slightly below the M2090's average score (13,075), even though the K20Xm wins the head-to-head OpenCL test. This discrepancy highlights that the K20Xm's performance can be more variable across different applications, whereas the M2090 shows more consistent results in its specific test.
The deltaPct of -24% for the M2090 in the head-to-head test is the single most important data point for this comparison. It shows that when directly competing on a modern compute benchmark, the newer Kepler architecture is clearly superior. The K20Xm's higher memory bandwidth also contributes to this win, as OpenCL workloads often involve significant data movement. The M2090's lower bandwidth would be a bottleneck, preventing it from fully utilizing its already lower compute throughput.
FAQ
Q: Which GPU has a higher FP32 compute performance?
A: The NVIDIA Tesla K20Xm has a significantly higher FP32 performance at 3.935 TFLOPS, compared to the NVIDIA Tesla M2090's 1,332.2 GFLOPS.
Q: Do both cards have the same amount of memory?
A: Yes, both the NVIDIA Tesla M2090 and the NVIDIA Tesla K20Xm come equipped with 6 GB of GDDR5 memory on a 384-bit bus.
Q: What is the difference in memory bandwidth between the two cards?
A: The NVIDIA Tesla K20Xm offers higher memory bandwidth at 249.6 GB/s, while the NVIDIA Tesla M2090 provides 177.4 GB/s.
Q: Which card supports the newer PCIe interface?
A: The NVIDIA Tesla K20Xm supports PCIe 3.0 x16, while the NVIDIA Tesla M2090 is limited to PCIe 2.0 x16.
Q: What is the TDP of each card?
A: The NVIDIA Tesla M2090 has a TDP of 250 W, and the NVIDIA Tesla K20Xm has a slightly lower TDP of 235 W.
Q: Are there any benchmark results where the M2090 beats the K20Xm?
A: Based on the data provided, no. The only head-to-head benchmark, Geekbench OpenCL, is won by the K20Xm with a score of 17,215 versus the M2090's 13,075.
The Verdict
The data clearly favors the NVIDIA Tesla K20Xm for modern compute workloads. It wins the only available head-to-head benchmark by a 24% margin, offers more than double the memory bandwidth (249.6 GB/s vs 177.4 GB/s), and has a substantially higher transistor density (12.6M / mm² vs 5.8M / mm²) enabled by its newer 28 nm process. The K20Xm's 2,688 shading units and 3.935 TFLOPS of FP32 performance make it the superior choice for tasks that are heavy in parallel math, such as scientific simulations and deep learning inference. Its support for PCIe 3.0 also reduces data transfer bottlenecks.
The NVIDIA Tesla M2090, while older, is not entirely obsolete. Its average benchmark score (13,075) is actually higher than the K20Xm's average score (12,625), which suggests that in some specific applications, the M2090 may perform more consistently relative to the broader GPU market. Its 53rd percentile ranking versus the K20Xm's 52nd percentile reinforces this point. However, this does not translate into a direct performance win. The M2090's lower power consumption is not a factor, as its TDP (250 W) is higher than the K20Xm's (235 W).
For a user choosing between these two end-of-life accelerators, the K20Xm is the logical pick for any task where raw compute throughput and memory bandwidth are the primary constraints. The M2090 should only be considered if there is a specific software requirement that is incompatible with the Kepler architecture, as its performance potential is lower. The K20Xm's launch MSRP was 7,699 USD, but that is a historical figure. Ultimately, the K20Xm is the more capable and efficient compute card.
Specification Differences
| Specification | NVIDIA Tesla M2090 | NVIDIA Tesla K20Xm |
| :--- | :--- | :--- |
| Chip | GF110 | GK110 |
| Architecture | Fermi 2.0 | Kepler |
| Process Node | 40 nm | 28 nm |
| Transistors | 3,000 million | 7,080 million |
| Die Size | 520 mm² | 561 mm² |
| Transistor Density | 5.8M / mm² | 12.6M / mm² |
| Memory Clock | 924 MHz (3.7 Gbps effective) | 1300 MHz (5.2 Gbps effective) |
| Memory Bandwidth | 177.4 GB/s | 249.6 GB/s |
| Shading Units | 512 | 2688 |
| TMUs | 64 | 224 |
| Pixel Rate | 20.83 GPixel/s | 40.99 GPixel/s |
| Texture Rate | 41.66 GTexel/s | 164.0 GTexel/s |
| FP32 | 1,332.2 GFLOPS | 3.935 TFLOPS |
| TDP | 250 W | 235 W |
| Power Connectors | 1x 6-pin + 1x 8-pin | None listed |
| Suggested PSU | 600 W | 550 W |
| Bus Interface | PCIe 2.0 x16 | PCIe 3.0 x16 |
| Vulkan API | None listed | 1.2.175 |
| Dimensions (Length) | 248 mm (9.8 inches) | 267 mm (10.5 inches) |
| Release Date | 2011-07-24 | 2012-11-11 |
| Predecessor | Tesla | Tesla Fermi |
| Successor | Tesla Kepler | Tesla Maxwell |