NVIDIA GeForce GTX 965M vs NVIDIA Tesla M2090 Comparison
NVIDIA GeForce GTX 965M
Tesla M2090
PERFORMANCE BENCHMARKS
Analysis: NVIDIA GeForce GTX 965M vs NVIDIA Tesla M2090
Where Each One Wins
The benchmark data splits cleanly along generational and architectural lines. The NVIDIA GeForce GTX 965M wins the only recorded head-to-head comparison, and it does so decisively. In the Geekbench OpenCL test, the GTX 965M scores 14,509 points against the Tesla M2090's 13,075 points, a margin of 11 percent. That is a substantial gap for two cards that occupy similar percentile positions in the database.
The GTX 965M's advantage is rooted in its modern Maxwell 2.0 design, which delivers superior compute efficiency per clock. Its 1.946 TFLOPS of FP32 throughput outpaces the Tesla M2090's 1,332.2 GFLOPS by roughly 46 percent. This explains why the mobile chip, despite its narrower memory bus and smaller frame buffer, outperforms the older workstation-class card in a raw compute workload like OpenCL. The GTX 965M also holds a slim lead in average benchmark score: 14,404 versus 13,075. That 1,329-point gap reflects the same underlying compute advantage.
However, the Tesla M2090 is not without its own strengths. Its 6 GB of GDDR5 memory, paired with a 384-bit bus, delivers 177.4 GB/s of bandwidth. The GTX 965M manages only 80.19 GB/s over a 128-bit interface. For memory-bound workloads, the Tesla's 2.2x bandwidth advantage could flip the outcome, though no such benchmark is recorded in the database. The Tesla also carries 48 ROPs against the GTX 965M's 32, and its 250 W TDP shows it was designed for sustained compute tasks rather than battery-conscious mobile use.
The wins are lopsided in the recorded data: the GTX 965M takes 1 win, the Tesla M2090 takes 0. But the database only covers OpenCL. The Tesla's architecture, with its 512 shading units and 64 TMUs, was optimized for double-precision and memory throughput in scientific computing, not consumer gaming benchmarks. The GTX 965M, by contrast, is a mobile gaming chip with DirectX 12 (12_1) support, which the Tesla lacks (it caps at DirectX 12 (11_0)). For real-world use cases, the Tesla would likely win any test that stresses memory capacity or bandwidth, while the GTX 965M excels at latency-sensitive, shader-heavy compute.
Architecture Differences
The two GPUs represent different eras of NVIDIA design. The GTX 965M uses the GM204 chip on TSMC's 28 nm process, packing 5,200 million transistors into a 398 mm² die. That yields a transistor density of 13.1 million per square millimeter. The Tesla M2090 uses the GF110 chip on TSMC's older 40 nm node, with 3,000 million transistors spread across a larger 520 mm² die, giving a density of just 5.8 million per square millimeter. The newer process node gives the GTX 965M a clear efficiency advantage: it achieves higher performance with fewer physical resources per area.
The compute architectures diverge significantly. Maxwell 2.0, found in the GTX 965M, was designed for high throughput per shader. Its 1,024 shading units operate at a base clock of 924 MHz with a boost of 950 MHz. The Fermi 2.0 architecture in the Tesla M2090 uses 512 shading units with no recorded base or boost clock in the database, but its memory clock is 924 MHz (3.7 Gbps effective). The GTX 965M's memory runs at 1,253 MHz (5 Gbps effective). Clock speeds alone do not tell the full story: Maxwell's shader efficiency is substantially higher per clock than Fermi's, which explains the 46 percent FP32 gap despite the Tesla having half the shading units.
Memory configurations could not be more different. The GTX 965M has 2 GB of GDDR5 on a 128-bit bus, yielding 80.19 GB/s. The Tesla M2090 has 6 GB of GDDR5 on a 384-bit bus, yielding 177.4 GB/s. That is a 121 percent bandwidth advantage for the Tesla. This makes the Tesla far better suited for large datasets that exceed 2 GB, while the GTX 965M's smaller pool forces frequent data swaps in memory-heavy workloads.
The physical designs reflect their intended roles. The GTX 965M is an MXM module, using the MXM-B (3.0) interface, with no power connectors and portable-device-dependent display outputs. The Tesla M2090 is a dual-slot card measuring 248 mm (9.8 inches) long, requiring a 1x 6-pin and 1x 8-pin power connector, with a 250 W TDP and a suggested 600 W power supply. It has no display outputs at all, confirming its role as a dedicated compute accelerator. The GTX 965M supports Vulkan 1.4, while the Tesla has no Vulkan support recorded. Both support OpenGL 4.6. The GTX 965M supports DirectX 12 (12_1), the Tesla only DirectX 12 (11_0).
The Verdict
The data points to a clear choice for different workloads. The GTX 965M wins the only recorded benchmark, the Geekbench OpenCL test, by 11 percent. Its 14,509 score versus 13,075 places it 1.1 percent ahead of the Tesla's nearest rival, the AMD Radeon RX 580, and 0.7 percent behind the NVIDIA GeForce GTX 1660 SUPER. The Tesla M2090 itself sits 0.9 percent behind the GTX 950 and 1 percent ahead of the RTX 3050 Ti Mobile. These percentile positions (56th for the GTX 965M, 53rd for the Tesla) indicate both cards are mid-pack, but the GTX 965M has the edge in raw compute.
For gaming or general-purpose consumer workloads, the GTX 965M is the obvious choice. It has modern API support (DirectX 12_1, Vulkan 1.4), a smaller footprint, lower power draw, and it wins the compute benchmark. For scientific or server-side compute tasks that require more than 2 GB of memory, the Tesla M2090's 6 GB frame buffer and 177.4 GB/s bandwidth make it the better fit, despite its older architecture. The Tesla has no display outputs, so it cannot drive a monitor; it is purely an accelerator.
The GTX 965M's nearest rivals show how competitive it is: it sits 0.1 percent ahead of the AMD Radeon RX Vega 11, 0.2 percent ahead of the NVIDIA GeForce GTX TITAN, 0.4 percent ahead of the AMD Radeon Vega 11, and 0.6 percent ahead of the Intel Iris Xe MAX Graphics. These are all modern or near-modern parts, and the 2015 mobile chip holds its own. The Tesla M2090, meanwhile, trades blows with the GTX 950 and GTX 1660 SUPER, showing that Fermi's compute cores remain relevant in raw throughput but lag in efficiency.
The verdict: pick the GTX 965M for balanced performance, modern APIs, and compute tasks that fit within 2 GB. Pick the Tesla M2090 only if your workload demands 6 GB of memory and high bandwidth, and you can tolerate its power draw and lack of display outputs.
FAQ
Q: Which GPU has a higher average benchmark score?
A: The NVIDIA GeForce GTX 965M has an average benchmark score of 14,404, while the NVIDIA Tesla M2090 scores 13,075. That is a 1,329-point gap in favor of the GTX 965M.
Q: How much faster is the GTX 965M in OpenCL?
A: The GTX 965M scores 14,509 in Geekbench OpenCL versus the Tesla M2090's 13,075, a delta of 11 percent.
Q: Does the Tesla M2090 have more memory?
A: Yes, the Tesla M2090 has 6 GB of GDDR5 memory on a 384-bit bus, while the GTX 965M has 2 GB on a 128-bit bus. The Tesla also has over twice the bandwidth: 177.4 GB/s versus 80.19 GB/s.
Q: Can the Tesla M2090 output to a display?
A: No, the Tesla M2090 has no display outputs. The GTX 965M's outputs are portable-device dependent, meaning they vary by laptop.
Q: Which card supports newer APIs?
A: The GTX 965M supports DirectX 12 (12_1) and Vulkan 1.4. The Tesla M2090 supports DirectX 12 (11_0) and has no Vulkan support recorded. Both support OpenGL 4.6.
Q: What are the transistor counts of each chip?
A: The GTX 965M's GM204 chip contains 5,200 million transistors on a 398 mm² die. The Tesla M2090's GF110 chip contains 3,000 million transistors on a 520 mm² die.
Head-to-Head Benchmarks
The only recorded head-to-head benchmark is Geekbench OpenCL, and it is a clear win for the GTX 965M. The GTX 965M posts 14,509 points, the Tesla M2090 posts 13,075 points, and the delta is 11 percent in favor of the mobile chip. This result is consistent with the raw FP32 throughput figures: the GTX 965M delivers 1.946 TFLOPS, while the Tesla M2090 delivers 1,332.2 GFLOPS. The 46 percent advantage in shader throughput translates into a smaller but still decisive 11 percent win in the actual benchmark, likely due to memory bandwidth constraints on the GTX 965M.
The Tesla M2090's nearest rivals put its score in context. It sits 0.7 percent behind the NVIDIA GeForce GTX 1660 SUPER (12,986), 0.9 percent ahead of the GTX 950 (13,189), 1 percent ahead of the RTX 3050 Ti Mobile (12,940), and 1.1 percent ahead of the AMD Radeon RX 580 (12,928). These are tight margins, showing that the Tesla M2090's Fermi architecture still trades blows with much newer cards in pure compute throughput.
The GTX 965M's nearest rivals are similarly tight: it is 0.1 percent ahead of the AMD Radeon RX Vega 11 (14,385), 0.2 percent ahead of the NVIDIA GeForce GTX TITAN (14,373), 0.4 percent ahead of the AMD Radeon Vega 11 (14,352), and 0.6 percent ahead of the Intel Iris Xe MAX Graphics (14,315). These margins are within noise, but the GTX 965M comes out on top against each one.
The data shows that the GTX 965M's win over the Tesla M2090 is not a fluke. It reflects a fundamentally more efficient architecture. Maxwell 2.0 achieves higher compute density per transistor and per clock than Fermi 2.0. The Tesla's 177.4 GB/s bandwidth is its saving grace, but in the recorded benchmark, that bandwidth advantage could not compensate for the GTX 965M's superior shader throughput. If a memory-intensive benchmark existed in the database, the Tesla would likely narrow or reverse the gap, but as recorded, the GTX 965M is the clear winner.
Both cards are end-of-life products. The GTX 965M was released in January 2015, the Tesla M2090 in July 2011. The GTX 965M's successor is the GeForce 10 Mobile series, while the Tesla M2090's successor is the Tesla Kepler. The 28 nm process node and Maxwell architecture give the GTX 965M a longevity advantage that the 40 nm Fermi chip cannot match, even four years later in release terms. The benchmark results confirm that newer process technology and architecture design matter more than raw memory capacity in compute workloads.