NVIDIA GeForce GTX 965M vs NVIDIA Tesla K20Xm Comparison
NVIDIA GeForce GTX 965M
Tesla K20Xm
PERFORMANCE BENCHMARKS
Analysis: NVIDIA GeForce GTX 965M vs NVIDIA Tesla K20Xm
# Head-to-Head Benchmarks
The only directly comparable benchmark in the database is Geekbench OpenCL, and it produces a decisive result. The NVIDIA Tesla K20Xm scores 17,215, while the NVIDIA GeForce GTX 965M scores 14,509. That is a delta of 15.7% in favor of the Tesla K20Xm. In practical terms, the K20Xm is roughly one-sixth faster in raw compute workloads, a meaningful margin for any task that scales well across shading units.
The GTX 965M’s average benchmark score sits at 14,404, placing it in the 56th percentile of all GPUs in the database. Its nearest rivals are all within a very tight band: the AMD Radeon RX Vega 11 at 14,385 (0.1% behind), the NVIDIA GeForce GTX TITAN at 14,373 (0.2% behind), the AMD Radeon Vega 11 at 14,352 (0.4% behind), and the Intel Iris Xe MAX Graphics at 14,315 (0.6% behind). The GTX 965M leads its closest competitor by only 0.1%, which is within measurement noise. The data suggests this is a tightly packed cluster of comparable performers.
The Tesla K20Xm, by contrast, has an average benchmark score of 12,625, which puts it in the 52nd percentile. Its nearest rivals are all slightly ahead: the AMD Radeon RX 7600M XT at 12,710 (0.7% ahead), the NVIDIA GeForce GTX 670 at 12,773 (1.2% ahead), the NVIDIA GeForce GTX 590 at 12,830 (1.6% ahead), and the AMD Radeon Pro 455 at 12,831 (1.6% ahead). Interestingly, the K20Xm’s average score is dragged down by its Geekbench Metal result of 8,035, which is far below its OpenCL score. The GTX 965M has no Metal result recorded, so the average comparison is not apples-to-apples across the same test suite.
The OpenCL head-to-head is the cleanest comparison available: the K20Xm wins that single benchmark outright. The GTX 965M records no wins in any head-to-head test, while the K20Xm takes one win. That is a 1-0 sweep in the Tesla’s favor, albeit based on a narrow set of data points.
# Architecture Differences
The two GPUs come from different NVIDIA architectures and serve different design goals. The GTX 965M uses the GM204 chip, built on Maxwell 2.0 architecture, and belongs to the GeForce 900M generation. The Tesla K20Xm uses the GK110 chip, built on Kepler architecture, and belongs to the Tesla Kepler (Kxx) generation. Both are fabricated on a 28 nm process at TSMC, but the similarities end there.
The GTX 965M packs 5,200 million transistors onto a 398 mm² die, giving it a transistor density of 13.1M per mm². The Tesla K20Xm is a much larger chip: 7,080 million transistors on a 561 mm² die, with a slightly lower density of 12.6M per mm². The K20Xm’s larger transistor budget translates directly into more compute resources: 2,688 shading units versus 1,024, 224 texture mapping units versus 64, and 48 raster output units versus 32. That is a 2.6x advantage in shading units and a 3.5x advantage in TMUs.
Clock behavior differs as well. The GTX 965M has a base clock of 924 MHz and a boost clock of 950 MHz. The K20Xm does not list base or boost clocks in the database, so the comparison cannot be made on frequency. Instead, the K20Xm’s memory clock is listed at 1,300 MHz (5.2 Gbps effective), while the GTX 965M runs its memory at 1,253 MHz (5 Gbps effective). Memory configuration is another major split: the GTX 965M has 2 GB of GDDR5 on a 128-bit bus, yielding 80.19 GB/s of bandwidth. The K20Xm has 6 GB of GDDR5 on a 384-bit bus, yielding 249.6 GB/s. That is more than 3x the bandwidth, which matters for large data sets and compute-heavy tasks.
The K20Xm also supports a higher feature set in some areas: it reaches DirectX 12 (11_0) and Vulkan 1.2.175, whereas the GTX 965M supports DirectX 12 (12_1) and Vulkan 1.4. Both support OpenGL 4.6. The GTX 965M has a higher DirectX feature level, which could benefit gaming workloads, while the K20Xm’s Vulkan version is older but still functional.
# FAQ
Q: Which GPU is faster in OpenCL compute?
A: The Tesla K20Xm scores 17,215 in Geekbench OpenCL, which is 15.7% higher than the GTX 965M’s 14,509. The K20Xm wins this benchmark outright.
Q: How do the memory configurations compare?
A: The GTX 965M has 2 GB of GDDR5 on a 128-bit bus, providing 80.19 GB/s of bandwidth. The K20Xm has 6 GB of GDDR5 on a 384-bit bus, providing 249.6 GB/s. The K20Xm has 3x the capacity and roughly 3.1x the bandwidth.
Q: Which GPU has more shading units?
A: The K20Xm has 2,688 shading units, while the GTX 965M has 1,024. The K20Xm also has 224 TMUs versus 64, and 48 ROPs versus 32.
Q: What are the power requirements for each card?
A: The K20Xm has a listed TDP of 235 W and a suggested PSU of 550 W. The GTX 965M has no TDP listed, and it draws power as an MXM module without dedicated power connectors.
Q: Which card supports newer graphics APIs?
A: The GTX 965M supports DirectX 12 (12_1) and Vulkan 1.4, while the K20Xm supports DirectX 12 (11_0) and Vulkan 1.2.175. Both support OpenGL 4.6.
Q: How do the average benchmark scores compare?
A: The GTX 965M has an average benchmark score of 14,404, placing it in the 56th percentile. The K20Xm has an average score of 12,625, placing it in the 52nd percentile. However, the K20Xm’s average includes a low Metal score of 8,035, which pulls it down.
# Specification Differences
- Chip: GM204 (GTX 965M) versus GK110 (K20Xm)
- Architecture: Maxwell 2.0 versus Kepler
- Generation: GeForce 900M versus Tesla Kepler (Kxx)
- Transistors: 5,200 million versus 7,080 million
- Die Size: 398 mm² versus 561 mm²
- Transistor Density: 13.1M / mm² versus 12.6M / mm²
- Base Clock: 924 MHz versus not listed
- Boost Clock: 950 MHz versus not listed
- Memory Clock: 1253 MHz (5 Gbps effective) versus 1300 MHz (5.2 Gbps effective)
- Memory Size: 2 GB versus 6 GB
- Memory Bus Width: 128 bit versus 384 bit
- Memory Bandwidth: 80.19 GB/s versus 249.6 GB/s
- Shading Units: 1,024 versus 2,688
- TMUs: 64 versus 224
- ROPs: 32 versus 48
- Pixel Rate: 30.40 GPixel/s versus 40.99 GPixel/s
- Texture Rate: 60.80 GTexel/s versus 164.0 GTexel/s
- FP32: 1.946 TFLOPS versus 3.935 TFLOPS
- TDP: not listed versus 235 W
- Slot Width: MXM Module versus Dual-slot
- Power Connectors: None versus not listed
- Suggested PSU: not listed versus 550 W
- Bus Interface: MXM-B (3.0) versus PCIe 3.0 x16
- Display Outputs: Portable Device Dependent versus No outputs
- DirectX: 12 (12_1) versus 12 (11_0)
- Vulkan: 1.4 versus 1.2.175
- Length: not listed versus 267 mm (10.5 inches)
- Release Date: 2015-01-08 versus 2012-11-11
- Predecessor: GeForce 800M versus Tesla Fermi
- Successor: GeForce 10 Mobile versus Tesla Maxwell
- Launch MSRP: not listed versus 7,699 USD
# Where Each One Wins
The Tesla K20Xm wins in raw compute throughput. Its FP32 output of 3.935 TFLOPS is more than double the GTX 965M’s 1.946 TFLOPS. Its texture rate of 164.0 GTexel/s versus 60.80 GTexel/s and pixel rate of 40.99 GPixel/s versus 30.40 GPixel/s reflect its larger execution engine. The K20Xm also wins on memory bandwidth by a wide margin, which is critical for data-heavy workloads like scientific simulation, large matrix operations, or any task that repeatedly streams large buffers.
The GTX 965M wins in practical system integration and API modernity. It supports DirectX 12 (12_1), a higher feature level than the K20Xm’s 12 (11_0), and Vulkan 1.4 versus 1.2.175. It is designed as an MXM module with portable device-dependent display outputs, meaning it can drive a screen directly, while the K20Xm has no display outputs at all. The GTX 965M also has no external power connector requirement, while the K20Xm needs a 550 W suggested PSU and uses a dual-slot form factor. For a laptop or compact system, the GTX 965M is the only realistic choice.
The benchmark data supports a split verdict: in OpenCL, the K20Xm leads by 15.7%. But in the overall average across all recorded tests, the GTX 965M sits at 14,404 versus 12,625, a 14% advantage in the GTX’s favor. That gap is explained by the K20Xm’s weak Metal result of 8,035, which is not a workload the GTX 965M even has recorded. The average is not a like-for-like comparison, so the OpenCL result is the more reliable indicator of relative compute performance.
# The Verdict
If the workload is compute-heavy and the system can accommodate a dual-slot, 235 W card with a 550 W PSU, the Tesla K20Xm is the clear pick. It offers 2.6x the shading units, 2x the FP32 throughput, and 3x the memory bandwidth compared to the GTX 965M. Its OpenCL score of 17,215 beats the GTX 965M by 15.7%, and its 6 GB frame buffer is far more accommodating for large datasets.
If the workload is gaming, general desktop use, or anything requiring display output, the GTX 965M is the only viable option. The K20Xm has no display outputs, so it cannot drive a monitor. The GTX 965M also supports a newer DirectX feature level (12_1 versus 11_0) and a newer Vulkan version (1.4 versus 1.2.175), which matters for modern game titles. Its MXM form factor and lack of power connectors make it far easier to integrate into a portable system.
The data does not support a single winner across all scenarios. The K20Xm wins the only head-to-head benchmark and dominates on paper specifications for compute, but the GTX 965M wins on system compatibility, API support, and overall average benchmark score. The choice depends entirely on whether the task requires display output and modern gaming features, or purely raw compute in a stationary chassis. The K20Xm is a compute accelerator with no display capability; the GTX 965M is a mobile graphics solution with broad software compatibility. Pick the former for server-style compute, pick the latter for anything with a screen attached.