NVIDIA GeForce GTX 965M vs NVIDIA Tesla K20Xm Comparison

NVIDIA
GEFORCE

NVIDIA GeForce GTX 965M

CORE STATE GM204
VRAM 2 GB
CLOCK SPEED 950 MHz
TDP —
BUS WIDTH 128 bit
ARCHITECTURE Maxwell 2.0
nm
PROCESS 28 nm
LAUNCH DATE 2015
VS
NVIDIA
GEFORCE

Tesla K20Xm

CORE STATE GK110
VRAM 6 GB
CLOCK SPEED —
TDP 235 W
BUS WIDTH 384 bit
ARCHITECTURE Kepler
nm
PROCESS 28 nm
LAUNCH DATE 2012

PERFORMANCE BENCHMARKS

geekbench_opencl
14,509
17,215
geekbench_vulkan
14,299
N/A
geekbench_metal
N/A
8,035

Analysis: NVIDIA GeForce GTX 965M vs NVIDIA Tesla K20Xm

# Head-to-Head Benchmarks

The only directly comparable benchmark in the database is Geekbench OpenCL, and it produces a decisive result. The NVIDIA Tesla K20Xm scores 17,215, while the NVIDIA GeForce GTX 965M scores 14,509. That is a delta of 15.7% in favor of the Tesla K20Xm. In practical terms, the K20Xm is roughly one-sixth faster in raw compute workloads, a meaningful margin for any task that scales well across shading units.

The GTX 965M’s average benchmark score sits at 14,404, placing it in the 56th percentile of all GPUs in the database. Its nearest rivals are all within a very tight band: the AMD Radeon RX Vega 11 at 14,385 (0.1% behind), the NVIDIA GeForce GTX TITAN at 14,373 (0.2% behind), the AMD Radeon Vega 11 at 14,352 (0.4% behind), and the Intel Iris Xe MAX Graphics at 14,315 (0.6% behind). The GTX 965M leads its closest competitor by only 0.1%, which is within measurement noise. The data suggests this is a tightly packed cluster of comparable performers.

The Tesla K20Xm, by contrast, has an average benchmark score of 12,625, which puts it in the 52nd percentile. Its nearest rivals are all slightly ahead: the AMD Radeon RX 7600M XT at 12,710 (0.7% ahead), the NVIDIA GeForce GTX 670 at 12,773 (1.2% ahead), the NVIDIA GeForce GTX 590 at 12,830 (1.6% ahead), and the AMD Radeon Pro 455 at 12,831 (1.6% ahead). Interestingly, the K20Xm’s average score is dragged down by its Geekbench Metal result of 8,035, which is far below its OpenCL score. The GTX 965M has no Metal result recorded, so the average comparison is not apples-to-apples across the same test suite.

The OpenCL head-to-head is the cleanest comparison available: the K20Xm wins that single benchmark outright. The GTX 965M records no wins in any head-to-head test, while the K20Xm takes one win. That is a 1-0 sweep in the Tesla’s favor, albeit based on a narrow set of data points.

# Architecture Differences

The two GPUs come from different NVIDIA architectures and serve different design goals. The GTX 965M uses the GM204 chip, built on Maxwell 2.0 architecture, and belongs to the GeForce 900M generation. The Tesla K20Xm uses the GK110 chip, built on Kepler architecture, and belongs to the Tesla Kepler (Kxx) generation. Both are fabricated on a 28 nm process at TSMC, but the similarities end there.

The GTX 965M packs 5,200 million transistors onto a 398 mm² die, giving it a transistor density of 13.1M per mm². The Tesla K20Xm is a much larger chip: 7,080 million transistors on a 561 mm² die, with a slightly lower density of 12.6M per mm². The K20Xm’s larger transistor budget translates directly into more compute resources: 2,688 shading units versus 1,024, 224 texture mapping units versus 64, and 48 raster output units versus 32. That is a 2.6x advantage in shading units and a 3.5x advantage in TMUs.

Clock behavior differs as well. The GTX 965M has a base clock of 924 MHz and a boost clock of 950 MHz. The K20Xm does not list base or boost clocks in the database, so the comparison cannot be made on frequency. Instead, the K20Xm’s memory clock is listed at 1,300 MHz (5.2 Gbps effective), while the GTX 965M runs its memory at 1,253 MHz (5 Gbps effective). Memory configuration is another major split: the GTX 965M has 2 GB of GDDR5 on a 128-bit bus, yielding 80.19 GB/s of bandwidth. The K20Xm has 6 GB of GDDR5 on a 384-bit bus, yielding 249.6 GB/s. That is more than 3x the bandwidth, which matters for large data sets and compute-heavy tasks.

The K20Xm also supports a higher feature set in some areas: it reaches DirectX 12 (11_0) and Vulkan 1.2.175, whereas the GTX 965M supports DirectX 12 (12_1) and Vulkan 1.4. Both support OpenGL 4.6. The GTX 965M has a higher DirectX feature level, which could benefit gaming workloads, while the K20Xm’s Vulkan version is older but still functional.

# FAQ

Q: Which GPU is faster in OpenCL compute?

A: The Tesla K20Xm scores 17,215 in Geekbench OpenCL, which is 15.7% higher than the GTX 965M’s 14,509. The K20Xm wins this benchmark outright.

Q: How do the memory configurations compare?

A: The GTX 965M has 2 GB of GDDR5 on a 128-bit bus, providing 80.19 GB/s of bandwidth. The K20Xm has 6 GB of GDDR5 on a 384-bit bus, providing 249.6 GB/s. The K20Xm has 3x the capacity and roughly 3.1x the bandwidth.

Q: Which GPU has more shading units?

A: The K20Xm has 2,688 shading units, while the GTX 965M has 1,024. The K20Xm also has 224 TMUs versus 64, and 48 ROPs versus 32.

Q: What are the power requirements for each card?

A: The K20Xm has a listed TDP of 235 W and a suggested PSU of 550 W. The GTX 965M has no TDP listed, and it draws power as an MXM module without dedicated power connectors.

Q: Which card supports newer graphics APIs?

A: The GTX 965M supports DirectX 12 (12_1) and Vulkan 1.4, while the K20Xm supports DirectX 12 (11_0) and Vulkan 1.2.175. Both support OpenGL 4.6.

Q: How do the average benchmark scores compare?

A: The GTX 965M has an average benchmark score of 14,404, placing it in the 56th percentile. The K20Xm has an average score of 12,625, placing it in the 52nd percentile. However, the K20Xm’s average includes a low Metal score of 8,035, which pulls it down.

# Specification Differences

  • Chip: GM204 (GTX 965M) versus GK110 (K20Xm)
  • Architecture: Maxwell 2.0 versus Kepler
  • Generation: GeForce 900M versus Tesla Kepler (Kxx)
  • Transistors: 5,200 million versus 7,080 million
  • Die Size: 398 mm² versus 561 mm²
  • Transistor Density: 13.1M / mm² versus 12.6M / mm²
  • Base Clock: 924 MHz versus not listed
  • Boost Clock: 950 MHz versus not listed
  • Memory Clock: 1253 MHz (5 Gbps effective) versus 1300 MHz (5.2 Gbps effective)
  • Memory Size: 2 GB versus 6 GB
  • Memory Bus Width: 128 bit versus 384 bit
  • Memory Bandwidth: 80.19 GB/s versus 249.6 GB/s
  • Shading Units: 1,024 versus 2,688
  • TMUs: 64 versus 224
  • ROPs: 32 versus 48
  • Pixel Rate: 30.40 GPixel/s versus 40.99 GPixel/s
  • Texture Rate: 60.80 GTexel/s versus 164.0 GTexel/s
  • FP32: 1.946 TFLOPS versus 3.935 TFLOPS
  • TDP: not listed versus 235 W
  • Slot Width: MXM Module versus Dual-slot
  • Power Connectors: None versus not listed
  • Suggested PSU: not listed versus 550 W
  • Bus Interface: MXM-B (3.0) versus PCIe 3.0 x16
  • Display Outputs: Portable Device Dependent versus No outputs
  • DirectX: 12 (12_1) versus 12 (11_0)
  • Vulkan: 1.4 versus 1.2.175
  • Length: not listed versus 267 mm (10.5 inches)
  • Release Date: 2015-01-08 versus 2012-11-11
  • Predecessor: GeForce 800M versus Tesla Fermi
  • Successor: GeForce 10 Mobile versus Tesla Maxwell
  • Launch MSRP: not listed versus 7,699 USD

# Where Each One Wins

The Tesla K20Xm wins in raw compute throughput. Its FP32 output of 3.935 TFLOPS is more than double the GTX 965M’s 1.946 TFLOPS. Its texture rate of 164.0 GTexel/s versus 60.80 GTexel/s and pixel rate of 40.99 GPixel/s versus 30.40 GPixel/s reflect its larger execution engine. The K20Xm also wins on memory bandwidth by a wide margin, which is critical for data-heavy workloads like scientific simulation, large matrix operations, or any task that repeatedly streams large buffers.

The GTX 965M wins in practical system integration and API modernity. It supports DirectX 12 (12_1), a higher feature level than the K20Xm’s 12 (11_0), and Vulkan 1.4 versus 1.2.175. It is designed as an MXM module with portable device-dependent display outputs, meaning it can drive a screen directly, while the K20Xm has no display outputs at all. The GTX 965M also has no external power connector requirement, while the K20Xm needs a 550 W suggested PSU and uses a dual-slot form factor. For a laptop or compact system, the GTX 965M is the only realistic choice.

The benchmark data supports a split verdict: in OpenCL, the K20Xm leads by 15.7%. But in the overall average across all recorded tests, the GTX 965M sits at 14,404 versus 12,625, a 14% advantage in the GTX’s favor. That gap is explained by the K20Xm’s weak Metal result of 8,035, which is not a workload the GTX 965M even has recorded. The average is not a like-for-like comparison, so the OpenCL result is the more reliable indicator of relative compute performance.

# The Verdict

If the workload is compute-heavy and the system can accommodate a dual-slot, 235 W card with a 550 W PSU, the Tesla K20Xm is the clear pick. It offers 2.6x the shading units, 2x the FP32 throughput, and 3x the memory bandwidth compared to the GTX 965M. Its OpenCL score of 17,215 beats the GTX 965M by 15.7%, and its 6 GB frame buffer is far more accommodating for large datasets.

If the workload is gaming, general desktop use, or anything requiring display output, the GTX 965M is the only viable option. The K20Xm has no display outputs, so it cannot drive a monitor. The GTX 965M also supports a newer DirectX feature level (12_1 versus 11_0) and a newer Vulkan version (1.4 versus 1.2.175), which matters for modern game titles. Its MXM form factor and lack of power connectors make it far easier to integrate into a portable system.

The data does not support a single winner across all scenarios. The K20Xm wins the only head-to-head benchmark and dominates on paper specifications for compute, but the GTX 965M wins on system compatibility, API support, and overall average benchmark score. The choice depends entirely on whether the task requires display output and modern gaming features, or purely raw compute in a stationary chassis. The K20Xm is a compute accelerator with no display capability; the GTX 965M is a mobile graphics solution with broad software compatibility. Pick the former for server-style compute, pick the latter for anything with a screen attached.

DETAILED SPECIFICATIONS

SPECIFICATION
GTX 965M
Tesla K20Xm
Core Specs
Shading Units
1,024
2,688 +162.5%
Shaders
1,024
2,688 +162.5%
TMUs
64
224 +250.0%
ROPs
32
48 +50.0%
Clocks
Base Clock
924 MHz
—
Boost Clock
950 MHz
—
GPU Clock
—
732 MHz
Memory Clock
1253 MHz 5 Gbps effective
1300 MHz 5.2 Gbps effective
Memory
Memory Size
2 GB
6 GB
VRAM (MB)
2,048
6,144 +200.0%
Memory Type
GDDR5
GDDR5
Memory Bus
128 bit
384 bit
Bandwidth
80.19 GB/s
249.6 GB/s
Cache
L1 Cache
48 KB (per SMM)
16 KB (per SMX)
L2 Cache
1024 KB
1536 KB
Performance
Pixel Rate
30.40 GPixel/s
40.99 GPixel/s
Texture Rate
60.80 GTexel/s
164.0 GTexel/s
FP32 (TFLOPS)
1.946 TFLOPS
3.935 TFLOPS
FP64 (TFLOPS)
60.80 GFLOPS (1:32)
1,311.7 GFLOPS (1:3)
Power
TDP
—
235 W
TDP (W)
—
235
Suggested PSU
—
550 W
Power Connectors
None
—
Architecture
Architecture
Maxwell 2.0
Kepler
GPU Name
GM204
GK110
Generation
GeForce 900M
Tesla Kepler (Kxx)
Process Size
28 nm
28 nm
Transistors
5,200 million
7,080 million
Die Size
398 mm²
561 mm²
Foundry
TSMC
TSMC
Density
13.1M / mm²
12.6M / mm²
API Support
DirectX
12 (12_1)
12 (11_0)
OpenGL
4.6
4.6
Vulkan
1.4
1.2.175
OpenCL
3.0
3.0
CUDA
5.2
3.5
Shader Model
6.8
6.5 (5.1)
Physical
Slot Width
MXM Module
Dual-slot
Length
—
267 mm 10.5 inches
Outputs
Portable Device Dependent
No outputs
Bus Interface
MXM-B (3.0)
PCIe 3.0 x16
Other
Launch Price
—
7,699 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 800M
Tesla Fermi
Successor
GeForce 10 Mobile
Tesla Maxwell
View GeForce GTX 965M Details View Tesla K20Xm Details