NVIDIA GeForce GTX 1070 vs NVIDIA GeForce GTX 960A Comparison

NVIDIA
GEFORCE

NVIDIA GeForce GTX 1070

CORE STATE GP104
VRAM 8 GB
CLOCK SPEED 1683 MHz
TDP 150 W
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016
VS
NVIDIA
GEFORCE

GeForce GTX 960A

CORE STATE GM107
VRAM 2 GB
CLOCK SPEED 1176 MHz
TDP 75 W
BUS WIDTH 128 bit
ARCHITECTURE Maxwell
nm
PROCESS 28 nm
LAUNCH DATE 2015

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
1,082
N/A
geekbench_metal
18,801
N/A
geekbench_opencl
44,700
11,998
geekbench_vulkan
22,121
N/A
passmark_directx_10
82
N/A
passmark_directx_11
100
N/A
passmark_directx_12
48
N/A
passmark_directx_9
197
N/A
passmark_g2d
846
N/A
passmark_g3d
13,498
N/A
passmark_gpu_compute
6,102
N/A

Analysis: NVIDIA GeForce GTX 1070 vs NVIDIA GeForce GTX 960A

Head-to-Head Benchmarks

The only directly comparable benchmark in the database is Geekbench OpenCL, and the result is decisive. The NVIDIA GeForce GTX 1070 scores 44700 points, while the NVIDIA GeForce GTX 960A scores 11998 points. This represents a 73.2% advantage for the GTX 1070, meaning the GTX 960A trails by that same margin. In raw compute terms, the GTX 1070 delivers nearly four times the OpenCL throughput of the GTX 960A.

Context from the nearest rivals reinforces the gap. The GTX 960A posts an average benchmark score of 11998, which places it in the 51st percentile of all GPUs. Its closest competitor in the database is the NVIDIA GeForce GTX 1080, which averages 11960, a delta of only 0.3%. The GTX 960A is essentially level with that card in this metric, and it sits 1.3% ahead of the AMD Radeon RX 6500 XT (11842), 2.7% ahead of the NVIDIA GeForce GTX 1660 (11680), and 3.2% ahead of the AMD Radeon RX 7800 XT (11627). This is a narrow cluster of scores, indicating that the GTX 960A occupies a specific performance tier despite its modest specifications.

The GTX 1070, by contrast, has an average benchmark score of 9780, which places it in the 47th percentile. Its nearest rivals are all professional or workstation-oriented cards: the AMD FirePro W5000 averages 9803 (0.2% ahead of the GTX 1070), the NVIDIA Quadro M2000M averages 9832 (0.5% ahead), the NVIDIA Tesla M10 averages 9724 (0.6% behind), and the NVIDIA Tesla C2070 averages 9716 (0.7% behind). Interestingly, the GTX 1070's average score across all recorded benchmarks is lower than its Geekbench OpenCL result alone, because its benchmark suite includes additional tests such as Passmark DirectX 9 (197), DirectX 10 (82), DirectX 11 (100), DirectX 12 (48), G2D (846), G3D (13498), and GPU Compute (6102), as well as Geekbench Metal (18801) and Geekbench Vulkan (22121). The GTX 960A has only one recorded benchmark, so its average equals its OpenCL score.

The head-to-head comparison is lopsided, but the two cards are from different eras and different market segments. The GTX 960A belongs to the GeForce 900A generation with a Maxwell architecture, while the GTX 1070 is a GeForce 10-series card built on Pascal. The OpenCL result alone does not capture the full range of differences, but it does establish that the GTX 1070 is the far stronger compute performer in this specific test.

Architecture Differences

The two GPUs are built on fundamentally different architectures and process nodes. The GTX 960A uses the GM107 chip, which is manufactured on a 28 nm process at TSMC. The GTX 1070 uses the GP104 chip, manufactured on a 16 nm process, also at TSMC. This process shrink is significant: the GTX 1070 packs 7,200 million transistors into a die size of 314 mm², while the GTX 960A contains 1,870 million transistors on a 148 mm² die. Transistor density tells the story clearly: the GTX 1070 achieves 22.9 million transistors per square millimeter, versus 12.6 million for the GTX 960A. The smaller process node allows the GTX 1070 to nearly quadruple the transistor count while only slightly more than doubling the die area.

Architecturally, the GTX 960A is Maxwell-based, while the GTX 1070 is Pascal-based. This generation jump brings with it a higher DirectX feature level: the GTX 960A supports DirectX 12 (11_0), while the GTX 1070 supports DirectX 12 (12_1). Both cards support OpenGL 4.6 and Vulkan 1.4, so those API capabilities are identical.

The compute resources differ dramatically. The GTX 960A has 640 shading units, 40 texture mapping units, and 16 raster output units. The GTX 1070 has 1920 shading units, 120 TMUs, and 64 ROPs. Each of these counts is exactly three or four times larger on the GTX 1070: shading units are 3 times larger, TMUs are 3 times larger, and ROPs are 4 times larger. These differences directly explain the large gap in pixel and texture throughput. The GTX 960A achieves a pixel rate of 18.82 GPixel/s and a texture rate of 47.04 GTexel/s. The GTX 1070 achieves 107.7 GPixel/s and 202.0 GTexel/s, which are 5.7 times and 4.3 times higher, respectively.

Clock speeds also favor the GTX 1070. The GTX 960A has a base clock of 1097 MHz and a boost clock of 1176 MHz. The GTX 1070 operates at a base clock of 1506 MHz and a boost clock of 1683 MHz. That is a 409 MHz higher base clock and a 507 MHz higher boost clock. Floating-point performance reflects the combination of higher clocks and more shaders: the GTX 960A delivers 1.505 TFLOPS of FP32 compute, while the GTX 1070 delivers 6.463 TFLOPS. The GTX 1070 also has a recorded FP16 figure of 101.0 GFLOPS (at a 1:64 ratio), while the GTX 960A has no FP16 figure in the database.

The GTX 1070 also supports a more capable memory subsystem. Its memory runs at 2002 MHz with an 8 Gbps effective rate, while the GTX 960A's memory runs at 1253 MHz with a 5 Gbps effective rate. The GTX 1070 offers 8 GB of GDDR5 memory on a 256 bit bus, yielding 256.3 GB/s of bandwidth. The GTX 960A offers 2 GB of GDDR5 memory on a 128 bit bus, yielding 80.19 GB/s. The bandwidth difference is 3.2 times in favor of the GTX 1070.

Where Each One Wins

The GTX 960A wins in exactly one recorded comparison category: its percentile rank. It sits at the 51st percentile of all GPUs, while the GTX 1070 sits at the 47th percentile. This is a narrow margin, but it reflects the fact that the GTX 960A's single recorded benchmark score (11998) is high relative to its specification tier, whereas the GTX 1070's average across multiple tests (9780) is pulled down by lower scores in legacy DirectX tests and compute workloads.

The GTX 960A also wins on power efficiency in one specific sense: its TDP is 75 W, which is half of the GTX 1070's 150 W. For systems with strict power limits or minimal cooling, the GTX 960A is the more manageable option. It draws no power connectors, while the GTX 1070 requires a single 8-pin connector. The GTX 1070's suggested PSU is 450 W, while the GTX 960A has no suggested PSU listed.

The GTX 1070 wins on every compute and graphics metric in the database. Its OpenCL score is 3.7 times higher. Its pixel rate, texture rate, FP32 throughput, memory bandwidth, and memory capacity are all substantially higher. It also supports more display outputs: 1x DVI, 1x HDMI 2.0, and 3x DisplayPort 1.4a, whereas the GTX 960A's display outputs are listed as "Portable Device Dependent" because it is an MXM module. For desktop use with multiple monitors, the GTX 1070 is clearly the more flexible option.

The GTX 1070 also wins on API support in one respect: it supports DirectX 12 (12_1), whereas the GTX 960A supports DirectX 12 (11_0). This difference matters for titles that use feature level 12_1 features, though both cards support the same OpenGL and Vulkan versions.

Specification Differences

The two cards differ in nearly every specification field. The GTX 960A uses the GM107 chip on a 28 nm process, while the GTX 1070 uses the GP104 chip on a 16 nm process. Transistor counts are 1,870 million versus 7,200 million, and die sizes are 148 mm² versus 314 mm². Transistor density is 12.6M per mm² versus 22.9M per mm².

Clock speeds: the GTX 960A runs at 1097 MHz base and 1176 MHz boost; the GTX 1070 runs at 1506 MHz base and 1683 MHz boost. Memory clocks are 1253 MHz (5 Gbps effective) versus 2002 MHz (8 Gbps effective). Memory capacity is 2 GB versus 8 GB, bus width is 128 bit versus 256 bit, and bandwidth is 80.19 GB/s versus 256.3 GB/s.

Compute resources: 640 shading units versus 1920, 40 TMUs versus 120, and 16 ROPs versus 64. Pixel rate is 18.82 GPixel/s versus 107.7 GPixel/s. Texture rate is 47.04 GTexel/s versus 202.0 GTexel/s. FP32 performance is 1.505 TFLOPS versus 6.463 TFLOPS. The GTX 1070 has an FP16 figure of 101.0 GFLOPS (1:64), while the GTX 960A has none.

Power and physical specifications: the GTX 960A has a TDP of 75 W, is an MXM Module with no power connectors, and uses an MXM-B (3.0) bus interface. The GTX 1070 has a TDP of 150 W, is a dual-slot card with one 8-pin connector, a suggested PSU of 450 W, and a PCIe 3.0 x16 interface. The GTX 1070 measures 267 mm in length, 112 mm in height, and 40 mm in width. The GTX 960A has no recorded dimensions.

DirectX support differs: 12 (11_0) versus 12 (12_1). Both support OpenGL 4.6 and Vulkan 1.4. The GTX 1070 has a launch MSRP of 379 USD. The GTX 960A has no launch MSRP recorded.

FAQ

Q: Which card has the higher Geekbench OpenCL score?

A: The NVIDIA GeForce GTX 1070 scores 44700, while the NVIDIA GeForce GTX 960A scores 11998. The GTX 1070 leads by 73.2%.

Q: What are the memory capacities of these two cards?

A: The GTX 960A has 2 GB of GDDR5 memory on a 128 bit bus with 80.19 GB/s bandwidth. The GTX 1070 has 8 GB of GDDR5 memory on a 256 bit bus with 256.3 GB/s bandwidth.

Q: Which card supports a higher DirectX feature level?

A: The GTX 1070 supports DirectX 12 (12_1), while the GTX 960A supports DirectX 12 (11_0). Both cards support OpenGL 4.6 and Vulkan 1.4.

Q: What is the TDP difference between the two cards?

A: The GTX 960A has a TDP of 75 W and requires no power connectors. The GTX 1070 has a TDP of 150 W and requires one 8-pin power connector, with a suggested PSU of 450 W.

Q: How do their compute resources compare?

A: The GTX 960A has 640 shading units, 40 TMUs, and 16 ROPs. The GTX 1070 has 1920 shading units, 120 TMUs, and 64 ROPs. The GTX 1070 delivers 6.463 TFLOPS of FP32 compute versus 1.505 TFLOPS for the GTX 960A.

Q: Which card has a higher percentile ranking among all GPUs?

A: The GTX 960A ranks at the 51st percentile, while the GTX 1070 ranks at the 47th percentile. The GTX 960A's average benchmark score is 11998, and the GTX 1070's average is 9780.

The Verdict

The data points to a clear split. The NVIDIA GeForce GTX 1070 is the superior performer in almost every measurable way: higher OpenCL score, higher compute throughput, higher memory bandwidth, more memory capacity, and support for a higher DirectX feature level. Its 73.2% lead in the head-to-head OpenCL benchmark is substantial, and its specification advantages in shading units, TMUs, ROPs, and clock speeds are consistent with that result.

The NVIDIA GeForce GTX 960A, however, holds a narrow advantage in percentile rank (51st versus 47th) and a significant advantage in power draw: 75 W versus 150 W. It also requires no power connectors and is an MXM module, which makes it suitable for compact or portable systems. Its single benchmark score of 11998 places it in a tight cluster with cards like the GTX 1080 and RX 6500 XT, suggesting that for its intended form factor, it is a capable performer.

Users who need maximum compute performance, larger memory capacity, and desktop connectivity should choose the GTX 1070. Users who require a low-power, connector-free module for a portable or space-constrained system should consider the GTX 960A. The GTX 1070 was released on June 9, 2016, while the GTX 960A was released on March 12, 2015, and both are end-of-life products. The GTX 1070's successor is the GeForce 20 series, while the GTX 960A has no recorded successor.

DETAILED SPECIFICATIONS

SPECIFICATION
GTX 1070
GTX 960A
Core Specs
Shading Units
1,920
640 -66.7%
Shaders
1,920
640 -66.7%
TMUs
120
40 -66.7%
ROPs
64
16 -75.0%
SM Count
15
—
Clocks
Base Clock
1506 MHz
1097 MHz
Boost Clock
1683 MHz
1176 MHz
Memory Clock
2002 MHz 8 Gbps effective
1253 MHz 5 Gbps effective
Memory
Memory Size
8 GB
2 GB
VRAM (MB)
8,192
2,048 -75.0%
Memory Type
GDDR5
GDDR5
Memory Bus
256 bit
128 bit
Bandwidth
256.3 GB/s
80.19 GB/s
Cache
L1 Cache
48 KB (per SM)
64 KB (per SMM)
L2 Cache
2 MB
2 MB
Performance
Pixel Rate
107.7 GPixel/s
18.82 GPixel/s
Texture Rate
202.0 GTexel/s
47.04 GTexel/s
FP32 (TFLOPS)
6.463 TFLOPS
1.505 TFLOPS
FP64 (TFLOPS)
202.0 GFLOPS (1:32)
47.04 GFLOPS (1:32)
FP16 (TFLOPS)
101.0 GFLOPS (1:64)
—
Power
TDP
150 W
75 W
TDP (W)
150
75 -50.0%
Suggested PSU
450 W
—
Power Connectors
1x 8-pin
None
Architecture
Architecture
Pascal
Maxwell
GPU Name
GP104
GM107
Generation
GeForce 10
GeForce 900A
Process Size
16 nm
28 nm
Transistors
7,200 million
1,870 million
Die Size
314 mm²
148 mm²
Foundry
TSMC
TSMC
Density
22.9M / mm²
12.6M / mm²
API Support
DirectX
12 (12_1)
12 (11_0)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
6.1
5.0
Shader Model
6.8
6.7 (5.1)
Physical
Slot Width
Dual-slot
MXM Module
Length
267 mm 10.5 inches
—
Height
112 mm 4.4 inches
—
Outputs
1x DVI1x HDMI 2.03x DisplayPort 1.4a
Portable Device Dependent
Bus Interface
PCIe 3.0 x16
MXM-B (3.0)
Other
Launch Price
379 USD
—
Production
End-of-life
End-of-life
Predecessor
GeForce 900
GeForce 800A
Successor
GeForce 20
—
View GeForce GTX 1070 Details View GeForce GTX 960A Details