NVIDIA A2 vs NVIDIA Tesla M60 Comparison

NVIDIA
GEFORCE

NVIDIA A2

CORE STATE GA107
VRAM 16 GB
CLOCK SPEED 1770 MHz
TDP 60 W
BUS WIDTH 128 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Tesla M60

CORE STATE GM204
VRAM 8 GB
CLOCK SPEED 1178 MHz
TDP 300 W
BUS WIDTH 256 bit
ARCHITECTURE Maxwell 2.0
nm
PROCESS 28 nm
LAUNCH DATE 2015

PERFORMANCE BENCHMARKS

geekbench_opencl
35,357
29,506
geekbench_vulkan
34,023
31,473

Analysis: NVIDIA A2 vs NVIDIA Tesla M60

The NVIDIA A2 and NVIDIA Tesla M60 represent two very different approaches to accelerator design, separated by nearly six years of GPU architecture evolution. The A2 is built on the Ampere architecture using Samsung’s 8 nm process, while the M60 uses the Maxwell 2.0 architecture on TSMC’s 28 nm node. This process gap alone explains much of the performance and efficiency divide between the two.

The A2’s chip, GA107, contains 8,700 million transistors on a 200 mm² die, yielding a transistor density of 43.5M per mm². The M60’s GM204 packs 5,200 million transistors onto a much larger 398 mm² die, resulting in a density of just 13.1M per mm². The A2 is a single-chip solution, whereas the M60 is a dual-GPU card, though the database records the M60’s specifications as a single GM204 unit. Clock speeds tell a similar story: the A2 runs at a base of 1440 MHz and boosts to 1770 MHz, while the M60 operates at a much lower base of 557 MHz but boosts to 1178 MHz. The A2’s memory clock of 1563 MHz (12.5 Gbps effective) far exceeds the M60’s 1253 MHz (5 Gbps effective).

Architecturally, the A2 supports features the M60 simply lacks. The A2 has 10 ray tracing cores and 40 tensor cores, both absent from the M60. The A2 also supports DirectX 12 Ultimate (12_2), while the M60 is limited to DirectX 12 (12_1). Both cards support OpenGL 4.6 and Vulkan 1.4. The A2’s FP16 throughput is listed as 4.531 TFLOPS (1:1 ratio), while the M60 has no recorded FP16 capability. Shader counts differ in interesting ways: the M60 has more shading units (2048 vs. 1280), more texture mapping units (128 vs. 40), and more render output units (64 vs. 32). Yet the A2 still achieves competitive raw throughput due to its higher clocks. The A2 delivers 4.531 TFLOPS FP32, while the M60 delivers 4.825 TFLOPS FP32, a small 6.5% edge for the older card. Pixel rates favor the M60 at 75.39 GPixel/s versus the A2’s 56.64 GPixel/s. Texture rates are dramatically different: the M60 hits 150.8 GTexel/s, more than double the A2’s 70.80 GTexel/s.

Power consumption is where the generation gap becomes stark. The A2 is rated at 60 W TDP with no power connectors and a suggested 250 W power supply. The M60 draws 300 W, requires a single 8-pin connector, and needs a 700 W power supply. The A2 is single-slot; the M60 is dual-slot and measures 267 mm (10.5 inches) in length. Both cards have no display outputs, making them strictly compute or virtualization accelerators. The A2 uses a PCIe 4.0 x8 interface, while the M60 uses PCIe 3.0 x16. The A2 offers 16 GB of GDDR6 memory on a 128-bit bus, delivering 200.1 GB/s bandwidth. The M60 provides 8 GB of GDDR5 on a 256-bit bus, achieving 160.4 GB/s. Despite the narrower bus, the A2’s faster memory more than compensates.

The Verdict

The benchmark data clearly favors the NVIDIA A2 in both recorded tests. In Geekbench OpenCL, the A2 scores 35357 against the M60’s 29506, a 19.8% advantage. In Geekbench Vulkan, the A2 scores 34023 against 31473, an 8.1% lead. The A2 wins both head-to-head tests, giving it a 2-0 record. Its average benchmark score of 34690 places it in the 79th percentile of all GPUs, while the M60’s average of 30490 sits in the 75th percentile.

Who should pick which depends on the workload. The A2 is the clear choice for tasks that leverage newer instruction sets, ray tracing, tensor operations, or FP16 compute. Its 16 GB memory capacity is double the M60’s 8 GB, which matters for large models or datasets. The A2 also offers superior memory bandwidth at 200.1 GB/s versus 160.4 GB/s. For power-constrained environments, the A2’s 60 W TDP is transformative compared to the M60’s 300 W requirement.

The M60 retains a niche for workloads that are heavily texture-bound or that benefit from its higher pixel and texture rates. Its 2048 shading units and 128 TMUs give it a structural advantage in certain rasterization-style compute patterns. The M60’s 4.825 TFLOPS FP32 is also slightly higher than the A2’s 4.531 TFLOPS, so pure FP32 compute that does not use newer features might see a marginal benefit from the older card. However, the M60’s lack of tensor cores, ray tracing cores, and FP16 support limits its relevance for modern AI or graphics workloads.

The database shows the M60’s nearest rivals include the NVIDIA CMP 70HX (average score 30476, delta 0%), the AMD Radeon RX 6700 (30433, delta 0.2%), the AMD Radeon RX 6800 (30095, delta 1.3%), and the NVIDIA GeForce RTX 3070 Ti (29945, delta 1.8%). These are all within 1.8% of the M60, indicating that the M60 sits in a crowded performance band. The A2’s nearest rivals are the NVIDIA T1000 8 GB (34561, delta 0.4%), the AMD Radeon HD 7970 (34541, delta 0.4%), the NVIDIA TITAN V (34355, delta 1%), and the NVIDIA RTX A1000 (34207, delta 1.4%). The A2 leads all of them, but by small margins.

For most users, the A2 is the better investment of engineering effort. It achieves higher benchmark scores, uses one-fifth the power, offers double the memory, and adds features that the M60 cannot emulate. The M60 is a legacy product from 2015, and the data reflects its age. Unless a specific workload benefits from its higher texture rate or wider memory bus, the A2 is the superior accelerator.

Head-to-Head Benchmarks

The two recorded benchmark tests both result in wins for the NVIDIA A2. In Geekbench OpenCL, the A2 scores 35357, while the M60 scores 29506. This represents a 19.8% lead for the A2. The gap is substantial and consistent with the architectural differences. The A2’s newer memory subsystem, higher clocks, and tensor core support likely contribute to this result. The M60’s 2048 shading units cannot overcome the clock speed disadvantage and older memory technology.

In Geekbench Vulkan, the gap narrows but still favors the A2. The A2 scores 34023, and the M60 scores 31473, a delta of 8.1%. Vulkan is a lower-level API that can sometimes benefit from raw shader counts, which may explain why the M60 closes some of the distance. The M60’s 128 TMUs and 64 ROPs give it more parallel rasterization hardware, which could matter in Vulkan workloads. Still, the A2’s advantage persists.

The average benchmark score difference is 4200 points (34690 for the A2 versus 30490 for the M60). This is a 13.8% gap in the A2’s favor. The percentile standings reinforce this: the A2 sits at the 79th percentile, four points above the M60’s 75th percentile. While both cards are above the median, the A2 is clearly the stronger performer in the database’s measurements.

Looking at the A2’s rival set, its closest competitor is the NVIDIA T1000 8 GB, which scores 34561, just 0.4% behind. The A2 also beats the AMD Radeon HD 7970 (34541, delta 0.4%), the NVIDIA TITAN V (34355, delta 1%), and the NVIDIA RTX A1000 (34207, delta 1.4%). These deltas are small, indicating that the A2 is at the top of a tight cluster. The M60’s nearest rival is the NVIDIA CMP 70HX at 30476, with a delta of 0%. The AMD Radeon RX 6700 trails by 0.2%, the RX 6800 by 1.3%, and the RTX 3070 Ti by 1.8%. The M60 is the leader of its own cluster, but that cluster sits well below the A2’s.

The biggest single-test win for the A2 is the OpenCL result at 19.8%. The smallest is Vulkan at 8.1%. Both are decisive. There is no benchmark in the database where the M60 wins.

Specification Differences

The two cards differ in nearly every measurable specification. The A2 uses the GA107 chip on the Ampere architecture, fabricated by Samsung on an 8 nm process. The M60 uses the GM204 chip on Maxwell 2.0, fabricated by TSMC on a 28 nm process. Transistor counts are 8,700 million for the A2 versus 5,200 million for the M60. Die sizes are 200 mm² for the A2 and 398 mm² for the M60. Transistor density is 43.5M per mm² for the A2, compared to 13.1M per mm² for the M60.

Clock speeds show the A2’s advantage: base 1440 MHz versus 557 MHz, boost 1770 MHz versus 1178 MHz. Memory clocks are 1563 MHz (12.5 Gbps effective) for the A2 and 1253 MHz (5 Gbps effective) for the M60. Memory capacity is 16 GB of GDDR6 on the A2 versus 8 GB of GDDR5 on the M60. Bus widths are 128-bit for the A2 and 256-bit for the M60. Bandwidth is 200.1 GB/s for the A2 and 160.4 GB/s for the M60.

Compute resources differ significantly. The A2 has 1280 shading units, 40 TMUs, 32 ROPs, 10 ray tracing cores, and 40 tensor cores. The M60 has 2048 shading units, 128 TMUs, 64 ROPs, and no ray tracing or tensor cores. Pixel rate is 56.64 GPixel/s for the A2 and 75.39 GPixel/s for the M60. Texture rate is 70.80 GTexel/s for the A2 and 150.8 GTexel/s for the M60. FP32 performance is 4.531 TFLOPS for the A2 and 4.825 TFLOPS for the M60. The A2 has FP16 at 4.531 TFLOPS (1:1), while the M60 has no recorded FP16.

Power and physical specifications are starkly different. The A2 has a 60 W TDP, single-slot width, no power connectors, and a suggested 250 W power supply. The M60 has a 300 W TDP, dual-slot width, one 8-pin power connector, and a suggested 700 W power supply. The M60 is 267 mm (10.5 inches) long. Both cards have no display outputs. The A2 uses PCIe 4.0 x8, the M60 uses PCIe 3.0 x16.

API support differs in DirectX: the A2 supports 12 Ultimate (12_2), while the M60 supports 12 (12_1). Both support OpenGL 4.6 and Vulkan 1.4. Release dates are 2021-11-09 for the A2 and 2015-08-29 for the M60. Both are end-of-life products. The A2’s predecessor is Quadro Turing and its successor is Workstation Ada. The M60’s predecessor is Tesla Kepler and its successor is Tesla Pascal.

FAQ

Q: Which card has higher average benchmark scores?

A: The NVIDIA A2 has an average benchmark score of 34690, while the NVIDIA Tesla M60 has an average score of 30490. The A2 leads by 4200 points.

Q: How much faster is the A2 in OpenCL?

A: The A2 scores 35357 in Geekbench OpenCL, compared to the M60’s 29506. This is a 19.8% advantage for the A2.

Q: Does the M60 have any advantages in raw compute resources?

A: Yes. The M60 has 2048 shading units, 128 TMUs, and 64 ROPs, compared to the A2’s 1280 shading units, 40 TMUs, and 32 ROPs. The M60 also has a higher FP32 rating at 4.825 TFLOPS versus the A2’s 4.531 TFLOPS.

Q: What memory configurations do the two cards use?

A: The A2 has 16 GB of GDDR6 on a 128-bit bus, delivering 200.1 GB/s. The M60 has 8 GB of GDDR5 on a 256-bit bus, delivering 160.4 GB/s.

Q: Which card supports ray tracing and tensor cores?

A: Only the NVIDIA A2. It has 10 ray tracing cores and 40 tensor cores. The Tesla M60 has neither.

Q: How do their power requirements compare?

A: The A2 has a 60 W TDP and requires no power connectors, with a suggested 250 W power supply. The M60 has a 300 W TDP, requires one 8-pin connector, and needs a 700 W power supply.

DETAILED SPECIFICATIONS

SPECIFICATION
A2
Tesla M60
Core Specs
Shading Units
1,280
2,048 +60.0%
Shaders
1,280
2,048 +60.0%
TMUs
40
128 +220.0%
ROPs
32
64 +100.0%
SM Count
10
Clocks
Base Clock
1440 MHz
557 MHz
Boost Clock
1770 MHz
1178 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
1253 MHz 5 Gbps effective
Memory
Memory Size
16 GB
8 GB
VRAM (MB)
16,384
8,192 -50.0%
Memory Type
GDDR6
GDDR5
Memory Bus
128 bit
256 bit
Bandwidth
200.1 GB/s
160.4 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SMM)
L2 Cache
2 MB
2 MB
Performance
Pixel Rate
56.64 GPixel/s
75.39 GPixel/s
Texture Rate
70.80 GTexel/s
150.8 GTexel/s
FP32 (TFLOPS)
4.531 TFLOPS
4.825 TFLOPS
FP64 (TFLOPS)
70.80 GFLOPS (1:64)
150.8 GFLOPS (1:32)
FP16 (TFLOPS)
4.531 TFLOPS (1:1)
AI/RT
RT Cores
10
Tensor Cores
40
Power
TDP
60 W
300 W
TDP (W)
60
300 +400.0%
Suggested PSU
250 W
700 W
Power Connectors
None
1x 8-pin
Architecture
Architecture
Ampere
Maxwell 2.0
GPU Name
GA107
GM204
Generation
Workstation Ampere (Ax000)
Tesla Maxwell (Mxx)
Process Size
8 nm
28 nm
Transistors
8,700 million
5,200 million
Die Size
200 mm²
398 mm²
Foundry
Samsung
TSMC
Density
43.5M / mm²
13.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
5.2
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Dual-slot
Length
267 mm 10.5 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x8
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Quadro Turing
Tesla Kepler
Successor
Workstation Ada
Tesla Pascal
View A2 Details View Tesla M60 Details