NVIDIA A2 vs NVIDIA Quadro M6000 Comparison
NVIDIA A2
Quadro M6000
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A2 vs NVIDIA Quadro M6000
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA Quadro M6000 leads with an average benchmark score of 43301, while the NVIDIA A2 scores 34690. The M6000 sits in the 84th percentile of all GPUs, whereas the A2 is in the 79th percentile.
Q: How large is the performance gap in the Vulkan benchmark?
A: The Quadro M6000 wins the Geekbench Vulkan test by a decisive 37.9% margin, scoring 46913 versus the A2's 34023.
Q: Does the A2 have any architectural advantage over the M6000?
A: Yes. The A2 is built on the Ampere architecture with a newer 8 nm process and includes 10 RT cores and 40 tensor cores. The M6000, based on Maxwell 2.0, has no RT or tensor cores at all.
Q: Which card has more memory?
A: The A2 has 16 GB of GDDR6, while the M6000 has 12 GB of GDDR5. However, the M6000 has a wider 384-bit bus and higher memory bandwidth at 317.4 GB/s versus the A2's 200.1 GB/s.
Q: What are the power requirements for each card?
A: The M6000 has a 250 W TDP and needs a 600 W suggested PSU with one 8-pin connector. The A2 is much more efficient with a 60 W TDP, a 250 W suggested PSU, and no power connectors required.
Q: Do both cards support the same graphics APIs?
A: Both support OpenGL 4.6 and Vulkan 1.4. They differ in DirectX support: the M6000 supports DirectX 12 (12_1), while the A2 supports DirectX 12 Ultimate (12_2).
Architecture Differences
The Quadro M6000 and A2 come from completely different eras of NVIDIA's GPU design. The M6000 uses the GM200 chip on the Maxwell 2.0 architecture, fabricated on TSMC's 28 nm process. It packs 8,000 million transistors across a 601 mm² die, giving a transistor density of 13.3M per mm². The A2 uses the GA107 chip on the Ampere architecture, built on Samsung's 8 nm process. It contains 8,700 million transistors on a much smaller 200 mm² die, achieving 43.5M transistors per mm², more than three times the density of the M6000.
The compute configurations differ sharply. The M6000 has 3072 shading units, 192 TMUs, and 96 ROPs. The A2 has 1280 shading units, 40 TMUs, and 32 ROPs. Critically, the A2 adds hardware that the M6000 lacks entirely: 10 RT cores for ray tracing and 40 tensor cores for AI workloads. The A2 also supports FP16 at a 1:1 ratio with FP32, while the M6000 has no listed FP16 capability.
Memory architecture reflects their different roles. The M6000 uses 12 GB of GDDR5 on a 384-bit bus, reaching 317.4 GB/s of bandwidth. The A2 uses 16 GB of GDDR6 on a 128-bit bus, delivering 200.1 GB/s. The M6000's wider bus gives it a substantial bandwidth advantage, but the A2's newer memory type and larger capacity suit different workloads.
Physical design and interface also diverge. The M6000 is a dual-slot card, 267 mm long, requiring a 600 W PSU and one 8-pin power connector. The A2 is a single-slot card with no power connectors, a 60 W TDP, and a suggested PSU of only 250 W. The M6000 uses PCIe 3.0 x16 and offers display outputs (1x DVI, 4x DisplayPort 1.2). The A2 uses PCIe 4.0 x8 and has no display outputs at all, clearly targeting compute-only server deployments.
Production timelines reinforce the generational gap. The M6000 launched on 2015-03-20, following the Quadro Kepler line and preceding Quadro Pascal. The A2 launched on 2021-11-09, succeeding Quadro Turing and preceding Workstation Ada. Both are now end-of-life products, but they represent very different design philosophies: the M6000 is a high-throughput workstation renderer, while the A2 is a low-power accelerator for inference and edge compute.
Head-to-Head Benchmarks
The recorded head-to-head data covers two Geekbench tests, and the M6000 wins both. In Geekbench OpenCL, the M6000 scores 39688 against the A2's 35357, a 12.2% advantage. In Geekbench Vulkan, the gap widens dramatically: the M6000 scores 46913, while the A2 manages 34023, a 37.9% margin. The M6000 takes both recorded wins, 2 to 0.
The OpenCL result is modest in comparison to the Vulkan result. A 12.2% lead suggests that in general compute workloads, the A2's newer architecture partially compensates for its smaller shader count. The A2's Ampere cores, running at higher clocks (1440 MHz base, 1770 MHz boost) versus the M6000's 988 MHz base and 1114 MHz boost, help it stay competitive despite having fewer than half the shading units.
The Vulkan result tells a different story. The M6000's 37.9% lead indicates that in this API, the older card's raw throughput advantages dominate. With over twice the shading units, nearly five times the TMUs, and triple the ROPs, the M6000 simply has more parallel hardware to feed. The A2's RT and tensor cores do not appear to help in this particular test, as Geekbench Vulkan does not heavily exercise those specialized units.
Looking at the broader database context, the M6000's average score of 43301 places it just above the NVIDIA GeForce RTX 5050 Mobile (43268, 0.1% ahead) and the Quadro M6000 24 GB (43262, 0.1% ahead). It trails the RTX 4090 Mobile by 0.8%. The A2's average of 34690 sits 0.4% above the NVIDIA T1000 8 GB and the AMD Radeon HD 7970, and 1% above the TITAN V. These relative positions confirm that while the M6000 competes with modern mid-range and mobile high-end parts, the A2 aligns with much older or lower-tier hardware in raw benchmark terms.
The Verdict
The data points to a clear winner for raw compute performance: the NVIDIA Quadro M6000. It wins both recorded benchmarks, holds a higher average score (43301 versus 34690), and sits in a higher percentile (84 versus 79). For any workload that relies on traditional rasterization, pixel throughput, or general OpenCL/Vulkan compute, the M6000 is the stronger card.
However, the A2 is not without its own case. Its 16 GB of memory exceeds the M6000's 12 GB, and its Ampere architecture brings RT cores and tensor cores that the M6000 cannot offer. The A2's 60 W TDP with no power connectors makes it suitable for systems where the M6000's 250 W requirement and dual-slot footprint would be prohibitive. The A2 also supports DirectX 12 Ultimate (12_2), a newer feature level than the M6000's DirectX 12 (12_1).
The choice depends entirely on use case. If the workload is traditional GPU compute, rendering, or any task that maps to the benchmarked OpenCL and Vulkan paths, the M6000's 12.2% to 37.9% lead makes it the logical pick. If the workload involves AI inference, ray tracing, or runs in a power-constrained server chassis, the A2's specialized hardware and minimal power draw are compelling despite its lower raw scores. The M6000 wins the performance crown; the A2 wins on efficiency and feature set.
Specification Differences
| Field | NVIDIA Quadro M6000 | NVIDIA A2 |
|---|---|---|
| Architecture | Maxwell 2.0 | Ampere |
| Chip | GM200 | GA107 |
| Process Node | 28 nm | 8 nm |
| Foundry | TSMC | Samsung |
| Transistors | 8,000 million | 8,700 million |
| Die Size | 601 mm² | 200 mm² |
| Transistor Density | 13.3M / mm² | 43.5M / mm² |
| Base Clock | 988 MHz | 1440 MHz |
| Boost Clock | 1114 MHz | 1770 MHz |
| Memory Size | 12 GB | 16 GB |
| Memory Type | GDDR5 | GDDR6 |
| Memory Bus Width | 384 bit | 128 bit |
| Memory Bandwidth | 317.4 GB/s | 200.1 GB/s |
| Shading Units | 3072 | 1280 |
| TMUs | 192 | 40 |
| ROPs | 96 | 32 |
| RT Cores | None | 10 |
| Tensor Cores | None | 40 |
| Pixel Rate | 106.9 GPixel/s | 56.64 GPixel/s |
| Texture Rate | 213.9 GTexel/s | 70.80 GTexel/s |
| FP32 | 6.844 TFLOPS | 4.531 TFLOPS |
| FP16 | Not listed | 4.531 TFLOPS (1:1) |
| TDP | 250 W | 60 W |
| Slot Width | Dual-slot | Single-slot |
| Power Connectors | 1x 8-pin | None |
| Suggested PSU | 600 W | 250 W |
| Bus Interface | PCIe 3.0 x16 | PCIe 4.0 x8 |
| Display Outputs | 1x DVI, 4x DisplayPort 1.2 | No outputs |
| DirectX | 12 (12_1) | 12 Ultimate (12_2) |
| Release Date | 2015-03-20 | 2021-11-09 |
Where Each One Wins
The Quadro M6000 wins in every measured benchmark category. Its pixel rate of 106.9 GPixel/s is nearly double the A2's 56.64 GPixel/s. Its texture rate of 213.9 GTexel/s dwarfs the A2's 70.80 GTexel/s. FP32 compute stands at 6.844 TFLOPS versus 4.531 TFLOPS. Memory bandwidth is 317.4 GB/s versus 200.1 GB/s. These are the metrics that matter for traditional 3D rendering, simulation, and compute workloads that rely on raw shader throughput.
The A2 wins in efficiency and specialization. Its 60 W TDP versus 250 W means it can run in systems where the M6000 would require significant power delivery and cooling. It has no power connectors, making installation simpler. Its 16 GB of GDDR6 memory exceeds the M6000's 12 GB, which matters for large model footprints. The 10 RT cores and 40 tensor cores open capabilities that the M6000 simply does not have, and the 1:1 FP16 support enables accelerated mixed-precision work. The A2's PCIe 4.0 x8 interface is also newer than the M6000's PCIe 3.0 x16.
For a builder assembling a workstation for rendering or compute, the M6000's benchmark dominance and higher percentile ranking make it the straightforward choice. For a builder populating a server with low-power inference cards or needing ray tracing support, the A2's feature set and minimal power footprint win. The recorded data shows the M6000 as the faster card; the A2 as the more versatile and efficient one.