NVIDIA A2 vs NVIDIA Quadro M6000 24 GB Comparison
NVIDIA A2
Quadro M6000 24 GB
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A2 vs NVIDIA Quadro M6000 24 GB
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA Quadro M6000 24 GB averages 43,262 points across recorded tests, placing it in the 83rd percentile of all GPUs. The NVIDIA A2 averages 34,690 points, sitting in the 79th percentile.
Q: How large is the performance gap in Vulkan workloads?
A: The Quadro M6000 24 GB scores 46,425 in Geekbench Vulkan, which is 36.5% higher than the A2's 34,023. This is the biggest delta between the two cards in any measured test.
Q: What are the memory differences between the two cards?
A: The Quadro M6000 24 GB has 24 GB of GDDR5 on a 384-bit bus delivering 317.4 GB/s of bandwidth. The A2 has 16 GB of GDDR6 on a 128-bit bus providing 200.1 GB/s.
Q: Which card has more shading units?
A: The Quadro M6000 24 GB packs 3,072 shading units, while the A2 has 1,280. The M6000 also has 192 texture mapping units and 96 ROPs, versus 40 TMUs and 32 ROPs on the A2.
Q: Does the A2 support any features the M6000 lacks?
A: Yes, the A2 includes 10 ray tracing cores and 40 tensor cores, while the M6000 has none. The A2 also supports DirectX 12 Ultimate (12_2), whereas the M6000 is limited to DirectX 12 (12_1).
Q: What is the power draw difference?
A: The Quadro M6000 24 GB has a 250 W TDP and requires a 600 W power supply with a single 8-pin connector. The A2 is rated at 60 W TDP, needing only a 250 W power supply, and draws power entirely from the PCIe slot with no external connectors.
The Verdict
The benchmark data paints a clear picture. The NVIDIA Quadro M6000 24 GB wins both recorded tests, taking 2 wins to 0 against the A2. Its average score of 43,262 is 24.7% higher than the A2 average of 34,690. Users who prioritize raw compute throughput should pick the M6000 24 GB without hesitation.
However, the A2 is the sensible choice in constrained environments. Its 60 W TDP, single-slot design, and lack of external power connectors make it dramatically easier to deploy in dense servers. The M6000's 250 W TDP and dual-slot footprint require substantial power and physical space. If the workload fits within 16 GB memory and can leverage the A2's tensor and RT cores, the newer architecture might be more future-proof for AI inference or ray tracing tasks despite its lower raw scores.
Neither card is currently in production, both are end-of-life. The M6000 has a launch MSRP of 4,999 USD, which historically reflects its professional workstation positioning. The A2's pricing is not listed in the database. The data suggests that the Quadro M6000 24 GB remains the stronger general-purpose compute card, while the A2 is the more efficient and specialized option.
Head-to-Head Benchmarks
The database records two direct comparisons. In Geekbench OpenCL, the Quadro M6000 24 GB scores 40,098 against the A2's 35,357, a 13.4% advantage. This test exercises general compute, and the M6000's larger shader array (3,072 vs 1,280) and wider memory bus (384-bit vs 128-bit) contribute to its lead.
The gap widens significantly in Geekbench Vulkan. The M6000 achieves 46,425, while the A2 manages only 34,023. That is a 36.5% swing in favor of the older card. The Vulkan result is striking because the A2 is built on the newer Ampere architecture with support for DirectX 12 Ultimate. The M6000's Maxwell architecture, despite being two generations older, delivers higher throughput in this API test.
The A2's average benchmark score of 34,690 places it within 1% of the NVIDIA T1000 8 GB (34,561) and AMD Radeon HD 7970 (34,541). Meanwhile, the M6000's 43,262 average sits 0.1% above the NVIDIA GeForce RTX 4070 SUPER (43,223) and just 0.9% below the GeForce RTX 4090 Mobile (43,667). These nearest-rival comparisons show the M6000 competing with modern mid-range cards, while the A2 aligns with older or lower-tier hardware.
The M6000 also dominates in pixel and texture rates. It outputs 106.9 GPixel/s and 213.9 GTexel/s, versus 56.64 GPixel/s and 70.80 GTexel/s for the A2. In FP32 compute, the M6000 delivers 6.844 TFLOPS against 4.531 TFLOPS for the A2, a 51% advantage.
Specification Differences
The two cards differ fundamentally in memory configuration. The M6000 offers 24 GB of GDDR5 across a 384-bit bus, yielding 317.4 GB/s bandwidth. The A2 provides 16 GB of GDDR6 across a 128-bit bus, yielding 200.1 GB/s. Despite the newer memory type, the A2's narrower bus reduces its bandwidth by 37%.
Clock speeds are reversed. The A2 runs at 1440 MHz base and 1770 MHz boost, while the M6000 is at 988 MHz base and 1114 MHz boost. The A2's higher clocks are typical of a newer process (8 nm vs 28 nm), but they do not overcome the M6000's massive shader count advantage.
Power requirements differ drastically. The M6000 consumes 250 W and needs a 600 W system PSU, drawing power from one 8-pin connector. The A2 uses only 60 W, requires a 250 W PSU, and draws power from the PCIe slot, requiring no external connector. Physical size also differs: the M6000 is dual-slot, 267 mm long (10.5 inches) and 111 mm tall (4.4 inches), while the A2 is single-slot with no dimensions recorded in the database.
The bus interface is different as well. The M6000 uses PCIe 3.0 x16, while the A2 uses PCIe 4.0 x8. Display outputs are another point of distinction: the M6000 has 1x DVI and 4x DisplayPort 1.2, while the A2 has no display outputs, indicating a compute-only orientation.
Architecture Differences
The M6000 is built on the Maxwell 2.0 architecture, using the GM200 chip. The A2 uses the Ampere architecture with the GA107 chip. The process nodes are generations apart: 28 nm at TSMC for the M6000 versus 8 nm at Samsung for the A2. This explains the transistor density gap: the A2 packs 8,700 million transistors into a 200 mm² die (43.5 million per mm²), while the M6000 fits 8,000 million into 601 mm² (13.3 million per mm²).
The A2's die is one-third the size of the M6000's die, yet it holds 700 million more transistors. That is the fundamental architectural advantage of the newer node. However, the M6000 uses its larger die to house 3,072 shading units, 192 TMUs, and 96 ROPs. The A2 has only 1,280 shading units, 40 TMUs, and 32 ROPs.
The A2 introduces hardware features absent from the M6000: 10 RT cores for ray tracing and 40 tensor cores for AI acceleration. It also supports half-precision FP16 at a 1:1 ratio with FP32, both rated at 4.531 TFLOPS. The M6000 has no FP16 support listed in the database. The A2's API support includes DirectX 12 Ultimate (12_2), while the M6000 reaches DirectX 12 (12_1). Both cards support OpenGL 4.6 and Vulkan 1.4.
The M6000 belongs to the Quadro Maxwell generation, with its predecessor in Quadro Kepler and successor in Quadro Pascal. The A2 belongs to the Workstation Ampere generation, following Quadro Turing and preceding Workstation Ada. Release dates differ by over five years: the M6000 launched in March 2016, the A2 in November 2021.
Where Each One Wins
The Quadro M6000 24 GB wins in general compute and graphics throughput. Its OpenCL lead of 13.4% and Vulkan lead of 36.5% show it excels in raw parallel workloads. The higher pixel rate (106.9 vs 56.64 GPixel/s) and texture rate (213.9 vs 70.80 GTexel/s) make it better suited for rendering tasks that scale with fill rates. Users with workloads that require 24 GB of memory will also prefer the M6000, since the A2 cannot hold nearly as much data in its 16 GB pool.
The A2 wins on efficiency and modern specialized features. Its 60 W TDP is 76% lower than the M6000's 250 W, making it far cheaper to operate from a power perspective and simpler to install. The single-slot design with no external power connectors means it can fit into space-constrained servers. The tensor cores and RT cores give it capabilities the M6000 cannot match in AI inference and ray-traced workloads. The FP16 throughput, equal to its FP32 rate, is a significant advantage for mixed-precision training.
For memory-bound tasks, the M6000's 317.4 GB/s bandwidth is uncut. For tasks that benefit from newer instructions and features, the A2's Ampere architecture and DirectX 12 Ultimate support provide a more complete feature set. The recorded data shows the M6000 leading in every benchmark, but the A2's advantages in power, size, and specialized cores are not captured by the Geekbench scores. The choice depends on whether the workload prioritizes raw throughput or modern efficiency and acceleration.