NVIDIA A10G vs NVIDIA Quadro P6000 Comparison
NVIDIA A10G
Quadro P6000
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A10G vs NVIDIA Quadro P6000
Where Each One Wins
The benchmark data presents a clear and unambiguous picture: the NVIDIA A10G wins in every recorded test. Across the two head-to-head benchmarks in the database, Geekbench OpenCL and Geekbench Vulkan, the A10G takes both victories. The Quadro P6000 records zero wins in this comparison. This is not a close contest where each card excels in a particular workload; the A10G dominates both compute-oriented and graphics-oriented API tests.
The Geekbench OpenCL benchmark, which typically reflects general-purpose compute throughput, shows the A10G at 158063 points versus the P6000's 66382 points. That is a delta of 138.1 percent in favor of the A10G. The Vulkan test, which leans on graphics pipeline performance and driver efficiency, tells a similar story: the A10G scores 145863, while the P6000 manages 73590. This represents a 98.2 percent advantage for the A10G.
What is notable is the consistency of the margin. The A10G does not merely edge ahead in one API and fall behind in another; it doubles or nearly doubles the P6000's performance in both. This suggests that the A10G is not tuned for a specific workload type but rather offers a broad generational uplift. For users running OpenCL-based compute tasks, the A10G is the clear choice. For users running Vulkan-based rendering or GPU-accelerated graphics, the A10G again takes the lead, though the margin shrinks compared to OpenCL.
The percentile data reinforces this split. The A10G sits in the 97th percentile of all GPUs in the database, while the P6000 sits in the 90th percentile. The seven-percentile gap indicates that the A10G belongs in a higher performance tier overall, even accounting for the fact that both cards are now listed as end-of-life products.
Architecture Differences
The architectural gap between these two cards spans two full generations of NVIDIA GPU design. The A10G uses the GA102 chip on the Ampere architecture, manufactured on an 8 nm process at Samsung. The P6000 uses the GP102 chip on the Pascal architecture, manufactured on a 16 nm process at TSMC. That process node difference is significant: the A10G packs 28,300 million transistors onto a 628 mm² die, while the P6000 contains 11,800 million transistors on a 471 mm² die. Transistor density tells the story clearly: the A10G achieves 45.1 million transistors per square millimeter, versus 25.1 million for the P6000.
The compute resources are equally divergent. The A10G carries 9216 shading units, 288 texture mapping units, and 96 raster output units. The P6000 has 3840 shading units, 240 TMUs, and 96 ROPs. The ROP count is identical, but the shading and texturing hardware is massively different. The A10G also includes 72 ray tracing cores and 288 tensor cores, hardware that is entirely absent from the P6000. The Pascal architecture predates both ray tracing acceleration and tensor core matrix math, so the P6000 has no equivalent features.
Clock speeds tell a nuanced story. The P6000 runs at a base clock of 1506 MHz and a boost of 1645 MHz, which is higher than the A10G's 1320 MHz base and 1710 MHz boost. The A10G boosts higher but idles lower. The memory clocks also differ: the A10G runs its GDDR6 memory at 1563 MHz (12.5 Gbps effective), while the P6000 runs GDDR5X at 1127 MHz (9 Gbps effective). Both cards have 24 GB of memory on a 384-bit bus, but the A10G's newer memory technology yields 600.2 GB/s of bandwidth versus 432.8 GB/s for the P6000.
The feature set diverges further. The A10G supports DirectX 12 Ultimate (12_2), while the P6000 only reaches DirectX 12 (12_1). Both cards support OpenGL 4.6 and Vulkan 1.4. The A10G uses a PCIe 4.0 x16 interface, while the P6000 is limited to PCIe 3.0 x16. The A10G has no display outputs, making it a pure compute accelerator, while the P6000 offers 1x DVI and 4x DisplayPort 1.4a outputs for direct display connectivity.
Power characteristics also separate the two. The A10G draws a 150 W TDP and uses an 8-pin EPS connector with a suggested 450 W power supply. The P6000 draws 250 W, uses a single 8-pin connector, and suggests a 600 W power supply. The A10G is single-slot, while the P6000 is dual-slot. Both cards measure 267 mm in length, with the A10G at 112 mm height and the P6000 at 111 mm.
Head-to-Head Benchmarks
The Geekbench OpenCL result is the single largest margin in this comparison. The A10G scores 158063, and the P6000 scores 66382. The delta is 138.1 percent, meaning the A10G more than doubles the P6000's output. This is consistent with the raw FP32 compute figures in the database: the A10G delivers 31.52 TFLOPS of FP32 performance, while the P6000 delivers 12.63 TFLOPS. The A10G also offers FP16 at a 1:1 ratio (31.52 TFLOPS), while the P6000's FP16 is severely limited at 197.4 GFLOPS with a 1:64 ratio. In compute-heavy OpenCL workloads, that FP16 capability alone can be decisive.
The Geekbench Vulkan result shows a narrower but still decisive gap. The A10G scores 145863 versus the P6000's 73590, a delta of 98.2 percent. The Vulkan test exercises the graphics pipeline, where the P6000's higher base clock (1506 MHz versus 1320 MHz) and identical ROP count help it remain competitive. The pixel rates are close: the A10G delivers 164.2 GPixel/s, and the P6000 delivers 157.9 GPixel/s. The texture rates differ more: 492.5 GTexel/s for the A10G versus 394.8 GTexel/s for the P6000. Yet even with those closer rasterization metrics, the A10G still nearly doubles the P6000's Vulkan score, likely due to the much larger shading unit count and the presence of dedicated ray tracing hardware.
Context from the nearest rivals strengthens the interpretation. The A10G's average benchmark score is 151963, placing it just 1.1 percent ahead of the Tesla V100 PCIe 32 GB (150305) and 9.3 percent ahead of the AMD Instinct MI100 (139035). It trails the AMD Radeon Pro W6800X by 5.4 percent and the A100 PCIe 40 GB by 6.5 percent. The P6000's average score is 69986, which sits within 0.2 percent of the AMD Radeon Pro WX 8200 (69870) and within 1.4 percent of the NVIDIA CMP 90HX (69000). The P6000's closest competitors are all within a 1.4 percent band, while the A10G's nearest rivals span a wider 15.8 percent range. This indicates that the A10G occupies a more competitive tier where small shifts in workload can change rankings, while the P6000 is firmly anchored in its performance class.
The Verdict
The data supports a straightforward conclusion: the NVIDIA A10G is the superior card in every measured dimension. It wins both head-to-head benchmarks by margins of 98.2 percent and 138.1 percent. It delivers more than double the FP32 compute throughput (31.52 TFLOPS versus 12.63 TFLOPS), more than double the FP16 throughput (31.52 TFLOPS versus 197.4 GFLOPS), and 38.7 percent more memory bandwidth (600.2 GB/s versus 432.8 GB/s). It adds ray tracing cores and tensor cores that the P6000 lacks entirely. It does all of this at a lower TDP (150 W versus 250 W) and in a single-slot form factor.
Users who need a compute accelerator for OpenCL-heavy workloads, machine learning inference, or any task that leverages FP16 or tensor operations should choose the A10G without hesitation. The 138.1 percent OpenCL margin is the strongest signal in the entire dataset. Users who need Vulkan performance for GPU-accelerated rendering will also find the A10G superior, though the 98.2 percent margin is slightly less extreme. The A10G's lack of display outputs means it cannot drive monitors directly, but the database does not record any headless compute penalty, and its PCIe 4.0 interface ensures modern platform compatibility.
The P6000 retains one practical advantage: display connectivity. It offers 1x DVI and 4x DisplayPort 1.4a outputs, making it usable as a workstation card that can drive multiple monitors directly. For a user who needs a single card to both compute and display, the P6000 has a functional edge that the A10G cannot match. The P6000 also has a higher base clock (1506 MHz versus 1320 MHz), which may help in latency-sensitive workloads that do not scale with parallel throughput. However, in every benchmark recorded in the database, the A10G wins by a wide margin.
FAQ
Q: Which card has more shading units?
A: The NVIDIA A10G has 9216 shading units, while the NVIDIA Quadro P6000 has 3840 shading units.
Q: What is the memory bandwidth difference between the two cards?
A: The A10G provides 600.2 GB/s of bandwidth from 24 GB of GDDR6 memory on a 384-bit bus. The P6000 provides 432.8 GB/s from 24 GB of GDDR5X memory on the same 384-bit bus. The A10G is 38.7 percent higher.
Q: Does the Quadro P6000 support ray tracing?
A: No. The P6000 has no ray tracing cores and no tensor cores. The A10G includes 72 ray tracing cores and 288 tensor cores.
Q: How do the two cards compare in the Geekbench Vulkan benchmark?
A: The A10G scores 145863, and the P6000 scores 73590. The A10G leads by 98.2 percent.
Q: Which card has a lower power draw?
A: The A10G has a 150 W TDP, while the P6000 has a 250 W TDP. The A10G also requires a lower suggested power supply (450 W versus 600 W).
Q: Can either card be used directly with a monitor?
A: Only the P6000 has display outputs, offering 1x DVI and 4x DisplayPort 1.4a. The A10G has no display outputs and requires a separate display adapter if visuals are needed.