NVIDIA A100 PCIe 40 GB vs NVIDIA Quadro P6000 Comparison
NVIDIA A100 PCIe 40 GB
Quadro P6000
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A100 PCIe 40 GB vs NVIDIA Quadro P6000
The NVIDIA A100 PCIe 40 GB and NVIDIA Quadro P6000 represent two distinct eras of NVIDIA’s professional GPU lineup, separated by architecture, memory technology, and intended workload focus. The A100, built on the Ampere architecture, is a server-oriented accelerator with a 97th percentile ranking among all GPUs in the database, while the Quadro P6000, a Pascal-generation workstation card, holds a 90th percentile ranking. The recorded benchmark data shows a substantial performance gap, with the A100 leading in both tested workloads. This analysis walks through the architectural differences, head-to-head results, and specification deltas to clarify where each card stands.
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA A100 PCIe 40 GB records an average benchmark score of 162,504, while the NVIDIA Quadro P6000 scores 69,986. This places the A100 in the 97th percentile of all GPUs, compared to the Quadro P6000’s 90th percentile.
Q: How much faster is the A100 in OpenCL compute?
A: In the geekbench_opencl test, the A100 scores 178,627 versus the Quadro P6000’s 66,382, a delta of 169.1% in favor of the A100. This is the largest margin recorded in the head-to-head data.
Q: What is the memory configuration difference between the two cards?
A: The A100 uses 40 GB of HBM2e memory with a 5120-bit bus and 1.56 TB/s bandwidth. The Quadro P6000 uses 24 GB of GDDR5X memory on a 384-bit bus, providing 432.8 GB/s bandwidth.
Q: Which card has a higher boost clock?
A: The Quadro P6000 has a boost clock of 1645 MHz, which is higher than the A100’s 1410 MHz boost clock. The Quadro also has a higher base clock at 1506 MHz versus 765 MHz.
Q: Does the A100 support display outputs?
A: No, the A100 lists “No outputs” for display outputs, whereas the Quadro P6000 includes 1x DVI and 4x DisplayPort 1.4a outputs.
Q: What is the difference in FP16 performance?
A: The A100 delivers 77.97 TFLOPS FP16 (4:1), while the Quadro P6000 provides 197.4 GFLOPS FP16 (1:64). The A100’s FP16 throughput is orders of magnitude higher.
Architecture Differences
The two GPUs come from different architectural generations, which explains most of their performance divergence. The A100 is based on the GA100 chip using the Ampere architecture, fabricated on a 7 nm process at TSMC. It integrates 54,200 million transistors on an 826 mm² die, yielding a transistor density of 65.6M per mm². In contrast, the Quadro P6000 uses the GP102 chip from the Pascal architecture, built on a 16 nm process, with 11,800 million transistors on a 471 mm² die, giving a density of 25.1M per mm².
The A100’s compute resources are substantially larger: it has 6912 shading units, 432 texture mapping units, and 160 ROPs. It also includes 432 tensor cores, which are absent from the Quadro P6000. The Quadro offers 3840 shading units, 240 TMUs, and 96 ROPs, with no tensor cores or ray tracing cores on either card. The A100’s FP32 throughput is rated at 19.49 TFLOPS, while the Quadro P6000 reaches 12.63 TFLOPS. FP16 performance shows an even larger gap: 77.97 TFLOPS on the A100 versus 197.4 GFLOPS on the Quadro, reflecting the A100’s dedicated tensor core path for mixed-precision work.
Memory architecture also diverges sharply. The A100 uses HBM2e with 40 GB capacity and a 5120-bit bus, achieving 1.56 TB/s bandwidth. The Quadro P6000 relies on GDDR5X with 24 GB capacity on a 384-bit bus, delivering 432.8 GB/s. The A100 supports PCIe 4.0 x16, while the Quadro is limited to PCIe 3.0 x16. The A100 has no display outputs, indicating a compute-first design, whereas the Quadro includes DVI and DisplayPort outputs for workstation visualization. Both cards are dual-slot, share a 250 W TDP, and list a 600 W suggested PSU, but the A100 uses an 8-pin EPS connector while the Quadro uses a single 8-pin.
Head-to-Head Benchmarks
The recorded head-to-head data contains two benchmarks, and the A100 wins both decisively. In geekbench_opencl, the A100 scores 178,627 against the Quadro P6000’s 66,382, a delta of 169.1%. This means the A100 more than doubles the Quadro’s OpenCL performance, reflecting the benefits of newer architecture, higher shading unit count, and vastly superior memory bandwidth. The A100’s 1.56 TB/s bandwidth allows it to feed its 6912 shading units far more effectively than the Quadro’s 432.8 GB/s can feed its 3840 units.
In geekbench_vulkan, the A100 scores 146,380 versus the Quadro P6000’s 73,590, a delta of 98.9%. While this margin is smaller than OpenCL, it still represents nearly double the Vulkan performance. The Vulkan test likely stresses graphics and compute pipelines, where the A100’s higher pixel rate (225.6 GPixel/s versus 157.9 GPixel/s) and texture rate (609.1 GTexel/s versus 394.8 GTexel/s) contribute to the lead. The A100’s FP32 throughput advantage of 19.49 versus 12.63 TFLOPS also plays a role in compute-heavy Vulkan workloads.
The nearest rivals for each card provide context. The A100’s closest competitor is the AMD Radeon Pro W6800X, with an average score of 160,671 and a delta of 1.1%, meaning the A100 leads by just over 1%. The NVIDIA RTX 4500 Ada Generation trails by 2.2%, while the AMD Radeon PRO W7800 and NVIDIA RTX A5500 sit 1.4% and 1.6% behind, respectively. For the Quadro P6000, the nearest rival is the AMD Radeon Pro WX 8200, with an average score of 69,870 and a delta of 0.2%, making them nearly identical. The NVIDIA CMP 90HX is 1.4% faster, while the NVIDIA RTX A3000 Mobile and AMD Radeon RX 6600 LE are 0.2% and 1.2% slower, respectively.
The Verdict
Based strictly on the recorded data, the NVIDIA A100 PCIe 40 GB is the superior performer in every measured benchmark. Its average score of 162,504 is more than double the Quadro P6000’s 69,986, and it holds a 97th percentile ranking versus 90th. The A100 wins both head-to-head tests by wide margins, with OpenCL showing a 169.1% advantage and Vulkan showing a 98.9% advantage. This is not a close contest; the A100 is a newer, more powerful accelerator designed for high-throughput compute, while the Quadro P6000 is an older workstation card with a more modest feature set.
The Quadro P6000’s strengths lie in its display outputs, which the A100 lacks entirely. For users who need a GPU for visualization or interactive graphics with monitor connectivity, the Quadro P6000 is the only option of the two, as the A100 cannot drive displays. However, for any compute-oriented task, the A100’s benchmark results indicate a clear and overwhelming advantage. The Quadro P6000’s higher clock speeds (1506 MHz base, 1645 MHz boost versus 765 MHz and 1410 MHz) do not compensate for its lower core count and slower memory. The verdict is straightforward: the A100 is for compute workloads, the Quadro P6000 for legacy workstation display use.
Specification Differences
The two cards differ across nearly every core specification field. The A100 uses a 7 nm process, while the Quadro P6000 uses 16 nm. Transistor count is 54,200 million versus 11,800 million, and die size is 826 mm² versus 471 mm², giving densities of 65.6M and 25.1M per mm², respectively. Base clocks are 765 MHz for the A100 and 1506 MHz for the Quadro; boost clocks are 1410 MHz and 1645 MHz. Memory speed is listed as 1215 MHz (2.4 Gbps effective) for the A100 and 1127 MHz (9 Gbps effective) for the Quadro.
Memory capacity differs: 40 GB HBM2e versus 24 GB GDDR5X. Bus width is 5120 bit versus 384 bit, and bandwidth is 1.56 TB/s versus 432.8 GB/s. Shading units are 6912 versus 3840, TMUs are 432 versus 240, and ROPs are 160 versus 96. The A100 has 432 tensor cores; the Quadro has none. Pixel rate is 225.6 GPixel/s versus 157.9 GPixel/s, and texture rate is 609.1 GTexel/s versus 394.8 GTexel/s. FP32 performance is 19.49 TFLOPS versus 12.63 TFLOPS, while FP16 is 77.97 TFLOPS versus 197.4 GFLOPS. The A100 uses PCIe 4.0 x16, the Quadro uses PCIe 3.0 x16. Display outputs are “No outputs” on the A100 and 1x DVI plus 4x DisplayPort 1.4a on the Quadro. The A100 supports no listed DirectX, OpenGL, or Vulkan versions, while the Quadro supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. Power connectors are 8-pin EPS for the A100 versus 1x 8-pin for the Quadro; both share a 250 W TDP and 600 W suggested PSU. Dimensions are identical at 267 mm length and 111 mm height.
Where Each One Wins
The A100 wins in every compute-focused category. Its FP32 and FP16 throughput are far higher, with 19.49 TFLOPS and 77.97 TFLOPS versus 12.63 TFLOPS and 197.4 GFLOPS. The A100’s memory bandwidth of 1.56 TB/s is 3.6 times the Quadro’s 432.8 GB/s, which directly benefits large data sets and memory-bound workloads. The A100’s 40 GB capacity also exceeds the Quadro’s 24 GB, allowing larger models or buffers to reside on the GPU. The A100’s tensor cores provide a dedicated path for AI and deep learning inference, a feature the Quadro entirely lacks. In the OpenCL and Vulkan benchmarks, the A100 leads by 169.1% and 98.9%, respectively, confirming its dominance in general-purpose compute and modern graphics APIs.
The Quadro P6000 wins in the area of display output. It offers 1x DVI and 4x DisplayPort 1.4a, enabling direct connection to monitors, while the A100 has no display outputs and requires a separate GPU for visualization. The Quadro also has higher base and boost clocks (1506 MHz and 1645 MHz versus 765 MHz and 1410 MHz), which can benefit lightly threaded workloads where clock speed matters more than core count. Its support for DirectX 12, OpenGL 4.6, and Vulkan 1.4 gives it a software compatibility advantage for traditional workstation applications, whereas the A100 lists no API support in the database. For legacy tasks, the Quadro’s 24 GB GDDR5X memory and 432.8 GB/s bandwidth remain adequate, but the A100’s raw compute power and memory subsystem make it the clear choice for any workload where display output is not required.