NVIDIA A100 PCIe 40 GB vs NVIDIA Quadro GP100 Comparison
NVIDIA A100 PCIe 40 GB
Quadro GP100
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A100 PCIe 40 GB vs NVIDIA Quadro GP100
Head-to-Head Benchmarks
The recorded database contains a single head-to-head benchmark between these two cards: Geekbench OpenCL. In this test, the NVIDIA A100 PCIe 40 GB scores 178627, while the NVIDIA Quadro GP100 scores 87445. The delta between them is 104.3%, meaning the A100 more than doubles the Quadro GP100’s OpenCL result. That is a decisive margin, not a narrow lead.
Looking at the broader benchmark context, the A100’s average benchmark score across all recorded tests is 162504, while the Quadro GP100’s average is 87445. The A100 sits at the 97th percentile among all GPUs in the database, whereas the Quadro GP100 sits at the 93rd percentile. The percentile gap is smaller than the raw score gap suggests, which indicates that the Quadro GP100 is still competitive relative to the wider GPU landscape, even though it is far behind the A100 in absolute terms.
Among the A100’s nearest rivals, the AMD Radeon Pro W6800X averages 160671, which puts it 1.1% behind the A100. The AMD Radeon PRO W7800 averages 164894, which is 1.4% ahead of the A100. The NVIDIA RTX A5500 averages 165217, 1.6% ahead, and the NVIDIA RTX 4500 Ada Generation averages 166094, 2.2% ahead. These are tight margins. The A100 is not the fastest card in its immediate peer group, but it is within a few percentage points of all of them. The Quadro GP100’s nearest rivals, by contrast, show a different pattern. The AMD Radeon PRO W7600 averages 87108, just 0.4% ahead of the Quadro GP100. The NVIDIA CMP 40HX averages 85637, which is 2.1% behind. The NVIDIA RTX A4500 Mobile averages 91134, 4.0% ahead, and the NVIDIA RTX A4500 averages 91671, 4.6% ahead. The Quadro GP100 is therefore near the middle of its own peer group, with rivals clustered close on either side.
The single benchmark result is consistent with the architectural gulf between the two products. The A100’s 104.3% lead in OpenCL is not a fluke of a single workload; it aligns with the massive differences in compute resources, memory subsystem, and fabrication process that separate the two cards. For any workload that scales with raw compute throughput, the A100 is the clear winner. The Quadro GP100, while respectable in its own generation, cannot match the newer card’s output.
Architecture Differences
The two GPUs come from different architectural generations entirely. The A100 uses the GA100 chip built on the Ampere architecture, fabricated on a 7 nm process at TSMC. The Quadro GP100 uses the GP100 chip on the Pascal architecture, fabricated on a 16 nm process, also at TSMC. The process shrink alone is dramatic: 7 nm versus 16 nm. Transistor counts reflect this. The GA100 packs 54,200 million transistors on a die size of 826 mm², yielding a transistor density of 65.6 million per square millimeter. The GP100 contains 15,300 million transistors on a 610 mm² die, for a density of 25.1 million per square millimeter. The A100 has over three and a half times the transistor count of the Quadro GP100.
Compute resources follow the same pattern. The A100 has 6912 shading units, 432 texture mapping units, and 160 render output units. The Quadro GP100 has 3584 shading units, 224 TMUs, and 96 ROPs. That is roughly double the shading units, double the TMUs, and a substantially higher ROP count. The A100 also includes 432 tensor cores, while the Quadro GP100 has no tensor cores at all, a reflection of the Pascal generation predating the tensor core era. Neither card has dedicated ray tracing cores.
Clock speeds tell a more nuanced story. The A100 has a base clock of 765 MHz and a boost clock of 1410 MHz. The Quadro GP100 has a base clock of 1304 MHz and a boost clock of 1443 MHz. The Quadro GP100’s base clock is much higher, and its boost clock is slightly higher as well. But the A100’s raw compute advantage comes from its far larger execution resource pool, not from clock speed. The A100 delivers 19.49 TFLOPS of FP32 performance and 77.97 TFLOPS of FP16 performance with a 4:1 ratio. The Quadro GP100 delivers 10.34 TFLOPS of FP32 and 20.69 TFLOPS of FP16 with a 2:1 ratio. In FP32, the A100 is nearly double. In FP16, the A100 is more than triple.
Memory architecture also differs fundamentally. The A100 has 40 GB of HBM2e memory on a 5120-bit bus, delivering 1.56 TB/s of bandwidth. The Quadro GP100 has 16 GB of HBM2 memory on a 4096-bit bus, delivering 732.2 GB/s. The A100 has more than twice the capacity and more than twice the bandwidth. The memory clock rates are not directly comparable because the A100 runs at 1215 MHz with 2.4 Gbps effective, while the Quadro GP100 runs at 715 MHz with 1430 Mbps effective, but the effective bandwidth figures already capture the practical difference.
Pixel and texture rates follow the resource counts. The A100 achieves 225.6 GPixel/s and 609.1 GTexel/s. The Quadro GP100 achieves 138.5 GPixel/s and 323.2 GTexel/s. The A100 leads by roughly 60% in pixel throughput and nearly 90% in texture throughput. Power draw is relatively close: the A100 is rated at 250 W with a suggested PSU of 600 W, while the Quadro GP100 is rated at 235 W with a suggested PSU of 550 W. Both are dual-slot cards with the same physical dimensions of 267 mm length and 111 mm height.
Where Each One Wins
The A100 wins in essentially every compute-heavy category. Its FP32 throughput is 19.49 TFLOPS versus 10.34 TFLOPS, a lead of about 88%. Its FP16 throughput is 77.97 TFLOPS versus 20.69 TFLOPS, a lead of about 277%. Its memory bandwidth of 1.56 TB/s versus 732.2 GB/s means workloads that saturate memory, such as large matrix operations, deep learning training, or data-parallel scientific codes, will run far faster on the A100. The presence of 432 tensor cores gives the A100 a dedicated path for tensor-heavy workloads that the Quadro GP100 cannot access at all. The A100’s 40 GB of HBM2e also allows larger datasets to remain resident on the GPU without spilling to host memory.
The Quadro GP100 does have one meaningful advantage: its base clock is 1304 MHz versus 765 MHz, and its boost clock is 1443 MHz versus 1410 MHz. For workloads that are latency-bound or that do not scale well with additional cores, the higher clock speeds could narrow the gap. The Quadro GP100 also has display outputs, specifically 1x DVI and 4x DisplayPort 1.4a, while the A100 has no display outputs at all. The Quadro GP100 is a Pascal-era workstation card that can drive monitors directly. The A100 is a server accelerator with no video output capability. The Quadro GP100 also supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.3, while the A100’s API support fields are null in the database, meaning those graphics APIs are not recorded for it.
For compute workloads, the A100 is the obvious choice. For any task that requires a display output, the Quadro GP100 is the only one of the two that can serve that role. The Quadro GP100’s 16 GB of HBM2 is still substantial, and its 732.2 GB/s bandwidth exceeds many newer cards. But the A100’s memory capacity and bandwidth are in a different class. The Quadro GP100’s nearest rival data shows it competing with mid-range workstation cards like the AMD Radeon PRO W7600 and the NVIDIA CMP 40HX, while the A100 competes with top-tier cards like the AMD Radeon PRO W7800 and the NVIDIA RTX 4500 Ada Generation. That positioning tells the story: the A100 belongs to a higher performance tier entirely.
The Verdict
The data points to a clear conclusion: the NVIDIA A100 PCIe 40 GB is the superior compute product in nearly every measurable way. Its OpenCL score of 178627 versus 87445 is a 104.3% lead. Its FP32 and FP16 throughput are roughly double and triple, respectively. Its memory bandwidth is more than double. It has tensor cores, the Quadro GP100 does not. It has 40 GB of memory versus 16 GB. It is built on a newer process node with more than three times the transistor count.
The Quadro GP100 is not a bad card. Its 93rd percentile ranking among all GPUs shows that it remains a capable performer, and its nearest rivals are all within about 5% of its average score. But it is a Pascal-generation product from 2016, and the A100 is an Ampere-generation server card from 2020. The four-year gap in release dates is visible throughout the specifications.
Who should pick the A100? Anyone running compute-heavy workloads: deep learning training, scientific simulations, large-scale data processing, or any task that benefits from tensor cores and massive memory bandwidth. The A100’s lack of display outputs is irrelevant in a server context. Who should pick the Quadro GP100? Anyone who needs a workstation GPU with direct display output, or anyone running legacy Pascal-optimized software that does not take advantage of tensor cores. The Quadro GP100 can still handle professional graphics workloads, and its higher clock speeds may help in some latency-sensitive tasks. But for raw performance, the A100 wins by a wide margin.
FAQ
Q: How much faster is the NVIDIA A100 PCIe 40 GB than the NVIDIA Quadro GP100 in OpenCL?
A: The A100 scores 178627 in Geekbench OpenCL, while the Quadro GP100 scores 87445. The A100 is 104.3% faster, meaning it more than doubles the Quadro GP100’s score.
Q: Does the Quadro GP100 have tensor cores?
A: No. The Quadro GP100 has no tensor cores listed in the database. The A100 has 432 tensor cores.
Q: Which card has more memory bandwidth?
A: The A100 has 1.56 TB/s of bandwidth from 40 GB of HBM2e on a 5120-bit bus. The Quadro GP100 has 732.2 GB/s from 16 GB of HBM2 on a 4096-bit bus.
Q: Can either card output video to a display?
A: Only the Quadro GP100. It has 1x DVI and 4x DisplayPort 1.4a outputs. The A100 has no display outputs.
Q: How do the two cards compare in FP32 performance?
A: The A100 delivers 19.49 TFLOPS of FP32, while the Quadro GP100 delivers 10.34 TFLOPS. The A100 is about 88% faster.
Q: What are the process nodes for each card?
A: The A100 is built on a 7 nm process at TSMC, while the Quadro GP100 is built on a 16 nm process, also at TSMC.
Specification Differences
| Specification | NVIDIA A100 PCIe 40 GB | NVIDIA Quadro GP100 |
|---|---|---|
| Architecture | Ampere | Pascal |
| Generation | Server Ampere (Axx) | Quadro Pascal (Px000) |
| Process Node | 7 nm | 16 nm |
| Transistors | 54,200 million | 15,300 million |
| Die Size | 826 mm² | 610 mm² |
| Transistor Density | 65.6M / mm² | 25.1M / mm² |
| Base Clock | 765 MHz | 1304 MHz |
| Boost Clock | 1410 MHz | 1443 MHz |
| Memory Clock | 1215 MHz, 2.4 Gbps effective | 715 MHz, 1430 Mbps effective |
| Memory Size | 40 GB | 16 GB |
| Memory Type | HBM2e | HBM2 |
| Memory Bus Width | 5120 bit | 4096 bit |
| Memory Bandwidth | 1.56 TB/s | 732.2 GB/s |
| Shading Units | 6912 | 3584 |
| TMUs | 432 | 224 |
| ROPs | 160 | 96 |
| Tensor Cores | 432 | None |
| Pixel Rate | 225.6 GPixel/s | 138.5 GPixel/s |
| Texture Rate | 609.1 GTexel/s | 323.2 GTexel/s |
| FP32 Performance | 19.49 TFLOPS | 10.34 TFLOPS |
| FP16 Performance | 77.97 TFLOPS (4:1) | 20.69 TFLOPS (2:1) |
| TDP | 250 W | 235 W |
| Power Connectors | 8-pin EPS | 1x 8-pin |
| Suggested PSU | 600 W | 550 W |
| Bus Interface | PCIe 4.0 x16 | PCIe 3.0 x16 |
| Display Outputs | No outputs | 1x DVI, 4x DisplayPort 1.4a |
| DirectX | Not recorded | 12 (12_1) |
| OpenGL | Not recorded | 4.6 |
| Vulkan | Not recorded | 1.3 |
| Release Date | 2020-06-21 | 2016-09-30 |
| Production Status | End-of-life | End-of-life |
| Predecessor | Tesla Turing | Quadro Maxwell |
| Successor | Server Ada | Quadro Volta |