NVIDIA A100 PCIe 40 GB vs NVIDIA GB10 Comparison
NVIDIA A100 PCIe 40 GB
GB10
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A100 PCIe 40 GB vs NVIDIA GB10
The Verdict
The benchmark data splits this comparison cleanly. The NVIDIA A100 PCIe 40 GB wins both recorded head-to-head tests, taking Geekbench OpenCL by 48.7% and Geekbench Vulkan by 27.7%. Its average benchmark score of 162,504 places it in the 97th percentile of all GPUs, while the NVIDIA GB10 sits at 117,393 and the 95th percentile. The A100 is the compute workhorse for legacy Ampere deployments, but the GB10 is the newer, more power-efficient design with vastly more memory. Choose the A100 if raw compute throughput in existing PCIe 4.0 server infrastructure is the priority. Choose the GB10 if you need 128 GB of memory, ray tracing support, a modern Blackwell architecture, and a much lower 140 W power draw in a compact IGP form factor. The A100’s 19.49 TFLOPS FP32 and 77.97 TFLOPS FP16 (4:1) are specialized for dense compute, while the GB10’s 29.71 TFLOPS FP32 and 29.71 TFLOPS FP16 (1:1) offer balanced performance across workloads. The data does not support a single winner; it supports two different use cases.
Architecture Differences
The architectural gap is generational. The A100 uses the GA100 chip on the Ampere architecture, built on a 7 nm process at TSMC with 54,200 million transistors on an 826 mm² die, yielding a transistor density of 65.6M per mm². The GB10 uses the GB20B chip on the Blackwell 2.0 architecture, built on a 5 nm process at TSMC with an unknown transistor count on a 382 mm² die. The node shrink and die size reduction are substantial, but the transistor density for the GB10 is not listed in the data.
Core configurations differ significantly. The A100 packs 6,912 shading units, 432 TMUs, 160 ROPs, and 432 tensor cores, with no RT cores listed. The GB10 has 6,144 shading units, 384 TMUs, just 48 ROPs, 384 tensor cores, and 48 RT cores. The A100 has more shading units, TMUs, ROPs, and tensor cores, but the GB10 adds dedicated ray tracing hardware that the A100 lacks entirely.
Clock behavior is a major differentiator. The A100 runs at a 765 MHz base and 1,410 MHz boost, while the GB10 runs at 1,665 MHz base and 2,418 MHz boost. The GB10’s boost clock is 71.5% higher than the A100’s, which helps explain how fewer shading units can produce higher FP32 throughput. Memory architecture is completely different: the A100 uses 40 GB of HBM2e on a 5,120-bit bus with 1.56 TB/s bandwidth, while the GB10 uses 128 GB of LPDDR5X on a 256-bit bus with 273.2 GB/s bandwidth. The A100’s bandwidth is 5.7 times higher, but the GB10 has 3.2 times more capacity. The A100’s memory clocks at 1,215 MHz (2.4 Gbps effective), while the GB10’s memory clocks at 1,067 MHz (8.5 Gbps effective).
Head-to-Head Benchmarks
The A100 dominates both recorded tests. In Geekbench OpenCL, the A100 scores 178,627 against the GB10’s 120,137, a 48.7% advantage. This is the largest margin in the comparison and reflects the A100’s superior memory bandwidth and higher shading unit count. In Geekbench Vulkan, the A100 scores 146,380 against the GB10’s 114,648, a 27.7% lead. The Vulkan gap is smaller, likely because the GB10’s higher clocks and ray tracing hardware narrow the compute gap in graphics-oriented workloads.
Context from nearest rivals clarifies these scores. The A100’s average of 162,504 sits 1.1% above the AMD Radeon Pro W6800X (160,671) and 1.4% below the AMD Radeon PRO W7800 (164,894). It also trails the NVIDIA RTX A5500 (165,217) by 1.6% and the NVIDIA RTX 4500 Ada Generation (166,094) by 2.2%. The A100 is competitive with these modern workstation cards despite being an older architecture. The GB10’s average of 117,393 is 0.3% above the NVIDIA RTX 4000 SFF Ada Generation (117,088) and 1.3% below the AMD Radeon PRO W7700 (118,976). It leads the NVIDIA Tesla V100 SXM2 16 GB (114,395) by 2.6% and the NVIDIA RTX A5500 Mobile (113,944) by 3.0%. The GB10’s nearest rivals are a mix of compact workstations and mobile parts, indicating its positioning as a lower-power, space-constrained solution.
The wins tally is 2-0 in favor of the A100, but the data shows the GB10 is not far behind in Vulkan and has advantages in other dimensions. The A100’s FP32 throughput of 19.49 TFLOPS is 34.4% lower than the GB10’s 29.71 TFLOPS, yet the A100 still wins both benchmarks. This suggests the A100’s memory bandwidth (1.56 TB/s vs 273.2 GB/s) is the decisive factor in these tests. The GB10’s pixel rate of 116.1 GPixel/s is about half the A100’s 225.6 GPixel/s, but its texture rate of 928.5 GTexel/s is 52.4% higher than the A100’s 609.1 GTexel/s.
Specification Differences
| Specification | NVIDIA A100 PCIe 40 GB | NVIDIA GB10 |
|---|---|---|
| Architecture | Ampere | Blackwell 2.0 |
| Process Node | 7 nm | 5 nm |
| Transistors | 54,200 million | unknown |
| Die Size | 826 mm² | 382 mm² |
| Transistor Density | 65.6M / mm² | null |
| Base Clock | 765 MHz | 1,665 MHz |
| Boost Clock | 1,410 MHz | 2,418 MHz |
| Memory Clock | 1,215 MHz (2.4 Gbps effective) | 1,067 MHz (8.5 Gbps effective) |
| Memory Size | 40 GB | 128 GB |
| Memory Type | HBM2e | LPDDR5X |
| Memory Bus Width | 5,120 bit | 256 bit |
| Memory Bandwidth | 1.56 TB/s | 273.2 GB/s |
| Shading Units | 6,912 | 6,144 |
| TMUs | 432 | 384 |
| ROPs | 160 | 48 |
| RT Cores | null | 48 |
| Tensor Cores | 432 | 384 |
| Pixel Rate | 225.6 GPixel/s | 116.1 GPixel/s |
| Texture Rate | 609.1 GTexel/s | 928.5 GTexel/s |
| FP32 | 19.49 TFLOPS | 29.71 TFLOPS |
| FP16 | 77.97 TFLOPS (4:1) | 29.71 TFLOPS (1:1) |
| TDP | 250 W | 140 W |
| Slot Width | Dual-slot | IGP |
| Power Connectors | 8-pin EPS | None |
| Suggested PSU | 600 W | 300 W |
| Bus Interface | PCIe 4.0 x16 | PCIe 5.0 x16 |
| Display Outputs | No outputs | 1x HDMI |
| Dimensions | 267 mm x 111 mm | 150 mm x 51 mm x 150 mm |
| Production Status | End-of-life | Active |
| Release Date | 2020-06-21 | 2025-10-14 |
| Predecessor | Tesla Turing | Server Hopper |
| Successor | Server Ada | Server Rubin |
| Launch MSRP | null | 3,999 USD |
FAQ
Q: Which GPU has higher FP32 performance?
A: The GB10 delivers 29.71 TFLOPS FP32, which is 52.4% higher than the A100’s 19.49 TFLOPS. This is despite the GB10 having fewer shading units (6,144 vs 6,912), thanks to its much higher boost clock of 2,418 MHz versus 1,410 MHz.
Q: How much more memory does the GB10 have?
A: The GB10 has 128 GB of LPDDR5X memory, which is 3.2 times the A100’s 40 GB of HBM2e. However, the A100’s memory bandwidth of 1.56 TB/s is 5.7 times higher than the GB10’s 273.2 GB/s due to the vastly wider 5,120-bit bus versus 256-bit.
Q: Which GPU wins the benchmark comparison?
A: The A100 wins both head-to-head tests. It leads by 48.7% in Geekbench OpenCL (178,627 vs 120,137) and by 27.7% in Geekbench Vulkan (146,380 vs 114,648). The A100’s average benchmark score of 162,504 also exceeds the GB10’s 117,393.
Q: Does the GB10 support ray tracing?
A: Yes. The GB10 has 48 RT cores, while the A100 has no RT cores listed in the data. This makes the GB10 the only one of the two with dedicated hardware for ray-traced workloads.
Q: What are the power requirements for each GPU?
A: The A100 has a TDP of 250 W with a suggested PSU of 600 W and requires an 8-pin EPS power connector. The GB10 has a TDP of 140 W with a suggested PSU of 300 W and requires no power connectors, as it is an IGP (integrated graphics processor) form factor.
Q: How do these GPUs compare to their nearest rivals?
A: The A100’s average score of 162,504 is 1.1% above the AMD Radeon Pro W6800X and 1.4% below the AMD Radeon PRO W7800. The GB10’s average of 117,393 is 0.3% above the NVIDIA RTX 4000 SFF Ada Generation and 1.3% below the AMD Radeon PRO W7700. Both sit near the top of their respective performance tiers.