GPU Comparison

NVIDIA
GEFORCE

NVIDIA A2

CORE STATE GA107
VRAM 16 GB
CLOCK SPEED 1770 MHz
TDP 60 W
BUS WIDTH 128 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Quadro GV100

CORE STATE GV100
VRAM 32 GB
CLOCK SPEED 1627 MHz
TDP 250 W
BUS WIDTH 4096 bit
ARCHITECTURE Volta
nm
PROCESS 12 nm
LAUNCH DATE 2018

PERFORMANCE BENCHMARKS

geekbench_opencl
35,357
150,004
geekbench_vulkan
34,023
139,526
passmark_directx_10
N/A
140
passmark_directx_11
N/A
168
passmark_directx_12
N/A
84
passmark_directx_9
N/A
207
passmark_g2d
N/A
836
passmark_g3d
N/A
19,650
passmark_gpu_compute
N/A
9,069

Analysis: NVIDIA A2 vs NVIDIA Quadro GV100

The NVIDIA A2 and NVIDIA Quadro GV100 represent two distinct eras of NVIDIA workstation hardware: the A2 is a compact, power-efficient Ampere-generation card aimed at dense compute environments, while the Quadro GV100 is a large Volta-based workstation GPU with massive memory and compute throughput. Benchmark data shows a striking contrast: in the only head-to-head test available, the Geekbench OpenCL score, the GV100 delivers 144393 points against the A2’s 34866, a 75.9% deficit for the A2. Yet the overall average benchmark scores tell a different story—the A2 averages 34866, while the GV100 averages 34677, a 0.5% difference. This divergence arises because the A2 has only one recorded benchmark, while the GV100 has nine, including low-scoring DirectX and PassMark entries that drag its average down. The data therefore demands careful interpretation: the GV100 dominates the common compute test, but the aggregate scores are nearly identical due to the dissimilar benchmark sets.

The Verdict

Based strictly on the available data, the NVIDIA Quadro GV100 is the clear winner in the only direct comparison. The Geekbench OpenCL benchmark shows the GV100 scoring 144393 against the A2’s 34866, a 75.9% margin. If the intended workload is OpenCL compute, the GV100 is overwhelmingly faster. However, the A2 has a higher average benchmark score (34866 vs. 34677), but that average is derived from a single test, whereas the GV100’s average includes several low DirectX and PassMark results. For users whose primary metric is the overall average across mixed workloads, the two cards are effectively tied—the A2 is 0.5% ahead per the nearestRivals data. But for any application that relies on OpenCL performance, the GV100 is the only rational choice. The A2 offers no other benchmark data to claim wins elsewhere. The GV100 also provides display outputs (4x DisplayPort 1.4a) and a much larger memory pool (32 GB vs. 16 GB), while the A2 has no display outputs and a lower TDP (60 W vs. 250 W). Thus, the verdict is straightforward: the GV100 is the performance leader in the tested compute scenario, while the A2’s advantages are limited to power efficiency and a slightly higher aggregate score that stems from incomplete benchmark coverage.

Architecture Differences

The two GPUs come from different architectural generations. The A2 uses the GA107 chip on an 8 nm Samsung process, while the GV100 uses the GV100 chip on a 12 nm TSMC process. The transistor counts differ substantially: the GV100 packs 21,100 million transistors on an 815 mm² die, whereas the A2 has 8,700 million transistors on a 200 mm² die. This results in a transistor density of 43.5M per mm² for the A2 versus 25.9M per mm² for the GV100. The A2 is built on the Ampere architecture, supporting DirectX 12 Ultimate (12_2), while the GV100 is Volta, limited to DirectX 12 (12_1). Both support OpenGL 4.6 and Vulkan 1.4.

Memory architecture is a major differentiator. The A2 uses 16 GB of GDDR6 on a 128-bit bus, yielding 200.1 GB/s bandwidth. The GV100 uses 32 GB of HBM2 on a 4096-bit bus, delivering 868.4 GB/s—over four times the bandwidth. The GV100’s memory clock is 848 MHz (1696 Mbps effective), while the A2’s memory runs at 1563 MHz (12.5 Gbps effective). The A2 has 1280 shading units, 40 TMUs, and 32 ROPs; the GV100 has 5120 shading units, 320 TMUs, and 128 ROPs. The A2 includes 10 RT cores and 40 tensor cores, while the GV100 has no RT cores but 640 tensor cores. Compute throughput also diverges: the GV100’s FP32 is 16.66 TFLOPS versus the A2’s 4.531 TFLOPS, and FP16 is 33.32 TFLOPS (2:1) versus the A2’s 4.531 TFLOPS (1:1). Clock speeds are lower on the GV100 (base 1132 MHz, boost 1627 MHz) compared to the A2 (1440 MHz base, 1770 MHz boost), but the massive difference in shading units and memory bandwidth compensates.

Power and physical design also differ. The A2 has a TDP of 60 W, is single-slot, requires no power connectors, and suggests a 250 W PSU. The GV100 has a TDP of 250 W, is dual-slot, needs a single 8-pin connector, and suggests a 600 W PSU. The bus interface is PCIe 4.0 x8 for the A2 versus PCIe 3.0 x16 for the GV100. The A2 has no display outputs; the GV100 has 4x DisplayPort 1.4a. The GV100 measures 267 mm in length and 111 mm in height, while the A2’s dimensions are not listed. Release dates differ by over three years: the A2 launched on 2021-11-09, the GV100 on 2018-03-26. The GV100 had a launch MSRP of 8,999 USD; the A2’s launch MSRP is not provided.

Where Each One Wins

In the single head-to-head benchmark—Geekbench OpenCL—the GV100 wins decisively with a score of 144393 versus the A2’s 34866, a 75.9% advantage. This is the only direct comparison available, and it leaves no ambiguity about OpenCL compute performance. The GV100 also holds advantages in raw specifications that translate to compute-heavy tasks: higher FP32 (16.66 TFLOPS), FP16 (33.32 TFLOPS), memory bandwidth (868.4 GB/s), and a larger memory pool (32 GB). These numbers suggest the GV100 is better suited for large-scale data processing, though the benchmark data only confirms OpenCL.

The A2, however, wins on the aggregate average benchmark score. Its average is 34866, while the GV100’s is 34677—a 0.5% lead per the nearestRivals delta. This is not a head-to-head win but a statistical artifact of the GV100’s inclusion of low DirectX scores (e.g., PassMark DirectX 10 at 140, DirectX 9 at 207). The A2’s single score is its OpenCL result, which is lower than the GV100’s OpenCL but higher than the GV100’s average because the GV100’s average is pulled down by other tests. Thus, the A2 “wins” only in the sense of having a higher mean across its recorded benchmarks—but that mean is based on one data point. The A2 also has a lower TDP (60 W vs. 250 W) and requires no external power connectors, making it more efficient per watt, though no efficiency metric is directly measured.

FAQ

Q: Which GPU has the higher Geekbench OpenCL score?

A: The NVIDIA Quadro GV100 scores 144393, while the NVIDIA A2 scores 34866. The GV100 is 75.9% higher, as indicated by the deltaPct of -75.9 for the A2.

Q: Do the two GPUs have similar average benchmark scores?

A: Yes. The A2 has an average benchmark score of 34866, and the GV100 has an average of 34677. The A2 leads by 0.5% according to the nearestRivals data, but this is based on the A2’s single benchmark versus the GV100’s nine benchmarks.

Q: Which card has more memory and bandwidth?

A: The Quadro GV100 has 32 GB of HBM2 memory on a 4096-bit bus, providing 868.4 GB/s bandwidth. The A2 has 16 GB of GDDR6 on a 128-bit bus, providing 200.1 GB/s.

Q: Does the A2 support ray tracing?

A: Yes, the A2 includes 10 RT cores. The Quadro GV100 has no RT cores.

Q: Which card has display outputs?

A: The Quadro GV100 has 4x DisplayPort 1.4a outputs. The A2 has no display outputs.

Q: What is the TDP difference?

A: The A2 has a TDP of 60 W, while the GV100 has a TDP of 250 W. The A2 requires no power connectors, while the GV100 needs a single 8-pin connector.

Head-to-Head Benchmarks

The only direct benchmark comparison in the data is the Geekbench OpenCL test. The results are unambiguous: the NVIDIA Quadro GV100 scores 144393, while the NVIDIA A2 scores 34866. The deltaPct of -75.9 indicates the A2 is 75.9% behind the GV100. This is a massive margin—the GV100’s score is more than four times that of the A2 (144393 / 34866 ≈ 4.14, though that derived ratio is not explicitly in the pack). The GV100’s advantage stems from its far larger compute configuration: 5120 shading units versus 1280, 640 tensor cores versus 40, and 868.4 GB/s memory bandwidth versus 200.1 GB/s. The A2’s higher boost clock (1770 MHz vs. 1627 MHz) cannot compensate for the GV100’s raw resource advantage.

Notably, the aggregate average benchmark scores are nearly identical: the A2 averages 34866, and the GV100 averages 34677, a 0.5% difference. This seems contradictory to the OpenCL result, but the GV100’s average includes many low-scoring DirectX and PassMark tests (e.g., PassMark DirectX 10 at 140, DirectX 11 at 168, DirectX 12 at 84, DirectX 9 at 207, G2D at 836, G3D at 19650, GPU compute at 9069) that pull its mean down. The A2’s average is exactly its OpenCL score, as it has no other recorded benchmarks. Therefore, the head-to-head OpenCL result is the most reliable indicator of relative compute performance; the average scores are misleading due to the disparate test sets.

Specification Differences

The following table highlights the key specification differences between the two GPUs, based solely on the provided data.

| Specification | NVIDIA A2 | NVIDIA Quadro GV100 |

|---------------|-----------|---------------------|

| Architecture | Ampere | Volta |

| Process Node | 8 nm | 12 nm |

| Foundry | Samsung | TSMC |

| Transistors | 8,700 million | 21,100 million |

| Die Size | 200 mm² | 815 mm² |

| Transistor Density | 43.5M / mm² | 25.9M / mm² |

| Base Clock | 1440 MHz | 1132 MHz |

| Boost Clock | 1770 MHz | 1627 MHz |

| Memory Size | 16 GB | 32 GB |

| Memory Type | GDDR6 | HBM2 |

| Memory Bus Width | 128 bit | 4096 bit |

| Memory Bandwidth | 200.1 GB/s | 868.4 GB/s |

| Shading Units | 1280 | 5120 |

| TMUs | 40 | 320 |

| ROPs | 32 | 128 |

| RT Cores | 10 | None |

| Tensor Cores | 40 | 640 |

| Pixel Rate | 56.64 GPixel/s | 208.3 GPixel/s |

| Texture Rate | 70.80 GTexel/s | 520.6 GTexel/s |

| FP32 Performance | 4.531 TFLOPS | 16.66 TFLOPS |

DETAILED SPECIFICATIONS

SPECIFICATION
A2
Quadro GV100
Core Specs
Shading Units
1,280
5,120 +300.0%
Shaders
1,280
5,120 +300.0%
TMUs
40
320 +700.0%
ROPs
32
128 +300.0%
SM Count
10
80 +700.0%
Clocks
Base Clock
1440 MHz
1132 MHz
Boost Clock
1770 MHz
1627 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
848 MHz 1696 Mbps effective
Memory
Memory Size
16 GB
32 GB
VRAM (MB)
16,384
32,768 +100.0%
Memory Type
GDDR6
HBM2
Memory Bus
128 bit
4096 bit
Bandwidth
200.1 GB/s
868.4 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
2 MB
6 MB
Performance
Pixel Rate
56.64 GPixel/s
208.3 GPixel/s
Texture Rate
70.80 GTexel/s
520.6 GTexel/s
FP32 (TFLOPS)
4.531 TFLOPS
16.66 TFLOPS
FP64 (TFLOPS)
70.80 GFLOPS (1:64)
8.330 TFLOPS (1:2)
FP16 (TFLOPS)
4.531 TFLOPS (1:1)
33.32 TFLOPS (2:1)
AI/RT
RT Cores
10
Tensor Cores
40
640 +1500.0%
Power
TDP
60 W
250 W
TDP (W)
60
250 +316.7%
Suggested PSU
250 W
600 W
Power Connectors
None
1x 8-pin
Architecture
Architecture
Ampere
Volta
GPU Name
GA107
GV100
Generation
Workstation Ampere (Ax000)
Quadro Volta (Vx000)
Process Size
8 nm
12 nm
Transistors
8,700 million
21,100 million
Die Size
200 mm²
815 mm²
Foundry
Samsung
TSMC
Density
43.5M / mm²
25.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
7.0
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Dual-slot
Length
267 mm 10.5 inches
Height
111 mm 4.4 inches
Outputs
No outputs
4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x8
PCIe 3.0 x16
Other
Launch Price
8,999 USD
Production
End-of-life
End-of-life
Predecessor
Quadro Turing
Quadro Pascal
Successor
Workstation Ada
Quadro Turing
View A2 Details View Quadro GV100 Details