NVIDIA A100 SXM4 40 GB vs NVIDIA A10G Comparison

NVIDIA
GEFORCE

NVIDIA A100 SXM4 40 GB

CORE STATE GA100
VRAM 40 GB
CLOCK SPEED 1410 MHz
TDP 400 W
BUS WIDTH 5120 bit
ARCHITECTURE Ampere
nm
PROCESS 7 nm
LAUNCH DATE 2020
VS
NVIDIA
GEFORCE

A10G

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1710 MHz
TDP 150 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
201,096
158,063
geekbench_vulkan
173,198
145,863

Analysis: NVIDIA A100 SXM4 40 GB vs NVIDIA A10G

# The Verdict

The NVIDIA A100 SXM4 40 GB and NVIDIA A10G are both Ampere-generation server parts, but they target different deployment profiles. The data is clear: the A100 SXM4 40 GB wins both recorded benchmark tests, taking geekbench_opencl by 27.2% and geekbench_vulkan by 18.7%. Its average benchmark score of 187147 places it in the 98th percentile of all GPUs, while the A10G sits at 151963 in the 97th percentile. The A100 SXM4 40 GB is the choice for maximum compute throughput when power and physical form factor are secondary concerns. The A10G, with its 150 W TDP versus the A100's 400 W, is the pick for dense, power-constrained environments where a single-slot card that draws less than half the power is more important than raw benchmark dominance. If your workload is memory-bandwidth bound or needs the largest possible frame buffer, the A100 SXM4 40 GB with its 40 GB of HBM2e and 1.56 TB/s bandwidth is the obvious pick. If you need a PCIe 4.0 x16 card that fits in a standard slot and can be powered by an 8-pin EPS connector, the A10G is the practical option, even though it trails in every measured benchmark.

# Architecture Differences

Both GPUs use the Ampere architecture, but they are built on different chips and different process nodes. The A100 SXM4 40 GB uses the GA100 chip fabricated by TSMC on a 7 nm process, while the A10G uses the GA102 chip fabricated by Samsung on an 8 nm process. The GA100 packs 54,200 million transistors on a 826 mm² die, yielding a transistor density of 65.6M per mm². The GA102 is smaller in both absolute and relative terms: 28,300 million transistors on a 628 mm² die, for a density of 45.1M per mm². The A100's chip has more than double the transistor count on a die that is only about 31% larger, which explains the substantial performance gap.

The memory subsystems are fundamentally different. The A100 SXM4 40 GB uses HBM2e across a 5120-bit bus, delivering 1.56 TB/s of bandwidth. The A10G uses GDDR6 across a 384-bit bus, delivering 600.2 GB/s. That is a 2.6x bandwidth advantage for the A100, which is critical for memory-intensive server workloads. The A100 has 40 GB of memory versus 24 GB on the A10G. Clock speeds also differ: the A100 runs at 1095 MHz base and 1410 MHz boost, while the A10G runs higher at 1320 MHz base and 1710 MHz boost. The A10G compensates for its narrower memory bus with higher clocks, but not enough to overcome the A100's bandwidth advantage.

Compute resources are distributed differently. The A100 SXM4 40 GB has 6912 shading units, 432 TMUs, 160 ROPs, and 432 tensor cores, but no dedicated RT cores listed. The A10G has more shading units at 9216, but fewer TMUs at 288, fewer ROPs at 96, and 288 tensor cores. The A10G does list 72 RT cores, while the A100's RT core count is not specified. The FP32 throughput tells the story: the A10G produces 31.52 TFLOPS of FP32 versus the A100's 19.49 TFLOPS, a 61.7% advantage for the A10G in raw FP32. However, the A100's FP16 throughput is 77.97 TFLOPS (with a 4:1 ratio), while the A10G's FP16 is 31.52 TFLOPS (1:1). The A100 is built for mixed-precision and reduced-precision compute.

# Head-to-Head Benchmarks

The Geekbench results show a consistent win for the A100 SXM4 40 GB, but the margins differ by test. In geekbench_opencl, the A100 scores 201096 versus the A10G's 158063, a 27.2% lead. In geekbench_vulkan, the A100 scores 173198 versus 145863, an 18.7% lead. The OpenCL gap is larger than the Vulkan gap, which suggests the A100's advantage is more pronounced in general compute workloads than in graphics-adjacent rendering paths. The A10G's higher FP32 throughput and RT cores do not translate into a benchmark win in either test.

The A100 SXM4 40 GB's average benchmark score of 187147 is 23.2% higher than the A10G's 151963. The A100 sits at the 98th percentile of all GPUs, while the A10G sits at the 97th percentile. That one-percentile difference understates the performance gap, but it does indicate both are high-end server parts. The A100's nearest rivals in the database include the NVIDIA RTX 5000 Ada Generation (184664, 1.3% behind), the NVIDIA A100 SXM4 80 GB (183725, 1.9% behind), and the NVIDIA RTX PRO 5000 Blackwell (182109, 2.8% behind). The only rival that beats the A100 is the NVIDIA Tesla V100S PCIe 32 GB at 194415, which is 3.7% ahead. The A10G's nearest rivals are the NVIDIA Tesla V100 PCIe 32 GB (150305, 1.1% behind), the AMD Radeon Pro W6800X (160671, 5.4% ahead of the A10G), and the NVIDIA A100 PCIe 40 GB (162504, 6.5% ahead). The A10G beats the AMD Instinct MI100 (139035) by 9.3%.

# Specification Differences

| Specification | NVIDIA A100 SXM4 40 GB | NVIDIA A10G |

|---|---|---|

| Chip | GA100 | GA102 |

| Process Node | 7 nm (TSMC) | 8 nm (Samsung) |

| Transistors | 54,200 million | 28,300 million |

| Die Size | 826 mm² | 628 mm² |

| Transistor Density | 65.6M / mm² | 45.1M / mm² |

| Base Clock | 1095 MHz | 1320 MHz |

| Boost Clock | 1410 MHz | 1710 MHz |

| Memory Size | 40 GB | 24 GB |

| Memory Type | HBM2e | GDDR6 |

| Memory Bus Width | 5120 bit | 384 bit |

| Memory Bandwidth | 1.56 TB/s | 600.2 GB/s |

| Memory Clock | 1215 MHz (2.4 Gbps effective) | 1563 MHz (12.5 Gbps effective) |

| Shading Units | 6912 | 9216 |

| TMUs | 432 | 288 |

| ROPs | 160 | 96 |

| RT Cores | Not specified | 72 |

| Tensor Cores | 432 | 288 |

| FP32 | 19.49 TFLOPS | 31.52 TFLOPS |

| FP16 | 77.97 TFLOPS (4:1) | 31.52 TFLOPS (1:1) |

| Pixel Rate | 225.6 GPixel/s | 164.2 GPixel/s |

| Texture Rate | 609.1 GTexel/s | 492.5 GTexel/s |

| TDP | 400 W | 150 W |

| Slot Width | SXM Module | Single-slot |

| Power Connectors | None | 8-pin EPS |

| Suggested PSU | 800 W | 450 W |

| Dimensions | Not specified | 267 mm (10.5 inches) length, 112 mm (4.4 inches) height |

| DirectX Support | Not specified | 12 Ultimate (12_2) |

| OpenGL Support | Not specified | 4.6 |

| Vulkan Support | Not specified | 1.4 |

| Release Date | 2020-05-13 | 2021-04-11 |

# FAQ

Q: Which GPU has higher raw FP32 compute throughput?

A: The NVIDIA A10G. It delivers 31.52 TFLOPS of FP32 versus the NVIDIA A100 SXM4 40 GB's 19.49 TFLOPS, a 61.7% advantage for the A10G.

Q: Which GPU has higher memory bandwidth?

A: The NVIDIA A100 SXM4 40 GB. It provides 1.56 TB/s of bandwidth via HBM2e on a 5120-bit bus, while the A10G provides 600.2 GB/s via GDDR6 on a 384-bit bus. The A100's bandwidth is 2.6x higher.

Q: Does the A10G support ray tracing?

A: The A10G lists 72 RT cores, while the A100 SXM4 40 GB does not list any RT cores in the provided data. The A10G also supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the A100's API support is not specified.

Q: How much power does each GPU consume?

A: The A100 SXM4 40 GB has a 400 W TDP and requires a suggested 800 W PSU. The A10G has a 150 W TDP and requires a suggested 450 W PSU. The A10G draws less than half the power of the A100.

Q: What is the physical form factor difference?

A: The A100 SXM4 40 GB is an SXM module with no power connectors and no specified dimensions. The A10G is a single-slot card with an 8-pin EPS power connector, measuring 267 mm (10.5 inches) in length and 112 mm (4.4 inches) in height.

Q: Which GPU has the higher average benchmark score?

A: The A100 SXM4 40 GB has an average benchmark score of 187147, placing it in the 98th percentile of all GPUs. The A10G has an average score of 151963, placing it in the 97th percentile. The A100 is 23.2% higher.

# Where Each One Wins

The NVIDIA A100 SXM4 40 GB wins in every recorded benchmark, but its strengths point to specific use cases. Its 1.56 TB/s memory bandwidth and 40 GB of HBM2e make it the pick for workloads that saturate memory, such as large model training and inference where the entire dataset or model must fit in fast memory. Its FP16 throughput of 77.97 TFLOPS (4:1) is more than double the A10G's 31.52 TFLOPS, making it the better choice for mixed-precision training where FP16 is the primary compute path. The A100's 98th percentile standing and its position within 1.9% of the A100 SXM4 80 GB in the average benchmark score show it is a top-tier compute part. The A100 also wins on texture rate (609.1 GTexel/s versus 492.5 GTexel/s) and pixel rate (225.6 GPixel/s versus 164.2 GPixel/s), despite having fewer shading units.

The NVIDIA A10G wins on the metrics that matter for different deployment scenarios. Its 31.52 TFLOPS of FP32 is 61.7% higher than the A100, so any workload that relies on single-precision compute will run faster on the A10G. Its higher boost clock of 1710 MHz versus 1410 MHz on the A100 supports this. The A10G's 150 W TDP and single-slot form factor make it the clear choice for servers where power density and physical space are constrained. It requires a suggested 450 W PSU versus the A100's 800 W, and it uses a standard 8-pin EPS power connector, while the A100 SXM module has no power connectors of its own. The A10G's 24 GB of GDDR6 is still a substantial memory pool, and its 600.2 GB/s bandwidth, while lower than the A100's, is sufficient for many inference and rendering tasks. The A10G also brings RT cores and full DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4 support, which the A100 does not list, making the A10G the better option for any workload that touches graphics APIs or ray-traced rendering. Its dimensions of 267 mm by 112 mm mean it fits in standard server chassis, whereas the A100 SXM module requires a specialized carrier board. For multi-GPU systems where every watt and every slot counts, the A10G's lower power draw and smaller footprint allow more cards per server. The A100 wins on absolute performance, but the A10G wins on efficiency per watt and deployment flexibility.

DETAILED SPECIFICATIONS

SPECIFICATION
A100 SXM4 40 GB
A10G
Core Specs
Shading Units
6,912
9,216 +33.3%
Shaders
6,912
9,216 +33.3%
TMUs
432
288 -33.3%
ROPs
160
96 -40.0%
SM Count
108
72 -33.3%
Clocks
Base Clock
1095 MHz
1320 MHz
Boost Clock
1410 MHz
1710 MHz
Memory Clock
1215 MHz 2.4 Gbps effective
1563 MHz 12.5 Gbps effective
Memory
Memory Size
40 GB
24 GB
VRAM (MB)
40,960
24,576 -40.0%
Memory Type
HBM2e
GDDR6
Memory Bus
5120 bit
384 bit
Bandwidth
1.56 TB/s
600.2 GB/s
Cache
L1 Cache
192 KB (per SM)
128 KB (per SM)
L2 Cache
40 MB
6 MB
Performance
Pixel Rate
225.6 GPixel/s
164.2 GPixel/s
Texture Rate
609.1 GTexel/s
492.5 GTexel/s
FP32 (TFLOPS)
19.49 TFLOPS
31.52 TFLOPS
FP64 (TFLOPS)
9.746 TFLOPS (1:2)
985.0 GFLOPS (1:32)
FP16 (TFLOPS)
77.97 TFLOPS (4:1)
31.52 TFLOPS (1:1)
AI/RT
RT Cores
—
72
Tensor Cores
432
288 -33.3%
BF16
311.84 TFLOPS (16:1)
—
TF32
155.92 TFLOPs (8:1)
—
Power
TDP
400 W
150 W
TDP (W)
400
150 -62.5%
Suggested PSU
800 W
450 W
Power Connectors
None
8-pin EPS
Architecture
Architecture
Ampere
Ampere
GPU Name
GA100
GA102
Generation
Server Ampere (Axx)
Server Ampere (Axx)
Process Size
7 nm
8 nm
Transistors
54,200 million
28,300 million
Die Size
826 mm²
628 mm²
Foundry
TSMC
Samsung
Density
65.6M / mm²
45.1M / mm²
API Support
DirectX
—
12 Ultimate (12_2)
OpenGL
—
4.6
Vulkan
—
1.4
OpenCL
3.0
3.0
CUDA
8.0
8.6
Shader Model
—
6.8
Physical
Slot Width
SXM Module
Single-slot
Length
—
267 mm 10.5 inches
Height
—
112 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Turing
Tesla Turing
Successor
Server Ada
Server Ada
View A100 SXM4 40 GB Details View A10G Details