NVIDIA A10G vs NVIDIA CMP 40HX Comparison

NVIDIA
GEFORCE

NVIDIA A10G

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1710 MHz
TDP 150 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
158,063
93,395
geekbench_vulkan
145,863
77,879

Analysis: NVIDIA A10G vs NVIDIA CMP 40HX

Where Each One Wins

The NVIDIA A10G and NVIDIA CMP 40HX occupy entirely different corners of the GPU landscape, and the benchmark data reflects that split clearly. The A10G wins both recorded head-to-head tests by a wide margin, making it the decisive choice for compute-heavy workloads. Its OpenCL score of 158,063 and Vulkan score of 145,863 place it in the 97th percentile of all GPUs in the database, a strong indicator of sustained throughput across general-purpose compute tasks.

The CMP 40HX, in contrast, was designed for a single purpose: mining. Its 93rd percentile ranking is respectable for what it is, but its OpenCL score of 93,395 and Vulkan score of 77,879 put it far behind the A10G. The CMP 40HX has no display outputs, and so does the A10G, meaning neither card is intended for desktop use. The difference is that the A10G is a server compute accelerator with a 24 GB memory pool, while the CMP 40HX is a mining-specific part with 8 GB. The data shows no scenario where the CMP 40HX outperforms the A10G in the recorded benchmarks.

Where each one wins is therefore less about workload type and more about capability tier. The A10G wins every measurable metric, and the CMP 40HX wins only in the sense that it exists as a lower-power, lower-cost alternative within the same generation family. For anyone choosing between the two, the A10G is the superior compute device, while the CMP 40HX would only be relevant for mining operations where the reduced memory and lower throughput are acceptable trade-offs.

Architecture Differences

The two GPUs come from different architectures, different foundries, and different process nodes. The A10G is built on the Ampere architecture using the GA102 chip, manufactured by Samsung on an 8 nm process. The CMP 40HX uses the older Turing architecture with the TU106 chip, produced by TSMC on a 12 nm process. This generation gap explains much of the performance difference.

The transistor counts tell the story. The A10G packs 28,300 million transistors on a 628 mm² die, yielding a transistor density of 45.1 million per mm². The CMP 40HX has 10,800 million transistors on a 445 mm² die, with a density of 24.3 million per mm². The A10G has nearly three times the transistor count, which translates directly into more shading units, more texture mapping units, and more render output units.

The A10G has 9,216 shading units, 288 TMUs, and 96 ROPs. The CMP 40HX has 2,304 shading units, 144 TMUs, and 64 ROPs. The A10G also carries 72 RT cores and 288 tensor cores, while the CMP 40HX has 36 RT cores and 288 tensor cores. The tensor core count is identical, but the RT core count is doubled on the A10G.

Memory architecture also diverges sharply. The A10G uses a 384 bit bus with 24 GB of GDDR6 running at 12.5 Gbps effective, producing 600.2 GB/s of bandwidth. The CMP 40HX uses a 256 bit bus with 8 GB of GDDR6 at 14 Gbps effective, yielding 448.0 GB/s. The A10G has three times the capacity and 34% more bandwidth.

The interface is another differentiator. The A10G uses PCIe 4.0 x16, while the CMP 40HX is limited to PCIe 1.0 x4. That is a massive gap in host communication bandwidth and would severely bottleneck any workload that requires frequent data transfer between CPU and GPU. The A10G also has a lower TDP of 150 W compared to the CMP 40HX at 185 W, despite delivering far more compute throughput.

Head-to-Head Benchmarks

The recorded head-to-head data contains two tests, and the A10G wins both by substantial margins. In Geekbench OpenCL, the A10G scores 158,063 against the CMP 40HX's 93,395, a 69.2% advantage. In Geekbench Vulkan, the A10G scores 145,863 against 77,879, a 87.3% advantage. The Vulkan gap is even larger than the OpenCL gap, suggesting the A10G scales better with the lower-level API.

The A10G's average benchmark score across the database is 151,963, while the CMP 40HX averages 85,637. That is a 77.4% difference in average score, consistent with the individual test results. The A10G sits in the 97th percentile of all GPUs, while the CMP 40HX sits in the 93rd percentile. Both are above average, but the A10G is in a different performance class.

Looking at the nearest rivals provides context for how each card sits relative to its competition. The A10G is 1.1% ahead of the NVIDIA Tesla V100 PCIe 32 GB in average score, and 9.3% ahead of the AMD Instinct MI100. It trails the AMD Radeon Pro W6800X by 5.4% and the NVIDIA A100 PCIe 40 GB by 6.5%. That places the A10G in the upper echelon of server accelerators, slightly below the A100 but competitive with the best of the previous generation.

The CMP 40HX, by comparison, is 1.7% behind the AMD Radeon PRO W7600 and 2.1% behind the NVIDIA Quadro GP100. It is 4.4% ahead of the AMD Radeon PRO W6600 and 5.8% ahead of the AMD Radeon Pro Vega 64X. The CMP 40HX is therefore a mid-tier performer that sits between workstation GPUs from the same era, but it is nowhere near the A10G in absolute terms.

The FP32 compute figures reinforce the benchmark results. The A10G delivers 31.52 TFLOPS of FP32 throughput, while the CMP 40HX delivers 7.603 TFLOPS. The A10G is roughly 4.1 times faster in raw single-precision compute. In FP16, the A10G also delivers 31.52 TFLOPS with a 1:1 ratio, while the CMP 40HX delivers 15.21 TFLOPS with a 2:1 ratio. The A10G maintains its lead in half-precision as well, though the margin shrinks to about 2.1 times.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA A10G has an average benchmark score of 151,963, while the NVIDIA CMP 40HX averages 85,637. The A10G is in the 97th percentile of all GPUs, while the CMP 40HX is in the 93rd percentile.

Q: How much faster is the A10G in OpenCL?

A: The A10G scores 158,063 in Geekbench OpenCL, compared to 93,395 for the CMP 40HX. That is a 69.2% advantage for the A10G.

Q: What is the memory capacity difference?

A: The A10G has 24 GB of GDDR6 memory on a 384 bit bus, while the CMP 40HX has 8 GB of GDDR6 on a 256 bit bus. The A10G also has higher memory bandwidth at 600.2 GB/s versus 448.0 GB/s.

Q: Do both cards have display outputs?

A: No. Neither the A10G nor the CMP 40HX has display outputs. Both are designed for compute or mining workloads, not desktop graphics.

Q: Which card has a higher TDP?

A: The CMP 40HX has a TDP of 185 W, while the A10G has a TDP of 150 W. The A10G delivers significantly more performance while consuming less power.

Q: What is the bus interface difference?

A: The A10G uses PCIe 4.0 x16, while the CMP 40HX uses PCIe 1.0 x4. This is a major difference in host communication bandwidth.

Specification Differences

The A10G and CMP 40HX differ across nearly every specification that matters for compute performance. The chip designs are from different generations, with the A10G using GA102 on Ampere and the CMP 40HX using TU106 on Turing. The process nodes differ as well, with the A10G at 8 nm and the CMP 40HX at 12 nm.

Transistor counts are 28,300 million for the A10G and 10,800 million for the CMP 40HX. Die sizes are 628 mm² and 445 mm² respectively. Transistor density is 45.1 million per mm² for the A10G and 24.3 million per mm² for the CMP 40HX.

Clock speeds show a mixed picture. The A10G has a base clock of 1320 MHz and a boost clock of 1710 MHz. The CMP 40HX has a higher base clock of 1470 MHz but a lower boost clock of 1650 MHz. Memory clocks differ as well, with the A10G at 1563 MHz (12.5 Gbps effective) and the CMP 40HX at 1750 MHz (14 Gbps effective).

Memory capacity is 24 GB for the A10G versus 8 GB for the CMP 40HX. Bus widths are 384 bit and 256 bit. Bandwidth is 600.2 GB/s versus 448.0 GB/s.

Shader resources are heavily skewed toward the A10G: 9,216 shading units versus 2,304, 288 TMUs versus 144, and 96 ROPs versus 64. RT cores are 72 versus 36, while tensor cores are 288 on both.

Rasterization and texture rates favor the A10G. Pixel rate is 164.2 GPixel/s versus 105.6 GPixel/s. Texture rate is 492.5 GTexel/s versus 237.6 GTexel/s.

FP32 compute is 31.52 TFLOPS for the A10G and 7.603 TFLOPS for the CMP 40HX. FP16 is 31.52 TFLOPS (1:1) for the A10G and 15.21 TFLOPS (2:1) for the CMP 40HX.

Power and physical dimensions differ. The A10G has a TDP of 150 W and is single-slot, using an 8-pin EPS connector. The CMP 40HX has a TDP of 185 W, is dual-slot, and uses a single 8-pin connector. Both suggest a 450 W power supply.

Physical size: the A10G is 267 mm long and 112 mm high. The CMP 40HX is 229 mm long, 111 mm high, and 35 mm wide.

The A10G uses PCIe 4.0 x16, while the CMP 40HX uses PCIe 1.0 x4. Both have no display outputs. Both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

The CMP 40HX has a launch MSRP of 699 USD. The A10G has no recorded launch MSRP in the database.

Release dates are close, with the A10G released on 2021-04-11 and the CMP 40HX on 2021-02-24. Both are end-of-life products. The A10G lists Tesla Turing as its predecessor and Server Ada as its successor, while the CMP 40HX has no recorded predecessor or successor.

DETAILED SPECIFICATIONS

SPECIFICATION
A10G
CMP 40HX
Core Specs
Shading Units
9,216
2,304 -75.0%
Shaders
9,216
2,304 -75.0%
TMUs
288
144 -50.0%
ROPs
96
64 -33.3%
SM Count
72
36 -50.0%
Clocks
Base Clock
1320 MHz
1470 MHz
Boost Clock
1710 MHz
1650 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
24 GB
8 GB
VRAM (MB)
24,576
8,192 -66.7%
Memory Type
GDDR6
GDDR6
Memory Bus
384 bit
256 bit
Bandwidth
600.2 GB/s
448.0 GB/s
Cache
L1 Cache
128 KB (per SM)
64 KB (per SM)
L2 Cache
6 MB
4 MB
Performance
Pixel Rate
164.2 GPixel/s
105.6 GPixel/s
Texture Rate
492.5 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
31.52 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
985.0 GFLOPS (1:32)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
31.52 TFLOPS (1:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
72
36 -50.0%
Tensor Cores
288
288 0.0%
Power
TDP
150 W
185 W
TDP (W)
150
185 +23.3%
Suggested PSU
450 W
450 W
Power Connectors
8-pin EPS
1x 8-pin
Architecture
Architecture
Ampere
Turing
GPU Name
GA102
TU106
Generation
Server Ampere (Axx)
Mining GPUs
Process Size
8 nm
12 nm
Transistors
28,300 million
10,800 million
Die Size
628 mm²
445 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
7.5
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Dual-slot
Length
267 mm 10.5 inches
229 mm 9 inches
Height
112 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 1.0 x4
Other
Launch Price
699 USD
Production
End-of-life
End-of-life
Predecessor
Tesla Turing
Successor
Server Ada
View A10G Details View CMP 40HX Details