NVIDIA A10G vs NVIDIA CMP 40HX Comparison
NVIDIA A10G
CMP 40HX
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A10G vs NVIDIA CMP 40HX
Where Each One Wins
The NVIDIA A10G and NVIDIA CMP 40HX occupy entirely different corners of the GPU landscape, and the benchmark data reflects that split clearly. The A10G wins both recorded head-to-head tests by a wide margin, making it the decisive choice for compute-heavy workloads. Its OpenCL score of 158,063 and Vulkan score of 145,863 place it in the 97th percentile of all GPUs in the database, a strong indicator of sustained throughput across general-purpose compute tasks.
The CMP 40HX, in contrast, was designed for a single purpose: mining. Its 93rd percentile ranking is respectable for what it is, but its OpenCL score of 93,395 and Vulkan score of 77,879 put it far behind the A10G. The CMP 40HX has no display outputs, and so does the A10G, meaning neither card is intended for desktop use. The difference is that the A10G is a server compute accelerator with a 24 GB memory pool, while the CMP 40HX is a mining-specific part with 8 GB. The data shows no scenario where the CMP 40HX outperforms the A10G in the recorded benchmarks.
Where each one wins is therefore less about workload type and more about capability tier. The A10G wins every measurable metric, and the CMP 40HX wins only in the sense that it exists as a lower-power, lower-cost alternative within the same generation family. For anyone choosing between the two, the A10G is the superior compute device, while the CMP 40HX would only be relevant for mining operations where the reduced memory and lower throughput are acceptable trade-offs.
Architecture Differences
The two GPUs come from different architectures, different foundries, and different process nodes. The A10G is built on the Ampere architecture using the GA102 chip, manufactured by Samsung on an 8 nm process. The CMP 40HX uses the older Turing architecture with the TU106 chip, produced by TSMC on a 12 nm process. This generation gap explains much of the performance difference.
The transistor counts tell the story. The A10G packs 28,300 million transistors on a 628 mm² die, yielding a transistor density of 45.1 million per mm². The CMP 40HX has 10,800 million transistors on a 445 mm² die, with a density of 24.3 million per mm². The A10G has nearly three times the transistor count, which translates directly into more shading units, more texture mapping units, and more render output units.
The A10G has 9,216 shading units, 288 TMUs, and 96 ROPs. The CMP 40HX has 2,304 shading units, 144 TMUs, and 64 ROPs. The A10G also carries 72 RT cores and 288 tensor cores, while the CMP 40HX has 36 RT cores and 288 tensor cores. The tensor core count is identical, but the RT core count is doubled on the A10G.
Memory architecture also diverges sharply. The A10G uses a 384 bit bus with 24 GB of GDDR6 running at 12.5 Gbps effective, producing 600.2 GB/s of bandwidth. The CMP 40HX uses a 256 bit bus with 8 GB of GDDR6 at 14 Gbps effective, yielding 448.0 GB/s. The A10G has three times the capacity and 34% more bandwidth.
The interface is another differentiator. The A10G uses PCIe 4.0 x16, while the CMP 40HX is limited to PCIe 1.0 x4. That is a massive gap in host communication bandwidth and would severely bottleneck any workload that requires frequent data transfer between CPU and GPU. The A10G also has a lower TDP of 150 W compared to the CMP 40HX at 185 W, despite delivering far more compute throughput.
Head-to-Head Benchmarks
The recorded head-to-head data contains two tests, and the A10G wins both by substantial margins. In Geekbench OpenCL, the A10G scores 158,063 against the CMP 40HX's 93,395, a 69.2% advantage. In Geekbench Vulkan, the A10G scores 145,863 against 77,879, a 87.3% advantage. The Vulkan gap is even larger than the OpenCL gap, suggesting the A10G scales better with the lower-level API.
The A10G's average benchmark score across the database is 151,963, while the CMP 40HX averages 85,637. That is a 77.4% difference in average score, consistent with the individual test results. The A10G sits in the 97th percentile of all GPUs, while the CMP 40HX sits in the 93rd percentile. Both are above average, but the A10G is in a different performance class.
Looking at the nearest rivals provides context for how each card sits relative to its competition. The A10G is 1.1% ahead of the NVIDIA Tesla V100 PCIe 32 GB in average score, and 9.3% ahead of the AMD Instinct MI100. It trails the AMD Radeon Pro W6800X by 5.4% and the NVIDIA A100 PCIe 40 GB by 6.5%. That places the A10G in the upper echelon of server accelerators, slightly below the A100 but competitive with the best of the previous generation.
The CMP 40HX, by comparison, is 1.7% behind the AMD Radeon PRO W7600 and 2.1% behind the NVIDIA Quadro GP100. It is 4.4% ahead of the AMD Radeon PRO W6600 and 5.8% ahead of the AMD Radeon Pro Vega 64X. The CMP 40HX is therefore a mid-tier performer that sits between workstation GPUs from the same era, but it is nowhere near the A10G in absolute terms.
The FP32 compute figures reinforce the benchmark results. The A10G delivers 31.52 TFLOPS of FP32 throughput, while the CMP 40HX delivers 7.603 TFLOPS. The A10G is roughly 4.1 times faster in raw single-precision compute. In FP16, the A10G also delivers 31.52 TFLOPS with a 1:1 ratio, while the CMP 40HX delivers 15.21 TFLOPS with a 2:1 ratio. The A10G maintains its lead in half-precision as well, though the margin shrinks to about 2.1 times.
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA A10G has an average benchmark score of 151,963, while the NVIDIA CMP 40HX averages 85,637. The A10G is in the 97th percentile of all GPUs, while the CMP 40HX is in the 93rd percentile.
Q: How much faster is the A10G in OpenCL?
A: The A10G scores 158,063 in Geekbench OpenCL, compared to 93,395 for the CMP 40HX. That is a 69.2% advantage for the A10G.
Q: What is the memory capacity difference?
A: The A10G has 24 GB of GDDR6 memory on a 384 bit bus, while the CMP 40HX has 8 GB of GDDR6 on a 256 bit bus. The A10G also has higher memory bandwidth at 600.2 GB/s versus 448.0 GB/s.
Q: Do both cards have display outputs?
A: No. Neither the A10G nor the CMP 40HX has display outputs. Both are designed for compute or mining workloads, not desktop graphics.
Q: Which card has a higher TDP?
A: The CMP 40HX has a TDP of 185 W, while the A10G has a TDP of 150 W. The A10G delivers significantly more performance while consuming less power.
Q: What is the bus interface difference?
A: The A10G uses PCIe 4.0 x16, while the CMP 40HX uses PCIe 1.0 x4. This is a major difference in host communication bandwidth.
Specification Differences
The A10G and CMP 40HX differ across nearly every specification that matters for compute performance. The chip designs are from different generations, with the A10G using GA102 on Ampere and the CMP 40HX using TU106 on Turing. The process nodes differ as well, with the A10G at 8 nm and the CMP 40HX at 12 nm.
Transistor counts are 28,300 million for the A10G and 10,800 million for the CMP 40HX. Die sizes are 628 mm² and 445 mm² respectively. Transistor density is 45.1 million per mm² for the A10G and 24.3 million per mm² for the CMP 40HX.
Clock speeds show a mixed picture. The A10G has a base clock of 1320 MHz and a boost clock of 1710 MHz. The CMP 40HX has a higher base clock of 1470 MHz but a lower boost clock of 1650 MHz. Memory clocks differ as well, with the A10G at 1563 MHz (12.5 Gbps effective) and the CMP 40HX at 1750 MHz (14 Gbps effective).
Memory capacity is 24 GB for the A10G versus 8 GB for the CMP 40HX. Bus widths are 384 bit and 256 bit. Bandwidth is 600.2 GB/s versus 448.0 GB/s.
Shader resources are heavily skewed toward the A10G: 9,216 shading units versus 2,304, 288 TMUs versus 144, and 96 ROPs versus 64. RT cores are 72 versus 36, while tensor cores are 288 on both.
Rasterization and texture rates favor the A10G. Pixel rate is 164.2 GPixel/s versus 105.6 GPixel/s. Texture rate is 492.5 GTexel/s versus 237.6 GTexel/s.
FP32 compute is 31.52 TFLOPS for the A10G and 7.603 TFLOPS for the CMP 40HX. FP16 is 31.52 TFLOPS (1:1) for the A10G and 15.21 TFLOPS (2:1) for the CMP 40HX.
Power and physical dimensions differ. The A10G has a TDP of 150 W and is single-slot, using an 8-pin EPS connector. The CMP 40HX has a TDP of 185 W, is dual-slot, and uses a single 8-pin connector. Both suggest a 450 W power supply.
Physical size: the A10G is 267 mm long and 112 mm high. The CMP 40HX is 229 mm long, 111 mm high, and 35 mm wide.
The A10G uses PCIe 4.0 x16, while the CMP 40HX uses PCIe 1.0 x4. Both have no display outputs. Both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.
The CMP 40HX has a launch MSRP of 699 USD. The A10G has no recorded launch MSRP in the database.
Release dates are close, with the A10G released on 2021-04-11 and the CMP 40HX on 2021-02-24. Both are end-of-life products. The A10G lists Tesla Turing as its predecessor and Server Ada as its successor, while the CMP 40HX has no recorded predecessor or successor.