NVIDIA A100 PCIe 40 GB vs NVIDIA CMP 40HX Comparison

NVIDIA
GEFORCE

NVIDIA A100 PCIe 40 GB

CORE STATE GA100
VRAM 40 GB
CLOCK SPEED 1410 MHz
TDP 250 W
BUS WIDTH 5120 bit
ARCHITECTURE Ampere
nm
PROCESS 7 nm
LAUNCH DATE 2020
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
178,627
93,395
geekbench_vulkan
146,380
77,879

Analysis: NVIDIA A100 PCIe 40 GB vs NVIDIA CMP 40HX

Head-to-Head Benchmarks

The recorded data shows a decisive performance advantage for the NVIDIA A100 PCIe 40 GB across both available benchmark tests. In Geekbench OpenCL, the A100 scores 178,627 against the CMP 40HX's 93,395, a delta of 91.3% in favor of the A100. That is not a marginal gap; it is nearly double the raw compute output in this workload. The Vulkan results follow the same pattern, with the A100 posting 146,380 versus 77,879 for the CMP 40HX, an 88% advantage. Both wins belong to the A100, and the head-to-head table records 2 wins for the A100 and 0 for the CMP 40HX.

Context from the nearest rivals reinforces how wide this gulf is. The A100's average benchmark score is 162,504, placing it in the 97th percentile of all GPUs in the database. Its closest competitor, the AMD Radeon PRO W7800, scores 164,894, which is 1.4% higher, while the NVIDIA RTX 4500 Ada Generation sits at 166,094, 2.2% higher. The AMD Radeon Pro W6800X trails by 1.1% at 160,671. In other words, the A100 is within a couple of percentage points of the fastest accelerators in its peer group, and its 91.3% lead over the CMP 40HX is roughly forty times larger than the gap to its nearest rival. The CMP 40HX, by contrast, averages 85,637, which places it in the 93rd percentile. Its nearest rival, the NVIDIA Quadro GP100, scores 87,445, 2.1% higher, and the AMD Radeon PRO W7600 is 1.7% higher at 87,108. The CMP 40HX does beat the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%, but those are the lower end of its comparison set. The percentile difference, 97 versus 93, understates the raw score gap because percentile ranks compress the tail, but the absolute numbers do not lie: the A100 delivers 89.8% more average benchmark score than the CMP 40HX.

Diving into the individual tests, the OpenCL result is the larger margin. A 91.3% delta means the A100 finishes the workload in roughly half the time, assuming linear scaling, which is a plausible interpretation for compute-bound tasks. The Vulkan delta of 88% is slightly smaller but still overwhelming. Both tests are synthetic compute workloads, so they reflect raw shader and tensor throughput rather than real-world gaming or rendering scenarios, but they are consistent with the hardware specifications. The CMP 40HX's FP32 throughput is 7.603 TFLOPS, while the A100 delivers 19.49 TFLOPS, a 2.56x ratio that closely tracks the benchmark delta. FP16 follows the same story: 77.97 TFLOPS (4:1) for the A100 versus 15.21 TFLOPS (2:1) for the CMP 40HX, a 5.1x gap that is even wider in mixed-precision work. The database does not include ray tracing or DLSS scores for these cards, so the analysis stops at compute metrics.

FAQ

Q: Which GPU wins in Geekbench OpenCL, and by how much?

A: The NVIDIA A100 PCIe 40 GB wins with a score of 178,627 against 93,395 for the NVIDIA CMP 40HX, a 91.3% advantage.

Q: Is the Vulkan result closer than OpenCL?

A: The Vulkan result is slightly closer but still lopsided. The A100 scores 146,380 and the CMP 40HX scores 77,879, an 88% delta. Both tests favor the A100 decisively.

Q: How does each card rank against all GPUs in the database?

A: The A100 sits in the 97th percentile with an average benchmark score of 162,504. The CMP 40HX sits in the 93rd percentile with an average score of 85,637.

Q: What are the closest rivals to the A100?

A: The AMD Radeon PRO W7800 is 1.4% higher at 164,894, the NVIDIA RTX 4500 Ada Generation is 2.2% higher at 166,094, the NVIDIA RTX A5500 is 1.6% higher at 165,217, and the AMD Radeon Pro W6800X is 1.1% lower at 160,671.

Q: What are the closest rivals to the CMP 40HX?

A: The NVIDIA Quadro GP100 is 2.1% higher at 87,445, the AMD Radeon PRO W7600 is 1.7% higher at 87,108, while the AMD Radeon PRO W6600 is 4.4% lower at 81,995 and the AMD Radeon Pro Vega 64X is 5.8% lower at 80,959.

Q: Does the CMP 40HX have any advantage in memory bandwidth per watt?

A: The data does not include a power-normalized metric. The CMP 40HX has a lower TDP of 185 W versus 250 W for the A100, and its memory bandwidth is 448.0 GB/s versus 1.56 TB/s, but the database does not calculate efficiency ratios.

Architecture Differences

The two GPUs come from different NVIDIA generations and foundry processes. The A100 uses the GA100 chip built on TSMC's 7 nm node, with 54,200 million transistors on an 826 mm² die, yielding a transistor density of 65.6 million per square millimeter. The CMP 40HX uses the TU106 chip on TSMC's 12 nm node, with 10,800 million transistors on a 445 mm² die, a density of 24.3 million per square millimeter. That is a 5x difference in transistor count and a 2.7x difference in density, which explains the compute gap more than any single clock or core count.

The memory subsystems are fundamentally different. The A100 has 40 GB of HBM2e on a 5120-bit bus, producing 1.56 TB/s of bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus, producing 448.0 GB/s. The A100's bandwidth is 3.5x higher, which matters for large matrix operations and data movement. The memory clock rates also differ: the A100 runs at 1215 MHz with 2.4 Gbps effective, while the CMP 40HX runs at 1750 MHz with 14 Gbps effective. The CMP 40HX uses a narrower bus but faster signaling, yet the total bandwidth still falls far short.

Shader and compute resources differ by a similar magnitude. The A100 has 6912 shading units, 432 TMUs, 160 ROPs, and 432 tensor cores. The CMP 40HX has 2304 shading units, 144 TMUs, 64 ROPs, and 288 tensor cores. The A100 also lists no RT cores in the database, while the CMP 40HX has 36 RT cores. The A100's pixel rate is 225.6 GPixel/s and its texture rate is 609.1 GTexel/s, versus 105.6 GPixel/s and 237.6 GTexel/s for the CMP 40HX. FP32 throughput is 19.49 TFLOPS versus 7.603 TFLOPS, and FP16 is 77.97 TFLOPS (4:1) versus 15.21 TFLOPS (2:1). The A100's FP16 figure uses a 4:1 ratio, indicating it is likely relying on tensor cores for that throughput, while the CMP 40HX's 2:1 ratio reflects its Turing-era FP16 path.

The CMP 40HX carries Turing's API feature set, including DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The A100 lists no API support in the database, which is consistent with its server-oriented profile and lack of display outputs. Both cards have no display outputs, so neither is suited for desktop use. The CMP 40HX is explicitly in the "Mining GPUs" generation, while the A100 is in "Server Ampere (Axx)". The A100's bus interface is PCIe 4.0 x16, while the CMP 40HX uses PCIe 1.0 x4, a severely limited connection that would bottleneck data transfer in any host system. The A100 also has a larger physical footprint: 267 mm length versus 229 mm, same 111 mm height, and the CMP 40HX adds a 35 mm width dimension. Both are dual-slot cards, but the A100 requires an 8-pin EPS power connector and a 600 W suggested PSU, while the CMP 40HX uses a single 8-pin and a 450 W suggested PSU.

The Verdict

The data points to a clear split by intended workload. The NVIDIA A100 PCIe 40 GB is the superior compute accelerator by every measured metric: 91.3% higher OpenCL score, 88% higher Vulkan score, 2.56x higher FP32 throughput, 3.5x higher memory bandwidth, and 5x more transistors. It belongs in the 97th percentile of all GPUs, within 2.2% of the top rivals in its peer group. Any workload that is compute-bound, memory-bound, or mixed-precision heavy should use the A100. The CMP 40HX cannot compete in raw performance, and its PCIe 1.0 x4 interface would throttle even its own lesser capabilities in a modern host system.

For the CMP 40HX, the case is narrower. It is a mining-focused Turing card with a 93rd percentile standing. It beats the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%, so it is not the weakest card in the database, but it is in the lower half of its own comparison set. Its 8 GB GDDR6 memory and 448.0 GB/s bandwidth are enough for tasks within its niche, and its 185 W TDP and dual-slot cooler, it fits in many systems. The launch MSRP was 699 USD. The absence of display outputs disqualifies it from any visual computing role. The A100 also has no display outputs, so the practical choice between the two is not about features but about scale. The A100 is for compute density; the CMP 40HX is for utility workloads. The A100 is end-of-life, released 2020-06-21, and the CMP 40HX is also end-of-life, released 2021-02-24. Both are no longer in production. If the task is inference training, scientific simulation, or any server-side acceleration, the A100 wins without qualification. The CMP 40HX is only relevant when the workload fits within its 7.603 TFLOPS FP32 and 448.0 GB/s envelope, and even then the gap to the A100 in every recorded benchmark is between 88% and 91.3%.

Specification Differences

The following table lists only the fields where the two GPUs differ in the database.

  • Chip: GA100 for the A100, TU106 for the CMP 40HX
  • Architecture: Ampere for the A100, Turing for the CMP 40HX
  • Generation: Server Ampere (Axx) for the A100, Mining GPUs for the CMP 40HX
  • Process node: 7 nm (TSMC) for the A100, 12 nm (TSMC) for the CMP 40HX
  • Transistors: 54,200 million for the A100, 10,800 million for the CMP 40HX
  • Die size: 826 mm² for the A100, 445 mm² for the CMP 40HX
  • Transistor density: 65.6M / mm² for the A100, 24.3M / mm² for the CMP 40HX
  • Base clock: 765 MHz for the A100, 1470 MHz for the CMP 40HX
  • Boost clock: 1410 MHz for the A100, 1650 MHz for the CMP 40HX
  • Memory clock: 1215 MHz (2.4 Gbps effective) for the A100, 1750 MHz (14 Gbps effective) for the CMP 40HX
  • Memory size: 40 GB for the A100, 8 GB for the CMP 40HX
  • Memory type: HBM2e for the A100, GDDR6 for the CMP 40HX
  • Memory bus width: 5120 bit for the A100, 256 bit for the CMP 40HX
  • Memory bandwidth: 1.56 TB/s for the A100, 448.0 GB/s for the CMP 40HX
  • Shading units: 6912 for the A100, 2304 for the CMP 40HX
  • TMUs: 432 for the A100, 144 for the CMP 40HX
  • ROPs: 160 for the A100, 64 for the CMP 40HX
  • RT cores: not listed for the A100, 36 for the CMP 40HX
  • Tensor cores: 432 for the A100, 288 for the CMP 40HX
  • Pixel rate: 225.6 GPixel/s for the A100, 105.6 GPixel/s for the CMP 40HX
  • Texture rate: 609.1 GTexel/s for the A100, 237.6 GTexel/s for the CMP 40HX
  • FP32: 19.49 TFLOPS for the A100, 7.603 TFLOPS for the CMP 40HX
  • FP16: 77.97 TFLOPS (4:1) for the A100, 15.21 TFLOPS (2:1) for the CMP 40HX
  • TDP: 250 W for the A100, 185 W for the CMP 40HX
  • Power connectors: 8-pin EPS for the A100, 1x 8-pin for the CMP 40HX
  • Suggested PSU: 600 W for the A100, 450 W for the CMP 40HX
  • Bus interface: PCIe 4.0 x16 for the A100, PCIe 1.0 x4 for the CMP 40HX
  • DirectX: not listed for the A100, 12 Ultimate (12_2) for the CMP 40HX
  • OpenGL: not listed for the A100, 4.6 for the CMP 40HX
  • Vulkan: not listed for the A100, 1.4 for the CMP 40HX
  • Length: 267 mm (10.5 inches) for the A100, 229 mm (9 inches) for the CMP 40HX
  • Width: not listed for the A100, 35 mm (1.4 inches) for the CMP 40HX
  • Release date: 2020-06-21 for the A100, 2021-02-24 for the CMP 40HX
  • Predecessor: Tesla Turing for the A100, not listed for the CMP 40HX
  • Successor: Server Ada for the A100, not listed for the CMP 40HX
  • Launch MSRP: not listed for the A100, 699 USD for the CMP 40HX
  • Geekbench OpenCL score: 178,627 for the A100, 93,395 for the CMP 40HX
  • Geekbench Vulkan score: 146,380 for the A100, 77,879 for the CMP 40HX
  • Average benchmark score: 162,504 for the A100, 85,637 for the CMP 40HX
  • Percentile vs all GPUs: 97 for the A100, 93 for the CMP 40HX

The cards share no identical specification fields other than manufacturer, foundry, slot width (dual-slot), height (111 mm), display outputs (none), and production status (end-of-life). The A100 is a server compute engine with a wide bus and massive memory pool; the CMP 40HX is a mining-oriented Turing card with a narrow PCIe link and modest memory. The specification sheet alone explains the benchmark results.

DETAILED SPECIFICATIONS

SPECIFICATION
A100 PCIe 40 GB
CMP 40HX
Core Specs
Shading Units
6,912
2,304 -66.7%
Shaders
6,912
2,304 -66.7%
TMUs
432
144 -66.7%
ROPs
160
64 -60.0%
SM Count
108
36 -66.7%
Clocks
Base Clock
765 MHz
1470 MHz
Boost Clock
1410 MHz
1650 MHz
Memory Clock
1215 MHz 2.4 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
40 GB
8 GB
VRAM (MB)
40,960
8,192 -80.0%
Memory Type
HBM2e
GDDR6
Memory Bus
5120 bit
256 bit
Bandwidth
1.56 TB/s
448.0 GB/s
Cache
L1 Cache
192 KB (per SM)
64 KB (per SM)
L2 Cache
40 MB
4 MB
Performance
Pixel Rate
225.6 GPixel/s
105.6 GPixel/s
Texture Rate
609.1 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
19.49 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
9.746 TFLOPS (1:2)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
77.97 TFLOPS (4:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
36
Tensor Cores
432
288 -33.3%
BF16
311.84 TFLOPS (16:1)
TF32
155.92 TFLOPs (8:1)
Power
TDP
250 W
185 W
TDP (W)
250
185 -26.0%
Suggested PSU
600 W
450 W
Power Connectors
8-pin EPS
1x 8-pin
Architecture
Architecture
Ampere
Turing
GPU Name
GA100
TU106
Generation
Server Ampere (Axx)
Mining GPUs
Process Size
7 nm
12 nm
Transistors
54,200 million
10,800 million
Die Size
826 mm²
445 mm²
Foundry
TSMC
TSMC
Density
65.6M / mm²
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.0
7.5
Shader Model
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
229 mm 9 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 1.0 x4
Other
Launch Price
699 USD
Production
End-of-life
End-of-life
Predecessor
Tesla Turing
Successor
Server Ada
View A100 PCIe 40 GB Details View CMP 40HX Details