NVIDIA A100 SXM4 40 GB vs NVIDIA CMP 40HX Comparison

NVIDIA
GEFORCE

NVIDIA A100 SXM4 40 GB

CORE STATE GA100
VRAM 40 GB
CLOCK SPEED 1410 MHz
TDP 400 W
BUS WIDTH 5120 bit
ARCHITECTURE Ampere
nm
PROCESS 7 nm
LAUNCH DATE 2020
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
201,096
93,395
geekbench_vulkan
173,198
77,879

Analysis: NVIDIA A100 SXM4 40 GB vs NVIDIA CMP 40HX

FAQ

Q: Which GPU is faster in the Geekbench OpenCL benchmark?

A: The NVIDIA A100 SXM4 40 GB scores 201096, which is 115.3% higher than the NVIDIA CMP 40HX's 93395.

Q: What is the difference in Geekbench Vulkan performance?

A: The A100 SXM4 40 GB records 173198, while the CMP 40HX achieves 77879, giving the A100 a 122.4% advantage.

Q: How do these GPUs compare to their nearest rivals in average benchmark scores?

A: The A100 SXM4 40 GB has an average benchmark score of 187147, placing it 1.3% ahead of the RTX 5000 Ada Generation (184664) and 1.9% ahead of the A100 SXM4 80 GB (183725). The CMP 40HX averages 85637, which is 1.7% behind the Radeon PRO W7600 (87108) and 4.4% ahead of the Radeon PRO W6600 (81995).

Q: What are the memory specifications of each card?

A: The A100 SXM4 40 GB uses 40 GB of HBM2e with a 5120-bit bus and 1.56 TB/s bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus with 448.0 GB/s bandwidth.

Q: What is the transistor count and die size for each chip?

A: The A100's GA100 chip contains 54,200 million transistors on an 826 mm² die. The CMP 40HX's TU106 chip has 10,800 million transistors on a 445 mm² die.

Q: Which GPU has a higher percentile ranking among all GPUs?

A: The A100 SXM4 40 GB sits at the 98th percentile, while the CMP 40HX is at the 93rd percentile.

Architecture Differences

The two GPUs represent entirely different product segments from NVIDIA. The A100 SXM4 40 GB belongs to the Server Ampere generation, built on the GA100 chip using TSMC's 7 nm process. The CMP 40HX is part of the Mining GPUs generation, using the TU106 chip from the Turing architecture on a 12 nm process. This generational gap shows up immediately in transistor density: the GA100 packs 65.6 million transistors per square millimeter, while the TU106 manages 24.3 million per square millimeter.

The compute resources differ dramatically. The A100 features 6912 shading units, 432 texture mapping units, 160 raster output units, and 432 tensor cores. The CMP 40HX offers 2304 shading units, 144 TMUs, 64 ROPs, 36 ray tracing cores, and 288 tensor cores. Notably, the CMP 40HX includes ray tracing hardware, while the A100 does not list RT cores in its specifications. The A100's tensor core count is higher, but the CMP 40HX still fields a substantial array.

Memory architecture diverges sharply. The A100 uses HBM2e with a massive 5120-bit bus, while the CMP 40HX relies on GDDR6 with a 256-bit bus. Clock speeds also differ: the A100 runs at 1095 MHz base and 1410 MHz boost, whereas the CMP 40HX operates at 1470 MHz base and 1650 MHz boost. The CMP's higher clocks reflect its smaller, denser compute layout. The A100's memory clock is listed as 1215 MHz with 2.4 Gbps effective, while the CMP's memory runs at 1750 MHz with 14 Gbps effective.

The A100 is a server module with SXM form factor, no power connectors, and no display outputs. The CMP 40HX is a dual-slot card with a single 8-pin connector, also lacking display outputs. The CMP 40HX does support DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the A100 lists no API support data. The CMP's PCIe interface is limited to PCIe 1.0 x4, a deliberate constraint for mining workloads, while the A100 uses PCIe 4.0 x16.

Head-to-Head Benchmarks

The recorded data shows a decisive performance gap across both benchmark tests. In Geekbench OpenCL, the A100 SXM4 40 GB scores 201096 against the CMP 40HX's 93395. That is a 115.3% advantage, meaning the A100 more than doubles the CMP's output. In Geekbench Vulkan, the margin widens slightly: the A100 reaches 173198, while the CMP manages 77879, a 122.4% lead.

These results align with the architectural disparity. The A100's FP32 throughput is listed at 19.49 TFLOPS, while the CMP 40HX delivers 7.603 TFLOPS. In FP16, the A100 achieves 77.97 TFLOPS using a 4:1 ratio, while the CMP reaches 15.21 TFLOPS at 2:1. The A100 also leads in pixel rate (225.6 GPixel/s vs 105.6 GPixel/s) and texture rate (609.1 GTexel/s vs 237.6 GTexel/s).

The average benchmark scores reinforce the gap. The A100's average of 187147 places it in the 98th percentile of all GPUs. The CMP 40HX averages 85637, sitting in the 93rd percentile. In terms of nearest rivals, the A100 edges out the RTX 5000 Ada Generation by 1.3% and the A100 SXM4 80 GB by 1.9%, while trailing the Tesla V100S PCIe 32 GB by 3.7%. The CMP 40HX trades closely with workstation cards: it is 1.7% behind the Radeon PRO W7600, 2.1% behind the Quadro GP100, but 4.4% ahead of the Radeon PRO W6600 and 5.8% ahead of the Radeon Pro Vega 64X.

The benchmark results show the A100 leading in every recorded test. The CMP 40HX posts no wins in the head-to-head comparison. This is not a close contest; it is a two-tier outcome where the server-grade Ampere part dominates the mining-oriented Turing part.

Specification Differences

The table below highlights only the fields where the two GPUs differ.

| Specification | NVIDIA A100 SXM4 40 GB | NVIDIA CMP 40HX |

|---|---|---|

| Architecture | Ampere | Turing |

| Generation | Server Ampere (Axx) | Mining GPUs |

| Process Node | 7 nm | 12 nm |

| Transistors | 54,200 million | 10,800 million |

| Die Size | 826 mm² | 445 mm² |

| Transistor Density | 65.6M / mm² | 24.3M / mm² |

| Base Clock | 1095 MHz | 1470 MHz |

| Boost Clock | 1410 MHz | 1650 MHz |

| Memory Clock | 1215 MHz, 2.4 Gbps effective | 1750 MHz, 14 Gbps effective |

| Memory Size | 40 GB | 8 GB |

| Memory Type | HBM2e | GDDR6 |

| Memory Bus | 5120 bit | 256 bit |

| Memory Bandwidth | 1.56 TB/s | 448.0 GB/s |

| Shading Units | 6912 | 2304 |

| TMUs | 432 | 144 |

| ROPs | 160 | 64 |

| RT Cores | None listed | 36 |

| Tensor Cores | 432 | 288 |

| Pixel Rate | 225.6 GPixel/s | 105.6 GPixel/s |

| Texture Rate | 609.1 GTexel/s | 237.6 GTexel/s |

| FP32 | 19.49 TFLOPS | 7.603 TFLOPS |

| FP16 | 77.97 TFLOPS (4:1) | 15.21 TFLOPS (2:1) |

| TDP | 400 W | 185 W |

| Slot Width | SXM Module | Dual-slot |

| Power Connectors | None | 1x 8-pin |

| Suggested PSU | 800 W | 450 W |

| Bus Interface | PCIe 4.0 x16 | PCIe 1.0 x4 |

| DirectX | None listed | 12 Ultimate (12_2) |

| OpenGL | None listed | 4.6 |

| Vulkan | None listed | 1.4 |

| Dimensions | Not listed | 229 mm x 111 mm x 35 mm |

| Release Date | 2020-05-13 | 2021-02-24 |

Both cards share the same manufacturer, use TSMC as the foundry, and have no display outputs. Both are also marked as end-of-life in production status.

The Verdict

The data points to a clear separation in intended use cases. The NVIDIA A100 SXM4 40 GB is a compute monster. Its 98th percentile ranking, 19.49 TFLOPS of FP32, and 1.56 TB/s of memory bandwidth make it suitable for heavy server workloads, scientific computing, and AI training. The benchmark results show it doubling the CMP 40HX's scores in both OpenCL and Vulkan tests. Its average score of 187147 places it among the top tier of GPUs, trading blows with the RTX 5000 Ada Generation and the A100 SXM4 80 GB.

The NVIDIA CMP 40HX serves a different purpose entirely. As a mining GPU, its design priorities are not compute throughput but efficiency and cost control. The 185 W TDP, 450 W suggested PSU, and PCIe 1.0 x4 interface all point to a card optimized for mining operations rather than general compute. Its 93rd percentile ranking is respectable, but its average score of 85637 places it firmly in the mid-range workstation class, competing with AMD's Radeon PRO W6600 and W7600 rather than server accelerators.

For users needing maximum compute performance, the A100 SXM4 40 GB is the obvious choice. The 115.3% OpenCL lead and 122.4% Vulkan lead are decisive. The 40 GB HBM2e memory with 1.56 TB/s bandwidth provides a massive advantage for memory-bound workloads. The 432 tensor cores and 77.97 TFLOPS FP16 performance make it suited for deep learning tasks.

For those prioritizing lower power draw and compact physical dimensions, the CMP 40HX offers a different trade-off. Its 185 W TDP is less than half the A100's 400 W, and its dual-slot 229 mm length fits in standard chassis. The inclusion of ray tracing cores and full DirectX 12 Ultimate support gives it broader API compatibility. However, the data shows it cannot match the A100's raw performance, and its 8 GB memory capacity is far more limited.

The verdict from the recorded measurements is straightforward: the A100 SXM4 40 GB wins every benchmark, holds a higher percentile ranking, and offers superior specifications across memory, compute, and bandwidth. The CMP 40HX is a specialized mining product, and its benchmark scores reflect that narrower focus. Buyers should choose based on workload: the A100 for serious compute, the CMP 40HX only if mining efficiency and low power draw are the primary concerns.

DETAILED SPECIFICATIONS

SPECIFICATION
A100 SXM4 40 GB
CMP 40HX
Core Specs
Shading Units
6,912
2,304 -66.7%
Shaders
6,912
2,304 -66.7%
TMUs
432
144 -66.7%
ROPs
160
64 -60.0%
SM Count
108
36 -66.7%
Clocks
Base Clock
1095 MHz
1470 MHz
Boost Clock
1410 MHz
1650 MHz
Memory Clock
1215 MHz 2.4 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
40 GB
8 GB
VRAM (MB)
40,960
8,192 -80.0%
Memory Type
HBM2e
GDDR6
Memory Bus
5120 bit
256 bit
Bandwidth
1.56 TB/s
448.0 GB/s
Cache
L1 Cache
192 KB (per SM)
64 KB (per SM)
L2 Cache
40 MB
4 MB
Performance
Pixel Rate
225.6 GPixel/s
105.6 GPixel/s
Texture Rate
609.1 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
19.49 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
9.746 TFLOPS (1:2)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
77.97 TFLOPS (4:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
36
Tensor Cores
432
288 -33.3%
BF16
311.84 TFLOPS (16:1)
TF32
155.92 TFLOPs (8:1)
Power
TDP
400 W
185 W
TDP (W)
400
185 -53.8%
Suggested PSU
800 W
450 W
Power Connectors
None
1x 8-pin
Architecture
Architecture
Ampere
Turing
GPU Name
GA100
TU106
Generation
Server Ampere (Axx)
Mining GPUs
Process Size
7 nm
12 nm
Transistors
54,200 million
10,800 million
Die Size
826 mm²
445 mm²
Foundry
TSMC
TSMC
Density
65.6M / mm²
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.0
7.5
Shader Model
6.8
Physical
Slot Width
SXM Module
Dual-slot
Length
229 mm 9 inches
Height
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 1.0 x4
Other
Launch Price
699 USD
Production
End-of-life
End-of-life
Predecessor
Tesla Turing
Successor
Server Ada
View A100 SXM4 40 GB Details View CMP 40HX Details