NVIDIA A100 PCIe 40 GB vs NVIDIA CMP 40HX Comparison
NVIDIA A100 PCIe 40 GB
CMP 40HX
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A100 PCIe 40 GB vs NVIDIA CMP 40HX
Head-to-Head Benchmarks
The recorded data shows a decisive performance advantage for the NVIDIA A100 PCIe 40 GB across both available benchmark tests. In Geekbench OpenCL, the A100 scores 178,627 against the CMP 40HX's 93,395, a delta of 91.3% in favor of the A100. That is not a marginal gap; it is nearly double the raw compute output in this workload. The Vulkan results follow the same pattern, with the A100 posting 146,380 versus 77,879 for the CMP 40HX, an 88% advantage. Both wins belong to the A100, and the head-to-head table records 2 wins for the A100 and 0 for the CMP 40HX.
Context from the nearest rivals reinforces how wide this gulf is. The A100's average benchmark score is 162,504, placing it in the 97th percentile of all GPUs in the database. Its closest competitor, the AMD Radeon PRO W7800, scores 164,894, which is 1.4% higher, while the NVIDIA RTX 4500 Ada Generation sits at 166,094, 2.2% higher. The AMD Radeon Pro W6800X trails by 1.1% at 160,671. In other words, the A100 is within a couple of percentage points of the fastest accelerators in its peer group, and its 91.3% lead over the CMP 40HX is roughly forty times larger than the gap to its nearest rival. The CMP 40HX, by contrast, averages 85,637, which places it in the 93rd percentile. Its nearest rival, the NVIDIA Quadro GP100, scores 87,445, 2.1% higher, and the AMD Radeon PRO W7600 is 1.7% higher at 87,108. The CMP 40HX does beat the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%, but those are the lower end of its comparison set. The percentile difference, 97 versus 93, understates the raw score gap because percentile ranks compress the tail, but the absolute numbers do not lie: the A100 delivers 89.8% more average benchmark score than the CMP 40HX.
Diving into the individual tests, the OpenCL result is the larger margin. A 91.3% delta means the A100 finishes the workload in roughly half the time, assuming linear scaling, which is a plausible interpretation for compute-bound tasks. The Vulkan delta of 88% is slightly smaller but still overwhelming. Both tests are synthetic compute workloads, so they reflect raw shader and tensor throughput rather than real-world gaming or rendering scenarios, but they are consistent with the hardware specifications. The CMP 40HX's FP32 throughput is 7.603 TFLOPS, while the A100 delivers 19.49 TFLOPS, a 2.56x ratio that closely tracks the benchmark delta. FP16 follows the same story: 77.97 TFLOPS (4:1) for the A100 versus 15.21 TFLOPS (2:1) for the CMP 40HX, a 5.1x gap that is even wider in mixed-precision work. The database does not include ray tracing or DLSS scores for these cards, so the analysis stops at compute metrics.
FAQ
Q: Which GPU wins in Geekbench OpenCL, and by how much?
A: The NVIDIA A100 PCIe 40 GB wins with a score of 178,627 against 93,395 for the NVIDIA CMP 40HX, a 91.3% advantage.
Q: Is the Vulkan result closer than OpenCL?
A: The Vulkan result is slightly closer but still lopsided. The A100 scores 146,380 and the CMP 40HX scores 77,879, an 88% delta. Both tests favor the A100 decisively.
Q: How does each card rank against all GPUs in the database?
A: The A100 sits in the 97th percentile with an average benchmark score of 162,504. The CMP 40HX sits in the 93rd percentile with an average score of 85,637.
Q: What are the closest rivals to the A100?
A: The AMD Radeon PRO W7800 is 1.4% higher at 164,894, the NVIDIA RTX 4500 Ada Generation is 2.2% higher at 166,094, the NVIDIA RTX A5500 is 1.6% higher at 165,217, and the AMD Radeon Pro W6800X is 1.1% lower at 160,671.
Q: What are the closest rivals to the CMP 40HX?
A: The NVIDIA Quadro GP100 is 2.1% higher at 87,445, the AMD Radeon PRO W7600 is 1.7% higher at 87,108, while the AMD Radeon PRO W6600 is 4.4% lower at 81,995 and the AMD Radeon Pro Vega 64X is 5.8% lower at 80,959.
Q: Does the CMP 40HX have any advantage in memory bandwidth per watt?
A: The data does not include a power-normalized metric. The CMP 40HX has a lower TDP of 185 W versus 250 W for the A100, and its memory bandwidth is 448.0 GB/s versus 1.56 TB/s, but the database does not calculate efficiency ratios.
Architecture Differences
The two GPUs come from different NVIDIA generations and foundry processes. The A100 uses the GA100 chip built on TSMC's 7 nm node, with 54,200 million transistors on an 826 mm² die, yielding a transistor density of 65.6 million per square millimeter. The CMP 40HX uses the TU106 chip on TSMC's 12 nm node, with 10,800 million transistors on a 445 mm² die, a density of 24.3 million per square millimeter. That is a 5x difference in transistor count and a 2.7x difference in density, which explains the compute gap more than any single clock or core count.
The memory subsystems are fundamentally different. The A100 has 40 GB of HBM2e on a 5120-bit bus, producing 1.56 TB/s of bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus, producing 448.0 GB/s. The A100's bandwidth is 3.5x higher, which matters for large matrix operations and data movement. The memory clock rates also differ: the A100 runs at 1215 MHz with 2.4 Gbps effective, while the CMP 40HX runs at 1750 MHz with 14 Gbps effective. The CMP 40HX uses a narrower bus but faster signaling, yet the total bandwidth still falls far short.
Shader and compute resources differ by a similar magnitude. The A100 has 6912 shading units, 432 TMUs, 160 ROPs, and 432 tensor cores. The CMP 40HX has 2304 shading units, 144 TMUs, 64 ROPs, and 288 tensor cores. The A100 also lists no RT cores in the database, while the CMP 40HX has 36 RT cores. The A100's pixel rate is 225.6 GPixel/s and its texture rate is 609.1 GTexel/s, versus 105.6 GPixel/s and 237.6 GTexel/s for the CMP 40HX. FP32 throughput is 19.49 TFLOPS versus 7.603 TFLOPS, and FP16 is 77.97 TFLOPS (4:1) versus 15.21 TFLOPS (2:1). The A100's FP16 figure uses a 4:1 ratio, indicating it is likely relying on tensor cores for that throughput, while the CMP 40HX's 2:1 ratio reflects its Turing-era FP16 path.
The CMP 40HX carries Turing's API feature set, including DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The A100 lists no API support in the database, which is consistent with its server-oriented profile and lack of display outputs. Both cards have no display outputs, so neither is suited for desktop use. The CMP 40HX is explicitly in the "Mining GPUs" generation, while the A100 is in "Server Ampere (Axx)". The A100's bus interface is PCIe 4.0 x16, while the CMP 40HX uses PCIe 1.0 x4, a severely limited connection that would bottleneck data transfer in any host system. The A100 also has a larger physical footprint: 267 mm length versus 229 mm, same 111 mm height, and the CMP 40HX adds a 35 mm width dimension. Both are dual-slot cards, but the A100 requires an 8-pin EPS power connector and a 600 W suggested PSU, while the CMP 40HX uses a single 8-pin and a 450 W suggested PSU.
The Verdict
The data points to a clear split by intended workload. The NVIDIA A100 PCIe 40 GB is the superior compute accelerator by every measured metric: 91.3% higher OpenCL score, 88% higher Vulkan score, 2.56x higher FP32 throughput, 3.5x higher memory bandwidth, and 5x more transistors. It belongs in the 97th percentile of all GPUs, within 2.2% of the top rivals in its peer group. Any workload that is compute-bound, memory-bound, or mixed-precision heavy should use the A100. The CMP 40HX cannot compete in raw performance, and its PCIe 1.0 x4 interface would throttle even its own lesser capabilities in a modern host system.
For the CMP 40HX, the case is narrower. It is a mining-focused Turing card with a 93rd percentile standing. It beats the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%, so it is not the weakest card in the database, but it is in the lower half of its own comparison set. Its 8 GB GDDR6 memory and 448.0 GB/s bandwidth are enough for tasks within its niche, and its 185 W TDP and dual-slot cooler, it fits in many systems. The launch MSRP was 699 USD. The absence of display outputs disqualifies it from any visual computing role. The A100 also has no display outputs, so the practical choice between the two is not about features but about scale. The A100 is for compute density; the CMP 40HX is for utility workloads. The A100 is end-of-life, released 2020-06-21, and the CMP 40HX is also end-of-life, released 2021-02-24. Both are no longer in production. If the task is inference training, scientific simulation, or any server-side acceleration, the A100 wins without qualification. The CMP 40HX is only relevant when the workload fits within its 7.603 TFLOPS FP32 and 448.0 GB/s envelope, and even then the gap to the A100 in every recorded benchmark is between 88% and 91.3%.
Specification Differences
The following table lists only the fields where the two GPUs differ in the database.
- Chip: GA100 for the A100, TU106 for the CMP 40HX
- Architecture: Ampere for the A100, Turing for the CMP 40HX
- Generation: Server Ampere (Axx) for the A100, Mining GPUs for the CMP 40HX
- Process node: 7 nm (TSMC) for the A100, 12 nm (TSMC) for the CMP 40HX
- Transistors: 54,200 million for the A100, 10,800 million for the CMP 40HX
- Die size: 826 mm² for the A100, 445 mm² for the CMP 40HX
- Transistor density: 65.6M / mm² for the A100, 24.3M / mm² for the CMP 40HX
- Base clock: 765 MHz for the A100, 1470 MHz for the CMP 40HX
- Boost clock: 1410 MHz for the A100, 1650 MHz for the CMP 40HX
- Memory clock: 1215 MHz (2.4 Gbps effective) for the A100, 1750 MHz (14 Gbps effective) for the CMP 40HX
- Memory size: 40 GB for the A100, 8 GB for the CMP 40HX
- Memory type: HBM2e for the A100, GDDR6 for the CMP 40HX
- Memory bus width: 5120 bit for the A100, 256 bit for the CMP 40HX
- Memory bandwidth: 1.56 TB/s for the A100, 448.0 GB/s for the CMP 40HX
- Shading units: 6912 for the A100, 2304 for the CMP 40HX
- TMUs: 432 for the A100, 144 for the CMP 40HX
- ROPs: 160 for the A100, 64 for the CMP 40HX
- RT cores: not listed for the A100, 36 for the CMP 40HX
- Tensor cores: 432 for the A100, 288 for the CMP 40HX
- Pixel rate: 225.6 GPixel/s for the A100, 105.6 GPixel/s for the CMP 40HX
- Texture rate: 609.1 GTexel/s for the A100, 237.6 GTexel/s for the CMP 40HX
- FP32: 19.49 TFLOPS for the A100, 7.603 TFLOPS for the CMP 40HX
- FP16: 77.97 TFLOPS (4:1) for the A100, 15.21 TFLOPS (2:1) for the CMP 40HX
- TDP: 250 W for the A100, 185 W for the CMP 40HX
- Power connectors: 8-pin EPS for the A100, 1x 8-pin for the CMP 40HX
- Suggested PSU: 600 W for the A100, 450 W for the CMP 40HX
- Bus interface: PCIe 4.0 x16 for the A100, PCIe 1.0 x4 for the CMP 40HX
- DirectX: not listed for the A100, 12 Ultimate (12_2) for the CMP 40HX
- OpenGL: not listed for the A100, 4.6 for the CMP 40HX
- Vulkan: not listed for the A100, 1.4 for the CMP 40HX
- Length: 267 mm (10.5 inches) for the A100, 229 mm (9 inches) for the CMP 40HX
- Width: not listed for the A100, 35 mm (1.4 inches) for the CMP 40HX
- Release date: 2020-06-21 for the A100, 2021-02-24 for the CMP 40HX
- Predecessor: Tesla Turing for the A100, not listed for the CMP 40HX
- Successor: Server Ada for the A100, not listed for the CMP 40HX
- Launch MSRP: not listed for the A100, 699 USD for the CMP 40HX
- Geekbench OpenCL score: 178,627 for the A100, 93,395 for the CMP 40HX
- Geekbench Vulkan score: 146,380 for the A100, 77,879 for the CMP 40HX
- Average benchmark score: 162,504 for the A100, 85,637 for the CMP 40HX
- Percentile vs all GPUs: 97 for the A100, 93 for the CMP 40HX
The cards share no identical specification fields other than manufacturer, foundry, slot width (dual-slot), height (111 mm), display outputs (none), and production status (end-of-life). The A100 is a server compute engine with a wide bus and massive memory pool; the CMP 40HX is a mining-oriented Turing card with a narrow PCIe link and modest memory. The specification sheet alone explains the benchmark results.