NVIDIA B200 vs NVIDIA CMP 40HX Comparison

NVIDIA
GEFORCE

NVIDIA B200

CORE STATE GB100
VRAM 90 GB
CLOCK SPEED 1965 MHz
TDP 1000 W
BUS WIDTH 4096 bit
ARCHITECTURE Blackwell
nm
PROCESS 5 nm
LAUNCH DATE
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
345,482
93,395
geekbench_vulkan
N/A
77,879

Analysis: NVIDIA B200 vs NVIDIA CMP 40HX

Head-to-Head Benchmarks

The benchmark data records a single head-to-head comparison between the NVIDIA B200 and the NVIDIA CMP 40HX, and the result is decisively one-sided. In the Geekbench OpenCL test, the B200 scores 345,482 against the CMP 40HX's 93,395, a delta of 269.9% in favor of the B200. This is not a marginal gap; it is a performance chasm that places the two products in entirely different performance strata.

To contextualize the B200's score, the database shows it sits at the 100th percentile among all GPUs, meaning no recorded GPU in the database scores higher. Its nearest rivals reinforce this position: it leads the NVIDIA H200 NVL by 3.2%, trails the NVIDIA B300 SXM6 AC by 6.6%, leads the AMD Instinct MI300X by 8.6%, and leads the NVIDIA L40S by 16.8%. These deltas, while smaller than the gap to the CMP 40HX, still show the B200 operating at the very top of the performance hierarchy.

The CMP 40HX, by contrast, sits at the 93rd percentile with an average benchmark score of 85,637 across its two recorded tests (OpenCL and Vulkan). Its OpenCL score of 93,395 is 269.9% lower than the B200's, and its Vulkan score of 77,879 is even lower. Against its own nearest rivals, the CMP 40HX is essentially at parity: it trails the AMD Radeon PRO W7600 by 1.7%, trails the NVIDIA Quadro GP100 by 2.1%, leads the AMD Radeon PRO W6600 by 4.4%, and leads the AMD Radeon Pro Vega 64X by 5.8%. This places the CMP 40HX in a completely different competitive bracket, one where single-digit percentage differences matter, rather than the multi-hundred-percent gaps seen at the top.

The verdict from the head-to-head data is unambiguous: the B200 outperforms the CMP 40HX by a factor of roughly 3.7 in raw OpenCL throughput. No recorded benchmark shows the CMP 40HX winning any test. The wins tally is 1 for the B200 and 0 for the CMP 40HX.

Where Each One Wins

Given the single recorded comparison, the use-case split is stark. The B200 wins the only benchmark test that both products share, and it does so by an enormous margin. Any workload that relies on OpenCL compute performance, which includes many scientific, AI, and data-center tasks, will see a massive advantage from the B200. The data shows no scenario in which the CMP 40HX leads.

However, the CMP 40HX does hold advantages in areas not captured by the compute benchmark. Its physical footprint is smaller: it is a dual-slot card measuring 229 mm in length, 111 mm in height, and 35 mm in width, whereas the B200 is an SXM module with no recorded board dimensions. The CMP 40HX also has a dramatically lower power draw, with a TDP of 185 W versus the B200's 1000 W, and it requires only a 450 W suggested PSU against the B200's 1400 W. The CMP 40HX uses a single 8-pin power connector, while the B200's power connector is not recorded.

For use cases that prioritize density, power efficiency, or integration into existing PCIe slots with modest power delivery, the CMP 40HX is the more practical choice despite its compute deficit. The B200, on the other hand, is the clear winner for any workload where raw compute throughput is the primary constraint. The data does not show the CMP 40HX winning any performance test, so its strengths lie entirely in system integration and operational overhead.

Architecture Differences

The two GPUs come from different architectural eras and are built for different purposes. The B200 uses the GB100 chip, based on the Blackwell architecture, and belongs to the Server Blackwell generation. It is fabricated on a 5 nm process at TSMC and packs 104,000 million transistors. The CMP 40HX uses the TU106 chip, based on the Turing architecture, and belongs to the Mining GPUs generation. It is fabricated on a 12 nm process at TSMC and contains 10,800 million transistors, with a die size of 445 mm² and a transistor density of 24.3M per mm².

The memory subsystems are fundamentally different. The B200 carries 90 GB of HBM3e memory on a 4096-bit bus, delivering 4.10 TB/s of bandwidth. The CMP 40HX carries 8 GB of GDDR6 memory on a 256-bit bus, delivering 448.0 GB/s of bandwidth. This is a 9.2x difference in memory bandwidth, which directly impacts any memory-bound workload.

Compute resources also differ enormously. The B200 has 18,944 shading units, 592 TMUs, and 24 ROPs, with 592 tensor cores. The CMP 40HX has 2,304 shading units, 144 TMUs, and 64 ROPs, with 36 ray-tracing cores and 288 tensor cores. The B200's pixel rate is 47.16 GPixel/s, and its texture rate is 1,163.3 GTexel/s. The CMP 40HX's pixel rate is 105.6 GPixel/s, and its texture rate is 237.6 GTexel/s. Interestingly, the CMP 40HX has a higher pixel rate, which reflects its higher ROP count and clock speeds.

Clock speeds tell a similar story. The B200 has a base clock of 700 MHz and a boost clock of 1965 MHz, while the CMP 40HX has a base clock of 1470 MHz and a boost clock of 1650 MHz. The CMP 40HX starts higher but boosts lower, while the B200 starts very low but boosts much higher. Memory clocks also differ: the B200 runs at 2000 MHz (8 Gbps effective), while the CMP 40HX runs at 1750 MHz (14 Gbps effective).

The B200 supports PCIe 5.0 x16, while the CMP 40HX supports PCIe 1.0 x4, a significant interface difference. The B200 has no display outputs, and the CMP 40HX also has no display outputs, which is consistent with the CMP 40HX's mining purpose. However, the CMP 40HX does report API support for DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the B200 reports no API data in the database.

FAQ

Q: Which GPU has a higher OpenCL benchmark score?

A: The NVIDIA B200 scores 345,482 in Geekbench OpenCL, while the NVIDIA CMP 40HX scores 93,395. The B200 leads by 269.9%.

Q: What is the memory capacity difference between the two?

A: The B200 has 90 GB of HBM3e memory on a 4096-bit bus with 4.10 TB/s bandwidth. The CMP 40HX has 8 GB of GDDR6 memory on a 256-bit bus with 448.0 GB/s bandwidth.

Q: How do their power requirements compare?

A: The B200 has a TDP of 1000 W and a suggested PSU of 1400 W. The CMP 40HX has a TDP of 185 W and a suggested PSU of 450 W.

Q: Which GPU is positioned higher in the overall performance percentile?

A: The B200 is at the 100th percentile among all GPUs, while the CMP 40HX is at the 93rd percentile.

Q: What are the process nodes for each GPU?

A: The B200 is fabricated on a 5 nm process at TSMC. The CMP 40HX is fabricated on a 12 nm process at TSMC.

Q: Does the CMP 40HX win any benchmark against the B200?

A: No. The recorded head-to-head data shows the B200 winning the only shared benchmark (Geekbench OpenCL) with a 269.9% advantage. The wins tally is 1 for the B200 and 0 for the CMP 40HX.

The Verdict

The data supports a clear conclusion: the NVIDIA B200 is the superior compute product by every measured performance metric. Its OpenCL score is 269.9% higher, it holds the 100th percentile ranking, and it leads its nearest rivals by margins ranging from 3.2% to 16.8%. For any workload where raw compute throughput, memory bandwidth, or tensor performance is the bottleneck, the B200 is the definitive choice.

The CMP 40HX, however, is not without merit. Its 93rd percentile ranking places it above the vast majority of GPUs, and its nearest-rival deltas show it trading blows with professional workstation cards like the AMD Radeon PRO W7600 and the NVIDIA Quadro GP100. Its 185 W TDP, dual-slot form factor, and standard PCIe mounting make it far easier to integrate into existing systems, and its 8 GB of GDDR6 memory is sufficient for many compute tasks, even if it is dwarfed by the B200's 90 GB HBM3e.

Who should pick which? The answer depends on the deployment context. If the goal is maximum compute performance in a server environment with adequate power and cooling, the B200 is the only rational choice. It delivers a 269.9% performance advantage in the recorded benchmark, and its memory bandwidth of 4.10 TB/s is in a different league from the CMP 40HX's 448.0 GB/s. The B200 is also an active product with a successor already recorded (Server Rubin), indicating ongoing platform support.

If the goal is a low-power, compact compute accelerator for edge deployment, mining, or legacy system integration, the CMP 40HX is a viable option. It is end-of-life, but its 185 W power draw and 450 W suggested PSU make it accessible where the B200's 1000 W TDP would be prohibitive. The CMP 40HX also supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, whereas the B200 records no API support, making the CMP 40HX more flexible for graphics-adjacent workloads.

The verdict is not a close call on performance, but it is a nuanced call on applicability. The B200 wins compute outright; the CMP 40HX wins on operational simplicity. Choose accordingly.

Specification Differences

| Field | NVIDIA B200 | NVIDIA CMP 40HX |

|-------|-------------|-----------------|

| Chip | GB100 | TU106 |

| Architecture | Blackwell | Turing |

| Generation | Server Blackwell (Bxx) | Mining GPUs |

| Process Node | 5 nm | 12 nm |

| Transistors | 104,000 million | 10,800 million |

| Die Size | Not recorded | 445 mm² |

| Transistor Density | Not recorded | 24.3M / mm² |

| Base Clock | 700 MHz | 1470 MHz |

| Boost Clock | 1965 MHz | 1650 MHz |

| Memory Clock | 2000 MHz (8 Gbps effective) | 1750 MHz (14 Gbps effective) |

| Memory Size | 90 GB | 8 GB |

| Memory Type | HBM3e | GDDR6 |

| Memory Bus Width | 4096 bit | 256 bit |

| Memory Bandwidth | 4.10 TB/s | 448.0 GB/s |

| Shading Units | 18,944 | 2,304 |

| TMUs | 592 | 144 |

| ROPs | 24 | 64 |

| RT Cores | Not recorded | 36 |

| Tensor Cores | 592 | 288 |

| Pixel Rate | 47.16 GPixel/s | 105.6 GPixel/s |

| Texture Rate | 1,163.3 GTexel/s | 237.6 GTexel/s |

| FP32 Performance | 74.45 TFLOPS | 7.603 TFLOPS |

| FP16 Performance | 1,191.2 TFLOPS (16:1) | 15.21 TFLOPS (2:1) |

| TDP | 1000 W | 185 W |

| Slot Width | SXM Module | Dual-slot |

| Power Connectors | Not recorded | 1x 8-pin |

| Suggested PSU | 1400 W | 450 W |

| Bus Interface | PCIe 5.0 x16 | PCIe 1.0 x4 |

| API Support | Not recorded | DirectX 12 Ultimate, OpenGL 4.6, Vulkan 1.4 |

| Dimensions | Not recorded | 229 mm x 111 mm x 35 mm |

| Production Status | Active | End-of-life |

| Release Date | Not recorded | 2021-02-24 |

| Predecessor | Server Hopper | Not recorded |

| Successor | Server Rubin | Not recorded |

| Launch MSRP | Not recorded | 699 USD |

DETAILED SPECIFICATIONS

SPECIFICATION
B200
CMP 40HX
Core Specs
Shading Units
18,944
2,304 -87.8%
Shaders
18,944
2,304 -87.8%
TMUs
592
144 -75.7%
ROPs
24
64 +166.7%
SM Count
148
36 -75.7%
Clocks
Base Clock
700 MHz
1470 MHz
Boost Clock
1965 MHz
1650 MHz
Memory Clock
2000 MHz 8 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
90 GB
8 GB
VRAM (MB)
92,160
8,192 -91.1%
Memory Type
HBM3e
GDDR6
Memory Bus
4096 bit
256 bit
Bandwidth
4.10 TB/s
448.0 GB/s
Cache
L1 Cache
256 KB (per SM)
64 KB (per SM)
L2 Cache
50 MB
4 MB
Performance
Pixel Rate
47.16 GPixel/s
105.6 GPixel/s
Texture Rate
1,163.3 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
74.45 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
37.22 TFLOPS (1:2)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
1,191.2 TFLOPS (16:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
36
Tensor Cores
592
288 -51.4%
Power
TDP
1000 W
185 W
TDP (W)
1,000
185 -81.5%
Suggested PSU
1400 W
450 W
Power Connectors
1x 8-pin
Architecture
Architecture
Blackwell
Turing
GPU Name
GB100
TU106
Generation
Server Blackwell (Bxx)
Mining GPUs
Process Size
5 nm
12 nm
Transistors
104,000 million
10,800 million
Die Size
445 mm²
Foundry
TSMC
TSMC
Density
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
10.0
7.5
Shader Model
6.8
Physical
Slot Width
SXM Module
Dual-slot
Length
229 mm 9 inches
Height
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 1.0 x4
Other
Launch Price
699 USD
Production
Active
End-of-life
Predecessor
Server Hopper
Successor
Server Rubin
View B200 Details View CMP 40HX Details