NVIDIA B300 SXM6 AC vs NVIDIA CMP 40HX Comparison
NVIDIA B300 SXM6 AC
CMP 40HX
PERFORMANCE BENCHMARKS
Analysis: NVIDIA B300 SXM6 AC vs NVIDIA CMP 40HX
Where Each One Wins
The benchmark landscape between these two NVIDIA products is remarkably one-sided, but the nature of that imbalance tells a story about their intended purposes. The NVIDIA B300 SXM6 AC dominates the single recorded benchmark in the database, the Geekbench OpenCL test, where it posts a score of 369,831 against the CMP 40HX's 93,395. That is a 296% advantage, a gap so wide that the CMP 40HX never wins a single benchmark in the head-to-head records.
The B300 SXM6 AC sits at the 100th percentile among all GPUs in the database, meaning no recorded GPU scores higher. Its nearest rivals, the NVIDIA B200 at 345,482 and the NVIDIA H200 NVL at 334,891, trail by 7% and 10.4% respectively. The AMD Instinct MI300X at 317,994 is 16.3% behind, and the NVIDIA L40S at 295,763 is a full 25% back. This is a compute monster designed for the absolute top of the server hierarchy.
The CMP 40HX, by contrast, occupies the 93rd percentile, which is still respectable but places it in a completely different competitive tier. Its nearest rivals include the AMD Radeon PRO W7600 at 87,108 (1.7% ahead), the NVIDIA Quadro GP100 at 87,445 (2.1% ahead), the AMD Radeon PRO W6600 at 81,995 (4.4% behind), and the AMD Radeon Pro Vega 64X at 80,959 (5.8% behind). The CMP 40HX was built for a specific mining workload, not general compute leadership, and its modest OpenCL showing reflects that narrow focus.
The use-case split is therefore stark. The B300 SXM6 AC wins every scenario where raw compute throughput matters: AI training, scientific simulation, massive data center workloads. The CMP 40HX, with its mining-specific lineage, wins no benchmark here, but its existence in the database with a 93rd percentile score suggests it still holds value in niche applications where power efficiency and compact size outweigh raw performance.
Architecture Differences
The architectural gap between these two GPUs is generational and foundational. The B300 SXM6 AC uses the GB110 chip on the Blackwell Ultra architecture, fabricated on a 5 nm process at TSMC. It packs 208,000 million transistors onto a 1628 mm² die, yielding a transistor density of 127.8 million per square millimeter. The CMP 40HX uses the TU106 chip on the older Turing architecture, also TSMC but on a 12 nm process, with just 10,800 million transistors on a 445 mm² die, for a density of 24.3 million per square millimeter.
The memory subsystems could not be more different. The B300 SXM6 AC carries 288 GB of HBM3e across an 8192-bit bus, delivering 8.19 TB/s of bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus, providing 448.0 GB/s. That is a 18-fold difference in capacity and a roughly 18-fold difference in bandwidth, both favoring the newer part.
Compute resources tell a similar tale. The B300 SXM6 AC fields 18,944 shading units, 592 TMUs, 592 tensor cores, and just 24 ROPs. The CMP 40HX has 2,304 shading units, 144 TMUs, 288 tensor cores, and 64 ROPs. The B300's pixel rate of 48.77 GPixel/s is actually lower than the CMP 40HX's 105.6 GPixel/s, a consequence of the B300's minimal ROP count, but its texture rate of 1,202.9 GTexel/s crushes the CMP's 237.6 GTexel/s.
Clock speeds differ as well. The B300 SXM6 AC runs at 1665 MHz base and 2032 MHz boost, with memory at 2000 MHz (8 Gbps effective). The CMP 40HX runs at 1470 MHz base and 1650 MHz boost, with memory at 1750 MHz (14 Gbps effective). Despite lower clock speeds, the B300's massive parallelism produces 76.99 TFLOPS of FP32 and FP16 (1:1), while the CMP 40HX delivers 7.603 TFLOPS FP32 and 15.21 TFLOPS FP16 (2:1). The B300 is over 10 times faster in FP32.
The interface and power profiles diverge sharply. The B300 SXM6 AC uses PCIe 6.0 x16, consumes up to 1100 W, and comes as an SXM module with no display outputs. The CMP 40HX uses PCIe 1.0 x4, draws 185 W, takes a single 8-pin connector, and is a dual-slot card with no display outputs. The B300 also supports a full API stack for graphics (DirectX 12 Ultimate, OpenGL 4.6, Vulkan 1.4), while the B300 lists N/A for all graphics APIs, confirming its compute-only orientation.
Head-to-Head Benchmarks
The sole recorded head-to-head benchmark is Geekbench OpenCL, and the result is categorical. The NVIDIA B300 SXM6 AC scores 369,831 against the CMP 40HX's 93,395, a delta of 296%. This means the B300 is roughly four times faster in this workload, not merely incrementally ahead.
Contextualizing within their rival groups sharpens the picture. The B300's 369,831 places it 7% above the B200's 345,482, 10.4% above the H200 NVL's 334,891, 16.3% above the MI300X's 317,994, and 25% above the L40S's 295,763. Each of those rivals is itself a flagship-class accelerator, so the B300's margin over them is significant. The CMP 40HX's 93,395, meanwhile, sits within 5.8% of its closest rivals, the W7600, GP100, W6600, and Vega 64X. The CMP is competitive within its own tier but utterly outclassed by the B300.
The FP32 compute gap reinforces the benchmark result. The B300's 76.99 TFLOPS is more than ten times the CMP 40HX's 7.603 TFLOPS. The B300's FP16 output matches its FP32 at 76.99 TFLOPS, whereas the CMP 40HX doubles its FP32 to reach 15.21 TFLOPS in FP16. Even the CMP's best-case FP16 number is less than a quarter of the B300's FP32 output.
Memory bandwidth tells a parallel story. The B300's 8.19 TB/s dwarfs the CMP's 448.0 GB/s by a factor of roughly 18. For workloads that are memory-bound, which many AI and HPC tasks are, that bandwidth advantage is decisive. The B300's 288 GB capacity also allows it to hold far larger models and datasets in memory without spilling to slower storage.
FAQ
Q: Which GPU has a higher Geekbench OpenCL score?
A: The NVIDIA B300 SXM6 AC scores 369,831, which is 296% higher than the CMP 40HX's 93,395.
Q: How do the memory capacities compare?
A: The B300 SXM6 AC has 288 GB of HBM3e, while the CMP 40HX has 8 GB of GDDR6. The B300's bandwidth is 8.19 TB/s versus 448.0 GB/s for the CMP.
Q: What are the transistor counts and process nodes?
A: The B300 SXM6 AC uses 208,000 million transistors on a 5 nm process. The CMP 40HX uses 10,800 million transistors on a 12 nm process.
Q: Is the CMP 40HX still in production?
A: No, the CMP 40HX is marked as end-of-life, while the B300 SXM6 AC is listed as active in production.
Q: What is the power draw difference?
A: The B300 SXM6 AC has a TDP of 1100 W with a suggested PSU of 1500 W. The CMP 40HX has a TDP of 185 W with a suggested PSU of 450 W.
Q: Does either GPU support graphics APIs?
A: The CMP 40HX supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4. The B300 SXM6 AC lists N/A for all graphics APIs, reflecting its compute-only design.
The Verdict
The data points to an unambiguous conclusion for different audiences. The NVIDIA B300 SXM6 AC is the choice for anyone running large-scale AI training, inference, or scientific computing where every TFLOPS and every gigabyte of bandwidth matters. Its 76.99 TFLOPS FP32, 8.19 TB/s memory bandwidth, and 288 GB capacity put it in a class of its own, confirmed by its 100th percentile standing and 296% lead over the CMP 40HX in the single benchmark. The 7% to 25% margins over its nearest rivals (B200, H200 NVL, MI300X, L40S) show it is not just better than the CMP; it is measurably ahead of the next-best accelerators in the database.
The NVIDIA CMP 40HX, by contrast, is a legacy mining part that has been end-of-life since its 2021 release. Its 93rd percentile score and tight clustering with professional workstation GPUs like the W7600 and Quadro GP100 indicate it remains a competent but unremarkable performer for its era. Its 185 W power draw and dual-slot form factor make it far more accessible in terms of power and physical space, but its 7.603 TFLOPS FP32 and 448.0 GB/s bandwidth are simply not in the same league.
For a builder assembling a server-grade compute node, the B300 SXM6 AC is the only rational choice from this pairing. For someone seeking a low-power, compact accelerator for legacy mining or light compute duties, the CMP 40HX still functions, but the performance gap is so vast that any modern workload would quickly expose its limits. The verdict is not a close call: the B300 SXM6 AC wins on every measured axis, and the CMP 40HX's only advantages are its lower power draw, smaller physical footprint, and support for graphics APIs that the B300 does not offer.
Specification Differences
| Specification | NVIDIA B300 SXM6 AC | NVIDIA CMP 40HX |
|----------------|---------------------|------------------|
| Architecture | Blackwell Ultra | Turing |
| Process Node | 5 nm | 12 nm |
| Transistors | 208,000 million | 10,800 million |
| Die Size | 1628 mm² | 445 mm² |
| Transistor Density | 127.8M / mm² | 24.3M / mm² |
| Base Clock | 1665 MHz | 1470 MHz |
| Boost Clock | 2032 MHz | 1650 MHz |
| Memory Clock | 2000 MHz (8 Gbps effective) | 1750 MHz (14 Gbps effective) |
| Memory Size | 288 GB | 8 GB |
| Memory Type | HBM3e | GDDR6 |
| Memory Bus | 8192 bit | 256 bit |
| Memory Bandwidth | 8.19 TB/s | 448.0 GB/s |
| Shading Units | 18,944 | 2,304 |
| TMUs | 592 | 144 |
| ROPs | 24 | 64 |
| Tensor Cores | 592 | 288 |
| RT Cores | Not listed | 36 |
| Pixel Rate | 48.77 GPixel/s | 105.6 GPixel/s |
| Texture Rate | 1,202.9 GTexel/s | 237.6 GTexel/s |
| FP32 Performance | 76.99 TFLOPS | 7.603 TFLOPS |
| FP16 Performance | 76.99 TFLOPS (1:1) | 15.21 TFLOPS (2:1) |
| TDP | 1100 W | 185 W |
| Slot Width | SXM Module | Dual-slot |
| Power Connectors | Not listed | 1x 8-pin |
| Suggested PSU | 1500 W | 450 W |
| Bus Interface | PCIe 6.0 x16 | PCIe 1.0 x4 |
| Graphics APIs | N/A | DirectX 12 Ultimate, OpenGL 4.6, Vulkan 1.4 |
| Production Status | Active | End-of-life |
| Release Date | 2025-09-10 | 2021-02-24 |
| Launch MSRP | Not listed | 699 USD |