NVIDIA CMP 40HX vs NVIDIA PG506-232 Comparison
NVIDIA CMP 40HX
PG506-232
PERFORMANCE BENCHMARKS
Analysis: NVIDIA CMP 40HX vs NVIDIA PG506-232
Head-to-Head Benchmarks
The only directly comparable benchmark recorded in the database is Geekbench OpenCL. In that test, the NVIDIA PG506-232 scores 225,124, while the NVIDIA CMP 40HX scores 93,395. This gives the PG506-232 a 141% advantage, a decisive margin that places the two cards in completely different performance tiers. The PG506-232 lands at the 99th percentile among all GPUs in the database, while the CMP 40HX sits at the 93rd percentile. That 6 percentile point gap might sound modest, but the raw score difference tells the real story: the PG506-232 delivers roughly two and a half times the OpenCL compute throughput of the CMP 40HX.
Looking at the nearest rivals for each card helps frame these numbers. The PG506-232 edges out the AMD Radeon PRO W7900D by 2.4% and the NVIDIA A100 PCIe 80 GB by 8.7%, while trailing the NVIDIA L20 by 10.4%. It sits comfortably ahead of the NVIDIA RTX 6000D by 14.9%. This places the PG506-232 in the upper echelon of workstation and server accelerators. The CMP 40HX, by contrast, trades blows with mid-range professional cards: it trails the AMD Radeon PRO W7600 by 1.7% and the NVIDIA Quadro GP100 by 2.1%, while beating the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%. In other words, the CMP 40HX is competitive with mid-tier workstation GPUs, but it is nowhere near the performance class of the PG506-232.
The CMP 40HX does have one additional recorded benchmark: Geekbench Vulkan, where it scores 77,879. The PG506-232 has no Vulkan score in the database, and its API support fields for DirectX, OpenGL, and Vulkan are all null, reflecting its compute-focused design with no display outputs. The CMP 40HX supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, so in any graphics API workload the CMP 40HX is the only one of the two that even has a path forward. But for raw compute, the PG506-232 is overwhelmingly faster.
FAQ
Q: Which card has the higher average benchmark score?
A: The NVIDIA PG506-232 has an average benchmark score of 225,124, while the NVIDIA CMP 40HX averages 85,637. That is a difference of roughly 163%.
Q: Does the CMP 40HX beat the PG506-232 in any benchmark?
A: No. The only shared benchmark is Geekbench OpenCL, and the PG506-232 wins that test outright. The CMP 40HX does have a Vulkan score of 77,879, but the PG506-232 has no Vulkan result recorded, so there is no direct comparison available in that API.
Q: How does the PG506-232 compare to the A100?
A: The PG506-232 scores 8.7% higher than the NVIDIA A100 PCIe 80 GB in the database. It also beats the RTX 6000D by 14.9% and the Radeon PRO W7900D by 2.4%, while the L20 outperforms it by 10.4%.
Q: Where does the CMP 40HX sit relative to its closest competitors?
A: The CMP 40HX is slightly behind the AMD Radeon PRO W7600 by 1.7% and the NVIDIA Quadro GP100 by 2.1%. It leads the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%.
Q: Can either card output video to a display?
A: No. Both cards have no display outputs. They are designed for compute or mining workloads, not for connecting monitors.
Q: What is the power requirement for each card?
A: Both cards have a suggested PSU rating of 450 W. The PG506-232 has a TDP of 165 W and uses an 8-pin EPS power connector, while the CMP 40HX has a TDP of 185 W and uses a single 8-pin connector.
Architecture Differences
The two cards come from different NVIDIA architectures and different process nodes. The PG506-232 is built on the Ampere architecture using the GA100 chip, fabricated on TSMC's 7 nm process. It packs 54,200 million transistors into a die size of 826 mm², which works out to a transistor density of 65.6 million transistors per square millimeter. The CMP 40HX, on the other hand, uses the older Turing architecture with the TU106 chip, built on TSMC's 12 nm process. It contains 10,800 million transistors on a 445 mm² die, giving it a density of 24.3 million transistors per square millimeter. The PG506-232 is not just bigger, it is more than twice as dense per square millimeter.
The memory subsystems are fundamentally different as well. The PG506-232 uses 24 GB of HBM2 memory on a 3072-bit bus, delivering 933.1 GB/s of bandwidth. The CMP 40HX uses 8 GB of GDDR6 on a 256-bit bus, with 448.0 GB/s of bandwidth. That is a massive gap: the PG506-232 has three times the memory capacity and more than double the bandwidth. The memory clock figures reflect this too, with the PG506-232 running at 1215 MHz (2.4 Gbps effective) and the CMP 40HX at 1750 MHz (14 Gbps effective). The GDDR6 runs at a much higher effective data rate, but the narrow bus and smaller capacity cannot compensate for the HBM2 advantage.
Compute resources differ sharply. The PG506-232 has 3584 shading units, 224 texture mapping units, 96 ROPs, and 224 tensor cores. It has no RT cores listed. The CMP 40HX has 2304 shading units, 144 TMUs, 64 ROPs, 288 tensor cores, and 36 RT cores. The PG506-232 has more raw shading and texture hardware, but the CMP 40HX actually has more tensor cores and includes ray tracing cores. The FP16 output tells an interesting story: the PG506-232 delivers 10.32 TFLOPS of FP16 at a 1:1 ratio with FP32, while the CMP 40HX delivers 15.21 TFLOPS of FP16 at a 2:1 ratio. The CMP 40HX is the faster card for FP16 work, despite being far slower in FP32.
The PG506-232 belongs to the Server Ampere generation with a release date of April 11, 2021, and its predecessor is Tesla Turing with a successor of Server Ada. The CMP 40HX belongs to the Mining GPUs generation, released February 24, 2021, with no predecessor or successor listed. Both are marked end-of-life in the database.
Specification Differences
The PG506-232 and CMP 40HX differ in nearly every specification category. Starting with the chip: the PG506-232 uses GA100 on 7 nm, while the CMP 40HX uses TU106 on 12 nm. Transistor counts are 54,200 million versus 10,800 million. Die sizes are 826 mm² versus 445 mm². Transistor density is 65.6M per mm² versus 24.3M per mm².
Clocks: the PG506-232 has a base clock of 930 MHz and a boost of 1440 MHz. The CMP 40HX has a base clock of 1470 MHz and a boost of 1650 MHz. The CMP 40HX runs at higher clock speeds, but the PG506-232 compensates with far more hardware. Memory clocks are 1215 MHz (2.4 Gbps effective) for the PG506-232 and 1750 MHz (14 Gbps effective) for the CMP 40HX.
Memory: 24 GB HBM2 on a 3072-bit bus with 933.1 GB/s bandwidth versus 8 GB GDDR6 on a 256-bit bus with 448.0 GB/s bandwidth. The PG506-232 has triple the capacity and more than double the bandwidth.
Compute units: the PG506-232 has 3584 shading units, 224 TMUs, 96 ROPs, and 224 tensor cores. The CMP 40HX has 2304 shading units, 144 TMUs, 64 ROPs, 288 tensor cores, and 36 RT cores. The PG506-232 has no RT cores listed, while the CMP 40HX includes them.
Pixel and texture rates: the PG506-232 achieves 138.2 GPixel/s and 322.6 GTexel/s. The CMP 40HX achieves 105.6 GPixel/s and 237.6 GTexel/s. The PG506-232 leads in both.
FP32 and FP16: the PG506-232 delivers 10.32 TFLOPS in both FP32 and FP16 (1:1). The CMP 40HX delivers 7.603 TFLOPS in FP32 and 15.21 TFLOPS in FP16 (2:1). The PG506-232 wins FP32, the CMP 40HX wins FP16.
Power and physical specs: the PG506-232 has a TDP of 165 W and uses an 8-pin EPS connector. The CMP 40HX has a TDP of 185 W and uses a single 8-pin connector. Both suggest a 450 W PSU. Both are dual-slot cards. The PG506-232 measures 267 mm long and 112 mm high. The CMP 40HX measures 229 mm long, 111 mm high, and 35 mm wide.
Bus interface: the PG506-232 uses PCIe 4.0 x16. The CMP 40HX uses PCIe 1.0 x4, which is an unusual and limiting choice that likely reflects its mining-oriented design, where bandwidth to the host is less critical.
API support: the PG506-232 has no recorded DirectX, OpenGL, or Vulkan support. The CMP 40HX supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.
Display outputs: neither card has any.
The CMP 40HX has a launch MSRP of 699 USD, stated once here for reference. The PG506-232 has no launch MSRP recorded.
Where Each One Wins
The PG506-232 wins decisively in OpenCL compute performance, with a 141% lead over the CMP 40HX. It also wins on memory capacity, bandwidth, FP32 throughput, pixel rate, texture rate, and transistor density. For any workload that depends on large memory allocations, high bandwidth, or raw FP32 compute, the PG506-232 is the obvious choice. Its 24 GB of HBM2 memory is well suited for large data sets, and its 933.1 GB/s bandwidth ensures that data can be fed to the compute units without bottlenecking. The card's 99th percentile ranking among all GPUs in the database reinforces this position. It is a top-tier compute accelerator, closer in performance to the L20 and A100 than to anything in the CMP 40HX's class.
The CMP 40HX wins in a few specific areas. It has higher base and boost clocks, more tensor cores, and includes 36 RT cores. Its FP16 output of 15.21 TFLOPS is significantly higher than the PG506-232's 10.32 TFLOPS, so any workload that relies on FP16 tensor operations might favor the CMP 40HX despite its lower FP32 performance. It also has full API support for DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, whereas the PG506-232 has no recorded API support. The CMP 40HX is physically smaller at 229 mm in length, making it easier to fit into tight chassis. Its PCIe 1.0 x4 interface is not a point in its favor, but it does not affect compute throughput that stays on the card.
For mining workloads, the CMP 40HX was purpose-built. Its Turing architecture, higher clocks, and 2:1 FP16 ratio make it a capable hashing card, and its 185 W TDP is modest. The PG506-232, with its 165 W TDP and far larger compute resources, is a server-class accelerator aimed at data center tasks like AI inference, scientific simulation, and high-performance computing. The two cards were designed for different markets, and the benchmark data reflects that split.
The Verdict
The data is unambiguous: the NVIDIA PG506-232 is the far more powerful card. It beats the CMP 40HX by 141% in the only shared benchmark, outscores it by roughly 163% in average benchmark score, and sits at the 99th percentile compared to the CMP 40HX's 93rd. It has triple the memory, more than double the bandwidth, and a commanding lead in FP32 compute. Anyone choosing between these two for compute-heavy work should pick the PG506-232 without hesitation.
The case for the CMP 40HX rests on narrower grounds. It is the only one of the two with graphics API support, the only one with RT cores, and the only one with a Vulkan benchmark score. Its FP16 output is higher, and its physical footprint is smaller. If the workload is FP16-centric, or if software requires Vulkan or DirectX support, the CMP 40HX becomes relevant. But those are edge cases. For general compute, memory-bound tasks, and any FP32-heavy workload, the PG506-232 is in a different league.
The CMP 40HX's nearest rivals are mid-range professional cards like the Radeon PRO W7600 and Quadro GP100, and it trades wins and losses with them within a few percentage points. The PG506-232's nearest rivals are the L20, Radeon PRO W7900D, and A100, all top-tier accelerators. That comparison alone tells the buyer which tier each card occupies.
Both cards are end-of-life, so availability will be on the secondary market. The PG506-232 has no launch MSRP recorded, while the CMP 40HX launched at 699 USD, but pricing history is not the deciding factor here. The deciding factor is performance, and the PG506-232 wins that contest by a wide margin. Choose the PG506-232 for compute density, memory capacity, and raw throughput. Choose the CMP 40HX only if FP16 performance, API compatibility, or physical size are the dominant constraints. In every other scenario, the PG506-232 is the correct pick.