AMD Radeon Pro W6900X vs NVIDIA B200 Comparison

AMD
RADEON

AMD Radeon Pro W6900X

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2171 MHz
TDP 300 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

B200

CORE STATE GB100
VRAM 90 GB
CLOCK SPEED 1965 MHz
TDP 1000 W
BUS WIDTH 4096 bit
ARCHITECTURE Blackwell
nm
PROCESS 5 nm
LAUNCH DATE —

PERFORMANCE BENCHMARKS

geekbench_metal
226,821
N/A
geekbench_opencl
130,035
345,482
geekbench_vulkan
148,865
N/A

Analysis: AMD Radeon Pro W6900X vs NVIDIA B200

The NVIDIA B200 and AMD Radeon Pro W6900X occupy vastly different corners of the GPU landscape. One is a server-class accelerator built for massive compute throughput, while the other is a workstation-oriented card designed for Apple Mac Pro systems. The benchmark data places them far apart, with the B200 delivering a single dominant result that underscores its performance class.

Head-to-Head Benchmarks

The only direct benchmark comparison available is the Geekbench OpenCL test, and the result is a decisive victory for the NVIDIA B200. The B200 scores 345,482 points, while the Radeon Pro W6900X manages 130,035 points. This translates to a delta of 165.7% in favor of the NVIDIA part. In practical terms, the B200 delivers more than 2.6 times the raw OpenCL compute performance of the AMD card.

This gap is not subtle. The B200’s score places it in the 100th percentile of all GPUs, meaning it outperforms every other graphics card in the database on this workload. The W6900X, by contrast, sits in the 97th percentile, which is still excellent but clearly a tier below. The difference is comparable to the entire performance envelope of a mid-range GPU. The B200’s nearest rival, the NVIDIA B300 SXM6 AC, scores 369,831 points, a 6.6% advantage over the B200. The B200, in turn, leads the NVIDIA H200 NVL by 3.2%, with that card scoring 334,891 points.

For the Radeon Pro W6900X, its competitive set is much closer. Its average benchmark score of 168,574 places it just 1.5% ahead of the NVIDIA RTX 4500 Ada Generation, which scores 166,094. It also edges out the NVIDIA RTX A5500 by 2% and the AMD Radeon PRO W7800 by 2.2%. The delta between the W6900X and its closest rivals is measured in single digits, while the delta between the B200 and the W6900X is measured in triple digits. This shows that the AMD card is competitive within its own tier, but that tier is fundamentally different from the one the B200 occupies.

The B200’s OpenCL score is also worth contextualizing against its other rivals. It leads the AMD Instinct MI300X by 8.6% and the NVIDIA L40S by 16.8%. These are substantial margins, indicating that the B200 is not just slightly faster than its competition but meaningfully ahead in this specific workload. The W6900X, meanwhile, is essentially neck-and-neck with its nearest competitors, with the largest gap being a 3.7% lead over the NVIDIA A100 PCIe 40 GB.

The Verdict

The data paints a clear picture: the NVIDIA B200 and AMD Radeon Pro W6900X are not direct competitors. The B200 is in a performance class that the W6900X cannot approach, with a 165.7% lead in the sole head-to-head benchmark. Anyone needing maximum OpenCL compute throughput should choose the NVIDIA B200 without hesitation. Its 100th percentile ranking and its ability to outpace the H200 NVL, MI300X, and L40S make it the obvious choice for server workloads where raw performance is the priority.

The Radeon Pro W6900X, however, has its own merits. It is a 97th percentile GPU, and its average benchmark score of 168,574 is respectable. It trades blows with the RTX 4500 Ada Generation and RTX A5500, coming out slightly ahead of both. For users within the Apple Mac ecosystem, it offers a specific feature set that the B200 cannot provide, including display outputs and a compact form factor. The B200, by contrast, has no display outputs at all, making it unsuitable for any workstation role that requires a monitor connection.

The choice comes down to workload context. If the task is headless server-side compute, the B200 is the only sensible option. If the task requires a GPU inside a Mac Pro with video output capabilities, the W6900X is the designated part. The benchmark gap is enormous, but it is also irrelevant if the B200 cannot physically fit into the system or connect to a display. The data does not suggest that the W6900X is a bad GPU; it suggests that it is a different kind of GPU.

FAQ

Q: How much faster is the NVIDIA B200 than the AMD Radeon Pro W6900X in OpenCL?

A: The B200 scores 345,482 points compared to 130,035 for the W6900X, a 165.7% advantage.

Q: What percentile does each GPU rank in globally?

A: The NVIDIA B200 is in the 100th percentile of all GPUs, while the AMD Radeon Pro W6900X is in the 97th percentile.

Q: How does the B200 compare to its closest rival, the NVIDIA B300 SXM6 AC?

A: The B300 SXM6 AC scores 369,831 points, which is 6.6% higher than the B200's 345,482.

Q: Which GPUs are closest in performance to the Radeon Pro W6900X?

A: The NVIDIA RTX 4500 Ada Generation is 1.5% slower, the NVIDIA RTX A5500 is 2% slower, and the AMD Radeon PRO W7800 is 2.2% slower than the W6900X.

Q: Does the Radeon Pro W6900X have any benchmark results other than OpenCL?

A: Yes, it also has a Geekbench Metal score of 226,821 and a Geekbench Vulkan score of 148,865.

Q: What is the average benchmark score for each GPU?

A: The NVIDIA B200 has an average score of 345,482, while the AMD Radeon Pro W6900X has an average score of 168,574.

Specification Differences

The two GPUs differ across nearly every major specification category. The NVIDIA B200 is built on a 5 nm process at TSMC, while the AMD Radeon Pro W6900X uses a 7 nm process, also at TSMC. Transistor counts reflect the generational gap: the B200 packs 104,000 million transistors, whereas the W6900X has 26,800 million. The B200 does not list a die size, but the W6900X measures 520 mm² with a transistor density of 51.5M per mm².

Memory configurations are starkly different. The B200 features 90 GB of HBM3e memory on a 4096-bit bus, delivering 4.10 TB/s of bandwidth. The W6900X has 32 GB of GDDR6 memory on a 256-bit bus, with a bandwidth of 512.0 GB/s. Clock speeds also diverge: the B200 has a base clock of 700 MHz and a boost clock of 1965 MHz, while the W6900X runs at 1825 MHz base and 2171 MHz boost. Memory clocks are listed as 2000 MHz with 8 Gbps effective for the B200, and 2000 MHz with 16 Gbps effective for the W6900X.

Compute unit counts favor the NVIDIA part heavily. The B200 has 18,944 shading units, 592 TMUs, and 24 ROPs, plus 592 tensor cores. The W6900X has 5,120 shading units, 320 TMUs, 128 ROPs, and 80 ray tracing cores. The B200 lists no RT cores, while the W6900X lists no tensor cores. Pixel and texture rates also differ: the B200 produces 47.16 GPixel/s and 1,163.3 GTexel/s, while the W6900X produces 277.9 GPixel/s and 694.7 GTexel/s.

Power and form factor are another major divide. The B200 has a TDP of 1000 W and requires a suggested PSU of 1400 W, mounted as an SXM Module. The W6900X has a TDP of 300 W with a suggested PSU of 700 W. The B200 uses a PCIe 5.0 x16 interface, while the W6900X uses an Apple MPX interface. The B200 has no display outputs, but the W6900X has 1x HDMI 2.1 and 4x Thunderbolt outputs. The W6900X measures 267 mm in length and 120 mm in height.

Architecture Differences

The architectural gap between these two GPUs is fundamental. The NVIDIA B200 is based on the Blackwell architecture, specifically the GB100 chip, and belongs to the Server Blackwell (Bxx) generation. The AMD Radeon Pro W6900X is built on RDNA 2.0, using the Navi 21 chip, and belongs to the Radeon Pro Mac (Navi II Series) generation. These are entirely different design philosophies: Blackwell is optimized for server-scale compute and AI workloads, while RDNA 2.0 is a graphics-first architecture adapted for professional Mac workstations.

The B200’s feature set is heavily weighted toward tensor operations, with 592 tensor cores enabling FP16 throughput of 1,191.2 TFLOPS at a 16:1 ratio. Its FP32 performance is 74.45 TFLOPS. The W6900X, lacking tensor cores, delivers FP16 at 44.46 TFLOPS with a 2:1 ratio and FP32 at 22.23 TFLOPS. The B200’s FP16 advantage is enormous, reflecting its design for deep learning and scientific computing. The W6900X instead includes 80 ray tracing cores, which the B200 does not list, catering to rendering and visualization tasks.

The B200 also has a different API support profile. It lists no DirectX, OpenGL, or Vulkan versions, whereas the W6900X supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. This reinforces the B200’s role as a compute accelerator without graphics output, while the W6900X is a full-featured graphics card. The B200’s predecessor is Server Hopper, and its successor is Server Rubin, indicating a clear roadmap in the server segment. The W6900X, now end-of-life, has no listed predecessor or successor. The B200 remains in active production, while the W6900X has been discontinued, a status that aligns with its 2021 release date.

DETAILED SPECIFICATIONS

SPECIFICATION
Pro W6900X
B200
Core Specs
Shading Units
5,120
18,944 +270.0%
Shaders
5,120
18,944 +270.0%
TMUs
320
592 +85.0%
ROPs
128
24 -81.3%
Compute Units
80
—
SM Count
—
148
Clocks
Base Clock
1825 MHz
700 MHz
Boost Clock
2171 MHz
1965 MHz
Memory Clock
2000 MHz 16 Gbps effective
2000 MHz 8 Gbps effective
Memory
Memory Size
32 GB
90 GB
VRAM (MB)
32,768
92,160 +181.3%
Memory Type
GDDR6
HBM3e
Memory Bus
256 bit
4096 bit
Bandwidth
512.0 GB/s
4.10 TB/s
Cache
L1 Cache
128 KB per Array
256 KB (per SM)
L2 Cache
4 MB
50 MB
L3 Cache
128 MB
—
L0 Cache
32 KB per WGP
—
Performance
Pixel Rate
277.9 GPixel/s
47.16 GPixel/s
Texture Rate
694.7 GTexel/s
1,163.3 GTexel/s
FP32 (TFLOPS)
22.23 TFLOPS
74.45 TFLOPS
FP64 (TFLOPS)
1,389.4 GFLOPS (1:16)
37.22 TFLOPS (1:2)
FP16 (TFLOPS)
44.46 TFLOPS (2:1)
1,191.2 TFLOPS (16:1)
AI/RT
RT Cores
80
—
Tensor Cores
—
592
Power
TDP
300 W
1000 W
TDP (W)
300
1,000 +233.3%
Suggested PSU
700 W
1400 W
Architecture
Architecture
RDNA 2.0
Blackwell
GPU Name
Navi 21
GB100
Generation
Radeon Pro Mac (Navi II Series)
Server Blackwell (Bxx)
Process Size
7 nm
5 nm
Transistors
26,800 million
104,000 million
Die Size
520 mm²
—
Foundry
TSMC
TSMC
Density
51.5M / mm²
—
API Support
DirectX
12 Ultimate (12_2)
—
OpenGL
4.6
—
Vulkan
1.4
—
OpenCL
2.1
3.0
CUDA
—
10.0
Shader Model
6.8
—
Physical
Slot Width
—
SXM Module
Length
267 mm 10.5 inches
—
Height
120 mm 4.7 inches
—
Outputs
1x HDMI 2.14x Thunderbolt
No outputs
Bus Interface
Apple MPX
PCIe 5.0 x16
Other
Launch Price
5,999 USD
—
Production
End-of-life
Active
Predecessor
—
Server Hopper
Successor
—
Server Rubin
View Radeon Pro W6900X Details View B200 Details