AMD Instinct MI350X vs NVIDIA B300 SXM6 AC Comparison

AMD
RADEON

AMD Instinct MI350X

CORE STATE MI350 256CU
VRAM 288 GB
CLOCK SPEED 2200 MHz
TDP 1000 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 4.0
nm
PROCESS 3 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

B300 SXM6 AC

CORE STATE GB110
VRAM 288 GB
CLOCK SPEED 2032 MHz
TDP 1100 W
BUS WIDTH 8192 bit
ARCHITECTURE Blackwell Ultra
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

geekbench_opencl
N/A
369,831

Analysis: AMD Instinct MI350X vs NVIDIA B300 SXM6 AC

Head-to-Head Benchmarks

The database contains one recorded benchmark for NVIDIA B300 SXM6 AC: a Geekbench OpenCL score of 369,831. AMD Instinct MI350X has no recorded benchmark scores in the database, so the head-to-head comparison relies on the B300 SXM6 AC's nearest rival data and the architectural specifications of both accelerators.

The NVIDIA B300 SXM6 AC delivers an OpenCL score that places it in the 100th percentile of all GPUs in the database. Its nearest documented rival, NVIDIA B200, averages 345,482 points, which puts the B300 SXM6 AC 7% ahead. The B300 SXM6 AC also leads the NVIDIA H200 NVL by 10.4%, with the H200 NVL averaging 334,891 points. Against AMD Instinct MI300X, the B300 SXM6 AC holds a 16.3% advantage, as the MI300X averages 317,994 points. The NVIDIA L40S trails by 25%, averaging 295,763 points.

In raw compute specifications, the B300 SXM6 AC posts an FP32 throughput of 76.99 TFLOPS, while the MI350X delivers 72.09 TFLOPS. That is a 6.8% gap in favor of the NVIDIA part. Both accelerators maintain a 1:1 FP16 to FP32 ratio, meaning the B300 SXM6 AC also delivers 76.99 TFLOPS for FP16 workloads, while the MI350X manages 72.09 TFLOPS. The MI350X does counter with a higher texture rate at 2,252.8 GTexel/s versus 1,202.9 GTexel/s for the B300 SXM6 AC, a difference of roughly 87%. The B300 SXM6 AC, however, is the only one with a documented pixel rate at 48.77 GPixel/s, while the MI350X lists 0 MPixel/s.

Memory configurations are identical: both use 288 GB of HBM3e with an 8192-bit bus and 8.19 TB/s of bandwidth. Clock behavior differs meaningfully. The B300 SXM6 AC runs a base clock of 1665 MHz and a boost clock of 2032 MHz. The MI350X has a 1000 MHz base and a 2200 MHz boost, so the AMD part has a higher peak clock but a much lower idle or sustained base. The B300 SXM6 AC also carries more shading units at 18,944 versus 16,384 for the MI350X, a 15.6% advantage. The MI350X has more texture mapping units at 1,024 versus 592 for the B300 SXM6 AC. The NVIDIA part includes 24 ROPs and 592 tensor cores, while the MI350X lists no ROPs and no tensor cores in the recorded data.

Power draw figures place the B300 SXM6 AC at 1100 W TDP, with a suggested PSU of 1500 W. The MI350X draws 1000 W TDP with a suggested PSU of 1400 W. The B300 SXM6 AC uses a PCIe 6.0 x16 interface, while the MI350X uses PCIe 5.0 x16. Form factors differ as well: the B300 SXM6 AC mounts in an SXM module, and the MI350X is an OAM module. Neither accelerator has display outputs, and both have no applicable DirectX, OpenGL, or Vulkan support in the database.

The Verdict

The recorded data favors the NVIDIA B300 SXM6 AC in every measurable performance category. Its Geekbench OpenCL score of 369,831 places it at the 100th percentile of all GPUs in the database, and it leads each of its four nearest rivals by margins ranging from 7% to 25%. The AMD Instinct MI350X, by contrast, has no benchmark score and sits at the 50th percentile, which in the database reflects the absence of recorded performance data rather than a measured result.

The B300 SXM6 AC also has a higher FP32 throughput at 76.99 TFLOPS versus 72.09 TFLOPS for the MI350X, a larger shading unit count at 18,944 versus 16,384, and a more advanced bus interface with PCIe 6.0 x16 versus PCIe 5.0 x16. The MI350X does have a higher boost clock at 2200 MHz versus 2032 MHz, but the B300 SXM6 AC compensates with a substantially higher base clock at 1665 MHz versus 1000 MHz. The MI350X also leads in texture fill rate at 2,252.8 GTexel/s versus 1,202.9 GTexel/s, but that does not translate into a higher shading or compute score in the available data.

For buyers selecting strictly on recorded benchmark performance, the B300 SXM6 AC is the clear choice. It is faster in the single measured test, has more compute cores, and delivers higher FP32 and FP16 throughput. The MI350X is not without merit, but its advantages are confined to texture throughput and peak boost clock, neither of which appears in a recorded benchmark result. The B300 SXM6 AC also holds a higher percentile ranking, which indicates that among all GPUs in the database, it sits at the top, while the MI350X sits at the median.

Where Each One Wins

The NVIDIA B300 SXM6 AC wins in general compute workloads as measured by Geekbench OpenCL. Its score of 369,831 puts it ahead of the nearest rival B200 by 7%, and ahead of the MI300X by 16.3%. The B300 SXM6 AC also wins in FP32 and FP16 compute throughput, delivering 76.99 TFLOPS in both, which surpasses the MI350X's 72.09 TFLOPS. The B300 SXM6 AC has more shading units at 18,944, which supports higher parallel throughput in shader-heavy tasks. It also includes tensor cores, with 592 available, which the MI350X does not list in the database, suggesting an advantage for workloads that utilize tensor operations.

The AMD Instinct MI350X wins in texture throughput, with 2,252.8 GTexel/s versus 1,202.9 GTexel/s for the B300 SXM6 AC. That makes the MI350X better suited for texture-bound operations, where the rate of texture mapping unit output is the limiting factor. The MI350X also has a higher boost clock at 2200 MHz versus 2032 MHz, which could provide an edge in single-threaded or lightly threaded workloads that scale with peak clock speed. The MI350X uses a 3 nm process node, while the B300 SXM6 AC uses 5 nm, both from TSMC. The smaller node may offer efficiency advantages, though the database does not record power efficiency metrics beyond TDP, where the MI350X draws 100 W less.

Power consumption is a differentiator. The MI350X has a TDP of 1000 W, while the B300 SXM6 AC draws 1100 W. The suggested PSU is 1400 W for the MI350X and 1500 W for the B300 SXM6 AC. For deployments with strict power limits, the MI350X requires less power headroom, which may allow denser packing or lower cooling demands. The MI350X also has a larger die at 2380 mm² versus 1628 mm² for the B300 SXM6 AC, though the B300 SXM6 AC has a higher transistor density at 127.8M per mm² versus 77.7M per mm² for the MI350X.

FAQ

Q: What is the recorded benchmark score for the NVIDIA B300 SXM6 AC?

A: The B300 SXM6 AC has a Geekbench OpenCL score of 369,831, which places it at the 100th percentile of all GPUs in the database.

Q: Does the AMD Instinct MI350X have any recorded benchmark scores?

A: No. The MI350X has an empty benchmark list and an average benchmark score of 0 in the database.

Q: How does the B300 SXM6 AC compare to its nearest rival in the database?

A: The B300 SXM6 AC is 7% ahead of the NVIDIA B200, 10.4% ahead of the NVIDIA H200 NVL, 16.3% ahead of the AMD Instinct MI300X, and 25% ahead of the NVIDIA L40S.

Q: What memory configuration do both accelerators use?

A: Both the MI350X and the B300 SXM6 AC use 288 GB of HBM3e memory with an 8192-bit bus and 8.19 TB/s bandwidth.

Q: Which accelerator has a higher FP32 throughput?

A: The B300 SXM6 AC has a higher FP32 throughput at 76.99 TFLOPS, compared to 72.09 TFLOPS for the MI350X.

Q: What are the power requirements for each accelerator?

A: The MI350X has a TDP of 1000 W with a suggested PSU of 1400 W. The B300 SXM6 AC has a TDP of 1100 W with a suggested PSU of 1500 W.

Architecture Differences

The two accelerators use different manufacturing processes. The AMD Instinct MI350X is built on a 3 nm node at TSMC, while the NVIDIA B300 SXM6 AC uses a 5 nm node, also at TSMC. The MI350X has a larger die at 2380 mm², while the B300 SXM6 AC has a smaller die at 1628 mm². Transistor counts differ as well: the B300 SXM6 AC has 208,000 million transistors, while the MI350X has 185,000 million. The B300 SXM6 AC achieves a higher transistor density at 127.8M per mm², versus 77.7M per mm² for the MI350X.

The core architectures are distinct. The MI350X uses AMD's CDNA 4.0 architecture with the MI350 256CU chip, while the B300 SXM6 AC uses NVIDIA's Blackwell Ultra architecture with the GB110 chip. The MI350X belongs to the Instinct (MIx) generation, while the B300 SXM6 AC is part of the Server Blackwell (Bxx) generation. The B300 SXM6 AC lists a production status of Active, and its predecessor is Server Hopper with a successor of Server Rubin. The MI350X lists Radeon Instinct as its predecessor.

Shading unit counts differ significantly. The B300 SXM6 AC has 18,944 shading units, while the MI350X has 16,384. Texture mapping units also differ, with the MI350X having 1,024 TMUs and the B300 SXM6 AC having 592. The B300 SXM6 AC has 24 ROPs and 592 tensor cores, while the MI350X lists no ROPs and no tensor cores. The MI350X has a pixel rate of 0 MPixel/s, while the B300 SXM6 AC has 48.77 GPixel/s.

Clock behavior shows a split. The MI350X has a base clock of 1000 MHz and a boost clock of 2200 MHz. The B300 SXM6 AC has a base clock of 1665 MHz and a boost clock of 2032 MHz. Memory clocks are identical at 2000 MHz with 8 Gbps effective. Both use HBM3e memory with identical capacity, bus width, and bandwidth.

The bus interface differs, with the B300 SXM6 AC using PCIe 6.0 x16 and the MI350X using PCIe 5.0 x16. Form factors also differ: the B300 SXM6 AC is an SXM module, while the MI350X is an OAM module. The MI350X has recorded dimensions of 102 mm in length and 165 mm in width, while the B300 SXM6 AC has no recorded dimensions. Neither accelerator has display outputs, and neither supports DirectX, OpenGL, or Vulkan in the database. The B300 SXM6 AC has a release date of September 2025, while the MI350X has a release date of June 2025.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI350X
B300 SXM6 AC
Core Specs
Shading Units
16,384
18,944 +15.6%
Shaders
16,384
18,944 +15.6%
TMUs
1,024
592 -42.2%
ROPs
0
24 +∞%
Compute Units
256
—
SM Count
—
148
Clocks
Base Clock
1000 MHz
1665 MHz
Boost Clock
2200 MHz
2032 MHz
Memory Clock
2000 MHz 8 Gbps effective
2000 MHz 8 Gbps effective
Memory
Memory Size
288 GB
288 GB
VRAM (MB)
294,912
294,912 0.0%
Memory Type
HBM3e
HBM3e
Memory Bus
8192 bit
8192 bit
Bandwidth
8.19 TB/s
8.19 TB/s
Cache
L1 Cache
16 KB (per CU)
256 KB (per SM)
L2 Cache
16 MB
126 MB
L3 Cache
256 MB
—
Performance
Pixel Rate
0 MPixel/s
48.77 GPixel/s
Texture Rate
2,252.8 GTexel/s
1,202.9 GTexel/s
FP32 (TFLOPS)
72.09 TFLOPS
76.99 TFLOPS
FP64 (TFLOPS)
36.04 TFLOPS (1:2)
1,202.9 GFLOPS (1:64)
FP16 (TFLOPS)
72.09 TFLOPS (1:1)
76.99 TFLOPS (1:1)
AI/RT
Tensor Cores
—
592
Matrix Cores
1,024
—
Power
TDP
1000 W
1100 W
TDP (W)
1,000
1,100 +10.0%
Suggested PSU
1400 W
1500 W
Power Connectors
None
—
Architecture
Architecture
CDNA 4.0
Blackwell Ultra
GPU Name
MI350 256CU
GB110
Generation
Instinct (MIx)
Server Blackwell (Bxx)
Process Size
3 nm
5 nm
Transistors
185,000 million
208,000 million
Die Size
2380 mm²
1628 mm²
Foundry
TSMC
TSMC
Density
77.7M / mm²
127.8M / mm²
AMD MCM
MCM
2
—
API Support
OpenCL
3.0
3.0
CUDA
—
10.3
Physical
Slot Width
OAM Module
SXM Module
Length
102 mm 4 inches
—
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 6.0 x16
Other
Production
—
Active
Predecessor
Radeon Instinct
Server Hopper
Successor
—
Server Rubin
View Instinct MI350X Details View B300 SXM6 AC Details