AMD Instinct MI355X vs NVIDIA B300 SXM6 AC Comparison

AMD
RADEON

AMD Instinct MI355X

CORE STATE MI350 256CU
VRAM 288 GB
CLOCK SPEED 2400 MHz
TDP 1400 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 4.0
nm
PROCESS 3 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

B300 SXM6 AC

CORE STATE GB110
VRAM 288 GB
CLOCK SPEED 2032 MHz
TDP 1100 W
BUS WIDTH 8192 bit
ARCHITECTURE Blackwell Ultra
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

geekbench_opencl
N/A
369,831

Analysis: AMD Instinct MI355X vs NVIDIA B300 SXM6 AC

AMD Instinct MI355X and NVIDIA B300 SXM6 AC are both enormous accelerator modules aimed at the same high-density compute segment. The recorded data shows two very different design philosophies, with AMD pushing a massive 2380 mm² die on a 3 nm process, while NVIDIA opts for a smaller 1628 mm² die on 5 nm with a denser transistor layout. Both deliver identical memory configurations, yet their compute characteristics and architectural choices separate them clearly in the database.

Where Each One Wins

The NVIDIA B300 SXM6 AC holds the only recorded benchmark score between the two, a Geekbench OpenCL result of 369831. That places it at the 100th percentile of all GPUs in the database, meaning it outranks every other recorded part. Its nearest rival, the NVIDIA B200, scores 345482, a 7% gap. The H200 NVL trails by 10.4% with 334891. The AMD Instinct MI300X, a previous-generation part in the same family, scores 317994, which is 16.3% behind. The L40S sits at 295763, 25% lower. The B300 SXM6 AC therefore wins outright in the only benchmark where both have recorded data.

The AMD Instinct MI355X has no benchmark entries and an average score of zero in the database. Its percentile ranking sits at 50, squarely in the middle of all recorded GPUs, which is a neutral placeholder rather than a measured result. The data does not show any test where the MI355X beats the B300. However, the MI355X has structural advantages in raw throughput metrics that do not appear in the benchmark record. Its texture rate is 2,457.6 GTexel/s, more than double the B300's 1,202.9 GTexel/s. Its FP32 and FP16 figures of 78.64 TFLOPS edge past the B300's 76.99 TFLOPS in both precisions. These are theoretical peak rates, not measured application results, so the benchmark gap remains in NVIDIA's favor.

The B300 also wins on shading unit count, with 18944 versus 16384 for the MI355X. That is a 15.6% advantage in raw shader hardware. The MI355X counters with 1024 texture mapping units against 592, a 73% lead in TMU count, and the texture rate reflects that heavily. The MI355X has no ROPs recorded at all, with a pixel rate of 0 MPixel/s, while the B300 has 24 ROPs and a 48.77 GPixel/s pixel rate. Any rasterization work, if it were possible on these compute modules, would favor NVIDIA decisively.

Architecture Differences

The MI355X uses the MI350 256CU chip built on CDNA 4.0 architecture, while the B300 uses the GB110 chip on Blackwell Ultra. Both are designed for server compute, with no display outputs and no DirectX, OpenGL, or Vulkan support recorded for either. The MI355X is fabricated on a 3 nm process at TSMC, the B300 on a 5 nm process, also TSMC. The die sizes differ substantially: 2380 mm² for AMD versus 1628 mm² for NVIDIA. Transistor counts also diverge, with NVIDIA packing 208,000 million transistors into the smaller die, yielding a density of 127.8M per mm². AMD fits 185,000 million transistors into the larger die, a density of 77.7M per mm². The density gap is 64.5% in NVIDIA's favor, reflecting the different transistor budgets.

Clock behavior is reversed between the two. The MI355X has a 1000 MHz base clock and a 2400 MHz boost, a 140% boost headroom. The B300 runs a 1665 MHz base and 2032 MHz boost, a 22% headroom. The MI355X boost clock is 18.1% higher than the B300's. Memory clocks are identical: 2000 MHz with 8 Gbps effective for both. Memory capacity, type, and bus width also match: 288 GB of HBM3e across an 8192 bit bus, delivering 8.19 TB/s bandwidth. The MI355X draws 1400 W with a suggested PSU of 1800 W. The B300 draws 1100 W with a suggested PSU of 1500 W. The B300 uses 300 W less power while producing a higher benchmark score. The MI355X carries the Radeon Instinct predecessor lineage, the B300 descends from Server Hopper and has a Server Rubin successor recorded.

The bus interface differs: the MI355X uses PCIe 5.0 x16, the B300 uses PCIe 6.0 x16. The MI355X is an OAM Module with 102 mm length and 165 mm width. The B300 is an SXM Module with no dimensions recorded. The B300 has 592 tensor cores and 592 TMUs, a 1:1 ratio. The MI355X lists no tensor core count. The B300 is marked Active in production status; the MI355X has no status recorded. Release dates are close, with the MI355X on 2025-06-11 and the B300 on 2025-09-10.

Head-to-Head Benchmarks

The database contains no direct head-to-head benchmark entries comparing the MI355X and B300. The only measured score belongs to the B300, a Geekbench OpenCL result of 369831. That score establishes a baseline. The nearest rival data puts the B300 7% ahead of the B200, which scores 345482. The H200 NVL at 334891 is 10.4% behind. The MI300X at 317994 is 16.3% behind. The L40S at 295763 is 25% behind. Those deltas show the B300's dominance within its own recorded peer group, but they do not directly quantify the MI355X deficit because the MI355X has no score.

Theoretical peak compute tells a different story. The MI355X records 78.64 TFLOPS FP32 and the same 78.64 TFLOPS FP16 at a 1:1 ratio. The B300 records 76.99 TFLOPS in both. The MI355X leads by 1.65 TFLOPS, roughly 2.1% higher. Texture throughput is the largest single metric gap: 2,457.6 GTexel/s for the MI355X versus 1,202.9 GTexel/s for the B300, a 104.3% advantage for AMD. Pixel rate is the largest reverse gap: the B300's 48.77 GPixel/s versus 0 for the MI355X. The MI355X has 16,384 shading units, the B300 has 18,944, a 15.6% NVIDIA lead. The MI355X has 1,024 TMUs, the B300 has 592, a 73% AMD lead.

The clock profiles explain some of this. The MI355X boosts to 2400 MHz, substantially higher than the B300's 2032 MHz. That higher boost enables the AMD part to push more texture operations per second despite a lower base clock. The B300's higher base clock of 1665 MHz versus 1000 MHz suggests a more consistent sustained operating point, while the MI355X relies on boost behavior to reach its peak figures. The identical memory systems, both 288 GB HBM3e at 8.19 TB/s, mean memory bandwidth is not a differentiator in any comparison between these two.

FAQ

Q: Which accelerator has the higher FP32 compute?

A: The AMD Instinct MI355X records 78.64 TFLOPS FP32, while the NVIDIA B300 SXM6 AC records 76.99 TFLOPS. The MI355X holds a 1.65 TFLOPS lead.

Q: What is the only recorded benchmark score?

A: The NVIDIA B300 SXM6 AC has a Geekbench OpenCL score of 369831. The AMD Instinct MI355X has no recorded benchmark scores.

Q: How does the B300 compare to its nearest recorded rivals?

A: The B300 scores 7% above the NVIDIA B200 at 345482, 10.4% above the H200 NVL at 334891, 16.3% above the MI300X at 317994, and 25% above the L40S at 295763.

Q: Do the two modules have the same memory?

A: Yes. Both use 288 GB of HBM3e with an 8192 bit bus and 8.19 TB/s bandwidth. Memory clock is also identical at 2000 MHz with 8 Gbps effective.

Q: Which module has more texture units?

A: The AMD Instinct MI355X has 1,024 TMUs and a texture rate of 2,457.6 GTexel/s. The NVIDIA B300 has 592 TMUs and a texture rate of 1,202.9 GTexel/s.

Q: What is the power draw difference?

A: The MI355X is rated at 1400 W with a suggested PSU of 1800 W. The B300 is rated at 1100 W with a suggested PSU of 1500 W.

The Verdict

The recorded data gives the clear performance crown to the NVIDIA B300 SXM6 AC. Its Geekbench OpenCL score of 369831 is a real measurement, and it sits at the 100th percentile of all GPUs in the database. The AMD Instinct MI355X has no measured score, a 50th percentile placeholder, and zero average benchmark score. In any comparison grounded strictly on recorded results, the B300 wins outright.

The theoretical metrics complicate that simple verdict. The MI355X leads in FP32, FP16, texture rate, TMU count, and boost clock. Its 78.64 TFLOPS FP32 tops the B300's 76.99 TFLOPS. Its 2,457.6 GTexel/s texture rate is more than double. Those numbers suggest workloads that saturate texture units or raw shader math could favor the AMD design. But the B300 counters with more shading units, tensor cores, ROPs, a higher base clock, a denser transistor layout at 127.8M per mm², and a 300 W lower power draw. The B300 also uses PCIe 6.0 x16, a newer bus standard than the MI355X's PCIe 5.0 x16.

The power efficiency picture favors NVIDIA. The B300 produces its benchmark score at 1100 W against the MI355X's 1400 W. The suggested PSU figures follow the same pattern: 1500 W for the B300, 1800 W for the MI355X. With identical memory capacity and bandwidth, neither side gains an edge from memory. The MI355X offers a higher peak clock and more texture hardware, but those advantages come with a higher power envelope and no measured performance evidence. The B300 has the only recorded win, the higher percentile rank, and the stronger rival deltas. For any user relying on database measurements, the B300 is the proven performer. The MI355X remains an unverified contender with strong theoretical specifications.

Specification Differences

| Specification | AMD Instinct MI355X | NVIDIA B300 SXM6 AC |

|---|---|---|

| Architecture | CDNA 4.0 | Blackwell Ultra |

| Process node | 3 nm | 5 nm |

| Transistors | 185,000 million | 208,000 million |

| Die size | 2380 mm² | 1628 mm² |

| Transistor density | 77.7M / mm² | 127.8M / mm² |

| Base clock | 1000 MHz | 1665 MHz |

| Boost clock | 2400 MHz | 2032 MHz |

| Shading units | 16384 | 18944 |

| TMUs | 1024 | 592 |

| ROPs | 0 | 24 |

| Tensor cores | None recorded | 592 |

| Pixel rate | 0 MPixel/s | 48.77 GPixel/s |

| Texture rate | 2,457.6 GTexel/s | 1,202.9 GTexel/s |

| FP32 | 78.64 TFLOPS | 76.99 TFLOPS |

| FP16 | 78.64 TFLOPS (1:1) | 76.99 TFLOPS (1:1) |

| TDP | 1400 W | 1100 W |

| Suggested PSU | 1800 W | 1500 W |

| Slot width | OAM Module | SXM Module |

| Bus interface | PCIe 5.0 x16 | PCIe 6.0 x16 |

| Release date | 2025-06-11 | 2025-09-10 |

| Production status | Not recorded | Active |

| Predecessor | Radeon Instinct | Server Hopper |

| Successor | None recorded | Server Rubin |

| Benchmark score | None recorded | 369831 |

| Percentile | 50 | 100 |

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI355X
B300 SXM6 AC
Core Specs
Shading Units
16,384
18,944 +15.6%
Shaders
16,384
18,944 +15.6%
TMUs
1,024
592 -42.2%
ROPs
0
24 +∞%
Compute Units
256
SM Count
148
Clocks
Base Clock
1000 MHz
1665 MHz
Boost Clock
2400 MHz
2032 MHz
Memory Clock
2000 MHz 8 Gbps effective
2000 MHz 8 Gbps effective
Memory
Memory Size
288 GB
288 GB
VRAM (MB)
294,912
294,912 0.0%
Memory Type
HBM3e
HBM3e
Memory Bus
8192 bit
8192 bit
Bandwidth
8.19 TB/s
8.19 TB/s
Cache
L1 Cache
32 KB (per CU)
256 KB (per SM)
L2 Cache
32 MB
126 MB
L3 Cache
256 MB
Performance
Pixel Rate
0 MPixel/s
48.77 GPixel/s
Texture Rate
2,457.6 GTexel/s
1,202.9 GTexel/s
FP32 (TFLOPS)
78.64 TFLOPS
76.99 TFLOPS
FP64 (TFLOPS)
39.32 TFLOPS (1:2)
1,202.9 GFLOPS (1:64)
FP16 (TFLOPS)
78.64 TFLOPS (1:1)
76.99 TFLOPS (1:1)
AI/RT
Tensor Cores
592
Matrix Cores
1,024
Power
TDP
1400 W
1100 W
TDP (W)
1,400
1,100 -21.4%
Suggested PSU
1800 W
1500 W
Power Connectors
None
Architecture
Architecture
CDNA 4.0
Blackwell Ultra
GPU Name
MI350 256CU
GB110
Generation
Instinct (MIx)
Server Blackwell (Bxx)
Process Size
3 nm
5 nm
Transistors
185,000 million
208,000 million
Die Size
2380 mm²
1628 mm²
Foundry
TSMC
TSMC
Density
77.7M / mm²
127.8M / mm²
API Support
OpenCL
3.0
3.0
CUDA
10.3
Physical
Slot Width
OAM Module
SXM Module
Length
102 mm 4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 6.0 x16
Other
Production
Active
Predecessor
Radeon Instinct
Server Hopper
Successor
Server Rubin
View Instinct MI355X Details View B300 SXM6 AC Details