AMD Instinct MI350P vs NVIDIA B300 SXM6 AC Comparison
AMD Instinct MI350P
B300 SXM6 AC
PERFORMANCE BENCHMARKS
Analysis: AMD Instinct MI350P vs NVIDIA B300 SXM6 AC
AMD Instinct MI350P and NVIDIA B300 SXM6 AC represent two heavyweight server accelerators from competing architectures. The recorded data shows a substantial performance gap in favor of the NVIDIA part, though the AMD offering brings its own architectural merits to the table. The B300 SXM6 AC delivers an average benchmark score of 369,831, placing it in the 100th percentile of all GPUs in the database, while the MI350P holds a 50th percentile position with an average score of zero, indicating no completed benchmark entries for the AMD part.
Head-to-Head Benchmarks
The only recorded benchmark result belongs to the NVIDIA B300 SXM6 AC, which scored 369,831 points in Geekbench OpenCL. This places it 7% ahead of the NVIDIA B200, which scored 345,482, and 10.4% ahead of the NVIDIA H200 NVL with its score of 334,891. The B300 SXM6 AC also outperforms the AMD Instinct MI300X by 16.3%, as that rival scored 317,994, and leads the NVIDIA L40S by 25%, which managed 295,763.
The AMD Instinct MI350P has no benchmarks recorded in the database, meaning there are no direct numerical comparisons between the two cards in this dataset. The MI350P's percentile ranking of 50 reflects its position among all GPUs, but without an average score, the head-to-head comparison relies entirely on the B300 SXM6 AC's measured performance and the architectural specifications of both parts.
The FP32 compute figures offer a clear point of reference. The B300 SXM6 AC delivers 76.99 TFLOPS of FP32 performance, while the MI350P provides 36.04 TFLOPS. This represents a 2.14x advantage for the NVIDIA part in single-precision throughput. Similarly, FP16 performance follows the same pattern, with the B300 SXM6 AC producing 76.99 TFLOPS compared to the MI350P's 36.04 TFLOPS, both operating at a 1:1 ratio.
Texture fill rates show a closer margin. The B300 SXM6 AC achieves 1,202.9 GTexel/s, while the MI350P reaches 1,126.4 GTexel/s, a difference of roughly 6.8% in favor of NVIDIA. Pixel rates diverge more significantly, with the B300 SXM6 AC producing 48.77 GPixel/s against the MI350P's 0 MPixel/s, though the AMD part's ROP count of zero makes this metric largely theoretical for its intended compute workload.
Where Each One Wins
The NVIDIA B300 SXM6 AC wins decisively in raw compute throughput, FP32 and FP16 performance, and pixel processing capabilities. Its 76.99 TFLOPS in both FP32 and FP16 makes it the stronger choice for workloads that depend heavily on single-precision math, such as large-scale matrix operations common in AI training and inference. The B300 SXM6 AC also holds an advantage in shading units, with 18,944 compared to the MI350P's 8,192, and texture mapping units, 592 versus 512.
The AMD Instinct MI350P wins in power efficiency per unit of compute. It consumes 600 W against the B300 SXM6 AC's 1100 W, while delivering 36.04 TFLOPS. The MI350P produces roughly 60.07 TFLOPS per kilowatt, while the B300 SXM6 AC achieves about 69.99 TFLOPS per kilowatt, so the NVIDIA part still leads in efficiency, but the AMD card requires far less total power and a smaller suggested PSU, 1000 W versus 1500 W.
The MI350P also offers a more conventional physical form factor. It is a dual-slot card measuring 267 mm in length, 111 mm in height, and 40 mm in width, with a single 16-pin power connector. The B300 SXM6 AC uses an SXM Module form factor with no listed power connector or dimensions, which means it requires a proprietary server chassis and power delivery system, whereas the AMD part can potentially fit into more standard server configurations with a PCIe 5.0 x16 interface.
Architecture Differences
The AMD Instinct MI350P uses the CDNA 4.0 architecture, built on a 3 nm process at TSMC with 73,000 million transistors on a 1,190 mm² die, yielding a transistor density of 61.3 million per square millimeter. The NVIDIA B300 SXM6 AC uses the Blackwell Ultra architecture on TSMC's 5 nm node, packing 208,000 million transistors onto a 1,628 mm² die, resulting in a density of 127.8 million per square millimeter.
The B300 SXM6 AC's transistor count is 2.85x higher than the MI350P's, and its die is 36.8% larger. The AMD part's 3 nm process node is smaller than NVIDIA's 5 nm node, but the NVIDIA chip compensates with a much higher transistor density, which suggests a more complex design with more functional blocks per area. The MI350P relies on its CDNA 4.0 architecture to deliver 8,192 shading units, while the B300 SXM6 AC fields 18,944 shading units and 592 tensor cores.
Both GPUs share the same memory configuration in key respects: 8192-bit bus width, HBM3e memory type, and 8.19 TB/s bandwidth. They also run memory at 2000 MHz with 8 Gbps effective speed. However, the B300 SXM6 AC carries 288 GB of memory, exactly double the MI350P's 144 GB. The B300 SXM6 AC also features 24 ROPs, while the MI350P lists zero ROPs, reflecting the AMD part's compute-focused design that omits rasterization hardware entirely.
The bus interfaces differ as well. The MI350P uses PCIe 5.0 x16, while the B300 SXM6 AC uses PCIe 6.0 x16, giving the NVIDIA part a newer interconnect standard. Neither card has display outputs, and both list N/A for DirectX, OpenGL, and Vulkan API support, confirming their server accelerator roles rather than graphics-oriented products.
Specification Differences
The two accelerators diverge on several key specifications. The MI350P uses the MI350 128CU chip, while the B300 SXM6 AC uses the GB110 chip. Clock speeds differ: the MI350P has a base clock of 1000 MHz and a boost of 2200 MHz, whereas the B300 SXM6 AC runs at 1665 MHz base and 2032 MHz boost. The AMD part boosts 168 MHz higher, but the NVIDIA part has a 665 MHz higher base clock.
Memory capacity is a major differentiator: 144 GB on the MI350P versus 288 GB on the B300 SXM6 AC. Both share the same 8192-bit bus width and 8.19 TB/s bandwidth, so the extra capacity on the NVIDIA card does not come with additional bandwidth. The B300 SXM6 AC has 18,944 shading units, 592 TMUs, and 24 ROPs, while the MI350P has 8,192 shading units, 512 TMUs, and 0 ROPs. The NVIDIA card also includes 592 tensor cores, a feature the AMD part does not list.
Power requirements show the B300 SXM6 AC consuming 1100 W with a suggested PSU of 1500 W, versus the MI350P's 600 W and 1000 W suggested PSU. The B300 SXM6 AC uses an SXM Module slot width, while the MI350P is dual-slot. The MI350P has a 267 mm length, 111 mm height, and 40 mm width, while the B300 SXM6 AC has no recorded dimensions. The MI350P uses a 1x 16-pin power connector, whereas the NVIDIA part has no listed power connector.
Release dates also differ. The MI350P is dated 2026-05-06, while the B300 SXM6 AC released 2025-09-10. The NVIDIA part has a production status of Active, while the AMD part has no recorded status. The B300 SXM6 AC lists a predecessor of Server Hopper and a successor of Server Rubin, while the MI350P lists a predecessor of Radeon Instinct and no successor.
FAQ
Q: Which GPU has higher FP32 performance?
A: The NVIDIA B300 SXM6 AC delivers 76.99 TFLOPS of FP32 compute, while the AMD Instinct MI350P provides 36.04 TFLOPS. The NVIDIA part offers 2.14x the single-precision throughput of the AMD card.
Q: How much memory does each accelerator have?
A: The AMD Instinct MI350P has 144 GB of HBM3e memory, while the NVIDIA B300 SXM6 AC has 288 GB of HBM3e memory. Both use an 8192-bit bus and provide 8.19 TB/s of bandwidth.
Q: What are the power consumption figures?
A: The MI350P consumes 600 W with a suggested PSU of 1000 W. The B300 SXM6 AC consumes 1100 W with a suggested PSU of 1500 W.
Q: Does either card support display outputs?
A: No. Both the MI350P and the B300 SXM6 AC list no display outputs, and both have N/A for DirectX, OpenGL, and Vulkan API support.
Q: What is the process node for each chip?
A: The AMD MI350P uses a 3 nm process at TSMC with 73,000 million transistors. The NVIDIA B300 SXM6 AC uses a 5 nm process at TSMC with 208,000 million transistors.
Q: How does the B300 SXM6 AC compare to other NVIDIA parts?
A: The B300 SXM6 AC scores 369,831, which is 7% higher than the B200's 345,482 and 10.4% higher than the H200 NVL's 334,891. It also leads the L40S by 25% with that card scoring 295,763.
The Verdict
The data points strongly toward the NVIDIA B300 SXM6 AC for workloads that demand maximum compute throughput. Its 76.99 TFLOPS in FP32 and FP16, combined with 288 GB of memory and 592 tensor cores, positions it as the higher-performance accelerator in this comparison. The measured benchmark score of 369,831 and 100th percentile ranking confirm its top-tier status relative to the entire GPU database. The 7% lead over the B200 and 16.3% lead over the MI300X show meaningful performance gains over established rivals.
The AMD Instinct MI350P suits scenarios where power draw and physical integration matter more than peak throughput. Its 600 W power envelope, dual-slot design, and standard PCIe 5.0 x16 interface make it a more flexible option for systems that cannot accommodate an SXM module or a 1100 W power draw. The 144 GB of HBM3e memory with identical 8.19 TB/s bandwidth to the NVIDIA part ensures that memory-intensive workloads still have ample capacity and speed.
For pure performance per watt, the B300 SXM6 AC still leads, producing roughly 70 TFLOPS per kilowatt against the MI350P's approximately 60 TFLOPS per kilowatt. However, the AMD part's lower absolute power requirement opens up deployment options in power-constrained environments. The MI350P's 3 nm process node and smaller transistor count suggest a more power-efficient design philosophy, but the B300 SXM6 AC's higher transistor density and larger die deliver superior raw capability.
Users prioritizing maximum FP32 or FP16 compute, the largest memory pool, or the highest benchmark scores should select the B300 SXM6 AC. Users needing a dual-slot accelerator with lower power draw and a conventional PCIe form factor should consider the MI350P. The absence of benchmark data for the MI350P means its real-world performance remains unverified in this database, while the B300 SXM6 AC has a confirmed score and percentile rank.