RADEON

AMD Instinct MI350P

AMD graphics card specifications and benchmark scores

144 GB
VRAM
2200
MHz Boost
600W
TDP
8192
Bus Width
MCM Design

At a Glance

AMD
VRAM 144 GB
Boost Clock 2,200 MHz
Shaders 8,192
Bus Width 8192-bit
TDP 600W
Memory Type HBM3e
Architecture CDNA 4.0
nm
Process 3 nm
Released May 2026

AMD Instinct MI350P Specifications

Instinct MI350P GPU Core

Shader units and compute resources

The AMD Instinct MI350P GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
8,192
Shaders
8,192
TMUs
512
Compute Units
128

Instinct MI350P Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Instinct MI350P's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Instinct MI350P by AMD dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1000 MHz
Base Clock
1,000 MHz
Boost Clock
2200 MHz
Boost Clock
2,200 MHz
Memory Clock
2000 MHz 8 Gbps effective
GDDR GDDR 6X 6X

AMD's Instinct MI350P Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Instinct MI350P's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
144 GB
VRAM
147,456 MB
Memory Type
HBM3e
VRAM Type
HBM3e
Memory Bus
8192 bit
Bus Width
8192-bit
Bandwidth
8.19 TB/s

Instinct MI350P by AMD Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Instinct MI350P, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
16 KB (per CU)
L2 Cache
16 MB
Infinity Cache
128 MB

Instinct MI350P Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the AMD Instinct MI350P against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
36.04 TFLOPS
FP64 (Double)
18.02 TFLOPS (1:2)
FP16 (Half)
36.04 TFLOPS (1:1)
Pixel Rate
0 MPixel/s
Texture Rate
1,126.4 GTexel/s

CDNA 4.0 Architecture & Process

Manufacturing and design details

The AMD Instinct MI350P is built on AMD's CDNA 4.0 architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Instinct MI350P will perform in GPU benchmarks compared to previous generations.

Architecture
CDNA 4.0
GPU Name
MI350 128CU
Process Node
3 nm
Foundry
TSMC
Transistors
73,000 million
Die Size
1190 mm²
Density
61.3M / mm²

AMD's Instinct MI350P Power & Thermal

TDP and power requirements

Power specifications for the AMD Instinct MI350P determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Instinct MI350P to maintain boost clocks without throttling.

TDP
600 W
TDP
600W
Power Connectors
1x 16-pin
Suggested PSU
1000 W

Instinct MI350P by AMD Physical & Connectivity

Dimensions and outputs

Physical dimensions of the AMD Instinct MI350P are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Dual-slot
Length
267 mm 10.5 inches
Height
111 mm 4.4 inches
Bus Interface
PCIe 5.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

AMD API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the AMD Instinct MI350P. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
N/A
DirectX
N/A
OpenGL
N/A
OpenGL
N/A
Vulkan
N/A
Vulkan
N/A
OpenCL
3.0
Shader Model
N/A

Instinct MI350P Product Information

Release and pricing details

The AMD Instinct MI350P is manufactured by AMD as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Instinct MI350P by AMD represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
AMD
Release Date
May 2026
Predecessor
Radeon Instinct

Instinct MI350P Benchmark Scores

No benchmark data available for this GPU.

About AMD Instinct MI350P

The AMD Instinct MI350P is a data-center compute accelerator built on the CDNA 4.0 architecture, targeting AI inference, scientific simulation, and high-throughput parallel workloads rather than consumer gaming. With a 50th percentile ranking among all GPUs, it sits in the middle of the database, but its raw memory capacity and compute throughput place it far ahead of typical desktop hardware. The MI350P is not a graphics card, it has no display outputs, no DirectX/OpenGL/Vulkan support, and a pixel rate of 0 MPixel/s, so it is exclusively for servers and workstations that run compute kernels. The data shows a 144 GB HBM3e memory pool, 36.04 TFLOPS FP32, and a 600 W TDP, making it a specialized tool for problems that exceed the VRAM limits of any consumer or prosumer GPU. This analysis covers who should deploy it, its ray tracing and feature set, benchmark performance, power and cooling, memory subsystem, and comparisons to nearest rivals.

Who Should Consider It

The MI350P is designed for users who need massive on-device memory capacity first and raw compute second. The 144 GB HBM3e frame buffer is the defining characteristic: it allows loading entire large language models, huge scientific datasets, or complex 3D render scenes without spilling to system RAM. For workloads that fit within 36.04 TFLOPS of FP32 throughput but require more than 24 GB or even 48 GB of VRAM, the MI350P is a necessity. This includes training or fine-tuning large transformer models, processing high-resolution volumetric data, or running multi-GPU distributed simulations where each card must hold a substantial slice of the problem. The 8192-bit memory bus and 8.19 TB/s bandwidth mean that memory-bound kernels, those that stream data rather than compute on cached values, will see substantial gains over a typical 256-bit or 512-bit consumer card.

Benchmark results place the MI350P at the 50th percentile of all GPUs, which might seem modest, but that percentile is driven by a database full of gaming cards with far smaller memory footprints. For compute tasks that leverage the full 144 GB, the effective performance is far higher than any gaming GPU because the workload simply cannot run on those cards. Resolution and settings recommendations are moot for a card with no display outputs, this is not a rasterization product. Instead, consider it for batch inference jobs where latency is less critical than throughput, and where the entire model fits in 144 GB, avoiding PCIe transfers. For scientific computing, the 1,126.4 GTexel/s texture rate and 36.04 TFLOPS FP16 (1:1 with FP32) suggest balanced performance for mixed-precision workloads. If your workload is small enough to fit in 32 GB or 64 GB, a smaller, lower-power card would be more efficient; the MI350P’s 600 W TDP demands serious cooling and power infrastructure that only pays off when the memory pool is fully utilized.

Ray Tracing and Feature Set

The MI350P has no ray tracing cores, no tensor cores, and no graphics APIs, DirectX, OpenGL, and Vulkan are all listed as N/A. This is a pure compute accelerator. The CDNA 4.0 architecture is AMD’s dedicated data-center design, distinct from the RDNA line used in gaming Radeon cards. The absence of RT cores means no hardware-accelerated ray tracing for rendering; any ray tracing would have to be done in software, which is impractical at scale. The lack of tensor cores means that matrix operations rely on the shader units (8192 shading units, 512 TMUs) for general-purpose compute. The FP16 performance is exactly equal to FP32 at 36.04 TFLOPS, which is a deliberate design choice for workloads that use mixed precision without a dedicated tensor path. This contrasts with NVIDIA data-center cards that often double or quadruple FP16 throughput via tensor cores; the MI350P takes a simpler approach, using the same 8192 shading units for both precisions.

The feature set is barebones: no display outputs, no API support, and no pixel rate. What it does offer is PCIe 5.0 x16 connectivity, ensuring high-bandwidth communication with the host CPU and other GPUs via NVLink or similar peer-to-peer mechanisms (though the pack does not specify interconnect details). The 73,000 million transistors on a 1190 mm² die, manufactured on TSMC’s 3 nm process, give a transistor density of 61.3M per mm². This is a massive chip, and the lack of RT/tensor cores means all that silicon is dedicated to FP32/FP16 math and the memory controller. For developers, the MI350P requires a compute stack that supports CDNA 4.0, likely ROCm or similar, since there is no graphics driver fallback. The card is a one-trick pony, but that trick is massive memory bandwidth and capacity.

Benchmark Performance

The MI350P has no benchmark scores in the database, its avgBenchmarkScore is 0, and the nearestRivals list is empty. The percentileVsAllGpus is 50, meaning that half of all GPUs in the database score higher and half score lower. This is an unusual situation: a card with 36.04 TFLOPS FP32 and 8.19 TB/s bandwidth should rank higher than most consumer cards, which typically have 10–20 TFLOPS and less than 1 TB/s bandwidth. The 50th percentile likely reflects the absence of benchmark entries rather than real-world parity. Without rival scores, we cannot compute deltaPct values or make direct percentage comparisons. However, we can infer relative performance from the hardware specifications. The 36.04 TFLOPS FP32 is roughly double a high-end consumer GPU like an RTX 4080 (which is not in the fact pack, so this is qualitative), but the MI350P’s 144 GB memory is 4–6 times larger than any gaming card. The 8.19 TB/s bandwidth is also multiple times higher than consumer HBM or GDDR6X implementations.

For FP16, the 1:1 ratio means the MI350P is not optimized for the half-precision workloads that dominate AI training on other architectures. Many accelerators offer 2:1 or 4:1 FP16/FP32 ratios, so the MI350P’s 36.04 TFLOPS FP16 is modest by modern data-center standards. The texture rate of 1,126.4 GTexel/s is high, but with no graphics pipeline, this is only useful for compute kernels that use texture sampling as a data access pattern. The pixel rate is 0, confirming no rasterization. In practice, the MI350P’s performance is best described as “memory-heavy compute”: it will excel at memory-bound operations, such as sparse matrix multiplication, graph analytics, or large database scans, where the 8.19 TB/s bandwidth is the bottleneck. For compute-bound kernels, the 36.04 TFLOPS FP32 is respectable but not class-leading. The 50th percentile suggests that in a mixed database of gaming and compute cards, the MI350P does not stand out for raw throughput, it stands out for capacity.

FAQ

Q: Does the MI350P support ray tracing?

A: No. The fact pack lists RT cores as null, and the API support for DirectX, OpenGL, and Vulkan is all N/A. This is a compute-only accelerator with no graphics pipeline.

Q: Can I use the MI350P for gaming?

A: No. It has no display outputs (listed as “No outputs”), no pixel rate (0 MPixel/s), and no graphics API support. It is designed exclusively for data-center compute workloads.

Q: What is the memory size and bandwidth?

A: The MI350P has 144 GB of HBM3e memory with an 8192-bit bus width and 8.19 TB/s bandwidth. This is far larger than any consumer GPU, which typically maxes out at 24 GB.

Q: What is the power consumption and PSU requirement?

A: The TDP is 600 W, and AMD recommends a 1000 W power supply. It uses a single 16-pin power connector and occupies a dual-slot form factor.

Q: Does the MI350P have tensor cores?

A: No. The tensorCores field is null. FP16 performance is 36.04 TFLOPS, which is identical to FP32 (1:1 ratio), meaning no dedicated tensor hardware.

Q: What is the release date and process node?

A: The release date is May 6, 2026, and it is built on TSMC’s 3 nm process with 73,000 million transistors on a 1190 mm² die.

Q: Can I install the MI350P in a standard PC case?

A: The card is 267 mm (10.5 inches) long, 111 mm (4.4 inches) high, and 40 mm (1.6 inches) wide, which is dual-slot. It fits in many full-tower cases, but the 600 W TDP and 1000 W PSU recommendation require a high-end power supply and adequate airflow.

Power and Cooling

The MI350P has a TDP of 600 W, which is among the highest in the database. This demands a power supply rated at 1000 W, as per the suggestedPsu field. The card uses a single 16-pin power connector (1x 16-pin), which is the modern PCIe 5.0 standard for high-wattage cards. The dual-slot form factor means it requires a case with sufficient clearance for a thick cooler, the width is 40 mm (1.6 inches), so adjacent PCIe slots may be blocked. Cooling is not specified in detail, but a 600 W TDP requires a robust air cooler or liquid cooling solution; the fact pack mentions only the dimensions and slot width. The 3 nm process node from TSMC helps with efficiency, but 600 W is still a massive thermal load. In a server environment, this card will need active cooling with high static pressure fans and proper chassis airflow. The 1000 W PSU recommendation is a minimum; if the system has multiple MI350P cards or a high-end CPU, the PSU should be sized accordingly. The power connector is a single 16-pin, which simplifies cabling compared to older cards that used two 8-pin connectors, but it means the PSU must have a native 16-pin cable or a high-quality adapter. The memory subsystem, 144 GB of HBM3e, also contributes to the thermal load, as HBM stacks sit close to the GPU die and require direct cooling. The card’s length of 267 mm is shorter than many flagship gaming GPUs, but the height of 111 mm may exceed standard ATX slot heights, so check case compatibility. For data-center racks, the dual-slot design allows for dense packing, but the 600 W TDP per card limits how many units can fit in a single chassis before power and cooling become prohibitive.

Memory Subsystem

The MI350P’s memory subsystem is its standout feature: 144 GB of HBM3e with an 8192-bit bus width and 8.19 TB/s bandwidth. This is the largest memory pool in the fact pack, far exceeding any gaming card’s 16–24 GB. The 8192-bit bus is double the width of most high-end cards (which typically use 384-bit or 512-bit buses), enabling the extraordinary 8.19 TB/s bandwidth. For high-resolution workloads, this means the card can hold entire datasets in VRAM, eliminating PCIe bottlenecks. For example, a 4K or 8K render scene with massive textures, or a large language model with billions of parameters, fits entirely on the card. The 144 GB capacity is also critical for scientific simulations that require storing intermediate results. The memory clock is 2000 MHz with 8 Gbps effective, but the sheer bus width compensates for the relatively modest clock speed. HBM3e is a stacked memory technology, which is why the die size is 1190 mm², the memory is packaged on the same substrate as the GPU. The 73,000 million transistors include the memory controller logic, but the HBM stacks themselves are separate dies.

The practical implication is that the MI350P can process datasets that would otherwise be split across multiple GPUs or spilled to system RAM. This reduces complexity and improves performance for memory-bound tasks. However, the 144 GB is not user-expandable, it is fixed on the board. For workloads that need more than 144 GB, you would need multiple MI350P cards, but the fact pack does not specify multi-GPU interconnect bandwidth. The 8.19 TB/s bandwidth is sufficient to feed the 36.04 TFLOPS FP32 compute, meaning the card is balanced in the sense that neither memory nor compute is a severe bottleneck for most kernels. The 1:1 FP16/FP32 ratio means that half-precision workloads will not see the typical 2x speedup found on other accelerators, but they also do not lose precision. For high-resolution (4K/8K) rendering, the MI350P is overkill since it has no display outputs, but for offline rendering, the 144 GB allows loading entire film-quality scenes. The 8192-bit bus is the key enabler: it provides 16 times the bandwidth of a 512-bit GDDR6 card (qualitative, since no rival specs are in the pack), making the MI350P ideal for data streaming.

How It Compares

The nearestRivals list is empty in the fact pack, so no direct percentage comparisons can be made. The percentileVsAllGpus is 50, meaning the MI350P sits at the median of the database. This is a remarkable position given its hardware: most GPUs in the database are consumer gaming cards, which have far less memory but often higher clock speeds. The 50th percentile suggests that the MI350P’s raw compute (36.04 TFLOPS FP32) is not exceptional, many gaming cards exceed this in FP32 with boost clocks around 2.5–3 GHz (qualitative, not from fact pack). However, the MI350P’s 144 GB memory is unmatched; no consumer card in the database is listed with more than 24 GB (again, not in fact pack, but the pack shows only the MI350P’s specs). In the absence of rival data, we must rely on the fact pack’s internal consistency. The MI350P has a 600 W TDP, which is high but not unprecedented for data-center accelerators. Its 3 nm process is leading-edge, but the 1190 mm² die is enormous, which explains the 73,000 million transistor count.

If we compare to the Radeon Instinct predecessor (qualitative, no specs given), the MI350P represents a generational shift to CDNA 4.0 and HBM3e. The lack of RT/tensor cores puts it in a different category from NVIDIA’s data-center GPUs, but the fact pack does not list any rivals. The 50th percentile might reflect that the database has many low-end and midrange GPUs, so a 36 TFLOPS card should rank higher. The absence of benchmark scores (avgBenchmarkScore 0) is puzzling, perhaps the card has not been tested yet, or the database does not have a suitable test suite for compute accelerators. In that context, the 50th percentile is a placeholder rather than a meaningful comparison. For users, the MI350P is best compared to itself: it is a 144 GB, 8.19 TB/s, 36 TFLOPS compute card that costs nothing (no launch MSRP) but demands 600 W and a 1000 W PSU. It sits in a niche where memory capacity is the primary constraint, and it delivers that capacity in spades. The 8192-bit bus is the widest in any GPU in the fact pack, and the 8.19 TB/s bandwidth is similarly unmatched. The card’s 267 mm length and dual-slot design are standard for data-center GPUs, but the 111 mm height may require specialized server chassis.

In summary, the MI350P is a memory-centric accelerator with no graphics capabilities, no ray tracing, and no tensor cores. Its 144 GB HBM3e and 8.19 TB/s bandwidth make it ideal for large-scale compute, but its 36.04 TFLOPS FP32 is mediocre for the 600 W TDP. The 50th percentile ranking reflects the lack of benchmark data rather than actual performance. Without nearest rivals, we cannot state deltas, but the card’s unique combination of capacity and bandwidth ensures it has no direct competitor in the database. The 3 nm process and 73,000 million transistors indicate a high-cost, high-performance chip, but the missing launch MSRP prevents any cost analysis. For a workload that fits in 144 GB, the MI350P is unmatched; for anything else, smaller and more efficient cards would suffice. The power and cooling requirements, 600 W TDP, 1000 W PSU, 16-pin connector, dual-slot, are substantial, but they are the price of 8.19 TB/s bandwidth. The card has no display outputs, so it is strictly for servers, and the lack of graphics APIs means it requires a compute stack that can program CDNA 4.0 directly. The FP16 1:1 ratio is a notable limitation for AI workloads that rely on half precision, but the memory capacity may compensate. Overall, the MI350P is a specialized tool for a specific problem: massive memory capacity with high bandwidth, at the expense of graphics features and raw compute efficiency.

The NVIDIA Equivalent of Instinct MI350P

Looking for a similar graphics card from NVIDIA? The NVIDIA GeForce RTX 5070 Mobile 12 GB offers comparable performance and features in the NVIDIA lineup.

NVIDIA GeForce RTX 5070 Mobile 12 GB

NVIDIA • 12 GB VRAM

View Specs Compare

Popular AMD Instinct MI350P Comparisons

See how the Instinct MI350P stacks up against similar graphics cards from the same generation and competing brands.

Compare Instinct MI350P with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs