AMD Instinct MI350P vs AMD Radeon 820M Comparison

AMD
RADEON

AMD Instinct MI350P

CORE STATE MI350 128CU
VRAM 144 GB
CLOCK SPEED 2200 MHz
TDP 600 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 4.0
nm
PROCESS 3 nm
LAUNCH DATE 2026
VS
AMD
RADEON

Radeon 820M

CORE STATE Krackan Point 2
VRAM System Shared
CLOCK SPEED 2800 MHz
TDP 15 W
BUS WIDTH System Shared
ARCHITECTURE RDNA 3.5
nm
PROCESS 4 nm
LAUNCH DATE 2025

Analysis: AMD Instinct MI350P vs AMD Radeon 820M

The AMD Instinct MI350P and AMD Radeon 820M occupy opposite ends of the GPU spectrum, one a 600 W accelerator with 144 GB of HBM3e memory, the other a 15 W integrated processor with system-shared memory. The database records no shared benchmark runs between these two parts, so the comparison rests on their recorded specifications and the architectural data available. The MI350P delivers 36.04 TFLOPS of FP32 compute, while the 820M delivers 716.8 GFLOPS, a factor of roughly 50 in raw floating-point throughput. The MI350P’s texture rate stands at 1,126.4 GTexel/s against the 820M’s 22.40 GTexel/s, and the MI350P’s pixel rate is recorded as 0 MPixel/s because it has no display outputs, whereas the 820M reaches 11.20 GPixel/s. Both parts sit at the 50th percentile in the database’s GPU rankings, but that percentile reflects their respective categories, not a head-to-head result.

Head-to-Head Benchmarks

The database contains no direct benchmark entries for this pairing. The head-to-head section is therefore based on the measured and recorded specifications that define each product’s capabilities. The MI350P’s FP32 output of 36.04 TFLOPS is the clearest point of dominance. That figure is 50 times the 820M’s 716.8 GFLOPS, and the same ratio applies to FP16, where both parts list 1:1 rates. The MI350P’s texture throughput of 1,126.4 GTexel/s exceeds the 820M’s 22.40 GTexel/s by a factor of 50.3. This disparity comes from the MI350P’s 512 texture mapping units versus the 820M’s 8 TMUs, combined with clock speeds that favor the larger chip.

Memory bandwidth tells a similar story. The MI350P accesses 144 GB of HBM3e over an 8192-bit bus, yielding 8.19 TB/s. The 820M uses system-shared memory with a bus width and type that are system dependent, and its bandwidth is recorded as system dependent as well, meaning no fixed number exists in the database. The MI350P’s memory clock is 2000 MHz with 8 Gbps effective, while the 820M’s memory clock is listed as system shared, again without a fixed value. The MI350P’s 8.19 TB/s is a concrete upper bound that the 820M cannot approach, given the latter’s reliance on shared system memory.

The 820M does win on pixel throughput. Its 11.20 GPixel/s comes from 4 ROPs, and the MI350P records 0 MPixel/s because it has no ROPs and no display outputs. This is not a performance comparison in the traditional sense; the MI350P is not designed to rasterize to a screen. The 820M also leads in API support, listing DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the MI350P lists N/A for all three APIs. The MI350P is a compute accelerator with no graphics pipeline, so the absence of those APIs is expected.

Clock speeds show a different kind of advantage. The 820M boosts to 2800 MHz from a 400 MHz base, while the MI350P boosts to 2200 MHz from a 1000 MHz base. The 820M’s higher boost clock reflects its integrated, power-constrained design, but the MI350P’s lower clock is offset by its massive shader count of 8192 shading units versus the 820M’s 128. The MI350P’s base clock is 1000 MHz, which is 2.5 times the 820M’s base, but the 820M’s boost clock is 27% higher than the MI350P’s boost. Neither number changes the overall compute picture.

Where Each One Wins

The MI350P wins in every compute-heavy metric recorded in the database. FP32 and FP16 performance are both 36.04 TFLOPS, which is the headline number for scientific simulation, machine learning training, and high-performance computing workloads. The 8192-bit memory bus and 144 GB capacity are designed for models and datasets that exceed the memory of any integrated GPU. The MI350P’s 73,000 million transistors on a 1190 mm² die, built on a 3 nm process at TSMC, give it a transistor density of 61.3M per square millimeter. That density supports the 512 TMUs and the 8.19 TB/s memory subsystem. The MI350P’s power connector is a single 16-pin, and its suggested power supply is 1000 W, which indicates the power envelope required for sustained compute workloads.

The 820M wins in portability and integration. Its 15 W TDP is 40 times lower than the MI350P’s 600 W, and it is an IGP with no slot width, no power connectors, and dimensions that are not recorded because it is integrated into a processor. The 820M’s 4 nm process at TSMC is less advanced than the MI350P’s 3 nm node, but the 820M’s transistor count and die size are unknown, so a direct density comparison is not possible. The 820M supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, making it suitable for graphics workloads and gaming on portable devices. Its display outputs are portable device dependent, meaning it can drive screens, while the MI350P has no outputs at all.

The 820M also holds an advantage in ray tracing. It has 2 ray tracing cores, while the MI350P lists null for RT cores. This is a functional difference: the 820M can accelerate ray-traced effects in supported applications, and the MI350P has no such capability because it is not a graphics part. The 820M’s 4 ROPs enable pixel output, and its 11.20 GPixel/s is a real number for display rendering. The MI350P’s 0 MPixel/s confirms that it never writes to a framebuffer.

The MI350P’s PCIe interface is PCIe 5.0 x16, while the 820M uses PCIe 4.0 x8. For the MI350P, that bandwidth supports data movement to and from the host CPU for large-scale compute jobs. For the 820M, PCIe 4.0 x8 is sufficient for an integrated part that shares system memory. The MI350P’s dimensions are 267 mm in length, 111 mm in height, and 40 mm in width, making it a dual-slot card. The 820M has no recorded dimensions because it is not a discrete card.

Architecture Differences

The MI350P is built on CDNA 4.0, AMD’s compute-focused architecture, while the 820M uses RDNA 3.5, the graphics and compute architecture for integrated and mobile parts. The MI350P’s chip is designated MI350 128CU, indicating 128 compute units, which aligns with its 8192 shading units at 64 shaders per compute unit. The 820M’s chip is Krackan Point 2, which is part of the Navi III IGP family for Strix Point Mobile. The MI350P belongs to the Instinct (MIx) generation, and its predecessor is listed as Radeon Instinct. The 820M’s generation is Navi III IGP (Strix Point Mobile), with a predecessor of Navi II IGP.

The process nodes differ: the MI350P is on 3 nm, and the 820M is on 4 nm, both at TSMC. The MI350P has 73,000 million transistors on a 1190 mm² die, giving a density of 61.3M per square millimeter. The 820M’s transistor count and die size are unknown, so no density can be calculated. The MI350P’s memory controller is 8192 bits wide, which is an order of magnitude wider than any typical graphics memory bus. The 820M’s memory bus is system shared, meaning it depends on the host processor’s memory controller.

The MI350P has no ROPs, no ray tracing cores, and no display outputs. Its API support is listed as N/A for DirectX, OpenGL, and Vulkan. This is a pure compute accelerator. The 820M has 4 ROPs, 2 ray tracing cores, and full API support for DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4. The MI350P’s shading units are 8192, versus 128 for the 820M, and its TMUs are 512 versus 8. The MI350P’s texture rate is 1,126.4 GTexel/s, while the 820M’s is 22.40 GTexel/s. The pixel rate is reversed: 0 MPixel/s for the MI350P and 11.20 GPixel/s for the 820M.

Memory type and capacity are fundamental architectural splits. The MI350P uses HBM3e with 144 GB capacity and a fixed bandwidth of 8.19 TB/s. The 820M uses system-shared memory, which means its capacity, type, bus width, and bandwidth are all system dependent. The MI350P’s memory clock is 2000 MHz with 8 Gbps effective, a fixed value. The 820M’s memory clock is listed as system shared. This difference defines their use cases: the MI350P is for data-center-scale workloads that require massive, fast memory, and the 820M is for thin-and-light laptops where memory is shared with the CPU.

Specification Differences

The two parts differ in nearly every recorded specification. The MI350P has a base clock of 1000 MHz and a boost clock of 2200 MHz, while the 820M has a base clock of 400 MHz and a boost clock of 2800 MHz. The MI350P’s memory clock is 2000 MHz with 8 Gbps effective, and the 820M’s is system shared. The MI350P has 144 GB of HBM3e memory on an 8192-bit bus with 8.19 TB/s bandwidth. The 820M has system-shared memory with a system-dependent bus width and bandwidth. The MI350P has 8192 shading units, 512 TMUs, and 0 ROPs. The 820M has 128 shading units, 8 TMUs, and 4 ROPs. The MI350P has no ray tracing cores, and the 820M has 2. The MI350P’s pixel rate is 0 MPixel/s, and the 820M’s is 11.20 GPixel/s. The MI350P’s texture rate is 1,126.4 GTexel/s, and the 820M’s is 22.40 GTexel/s. FP32 and FP16 are 36.04 TFLOPS for the MI350P and 716.8 GFLOPS for the 820M. The MI350P has a TDP of 600 W, and the 820M has a TDP of 15 W. The MI350P is dual-slot, and the 820M is an IGP. The MI350P uses one 16-pin power connector, and the 820M has no power connectors. The MI350P’s suggested PSU is 1000 W, and the 820M has no suggested PSU. The MI350P uses PCIe 5.0 x16, and the 820M uses PCIe 4.0 x8. The MI350P has no display outputs, and the 820M’s outputs are portable device dependent. The MI350P’s APIs are all N/A, and the 820M supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4. The MI350P is 267 mm long, 111 mm high, and 40 mm wide. The 820M has no recorded dimensions.

The MI350P is on a 3 nm process with 73,000 million transistors and a 1190 mm² die. The 820M is on a 4 nm process with unknown transistor count and die size. The MI350P’s architecture is CDNA 4.0, and the 820M’s is RDNA 3.5. The MI350P’s chip is MI350 128CU, and the 820M’s is Krackan Point 2. The MI350P’s generation is Instinct (MIx), and the 820M’s is Navi III IGP (Strix Point Mobile). The MI350P’s release date is 2026-05-06, and the 820M’s is 2025-02-28. The MI350P’s predecessor is Radeon Instinct, and the 820M’s is Navi II IGP. The 820M’s production status is active, and the MI350P has no recorded production status. Neither part has a launch MSRP in the database.

FAQ

Q: Which GPU has more FP32 compute power?

A: The AMD Instinct MI350P delivers 36.04 TFLOPS of FP32, while the AMD Radeon 820M delivers 716.8 GFLOPS. The MI350P is 50 times higher in FP32 throughput.

Q: How much memory does each GPU have?

A: The MI350P has 144 GB of HBM3e memory with an 8192-bit bus and 8.19 TB/s bandwidth. The 820M uses system-shared memory, and its capacity, type, bus width, and bandwidth are all system dependent.

Q: Does the Radeon 820M support ray tracing?

A: Yes, the 820M has 2 ray tracing cores. The Instinct MI350P lists no ray tracing cores, as it is a compute accelerator without a graphics pipeline.

Q: What is the power consumption difference?

A: The MI350P has a TDP of 600 W and a suggested PSU of 1000 W. The 820M has a TDP of 15 W and no suggested PSU, as it is an integrated part with no power connectors.

Q: Can the Instinct MI350P output video?

A: No. The MI350P has no display outputs and records a pixel rate of 0 MPixel/s. The 820M has portable device dependent display outputs and a pixel rate of 11.20 GPixel/s.

Q: Which APIs does each GPU support?

A: The MI350P lists N/A for DirectX, OpenGL, and Vulkan. The 820M supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI350P
820M
Core Specs
Shading Units
8,192
128 -98.4%
Shaders
8,192
128 -98.4%
TMUs
512
8 -98.4%
ROPs
0
4 +∞%
Compute Units
128
2 -98.4%
Clocks
Base Clock
1000 MHz
400 MHz
Boost Clock
2200 MHz
2800 MHz
Memory Clock
2000 MHz 8 Gbps effective
System Shared
Memory
Memory Size
144 GB
System Shared
VRAM (MB)
147,456
Memory Type
HBM3e
System Shared
Memory Bus
8192 bit
System Shared
Bandwidth
8.19 TB/s
System Dependent
Cache
L1 Cache
16 KB (per CU)
128 KB per Array
L2 Cache
16 MB
1024 KB
L3 Cache
128 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
0 MPixel/s
11.20 GPixel/s
Texture Rate
1,126.4 GTexel/s
22.40 GTexel/s
FP32 (TFLOPS)
36.04 TFLOPS
716.8 GFLOPS
FP64 (TFLOPS)
18.02 TFLOPS (1:2)
44.80 GFLOPS (1:16)
FP16 (TFLOPS)
36.04 TFLOPS (1:1)
716.8 GFLOPS (1:1)
AI/RT
RT Cores
2
Matrix Cores
512
Power
TDP
600 W
15 W
TDP (W)
600
15 -97.5%
Suggested PSU
1000 W
Power Connectors
1x 16-pin
None
Architecture
Architecture
CDNA 4.0
RDNA 3.5
GPU Name
MI350 128CU
Krackan Point 2
Generation
Instinct (MIx)
Navi III IGP (Strix Point Mobile)
Process Size
3 nm
4 nm
Transistors
73,000 million
unknown
Die Size
1190 mm²
unknown
Foundry
TSMC
TSMC
Density
61.3M / mm²
AMD MCM
MCM
2
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
2.1
Shader Model
6.8
Physical
Slot Width
Dual-slot
IGP
Length
267 mm 10.5 inches
Height
111 mm 4.4 inches
Outputs
No outputs
Portable Device Dependent
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x8
Other
Production
Active
Predecessor
Radeon Instinct
Navi II IGP
View Instinct MI350P Details View Radeon 820M Details