AMD Instinct MI355X vs NVIDIA L20 Comparison

AMD
RADEON

AMD Instinct MI355X

CORE STATE MI350 256CU
VRAM 288 GB
CLOCK SPEED 2400 MHz
TDP 1400 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 4.0
nm
PROCESS 3 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

L20

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 275 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

geekbench_opencl
N/A
274,276
geekbench_vulkan
N/A
228,018

Analysis: AMD Instinct MI355X vs NVIDIA L20

FAQ

Q: What are the two accelerators compared here?

A: The AMD Instinct MI355X and the NVIDIA L20. The MI355X is an AMD Instinct (MIx) generation part built on the CDNA 4.0 architecture, while the L20 is an NVIDIA Server Ada (Lxx) generation part using the Ada Lovelace architecture.

Q: Which card has more memory and bandwidth?

A: The AMD Instinct MI355X has 288 GB of HBM3e memory on an 8192-bit bus, delivering 8.19 TB/s of bandwidth. The NVIDIA L20 has 48 GB of GDDR6 memory on a 384-bit bus, delivering 864.0 GB/s.

Q: How do the FP32 compute figures compare?

A: The MI355X delivers 78.64 TFLOPS of FP32 throughput, while the L20 delivers 59.35 TFLOPS. The MI355X holds a 32.5% advantage in raw FP32 throughput.

Q: What are the thermal and power characteristics?

A: The MI355X has a TDP of 1400 W and requires a suggested 1800 W PSU, while the L20 has a TDP of 275 W and a suggested 600 W PSU. The MI355X is an OAM module with no power connectors listed, whereas the L20 is a dual-slot card with a 1x 16-pin connector.

Q: Which card has benchmark scores in the database?

A: Only the NVIDIA L20 has recorded benchmark scores. It shows a geekbench_opencl score of 274276 and a geekbench_vulkan score of 228018, with an average benchmark score of 251147. The MI355X has no entries in the benchmarks array.

Q: What is the L20's percentile ranking?

A: The L20 sits at the 99th percentile against all GPUs in the database. Its nearest rivals include the NVIDIA PG506-232 (11.6% behind), the AMD Radeon PRO W7900D (14.2% behind), the NVIDIA L40 (11.6% ahead), and the NVIDIA RTX 6000 Ada Generation (12.6% ahead).

The Verdict

The data presents two very different accelerators. The AMD Instinct MI355X is a massive compute-oriented module with a 1400 W TDP, 288 GB of HBM3e memory, and a 3 nm process node. It is designed for high-end compute workloads where memory capacity and bandwidth are critical. The NVIDIA L20 is a more conventional dual-slot PCIe card with a 275 W TDP, 48 GB of GDDR6 memory, and a 99th percentile benchmark ranking.

The L20 is the only one with actual benchmark scores. Its average score of 251147 places it well above the NVIDIA PG506-232 (225124, 11.6% behind) and the AMD Radeon PRO W7900D (219827, 14.2% behind). It trails the NVIDIA L40 (284111, 11.6% ahead) and the NVIDIA RTX 6000 Ada Generation (287237, 12.6% ahead). The MI355X has no recorded benchmarks, so its performance cannot be quantified from the database.

For users who need a validated, high-performance accelerator with DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4 API support, the L20 is the proven choice. It also offers display outputs (4x DisplayPort 1.4a), which the MI355X lacks entirely.

For workloads that demand enormous memory capacity and bandwidth, the MI355X is the clear specification winner. Its 288 GB of HBM3e and 8.19 TB/s bandwidth dwarf the L20's 48 GB and 864.0 GB/s. The MI355X also has a higher transistor count (185,000 million vs 76,300 million) and a larger die (2380 mm² vs 609 mm²), though its density is lower (77.7M / mm² vs 125.3M / mm²).

The choice depends on the workload. The L20 is a proven, efficient, API-complete accelerator. The MI355X is a specialized compute module that prioritizes memory and throughput over everything else.

Head-to-Head Benchmarks

The head-to-head benchmark array is empty, so no direct comparison scores exist. The only recorded data comes from the L20's individual benchmarks. Its geekbench_opencl score of 274276 and geekbench_vulkan score of 228018 produce an average of 251147.

Against its nearest rivals, the L20's performance is well documented. It beats the NVIDIA PG506-232 by 11.6% and the AMD Radeon PRO W7900D by 14.2%. It is beaten by the NVIDIA L40 by 11.6% and the NVIDIA RTX 6000 Ada Generation by 12.6%.

The MI355X has no benchmark scores in the database. Its average benchmark score is 0 and its percentile is 50, which is the median value but not a reflection of actual performance. The data cannot confirm how the MI355X stacks up against the L20 in real-world tests.

What can be inferred comes from raw specifications. The MI355X's FP32 throughput of 78.64 TFLOPS is 32.5% higher than the L20's 59.35 TFLOPS. Its texture rate of 2,457.6 GTexel/s is 2.65 times the L20's 927.4 GTexel/s. The MI355X's memory bandwidth of 8.19 TB/s is 9.48 times the L20's 864.0 GB/s. These figures suggest the MI355X would dominate memory-bound workloads, but without benchmark data, this remains projection.

The L20 has a pixel rate of 322.6 GPixel/s, whereas the MI355X has a pixel rate of 0 MPixel/s. This reflects the MI355X's compute-only design. The L20 also has 128 ROPs and 92 RT cores; the MI355X lists 0 ROPs and no RT cores. The L20's shading units number 11776, while the MI355X has 16384. The L20 has 368 TMUs and 368 tensor cores; the MI355X has 1024 TMUs and no tensor core count listed.

The L20's higher clock speeds (1440 MHz base, 2520 MHz boost) versus the MI355X (1000 MHz base, 2400 MHz boost) partially explain the L20's efficiency. The L20 achieves 59.35 TFLOPS at 275 W, while the MI355X needs 1400 W to reach 78.64 TFLOPS. The L20 is more power-efficient per FLOP, though the MI355X delivers more absolute throughput.

Specification Differences

The two accelerators differ on nearly every measurable specification. The MI355X uses a 3 nm process from TSMC with 185,000 million transistors on a 2380 mm² die. The L20 uses a 5 nm process from TSMC with 76,300 million transistors on a 609 mm² die. Transistor density favors the L20 at 125.3M / mm² versus the MI355X's 77.7M / mm².

Memory configurations are starkly different. The MI355X has 288 GB of HBM3e with an 8192-bit bus and 8.19 TB/s bandwidth. The L20 has 48 GB of GDDR6 with a 384-bit bus and 864.0 GB/s bandwidth. Memory clock speeds also differ: the MI355X runs at 2000 MHz (8 Gbps effective), the L20 at 2250 MHz (18 Gbps effective).

Compute resources differ significantly. The MI355X has 16384 shading units and 1024 TMUs, against the L20's 11776 shading units and 368 TMUs. The MI355X has 0 ROPs and no RT cores or tensor cores listed; the L20 has 128 ROPs, 92 RT cores, and 368 tensor cores. The MI355X's texture rate is 2,457.6 GTexel/s versus the L20's 927.4 GTexel/s. The MI355X's pixel rate is 0 MPixel/s; the L20's is 322.6 GPixel/s.

FP32 and FP16 throughput are both 1:1 on both cards. The MI355X delivers 78.64 TFLOPS in both, the L20 delivers 59.35 TFLOPS in both.

Power and physical design diverge completely. The MI355X has a TDP of 1400 W, is an OAM Module, has no power connectors, and needs a suggested 1800 W PSU. The L20 has a TDP of 275 W, is dual-slot, uses a 1x 16-pin connector, and needs a suggested 600 W PSU. The MI355X's dimensions are 102 mm by 165 mm; the L20 is 267 mm long and 111 mm high.

Bus interfaces differ: the MI355X uses PCIe 5.0 x16, the L20 uses PCIe 4.0 x16. Display outputs are absent on the MI355X, while the L20 has 4x DisplayPort 1.4a. The MI355X lists no API support (DirectX, OpenGL, Vulkan all N/A), while the L20 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Release dates are far apart. The MI355X was released on 2025-06-11, the L20 on 2023-11-15. The MI355X's predecessor is Radeon Instinct; the L20's predecessor is Server Ampere and its successor is Server Hopper. The L20 has an active production status; the MI355X's production status is not listed.

Architecture Differences

The MI355X is built on CDNA 4.0, AMD's compute-focused architecture. It uses the MI350 256CU chip. The L20 uses the AD102 chip based on Ada Lovelace, NVIDIA's architecture that also powers consumer and workstation graphics cards.

The process nodes differ: the MI355X uses a 3 nm process, the L20 uses a 5 nm process, both from TSMC. The MI355X's transistor count of 185,000 million is 2.42 times the L20's 76,300 million, but the MI355X's die is 3.91 times larger at 2380 mm² versus 609 mm². This results in the L20 having a higher transistor density.

Memory architecture reflects different design philosophies. The MI355X uses HBM3e, a high-bandwidth memory stacked on the package, with an 8192-bit bus. The L20 uses GDDR6, a conventional discrete memory, with a 384-bit bus. The MI355X's 8.19 TB/s bandwidth is designed for memory-bound compute, while the L20's 864.0 GB/s is more typical of a general-purpose GPU.

The MI355X has no ROPs, no RT cores, and no tensor core count, and its pixel rate is 0 MPixel/s. This indicates a pure compute design with no rasterization or ray tracing hardware. The L20 has 128 ROPs, 92 RT cores, and 368 tensor cores, making it a full-featured GPU with graphics capabilities. The L20's API support confirms this: it supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the MI355X lists N/A for all APIs.

The MI355X's clock speeds are lower (1000 MHz base, 2400 MHz boost) than the L20's (1440 MHz base, 2520 MHz boost). Despite this, the MI355X's higher shading unit count (16384 vs 11776) and TMU count (1024 vs 368) allow it to achieve higher FP32 throughput and texture rate.

The MI355X's 1400 W TDP versus the L20's 275 W TDP reflects the scale of the compute module. The MI355X is an OAM Module, designed for server racks with dedicated cooling and power delivery. The L20 is a dual-slot PCIe card that fits into standard server or workstation slots.

The L20's release date of 2023-11-15 places it in the Ada Lovelace generation, with a successor already named (Server Hopper). The MI355X's release date of 2025-06-11 is newer, and its predecessor is Radeon Instinct. The MI355X's CDNA 4.0 architecture represents AMD's latest compute-focused design, while the L20's Ada Lovelace is a mature architecture from NVIDIA.

The L20's transistor density advantage (125.3M / mm² vs 77.7M / mm²) suggests a more compact design, while the MI355X's massive die and HBM3e memory point to a design optimized for memory capacity and bandwidth over density. These architectural choices make the MI355X a specialist for memory-intensive AI and scientific workloads, while the L20 is a versatile accelerator with graphics capabilities and proven benchmark performance.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI355X
L20
Core Specs
Shading Units
16,384
11,776 -28.1%
Shaders
16,384
11,776 -28.1%
TMUs
1,024
368 -64.1%
ROPs
0
128 +∞%
Compute Units
256
SM Count
92
Clocks
Base Clock
1000 MHz
1440 MHz
Boost Clock
2400 MHz
2520 MHz
Memory Clock
2000 MHz 8 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
288 GB
48 GB
VRAM (MB)
294,912
49,152 -83.3%
Memory Type
HBM3e
GDDR6
Memory Bus
8192 bit
384 bit
Bandwidth
8.19 TB/s
864.0 GB/s
Cache
L1 Cache
32 KB (per CU)
128 KB (per SM)
L2 Cache
32 MB
96 MB
L3 Cache
256 MB
Performance
Pixel Rate
0 MPixel/s
322.6 GPixel/s
Texture Rate
2,457.6 GTexel/s
927.4 GTexel/s
FP32 (TFLOPS)
78.64 TFLOPS
59.35 TFLOPS
FP64 (TFLOPS)
39.32 TFLOPS (1:2)
927.4 GFLOPS (1:64)
FP16 (TFLOPS)
78.64 TFLOPS (1:1)
59.35 TFLOPS (1:1)
AI/RT
RT Cores
92
Tensor Cores
368
Matrix Cores
1,024
Power
TDP
1400 W
275 W
TDP (W)
1,400
275 -80.4%
Suggested PSU
1800 W
600 W
Power Connectors
None
1x 16-pin
Architecture
Architecture
CDNA 4.0
Ada Lovelace
GPU Name
MI350 256CU
AD102
Generation
Instinct (MIx)
Server Ada (Lxx)
Process Size
3 nm
5 nm
Transistors
185,000 million
76,300 million
Die Size
2380 mm²
609 mm²
Foundry
TSMC
TSMC
Density
77.7M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.9
Shader Model
6.8
Physical
Slot Width
OAM Module
Dual-slot
Length
102 mm 4 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
Outputs
No outputs
4x DisplayPort 1.4a
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x16
Other
Production
Active
Predecessor
Radeon Instinct
Server Ampere
Successor
Server Hopper
View Instinct MI355X Details View L20 Details