AMD Radeon Instinct MI60 vs NVIDIA L40 Comparison

AMD
RADEON

AMD Radeon Instinct MI60

CORE STATE Vega 20
VRAM 32 GB
CLOCK SPEED 1800 MHz
TDP 300 W
BUS WIDTH 4096 bit
ARCHITECTURE GCN 5.1
nm
PROCESS 7 nm
LAUNCH DATE 2018
VS
NVIDIA
GEFORCE

L40

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2490 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
92,488
330,926
geekbench_vulkan
92,444
237,295

Analysis: AMD Radeon Instinct MI60 vs NVIDIA L40

Where Each One Wins

The recorded benchmark data splits cleanly between these two accelerators, with the NVIDIA L40 taking every measured workload. Across the two Geekbench tests in the database, the L40 wins both: OpenCL and Vulkan. The AMD Radeon Instinct MI60 records no winning benchmark in this comparison.

The L40's OpenCL result of 330,926 places it in the 99th percentile of all GPUs tracked, while the MI60's 92,488 OpenCL score sits in the 93rd percentile. That percentile gap matters: the L40 is positioned among the top tier of every recorded GPU, while the MI60 is merely above average.

For Vulkan, the L40 scores 237,295, again at the 99th percentile overall. The MI60's Vulkan result is 92,444, nearly identical to its OpenCL figure. This suggests the MI60's compute throughput is consistent across both APIs, but at a fundamentally lower performance level than the L40.

The use-case split is therefore one-sided. Workloads that favor raw compute throughput, whether through OpenCL or Vulkan, will see the L40 dominate. The MI60 does offer one structural advantage: its memory bandwidth of 1.02 TB/s exceeds the L40's 864.0 GB/s. That could matter for bandwidth-bound tasks, though the recorded benchmarks do not show any workload where the MI60's higher bandwidth translates into a win.

Architecture Differences

The two cards come from different architectural generations and design philosophies. The L40 uses the AD102 chip built on Ada Lovelace, manufactured by TSMC on a 5 nm process. The MI60 uses the Vega 20 chip on GCN 5.1, also TSMC-made but on a 7 nm node.

Transistor counts reveal the scale gap. The L40 packs 76,300 million transistors on a 609 mm² die, yielding a density of 125.3 million per mm². The MI60 has 13,230 million transistors on a 331 mm² die, for 40.0 million per mm². The L40 has over 5.7 times the transistor count on a die that is not even twice as large.

Compute resources differ by an order of magnitude. The L40 has 18,176 shading units, 568 TMUs, and 192 ROPs. The MI60 has 4,096 shading units, 256 TMUs, and 64 ROPs. The L40 also includes 142 RT cores and 568 tensor cores; the MI60 has neither, reflecting its pre-ray-tracing GCN design.

Memory configurations diverge sharply. The L40 uses 48 GB of GDDR6 on a 384-bit bus. The MI60 uses 32 GB of HBM2 on a 4096-bit bus. The MI60's wider bus gives it the bandwidth edge (1.02 TB/s versus 864.0 GB/s), but the L40 has 50% more capacity.

Clock behavior also differs. The L40 has a low base clock of 735 MHz but boosts to 2490 MHz. The MI60 starts at 1200 MHz and boosts to 1800 MHz. The L40's boost clock is 690 MHz higher, which helps explain its compute advantage. Both cards draw 300 W under the database's TDP field and require a 700 W power supply.

API support reflects their eras. The L40 supports DirectX 12 Ultimate (12_2), Vulkan 1.4, and OpenGL 4.6. The MI60 supports DirectX 12 (12_1), Vulkan 1.3, and OpenGL 4.6. The L40 also has four DisplayPort 1.4a outputs; the MI60 has just one mini-DisplayPort 1.4a.

Head-to-Head Benchmarks

The Geekbench OpenCL test shows the largest absolute gap. The L40 scores 330,926, while the MI60 scores 92,488. That is a delta of 257.8% in the L40's favor. In relative terms, the L40 delivers roughly 3.58 times the OpenCL score of the MI60. This aligns with the raw compute figures: the L40's FP32 throughput is 90.52 TFLOPS versus the MI60's 14.75 TFLOPS, a 6.1x difference, though the benchmark does not scale linearly with theoretical peak.

The Vulkan result narrows the gap somewhat but still favors the L40 decisively. The L40 scores 237,295 against the MI60's 92,444, a delta of 156.7%. The L40 still more than doubles the MI60's Vulkan performance. Notably, the MI60's Vulkan score (92,444) is nearly identical to its OpenCL score (92,488), suggesting its compute throughput is API-independent. The L40's Vulkan score is about 28% lower than its OpenCL score, indicating some API-specific overhead on the Ada architecture.

For context, the L40's average benchmark score across all recorded tests is 284,111. Its nearest rival, the NVIDIA RTX 6000 Ada Generation, averages 287,237, which is 1.1% higher. The L40 also trails the NVIDIA L40S (295,763, 3.9% higher) and the AMD Instinct MI300X (317,994, 10.7% higher). Against the NVIDIA L20, the L40 leads by 13.1%.

The MI60's average benchmark score is 92,466. Its nearest rivals cluster close: the NVIDIA RTX A4500 averages 91,671 (0.9% lower), the RTX A4500 Mobile averages 91,134 (1.5% lower), the AMD Radeon Pro VII averages 97,131 (4.8% higher), and the AMD Radeon RX 7900M averages 97,487 (5.2% higher). The MI60 sits in a tight competitive band, whereas the L40 operates in a higher performance tier with wider separation between rivals.

FAQ

Q: Which GPU has higher raw FP32 compute?

A: The NVIDIA L40 delivers 90.52 TFLOPS FP32, compared to the AMD MI60's 14.75 TFLOPS. That is a 6.1x theoretical peak advantage for the L40.

Q: Does the AMD MI60 have any advantage over the NVIDIA L40?

A: Yes, in memory bandwidth. The MI60's HBM2 memory provides 1.02 TB/s bandwidth, which is higher than the L40's 864.0 GB/s GDDR6 bandwidth. The MI60 also has a wider 4096-bit memory bus versus the L40's 384-bit bus.

Q: How do their benchmark scores compare?

A: In Geekbench OpenCL, the L40 scores 330,926 versus the MI60's 92,488, a 257.8% difference. In Vulkan, the L40 scores 237,295 versus 92,444, a 156.7% difference. The L40 wins both recorded benchmarks.

Q: Do both cards support ray tracing?

A: No. The L40 includes 142 RT cores and 568 tensor cores. The MI60 has neither RT cores nor tensor cores, as its GCN 5.1 architecture predates dedicated ray tracing hardware.

Q: What is the memory capacity difference?

A: The L40 has 48 GB of GDDR6 memory. The MI60 has 32 GB of HBM2. The L40 offers 50% more capacity, which can be critical for large model inference or rendering workloads.

Q: How do their process nodes compare?

A: The L40 is built on TSMC's 5 nm process, while the MI60 uses TSMC's 7 nm node. The L40's die is 609 mm² with 76,300 million transistors; the MI60's die is 331 mm² with 13,230 million transistors.

Q: Which card has higher clock speeds?

A: The L40 has a base clock of 735 MHz and a boost clock of 2490 MHz. The MI60 has a base clock of 1200 MHz and a boost clock of 1800 MHz. The L40's boost clock is 690 MHz higher.

Specification Differences

| Field | NVIDIA L40 | AMD Radeon Instinct MI60 |

|---|---|---|

| Architecture | Ada Lovelace | GCN 5.1 |

| Chip | AD102 | Vega 20 |

| Process Node | 5 nm | 7 nm |

| Transistors | 76,300 million | 13,230 million |

| Die Size | 609 mm² | 331 mm² |

| Transistor Density | 125.3M / mm² | 40.0M / mm² |

| Base Clock | 735 MHz | 1200 MHz |

| Boost Clock | 2490 MHz | 1800 MHz |

| Memory Size | 48 GB | 32 GB |

| Memory Type | GDDR6 | HBM2 |

| Memory Bus Width | 384 bit | 4096 bit |

| Memory Bandwidth | 864.0 GB/s | 1.02 TB/s |

| Memory Clock | 2250 MHz (18 Gbps effective) | 1000 MHz (2 Gbps effective) |

| Shading Units | 18,176 | 4,096 |

| TMUs | 568 | 256 |

| ROPs | 192 | 64 |

| RT Cores | 142 | None |

| Tensor Cores | 568 | None |

| Pixel Rate | 478.1 GPixel/s | 115.2 GPixel/s |

| Texture Rate | 1,414.3 GTexel/s | 460.8 GTexel/s |

| FP32 Performance | 90.52 TFLOPS | 14.75 TFLOPS |

| FP16 Performance | 90.52 TFLOPS (1:1) | 29.49 TFLOPS (2:1) |

| TDP | 300 W | 300 W |

| Power Connectors | 1x 16-pin | 1x 6-pin + 1x 8-pin |

| Suggested PSU | 700 W | 700 W |

| Slot Width | Dual-slot | Dual-slot |

| Display Outputs | 4x DisplayPort 1.4a | 1x mini-DisplayPort 1.4a |

| DirectX Support | 12 Ultimate (12_2) | 12 (12_1) |

| Vulkan Support | 1.4 | 1.3 |

| OpenGL Support | 4.6 | 4.6 |

| Release Date | 2022-10-12 | 2018-11-17 |

| Production Status | End-of-life | End-of-life |

The specification table confirms the generational divide. The L40 is a newer, denser design with far more compute resources and a higher boost clock. The MI60's only clear specification wins are memory bandwidth and the width of its memory bus. Both cards share the same TDP (300 W), same dual-slot form factor, same physical dimensions (267 mm length, 111 mm height), and same PCIe 4.0 x16 bus interface.

The L40's release date is nearly four years after the MI60's, and the database marks both as end-of-life. The L40's predecessor is listed as Server Ampere with successor Server Hopper; the MI60's predecessor is FirePro Data Center with no successor listed.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI60
L40
Core Specs
Shading Units
4,096
18,176 +343.8%
Shaders
4,096
18,176 +343.8%
TMUs
256
568 +121.9%
ROPs
64
192 +200.0%
Compute Units
64
—
SM Count
—
142
Clocks
Base Clock
1200 MHz
735 MHz
Boost Clock
1800 MHz
2490 MHz
Memory Clock
1000 MHz 2 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
32 GB
48 GB
VRAM (MB)
32,768
49,152 +50.0%
Memory Type
HBM2
GDDR6
Memory Bus
4096 bit
384 bit
Bandwidth
1.02 TB/s
864.0 GB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
4 MB
96 MB
Performance
Pixel Rate
115.2 GPixel/s
478.1 GPixel/s
Texture Rate
460.8 GTexel/s
1,414.3 GTexel/s
FP32 (TFLOPS)
14.75 TFLOPS
90.52 TFLOPS
FP64 (TFLOPS)
7.373 TFLOPS (1:2)
1,414.3 GFLOPS (1:64)
FP16 (TFLOPS)
29.49 TFLOPS (2:1)
90.52 TFLOPS (1:1)
AI/RT
RT Cores
—
142
Tensor Cores
—
568
Power
TDP
300 W
300 W
TDP (W)
300
300 0.0%
Suggested PSU
700 W
700 W
Power Connectors
1x 6-pin + 1x 8-pin
1x 16-pin
Architecture
Architecture
GCN 5.1
Ada Lovelace
GPU Name
Vega 20
AD102
Generation
Radeon Instinct (MIx)
Server Ada (Lxx)
Process Size
7 nm
5 nm
Transistors
13,230 million
76,300 million
Die Size
331 mm²
609 mm²
Foundry
TSMC
TSMC
Density
40.0M / mm²
125.3M / mm²
API Support
DirectX
12 (12_1)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
—
8.9
Shader Model
6.7
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
1x mini-DisplayPort 1.4a
4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
FirePro Data Center
Server Ampere
Successor
—
Server Hopper
View Radeon Instinct MI60 Details View L40 Details