AMD Instinct MI100 vs NVIDIA L40 Comparison

AMD
RADEON

AMD Instinct MI100

CORE STATE Arcturus
VRAM 32 GB
CLOCK SPEED 1502 MHz
TDP 300 W
BUS WIDTH 4096 bit
ARCHITECTURE CDNA 1.0
nm
PROCESS 7 nm
LAUNCH DATE 2020
VS
NVIDIA
GEFORCE

L40

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2490 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
139,035
330,926
geekbench_vulkan
N/A
237,295

Analysis: AMD Instinct MI100 vs NVIDIA L40

# Head-to-Head Benchmarks

The benchmark data shows a decisive victory for the NVIDIA L40 in the single available head-to-head comparison. In Geekbench OpenCL, the L40 scores 330,926 points against the AMD Instinct MI100’s 139,035 points, a delta of 138% in favor of the NVIDIA part. This is not a marginal win; the L40 more than doubles the MI100’s raw compute output in this workload.

Context from the nearest-rivals data reinforces the gap. The L40’s average benchmark score of 284,111 places it at the 99th percentile of all GPUs, while the MI100’s average of 139,035 sits at the 96th percentile. The MI100’s nearest rivals — the NVIDIA Tesla V100 PCIe 16 GB (138,063), Tesla V100 SXM2 32 GB (137,731), AMD Radeon PRO V620 (136,472), and Radeon Pro W6800X Duo (135,774) — all cluster within a 2.4% delta of the MI100. That suggests the MI100 is competitive with an older generation of accelerators, but the L40 is operating in a different tier entirely.

For the L40, its nearest rivals include the NVIDIA RTX 6000 Ada Generation (287,237, delta -1.1%), the NVIDIA L40S (295,763, delta -3.9%), the NVIDIA L20 (251,147, delta 13.1%), and the AMD Instinct MI300X (317,994, delta -10.7%). Notably, the L40 trails the MI300X by 10.7%, and the L40S by 3.9%, but it leads the L20 by 13.1% and essentially trades blows with the RTX 6000 Ada. The MI100, by contrast, does not appear anywhere near the L40’s rival list — the closest AMD accelerator in that group is the MI300X, which is 10.7% ahead of the L40 but still nowhere close to the MI100’s score.

The wins tally reflects this asymmetry: the L40 wins 1 head-to-head benchmark, the MI100 wins 0. There is no Vulkan score for the MI100 in the data, so the comparison rests entirely on OpenCL. Even so, the result is unambiguous — the L40 delivers over twice the performance in the tested workload.

# FAQ

Q: How much faster is the NVIDIA L40 than the AMD Instinct MI100 in OpenCL?

A: The L40 scores 330,926 in Geekbench OpenCL, while the MI100 scores 139,035. That is a 138% delta in favor of the L40, meaning the L40 more than doubles the MI100’s performance in this specific test.

Q: Where does the MI100 rank relative to its closest competitors?

A: The MI100’s average score of 139,035 puts it 0.7% ahead of the Tesla V100 PCIe 16 GB, 0.9% ahead of the Tesla V100 SXM2 32 GB, 1.9% ahead of the Radeon PRO V620, and 2.4% ahead of the Radeon Pro W6800X Duo. All four rivals are within a narrow band around the MI100’s score.

Q: How does the L40 compare to the next-fastest NVIDIA accelerators?

A: The L40’s average score of 284,111 is 1.1% behind the RTX 6000 Ada Generation, 3.9% behind the L40S, and 13.1% ahead of the L20. It also trails the AMD Instinct MI300X by 10.7%.

Q: Does the MI100 have any benchmark advantage over the L40?

A: No. In the only head-to-head benchmark available (Geekbench OpenCL), the L40 wins outright. The MI100 has no winning entries in the data.

Q: What percentile do these GPUs occupy relative to all GPUs?

A: The L40 is at the 99th percentile, while the MI100 is at the 96th percentile. Despite the large performance gap, both are high-ranking accelerators in the overall distribution.

Q: Are there any other benchmark tests in the data?

A: The L40 has both a Geekbench OpenCL score (330,926) and a Geekbench Vulkan score (237,295). The MI100 only has a Geekbench OpenCL score (139,035). No Vulkan result is listed for the MI100, so cross-API comparison is not possible.

# Architecture Differences

The two accelerators are built on fundamentally different architectures. The NVIDIA L40 uses the AD102 chip with the Ada Lovelace architecture, manufactured on a 5 nm process at TSMC. The AMD Instinct MI100 uses the Arcturus chip with the CDNA 1.0 architecture, also fabbed at TSMC but on a 7 nm node. This process gap is significant: the L40 packs 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3 million per mm². The MI100, by contrast, has 25,600 million transistors on a larger 750 mm² die, for a density of just 34.1 million per mm². The L40 achieves roughly 3.7 times the transistor density of the MI100.

Core configuration diverges sharply as well. The L40 has 18,176 shading units, 568 texture mapping units, and 192 ROPs. It also includes 142 RT cores and 568 tensor cores, reflecting its Ada Lovelace heritage with dedicated ray tracing and AI acceleration hardware. The MI100 has 7,680 shading units, 480 TMUs, and 64 ROPs, with no RT cores and no tensor cores listed. The MI100’s compute paradigm relies on its CDNA design, which prioritizes raw FP32 and FP16 throughput over graphics-oriented features.

The L40’s FP32 throughput is 90.52 TFLOPS, with FP16 at the same 90.52 TFLOPS (1:1 ratio). The MI100 delivers 23.07 TFLOPS FP32 and 46.14 TFLOPS FP16 (2:1 ratio). This means the L40 has nearly four times the FP32 capability and roughly double the FP16 capability of the MI100.

Pixel and texture rates follow the same pattern. The L40 hits 478.1 GPixel/s and 1,414.3 GTexel/s, while the MI100 manages 96.13 GPixel/s and 721.0 GTexel/s. The L40’s pixel rate is nearly 5 times higher, and its texture rate is about 1.96 times higher.

API support is another differentiator. The L40 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The MI100 lists N/A for DirectX, OpenGL, and Vulkan — it is a compute-focused accelerator without graphics API support. Display outputs also reflect this: the L40 has 4x DisplayPort 1.4a, while the MI100 has no outputs at all.

# Specification Differences

Clock speeds differ notably. The L40 has a base clock of 735 MHz and a boost clock of 2490 MHz. The MI100 runs at a 1000 MHz base and 1502 MHz boost. Despite the L40’s lower base clock, its boost clock is significantly higher, which contributes to its performance lead.

Memory configurations are also different. The L40 comes with 48 GB of GDDR6 on a 384-bit bus, delivering 864.0 GB/s of bandwidth. The MI100 has 32 GB of HBM2 on a 4096-bit bus, achieving 1.23 TB/s of bandwidth. The MI100’s memory bandwidth is higher, but the L40 has more capacity and a newer memory type. Memory clocks are listed at 2250 MHz (18 Gbps effective) for the L40 and 1200 MHz (2.4 Gbps effective) for the MI100.

Power delivery differs in connector type. The L40 uses a single 16-pin connector, while the MI100 uses two 8-pin connectors. Both have the same 300 W TDP and the same 700 W suggested PSU. Both are dual-slot cards with identical physical dimensions: 267 mm in length and 111 mm in height.

The L40 supports PCIe 4.0 x16, as does the MI100. Release dates are far apart: the L40 launched on 2022-10-12, while the MI100 came out on 2020-11-15. Both are end-of-life products. The L40’s predecessor is Server Ampere and its successor is Server Hopper; the MI100’s predecessor is Radeon Instinct, with no successor listed.

# The Verdict

The data paints a stark picture. The NVIDIA L40 is the clear performance winner, with an OpenCL score that is 138% higher than the AMD Instinct MI100’s. This is not a close contest — the L40 sits at the 99th percentile of all GPUs, while the MI100 sits at the 96th, and the gap between them in raw score is larger than the gap between the MI100 and its nearest rivals.

For users prioritizing raw compute throughput, the L40’s 90.52 TFLOPS FP32 and 90.52 TFLOPS FP16 are more than double and nearly double the MI100’s respective figures. The L40 also brings modern features that the MI100 lacks entirely: RT cores, tensor cores, and full graphics API support with DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4. The MI100’s N/A entries for all graphics APIs and its lack of display outputs mark it as a pure compute accelerator, whereas the L40 can also drive displays and handle graphics workloads.

The MI100’s advantages are narrower. Its HBM2 memory delivers 1.23 TB/s of bandwidth versus the L40’s 864.0 GB/s, and its 4096-bit bus is wider than the L40’s 384-bit bus. But in terms of memory capacity, the L40 offers 48 GB versus the MI100’s 32 GB. The MI100 also has a higher base clock (1000 MHz vs 735 MHz), though the L40’s boost clock (2490 MHz vs 1502 MHz) more than compensates.

In terms of nearest-rival positioning, the MI100’s closest competitors are all older NVIDIA Tesla V100 variants and AMD Radeon Pro cards, all within a 2.4% delta. The L40, by contrast, competes with the RTX 6000 Ada Generation and the L40S, and even the AMD Instinct MI300X is only 10.7% ahead of it. The MI100 simply does not appear in the L40’s competitive set.

Who should pick which? Based strictly on the data, the NVIDIA L40 is the superior accelerator for anyone needing maximum compute performance, modern feature support, or graphics capability. The AMD Instinct MI100 remains a viable option only for workloads that specifically benefit from its higher memory bandwidth or wider bus, and even then, its lower FP32 and FP16 throughput and lack of graphics APIs make it a niche choice. The benchmark results are unambiguous: the L40 wins the only head-to-head test, and the architecture and specification differences reinforce that outcome.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI100
L40
Core Specs
Shading Units
7,680
18,176 +136.7%
Shaders
7,680
18,176 +136.7%
TMUs
480
568 +18.3%
ROPs
64
192 +200.0%
Compute Units
120
—
SM Count
—
142
Clocks
Base Clock
1000 MHz
735 MHz
Boost Clock
1502 MHz
2490 MHz
Memory Clock
1200 MHz 2.4 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
32 GB
48 GB
VRAM (MB)
32,768
49,152 +50.0%
Memory Type
HBM2
GDDR6
Memory Bus
4096 bit
384 bit
Bandwidth
1.23 TB/s
864.0 GB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
8 MB
96 MB
Performance
Pixel Rate
96.13 GPixel/s
478.1 GPixel/s
Texture Rate
721.0 GTexel/s
1,414.3 GTexel/s
FP32 (TFLOPS)
23.07 TFLOPS
90.52 TFLOPS
FP64 (TFLOPS)
11.54 TFLOPS (1:2)
1,414.3 GFLOPS (1:64)
FP16 (TFLOPS)
46.14 TFLOPS (2:1)
90.52 TFLOPS (1:1)
AI/RT
RT Cores
—
142
Tensor Cores
—
568
Power
TDP
300 W
300 W
TDP (W)
300
300 0.0%
Suggested PSU
700 W
700 W
Power Connectors
2x 8-pin
1x 16-pin
Architecture
Architecture
CDNA 1.0
Ada Lovelace
GPU Name
Arcturus
AD102
Generation
Instinct (MIx)
Server Ada (Lxx)
Process Size
7 nm
5 nm
Transistors
25,600 million
76,300 million
Die Size
750 mm²
609 mm²
Foundry
TSMC
TSMC
Density
34.1M / mm²
125.3M / mm²
API Support
DirectX
—
12 Ultimate (12_2)
OpenGL
—
4.6
Vulkan
—
1.4
OpenCL
2.1
3.0
CUDA
—
8.9
Shader Model
—
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Radeon Instinct
Server Ampere
Successor
—
Server Hopper
View Instinct MI100 Details View L40 Details