GPU Comparison

AMD
RADEON

AMD Instinct MI300X

CORE STATE Aqua Vanjaram
VRAM 192 GB
CLOCK SPEED 2100 MHz
TDP 750 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 3.0
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
317,994
330,727
geekbench_vulkan
N/A
260,799

Analysis: AMD Instinct MI300X vs NVIDIA L40S

Head-to-Head Benchmarks

The single direct comparison available between the AMD Instinct MI300X and the NVIDIA L40S is the Geekbench OpenCL compute test, and the result is close. The NVIDIA L40S edges out the AMD part with a score of 330,727 against the MI300X’s 317,994, a delta of -3.9% for the AMD accelerator. That is not a wide margin, it is within the kind of run-to-run variance one might expect from a single benchmark, but it is a clear win for the L40S on raw OpenCL throughput.

Context from the nearest-rival data reinforces how tight this pairing is. The MI300X’s average benchmark score sits at 317,994, while the L40S’s average is 295,763. The L40S’s single OpenCL run is actually 11.8% above its own average (330,727 vs. 295,763), which suggests that the direct head-to-head result may favor the L40S more than its typical performance profile would indicate. Conversely, the MI300X’s OpenCL score exactly equals its average, so there is no similar upward deviation for AMD.

Looking at the wider rival field, the MI300X is 7.5% ahead of the L40S based on average scores (317,994 vs. 295,763). That is a meaningful gap in the other direction. The discrepancy between the head-to-head delta (-3.9%) and the average-score delta (+7.5%) is explained by the L40S’s strong OpenCL showing relative to its other benchmark results. For the L40S, the data lists a second benchmark, Geekbench Vulkan at 260,799, which drags its average down considerably. The MI300X has no Vulkan result listed, so its average is purely OpenCL-driven.

The L40S also holds a percentile advantage, sitting at the 99th percentile versus the MI300X’s 100th percentile among all GPUs. That is a narrow gap in percentile terms, but it indicates both parts are at the very top of the database. In the nearest-rivals table, the MI300X’s deltaPct versus the L40S is +7.5, while the L40S’s deltaPct versus the MI300X is -7. That reciprocal relationship confirms that, on average, the AMD part is the stronger compute performer, even though the single OpenCL test favors NVIDIA.

The most lopsided wins in the rival data are not between these two cards but against other hardware. The MI300X is 10.7% ahead of the RTX 6000 Ada Generation, while the L40S is only 3% ahead of that same RTX 6000. Meanwhile, the L40S is 4.1% ahead of the NVIDIA L40, and the MI300X trails the NVIDIA H200 NVL by 5% and the B200 by 8%. Neither card tops the absolute best in the database, but both are clearly in the top tier.

Architecture Differences

The two accelerators come from fundamentally different design philosophies, and the data shows it clearly. The AMD Instinct MI300X uses the CDNA 3.0 architecture on the Aqua Vanjaram chip, built on a 5 nm TSMC process with 153,000 million transistors on a 1017 mm² die. The NVIDIA L40S uses the Ada Lovelace architecture with the AD102 chip, also on a 5 nm TSMC process, but with 76,300 million transistors on a 609 mm² die. The MI300X packs more than twice the transistor count and a 67% larger die, giving it a transistor density of 150.4M per mm² versus the L40S’s 125.3M per mm².

Memory is where the divergence becomes stark. The MI300X carries 192 GB of HBM3 across an 8192-bit bus, delivering 5.32 TB/s of bandwidth. The L40S has 48 GB of GDDR6 on a 384-bit bus, with 864.0 GB/s of bandwidth. That is a 4x capacity advantage and a 6.2x bandwidth advantage for AMD. For workloads that are memory-bound, large model inference, massive datasets, the MI300X has an enormous structural edge. The L40S’s memory clock is listed at 2250 MHz with 18 Gbps effective, while the MI300X’s memory runs at 1300 MHz with 5.2 Gbps effective, but the bus width difference renders those clock numbers almost irrelevant.

Compute resources also differ significantly. The MI300X has 19,456 shading units and 1,216 TMUs, but zero ROPs and a pixel rate of 0 MPixel/s. The L40S has 18,176 shading units, 568 TMUs, and 192 ROPs, with a pixel rate of 483.8 GPixel/s. The MI300X has no RT cores or tensor cores listed, while the L40S has 142 RT cores and 568 tensor cores. The MI300X’s texture rate is 2,553.6 GTexel/s, versus 1,431.4 GTexel/s for the L40S. In FP32 and FP16, the L40S is nominally higher at 91.61 TFLOPS for both, compared to the MI300X’s 81.72 TFLOPS for both.

Clock speeds tell a similar story. The MI300X has a base clock of 1000 MHz and a boost of 2100 MHz. The L40S boosts to 2520 MHz with a 1110 MHz base. The L40S’s higher clocks partially compensate for its smaller transistor budget, but the MI300X’s raw scale still gives it the average-score advantage.

Power and physical design are also divergent. The MI300X has a TDP of 750 W, uses an OAM module form factor, has no power connectors listed, and requires a suggested PSU of 1150 W. The L40S has a TDP of 300 W, is dual-slot, uses a single 16-pin connector, and has a suggested PSU of 700 W. The L40S also has display outputs, 1x HDMI 2.1 and 3x DisplayPort 1.4a, whereas the MI300X has no display outputs. The L40S supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4; the MI300X lists N/A for all graphics APIs. The L40S is 267 mm long and 111 mm tall, and its production status is end-of-life. The MI300X’s production status is not listed.

The Verdict

Based strictly on the data, the AMD Instinct MI300X is the stronger compute accelerator on average. Its average benchmark score of 317,994 places it 7.5% ahead of the L40S’s 295,763. The MI300X also holds the 100th percentile versus the L40S’s 99th, and it leads the RTX 6000 Ada Generation by 10.7%, while the L40S leads that same card by only 3%. For anyone prioritizing raw throughput in a compute-heavy environment, especially one that can tolerate a 750 W TDP and OAM module form factor, the MI300X is the data-backed choice.

However, the single head-to-head OpenCL test goes to the L40S by 3.9%, and the L40S’s Vulkan result at 260,799 shows it has broader API support. The L40S also has a dramatically lower power draw (300 W vs. 750 W), a standard dual-slot PCIe 4.0 form factor, display outputs, and full graphics API support. The MI300X has no graphics APIs, no ROPs, and no display outputs, it is purely a compute accelerator. The L40S is end-of-life, whereas the MI300X has no end-of-life flag.

The verdict from the data: pick the MI300X for maximum average compute performance, massive memory capacity (192 GB), and bandwidth (5.32 TB/s), provided the power and form factor constraints are acceptable. Pick the L40S for a balanced accelerator that also handles graphics, supports Vulkan and DirectX, draws 450 W less, and still delivers competitive OpenCL performance. The MI300X wins on scale; the L40S wins on versatility and efficiency.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The AMD Instinct MI300X, with an average score of 317,994, which is 7.5% higher than the NVIDIA L40S’s 295,763.

Q: What was the result of the direct head-to-head OpenCL test?

A: The NVIDIA L40S won with a score of 330,727 against the MI300X’s 317,994, a delta of -3.9% for the AMD part.

Q: How much memory does each card have, and what is the bandwidth?

A: The MI300X has 192 GB of HBM3 with 5.32 TB/s bandwidth. The L40S has 48 GB of GDDR6 with 864.0 GB/s bandwidth.

Q: Which card has a lower power draw?

A: The NVIDIA L40S has a TDP of 300 W, versus the MI300X’s 750 W. The L40S also has a suggested PSU of 700 W, compared to 1150 W for the MI300X.

Q: Does either card support graphics APIs?

A: The L40S supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4. The MI300X lists N/A for DirectX, OpenGL, and Vulkan.

Q: What is the transistor count difference?

A: The MI300X has 153,000 million transistors, while the L40S has 76,300 million. The MI300X’s die is 1017 mm² versus the L40S’s 609 mm².

Where Each One Wins

The AMD Instinct MI300X wins in scenarios that benefit from its massive memory subsystem. The 192 GB of HBM3 with 5.32 TB/s bandwidth is a 4x capacity and 6.2x bandwidth advantage over the L40S. Workloads that require holding large models or datasets in memory, without constant host transfers, will see a structural benefit from the MI300X. Its 2,553.6 GTexel/s texture rate is also 78% higher than the L40S’s 1,431.4 GTexel/s, which matters for texture-heavy compute tasks. The MI300X’s 7.5% average-score lead over the L40S, and its 10.7% lead over the RTX 6000 Ada Generation, make it the pick for pure compute throughput. Its 100th percentile ranking and absence of an end-of-life flag further support its position as a current-generation compute workhorse.

The NVIDIA L40S wins in efficiency and versatility. Its 300 W TDP is less than half the MI300X’s 750 W, and its dual-slot PCIe 4.0 form factor with a single 16-pin connector is far easier to integrate into standard servers. The L40S’s display outputs, 1x HDMI 2.1 and 3x DisplayPort 1.4a, and full graphics API support (DirectX 12 Ultimate, OpenGL 4.6, Vulkan 1.4) mean it can handle visualization and graphics workloads that the MI300X cannot. The L40S’s 483.8 GPixel/s pixel rate and 192 ROPs give it actual rasterization capability, which the MI300X lacks entirely (0 ROPs, 0 MPixel/s). The L40S also wins the single direct OpenCL comparison by 3.9%, and its Vulkan score of 260,799 demonstrates API breadth. For environments where power draw, graphics support, or standard mounting matter more than raw memory capacity, the L40S is the data-backed selection.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI300X
L40S
Core Specs
Shading Units
19,456
18,176 -6.6%
Shaders
19,456
18,176 -6.6%
TMUs
1,216
568 -53.3%
ROPs
0
192 +∞%
Compute Units
304
SM Count
142
Clocks
Base Clock
1000 MHz
1110 MHz
Boost Clock
2100 MHz
2520 MHz
Memory Clock
1300 MHz 5.2 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
192 GB
48 GB
VRAM (MB)
196,608
49,152 -75.0%
Memory Type
HBM3
GDDR6
Memory Bus
8192 bit
384 bit
Bandwidth
5.32 TB/s
864.0 GB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
16 MB
48 MB
L3 Cache
256 MB
Performance
Pixel Rate
0 MPixel/s
483.8 GPixel/s
Texture Rate
2,553.6 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
81.72 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
40.86 TFLOPS (1:2)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
81.72 TFLOPS (1:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
142
Tensor Cores
568
Matrix Cores
1,216
Power
TDP
750 W
300 W
TDP (W)
750
300 -60.0%
Suggested PSU
1150 W
700 W
Power Connectors
None
1x 16-pin
Architecture
Architecture
CDNA 3.0
Ada Lovelace
GPU Name
Aqua Vanjaram
AD102
Generation
Instinct (MIx)
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
153,000 million
76,300 million
Die Size
1017 mm²
609 mm²
Foundry
TSMC
TSMC
Density
150.4M / mm²
125.3M / mm²
AMD MCM
MCM
2
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.9
Shader Model
6.8
Physical
Slot Width
OAM Module
Dual-slot
Length
267 mm 10.5 inches
Height
111 mm 4.4 inches
Outputs
No outputs
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
Predecessor
Radeon Instinct
Server Ampere
Successor
Server Hopper
View Instinct MI300X Details View L40S Details