GPU Comparison
AMD Instinct MI300X
L40S
PERFORMANCE BENCHMARKS
Analysis: AMD Instinct MI300X vs NVIDIA L40S
Head-to-Head Benchmarks
The single direct comparison available between the AMD Instinct MI300X and the NVIDIA L40S is the Geekbench OpenCL compute test, and the result is close. The NVIDIA L40S edges out the AMD part with a score of 330,727 against the MI300X’s 317,994, a delta of -3.9% for the AMD accelerator. That is not a wide margin, it is within the kind of run-to-run variance one might expect from a single benchmark, but it is a clear win for the L40S on raw OpenCL throughput.
Context from the nearest-rival data reinforces how tight this pairing is. The MI300X’s average benchmark score sits at 317,994, while the L40S’s average is 295,763. The L40S’s single OpenCL run is actually 11.8% above its own average (330,727 vs. 295,763), which suggests that the direct head-to-head result may favor the L40S more than its typical performance profile would indicate. Conversely, the MI300X’s OpenCL score exactly equals its average, so there is no similar upward deviation for AMD.
Looking at the wider rival field, the MI300X is 7.5% ahead of the L40S based on average scores (317,994 vs. 295,763). That is a meaningful gap in the other direction. The discrepancy between the head-to-head delta (-3.9%) and the average-score delta (+7.5%) is explained by the L40S’s strong OpenCL showing relative to its other benchmark results. For the L40S, the data lists a second benchmark, Geekbench Vulkan at 260,799, which drags its average down considerably. The MI300X has no Vulkan result listed, so its average is purely OpenCL-driven.
The L40S also holds a percentile advantage, sitting at the 99th percentile versus the MI300X’s 100th percentile among all GPUs. That is a narrow gap in percentile terms, but it indicates both parts are at the very top of the database. In the nearest-rivals table, the MI300X’s deltaPct versus the L40S is +7.5, while the L40S’s deltaPct versus the MI300X is -7. That reciprocal relationship confirms that, on average, the AMD part is the stronger compute performer, even though the single OpenCL test favors NVIDIA.
The most lopsided wins in the rival data are not between these two cards but against other hardware. The MI300X is 10.7% ahead of the RTX 6000 Ada Generation, while the L40S is only 3% ahead of that same RTX 6000. Meanwhile, the L40S is 4.1% ahead of the NVIDIA L40, and the MI300X trails the NVIDIA H200 NVL by 5% and the B200 by 8%. Neither card tops the absolute best in the database, but both are clearly in the top tier.
Architecture Differences
The two accelerators come from fundamentally different design philosophies, and the data shows it clearly. The AMD Instinct MI300X uses the CDNA 3.0 architecture on the Aqua Vanjaram chip, built on a 5 nm TSMC process with 153,000 million transistors on a 1017 mm² die. The NVIDIA L40S uses the Ada Lovelace architecture with the AD102 chip, also on a 5 nm TSMC process, but with 76,300 million transistors on a 609 mm² die. The MI300X packs more than twice the transistor count and a 67% larger die, giving it a transistor density of 150.4M per mm² versus the L40S’s 125.3M per mm².
Memory is where the divergence becomes stark. The MI300X carries 192 GB of HBM3 across an 8192-bit bus, delivering 5.32 TB/s of bandwidth. The L40S has 48 GB of GDDR6 on a 384-bit bus, with 864.0 GB/s of bandwidth. That is a 4x capacity advantage and a 6.2x bandwidth advantage for AMD. For workloads that are memory-bound, large model inference, massive datasets, the MI300X has an enormous structural edge. The L40S’s memory clock is listed at 2250 MHz with 18 Gbps effective, while the MI300X’s memory runs at 1300 MHz with 5.2 Gbps effective, but the bus width difference renders those clock numbers almost irrelevant.
Compute resources also differ significantly. The MI300X has 19,456 shading units and 1,216 TMUs, but zero ROPs and a pixel rate of 0 MPixel/s. The L40S has 18,176 shading units, 568 TMUs, and 192 ROPs, with a pixel rate of 483.8 GPixel/s. The MI300X has no RT cores or tensor cores listed, while the L40S has 142 RT cores and 568 tensor cores. The MI300X’s texture rate is 2,553.6 GTexel/s, versus 1,431.4 GTexel/s for the L40S. In FP32 and FP16, the L40S is nominally higher at 91.61 TFLOPS for both, compared to the MI300X’s 81.72 TFLOPS for both.
Clock speeds tell a similar story. The MI300X has a base clock of 1000 MHz and a boost of 2100 MHz. The L40S boosts to 2520 MHz with a 1110 MHz base. The L40S’s higher clocks partially compensate for its smaller transistor budget, but the MI300X’s raw scale still gives it the average-score advantage.
Power and physical design are also divergent. The MI300X has a TDP of 750 W, uses an OAM module form factor, has no power connectors listed, and requires a suggested PSU of 1150 W. The L40S has a TDP of 300 W, is dual-slot, uses a single 16-pin connector, and has a suggested PSU of 700 W. The L40S also has display outputs, 1x HDMI 2.1 and 3x DisplayPort 1.4a, whereas the MI300X has no display outputs. The L40S supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4; the MI300X lists N/A for all graphics APIs. The L40S is 267 mm long and 111 mm tall, and its production status is end-of-life. The MI300X’s production status is not listed.
The Verdict
Based strictly on the data, the AMD Instinct MI300X is the stronger compute accelerator on average. Its average benchmark score of 317,994 places it 7.5% ahead of the L40S’s 295,763. The MI300X also holds the 100th percentile versus the L40S’s 99th, and it leads the RTX 6000 Ada Generation by 10.7%, while the L40S leads that same card by only 3%. For anyone prioritizing raw throughput in a compute-heavy environment, especially one that can tolerate a 750 W TDP and OAM module form factor, the MI300X is the data-backed choice.
However, the single head-to-head OpenCL test goes to the L40S by 3.9%, and the L40S’s Vulkan result at 260,799 shows it has broader API support. The L40S also has a dramatically lower power draw (300 W vs. 750 W), a standard dual-slot PCIe 4.0 form factor, display outputs, and full graphics API support. The MI300X has no graphics APIs, no ROPs, and no display outputs, it is purely a compute accelerator. The L40S is end-of-life, whereas the MI300X has no end-of-life flag.
The verdict from the data: pick the MI300X for maximum average compute performance, massive memory capacity (192 GB), and bandwidth (5.32 TB/s), provided the power and form factor constraints are acceptable. Pick the L40S for a balanced accelerator that also handles graphics, supports Vulkan and DirectX, draws 450 W less, and still delivers competitive OpenCL performance. The MI300X wins on scale; the L40S wins on versatility and efficiency.
FAQ
Q: Which GPU has the higher average benchmark score?
A: The AMD Instinct MI300X, with an average score of 317,994, which is 7.5% higher than the NVIDIA L40S’s 295,763.
Q: What was the result of the direct head-to-head OpenCL test?
A: The NVIDIA L40S won with a score of 330,727 against the MI300X’s 317,994, a delta of -3.9% for the AMD part.
Q: How much memory does each card have, and what is the bandwidth?
A: The MI300X has 192 GB of HBM3 with 5.32 TB/s bandwidth. The L40S has 48 GB of GDDR6 with 864.0 GB/s bandwidth.
Q: Which card has a lower power draw?
A: The NVIDIA L40S has a TDP of 300 W, versus the MI300X’s 750 W. The L40S also has a suggested PSU of 700 W, compared to 1150 W for the MI300X.
Q: Does either card support graphics APIs?
A: The L40S supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4. The MI300X lists N/A for DirectX, OpenGL, and Vulkan.
Q: What is the transistor count difference?
A: The MI300X has 153,000 million transistors, while the L40S has 76,300 million. The MI300X’s die is 1017 mm² versus the L40S’s 609 mm².
Where Each One Wins
The AMD Instinct MI300X wins in scenarios that benefit from its massive memory subsystem. The 192 GB of HBM3 with 5.32 TB/s bandwidth is a 4x capacity and 6.2x bandwidth advantage over the L40S. Workloads that require holding large models or datasets in memory, without constant host transfers, will see a structural benefit from the MI300X. Its 2,553.6 GTexel/s texture rate is also 78% higher than the L40S’s 1,431.4 GTexel/s, which matters for texture-heavy compute tasks. The MI300X’s 7.5% average-score lead over the L40S, and its 10.7% lead over the RTX 6000 Ada Generation, make it the pick for pure compute throughput. Its 100th percentile ranking and absence of an end-of-life flag further support its position as a current-generation compute workhorse.
The NVIDIA L40S wins in efficiency and versatility. Its 300 W TDP is less than half the MI300X’s 750 W, and its dual-slot PCIe 4.0 form factor with a single 16-pin connector is far easier to integrate into standard servers. The L40S’s display outputs, 1x HDMI 2.1 and 3x DisplayPort 1.4a, and full graphics API support (DirectX 12 Ultimate, OpenGL 4.6, Vulkan 1.4) mean it can handle visualization and graphics workloads that the MI300X cannot. The L40S’s 483.8 GPixel/s pixel rate and 192 ROPs give it actual rasterization capability, which the MI300X lacks entirely (0 ROPs, 0 MPixel/s). The L40S also wins the single direct OpenCL comparison by 3.9%, and its Vulkan score of 260,799 demonstrates API breadth. For environments where power draw, graphics support, or standard mounting matter more than raw memory capacity, the L40S is the data-backed selection.