AMD Radeon Instinct MI60 vs NVIDIA L40S Comparison
AMD Radeon Instinct MI60
L40S
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Instinct MI60 vs NVIDIA L40S
The Verdict
The recorded data shows a decisive performance separation between these two server accelerators. The NVIDIA L40S dominates the AMD Radeon Instinct MI60 in both recorded benchmark tests, with an average benchmark score of 295,763 compared to 92,466 for the AMD part. That is a 219.8% difference in average score, and the L40S ranks in the 99th percentile of all GPUs while the MI60 sits in the 93rd percentile.
Choose the NVIDIA L40S if your workloads are built around modern compute features, larger memory capacity, and the highest possible raw throughput. The database shows it winning both OpenCL and Vulkan tests by massive margins, and its 48 GB GDDR6 memory doubles the MI60's 32 GB HBM2 capacity. This is the card for current generation server deployments, AI inference, and rendering tasks that can exploit Ada Lovelace features.
Choose the AMD Radeon Instinct MI60 only if you have a specific legacy workload that was tuned for GCN 5.1 architecture, or if your software stack predates the L40S era. Its 93rd percentile ranking shows it remains a capable accelerator, and its nearest rivals include the NVIDIA RTX A4500 at only 0.9% behind. But the data does not support picking it over the L40S on raw performance grounds.
The L40S's nearest rivals in the database are telling. It sits 3% ahead of the NVIDIA RTX 6000 Ada Generation and 4.1% ahead of the NVIDIA L40, while trailing the AMD Instinct MI300X by 7% and the NVIDIA H200 NVL by 11.7%. The MI60, by contrast, is effectively matched with the RTX A4500 and sits 4.8% behind the AMD Radeon Pro VII. These are different performance tiers entirely.
Where Each One Wins
The head-to-head data records two wins for the NVIDIA L40S and zero for the AMD MI60. In Geekbench OpenCL, the L40S scores 330,727 against 92,488 for the MI60, a 257.6% advantage. In Geekbench Vulkan, the L40S scores 260,799 against 92,444, a 182.1% advantage.
The OpenCL result is the larger gap. A 257.6% delta means the L40S delivers more than three and a half times the MI60's score in that test. This suggests the L40S is disproportionately strong in compute-heavy OpenCL workloads, which often involve scientific simulation, machine learning kernels, and general purpose GPU compute.
The Vulkan gap, while still enormous at 182.1%, is comparatively smaller. The L40S still more than doubles the MI60, but the relative difference narrows. This could indicate that the MI60's GCN architecture handles graphics-style workloads slightly better relative to its compute performance, though it remains thoroughly outclassed.
For the MI60, the closest competitor in its own rival group is the NVIDIA RTX A4500, which trails by only 0.9%, and the RTX A4500 Mobile at 1.5% behind. The MI60 is not without merit in its own tier, but that tier is far below where the L40S operates.
Architecture Differences
The NVIDIA L40S uses the AD102 chip built on Ada Lovelace architecture, produced on a 5 nm process at TSMC. It packs 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3 million per mm². The AMD MI60 uses the Vega 20 chip on GCN 5.1 architecture, also from TSMC but on a larger 7 nm process. It contains 13,230 million transistors on a 331 mm² die with a density of 40.0 million per mm².
The L40S has 18,176 shading units and 568 TMUs. The MI60 has 4,096 shading units and 256 TMUs. The L40S also carries 142 RT cores and 568 tensor cores, while the MI60 records null entries for both RT cores and tensor cores. The MI60's GCN 5.1 design simply has no hardware ray tracing acceleration or tensor core equivalents listed in the database.
Memory subsystem differences are stark. The L40S uses 48 GB GDDR6 memory on a 384 bit bus with 864.0 GB/s bandwidth. The MI60 uses 32 GB HBM2 memory on a 4096 bit bus with 1.02 TB/s bandwidth. Despite the MI60's wider bus and higher bandwidth, the L40S compensates with its faster memory clock of 2250 MHz (18 Gbps effective, 2 Gbps for the MI60) and newer GDDR6 memory type.
The L40S boosts to 2520 MHz against the MI60's 1800 MHz boost. Base clocks are closer at 1110 MHz for NVIDIA and 1200 MHz for AMD, meaning the L40S has a larger relative clock headroom. API support differs as well, with the L40S supporting DirectX 12 Ultimate (12_2) and Vulkan 1.4, while the MI60 supports DirectX 12 (12_1) and Vulkan 1.3.
The MI60 predates the L40S by roughly four years, with a release date of November 2018 against October 2022 for the L40S. The L40S lists a predecessor of Server Ampere and a successor of Server Hopper. The MI60 lists FirePro Data Center as its predecessor and no successor in the database.
FAQ
Q: Which GPU has more memory and does it matter?
A: The NVIDIA L40S has 48 GB GDDR6 memory, while the AMD MI60 has 32 GB HBM2. The L40S also has a higher memory bandwidth of 864.0 GB/s, though the MI60's HBM2 achieves 1.02 TB/s. The L40S's larger capacity matters for models and datasets that exceed 32 GB.
Q: Does the AMD MI60 support ray tracing or tensor operations?
A: The database records null values for both RT cores and tensor cores on the MI60. The NVIDIA L40S has 142 RT cores and 568 tensor cores. If your workload depends on these features, the MI60 cannot accelerate them.
Q: How does the L40S compare to its closest rival, the RTX 6000 Ada Generation?
A: The L40S average benchmark score is 295,763 against 287,237 for the RTX 6000 Ada Generation, a 3% delta in favor of the L40S. It also leads the NVIDIA L40 by 4.1%.
Q: What GPUs are in the MI60's performance tier?
A: The MI60's nearest rivals are the NVIDIA RTX A4500 at 0.9% behind, the RTX A4500 Mobile at 1.5% behind, the AMD Radeon Pro VII at 4.8% ahead, and the AMD Radeon RX 7900M at 5.2% ahead. Its average score is 92,466.
Q: Which GPU has better Vulkan performance?
A: The NVIDIA L40S scores 260,799 in Geekbench Vulkan versus 92,444 for the MI60, a 182.1% advantage. The L40S supports Vulkan 1.4 while the MI60 supports Vulkan 1.3.
Q: Are both cards still in production?
A: No. The database lists both the NVIDIA L40S and the AMD Radeon Instinct MI60 as end-of-life products.
Head-to-Head Benchmarks
The Geekbench OpenCL result is the single largest recorded gap between these two cards. The NVIDIA L40S posts 330,727 points, the AMD MI60 posts 92,488 points, and the delta is 257.6%. This means the L40S delivers roughly 3.6 times the OpenCL score of the MI60. For compute workloads that rely on OpenCL, this is not a close contest.
The Geekbench Vulkan result shows a similar but slightly narrower margin. The L40S scores 260,799 against 92,444 for the MI60, a 182.1% advantage. The L40S still more than doubles the MI60's Vulkan output, but the relative difference is smaller than in OpenCL.
Looking at the L40S's position against its own rivals, the data shows it leads the RTX 6000 Ada Generation by 3% and the L40 by 4.1%. It trails the AMD Instinct MI300X by 7% and the NVIDIA H200 NVL by 11.7%. The L40S is firmly a top-tier accelerator, sitting just below the absolute fastest options in the database.
The MI60's rival group shows a different story. It is only 0.9% ahead of the RTX A4500 and 1.5% ahead of the RTX A4500 Mobile. It trails the Radeon Pro VII by 4.8% and the Radeon RX 7900M by 5.2%. The MI60 is competitive within its own generation, but that generation is far behind the L40S.
Both cards share a 300 W TDP and a suggested PSU of 700 W. Both are dual-slot designs with identical physical dimensions of 267 mm length and 111 mm height. Both use PCIe 4.0 x16 interfaces. The power connectors differ, with the L40S using a single 16-pin connector and the MI60 using one 6-pin plus one 8-pin connector.
The L40S offers 1x HDMI 2.1 and 3x DisplayPort 1.4a outputs, while the MI60 has a single mini-DisplayPort 1.4a output. For a server accelerator, display outputs are rarely the deciding factor, but the L40S offers more flexibility.
Specification Differences
The process node differs, with the L40S on 5 nm and the MI60 on 7 nm, both from TSMC. Transistor count is 76,300 million for the L40S versus 13,230 million for the MI60. Die size is 609 mm² against 331 mm². Transistor density is 125.3M per mm² versus 40.0M per mm².
Clock speeds: the L40S runs at 1110 MHz base and 2520 MHz boost. The MI60 runs at 1200 MHz base and 1800 MHz boost. The L40S has a lower base clock but a substantially higher boost clock.
Memory: the L40S has 48 GB GDDR6 on a 384 bit bus with 864.0 GB/s bandwidth and 2250 MHz memory clock (18 Gbps effective). The MI60 has 32 GB HBM2 on a 4096 bit bus with 1.02 TB/s bandwidth and 1000 MHz memory clock (2 Gbps effective).
Compute units: the L40S has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The MI60 has 4,096 shading units, 256 TMUs, 64 ROPs, and no recorded RT or tensor cores. The L40S delivers 91.61 TFLOPS FP32 and 91.61 TFLOPS FP16 (1:1 ratio). The MI60 delivers 14.75 TFLOPS FP32 and 29.49 TFLOPS FP16 (2:1 ratio).
Pixel and texture rates: the L40S outputs 483.8 GPixel/s and 1,431.4 GTexel/s. The MI60 outputs 115.2 GPixel/s and 460.8 GTexel/s.
API support: the L40S supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The MI60 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.3.
Power and physical: both are 300 W TDP with 700 W suggested PSU and dual-slot designs. The L40S uses a 1x 16-pin power connector, the MI60 uses 1x 6-pin plus 1x 8-pin. Both are 267 mm long and 111 mm tall. The L40S has four display outputs, the MI60 has one. Both are PCIe 4.0 x16. Both are end-of-life products. The L40S released in October 2022, the MI60 in November 2018.