AMD Instinct MI100 vs NVIDIA L40S Comparison
AMD Instinct MI100
L40S
PERFORMANCE BENCHMARKS
Analysis: AMD Instinct MI100 vs NVIDIA L40S
# NVIDIA L40S vs AMD Instinct MI100
The NVIDIA L40S and AMD Instinct MI100 are both end-of-life server accelerators aimed at compute workloads, but they represent vastly different generations and design philosophies. The L40S, built on NVIDIA's Ada Lovelace architecture with a 5 nm TSMC process, delivers an average benchmark score of 295,763, placing it in the 99th percentile of all GPUs. The MI100, based on AMD's CDNA 1.0 architecture on a 7 nm TSMC process, averages 139,035, sitting in the 96th percentile. The data shows a decisive performance gap: in the sole head-to-head Geekbench OpenCL test, the L40S scores 330,727 against the MI100's 139,035, a 137.9% advantage. This page breaks down where each card stands across architecture, specifications, and benchmark results.
FAQ
Q: Which GPU has the higher raw compute throughput in FP32?
A: The NVIDIA L40S delivers 91.61 TFLOPS of FP32 performance, while the AMD Instinct MI100 provides 23.07 TFLOPS. The L40S is roughly 4x faster in single-precision compute based on these figures.
Q: How do the memory subsystems compare between the two cards?
A: The L40S offers 48 GB of GDDR6 memory on a 384-bit bus with 864.0 GB/s bandwidth. The MI100 features 32 GB of HBM2 on a 4096-bit bus with 1.23 TB/s bandwidth. The MI100 has higher raw bandwidth, but the L40S has 50% more capacity.
Q: What is the performance percentile ranking for each GPU?
A: The L40S ranks in the 99th percentile of all GPUs, whereas the MI100 ranks in the 96th percentile. This reflects the L40S's position among the top 1% of all accelerators tracked.
Q: Which card supports ray tracing and tensor operations?
A: The NVIDIA L40S includes 142 ray tracing cores and 568 tensor cores. The AMD Instinct MI100 has no ray tracing cores and no tensor cores listed, reflecting its compute-focused CDNA architecture.
Q: What are the closest rivals for each card based on average benchmark score?
A: For the L40S, the nearest rivals are the NVIDIA H200 NVL (334,891 average, 11.7% higher), AMD Instinct MI300X (317,994, 7% higher), NVIDIA RTX 6000 Ada Generation (287,237, 3% lower), and NVIDIA L40 (284,111, 4.1% lower). For the MI100, the closest rivals are the NVIDIA Tesla V100 PCIe 16 GB (138,063, 0.7% lower), Tesla V100 SXM2 32 GB (137,731, 0.9% lower), AMD Radeon PRO V620 (136,472, 1.9% lower), and Radeon Pro W6800X Duo (135,774, 2.4% lower).
Q: Do both cards have the same power consumption?
A: Yes, both the L40S and MI100 have a TDP of 300 W and share a suggested PSU rating of 700 W. However, the L40S uses a single 16-pin power connector, while the MI100 uses two 8-pin connectors.
Architecture Differences
The architectural divide between the L40S and MI100 is profound. The L40S is built on NVIDIA's Ada Lovelace architecture, fabricated on TSMC's 5 nm process, incorporating 76,300 million transistors on a 609 mm² die. This yields a transistor density of 125.3 million per mm². The MI100 uses AMD's CDNA 1.0 architecture, made on TSMC's 7 nm process, with 25,600 million transistors on a larger 750 mm² die, giving a density of only 34.1 million per mm². The L40S packs nearly three times more transistors into a smaller physical area.
The compute feature sets diverge sharply. The L40S includes 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The MI100 has 7,680 shading units, 480 TMUs, and 64 ROPs, with no dedicated RT or tensor cores. This means the L40S can accelerate ray-traced workloads and tensor operations natively, while the MI100 relies purely on shader-based compute. The MI100's FP16 throughput is 46.14 TFLOPS at a 2:1 ratio relative to FP32, whereas the L40S achieves 91.61 TFLOPS FP16 at a 1:1 ratio—meaning the L40S sustains full FP16 performance without clock penalties.
Memory architecture also reflects different priorities. The MI100's HBM2 memory with a 4096-bit bus provides 1.23 TB/s bandwidth, which is 42% higher than the L40S's 864.0 GB/s GDDR6. However, the L40S counters with 48 GB capacity versus 32 GB, which is critical for large model residency. The L40S also supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the MI100 lists N/A for all graphics APIs—it is a pure compute accelerator with no display outputs, whereas the L40S offers 1x HDMI 2.1 and 3x DisplayPort 1.4a.
Where Each One Wins
The L40S wins decisively in general compute workloads as measured by Geekbench OpenCL. With a score of 330,727 versus 139,035, it achieves a 137.9% improvement. This advantage stems from its higher FP32 throughput (91.61 vs 23.07 TFLOPS), more shading units (18,176 vs 7,680), and faster clock speeds (boost 2520 MHz vs 1502 MHz). For FP16 workloads, the L40S's 1:1 ratio means it can double the MI100's rated FP16 performance without any architectural workaround.
The MI100 does have specific strengths. Its 1.23 TB/s memory bandwidth exceeds the L40S's 864.0 GB/s, which could benefit bandwidth-bound kernels that fit within its 32 GB capacity. The 4096-bit HBM2 bus is a legacy of AMD's compute-first design, and the MI100's 750 mm² die size suggests a focus on memory subsystem density. Additionally, the MI100's lower transistor count (25,600 million) and simpler feature set may imply easier software porting for pure compute codes that don't utilize RT or tensor paths.
The L40S also wins on compatibility and flexibility. It supports modern graphics APIs, includes display outputs, and has tensor cores that accelerate AI inference and training workloads—areas where the MI100 has no dedicated hardware. The L40S's 48 GB memory capacity is 50% larger, enabling larger batch sizes or model parameters. In the percentile rankings, the L40S's 99th percentile versus the MI100's 96th percentile reflects its broader applicability across benchmark suites.
Specification Differences
The two cards differ across nearly every major specification:
| Specification | NVIDIA L40S | AMD Instinct MI100 |
|---|---|---|
| Architecture | Ada Lovelace | CDNA 1.0 |
| Process Node | 5 nm | 7 nm |
| Transistors | 76,300 million | 25,600 million |
| Die Size | 609 mm² | 750 mm² |
| Transistor Density | 125.3M / mm² | 34.1M / mm² |
| Base Clock | 1110 MHz | 1000 MHz |
| Boost Clock | 2520 MHz | 1502 MHz |
| Memory Size | 48 GB | 32 GB |
| Memory Type | GDDR6 | HBM2 |
| Memory Bus | 384 bit | 4096 bit |
| Memory Bandwidth | 864.0 GB/s | 1.23 TB/s |
| Memory Clock | 2250 MHz (18 Gbps effective) | 1200 MHz (2.4 Gbps effective) |
| Shading Units | 18,176 | 7,680 |
| TMUs | 568 | 480 |
| ROPs | 192 | 64 |
| RT Cores | 142 | None |
| Tensor Cores | 568 | None |
| Pixel Rate | 483.8 GPixel/s | 96.13 GPixel/s |
| Texture Rate | 1,431.4 GTexel/s | 721.0 GTexel/s |
| FP32 | 91.61 TFLOPS | 23.07 TFLOPS |
| FP16 | 91.61 TFLOPS (1:1) | 46.14 TFLOPS (2:1) |
| Power Connectors | 1x 16-pin | 2x 8-pin |
| Display Outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a | None |
| DirectX | 12 Ultimate (12_2) | N/A |
| OpenGL | 4.6 | N/A |
| Vulkan | 1.4 | N/A |
The release dates also differ: the L40S launched on October 12, 2022, while the MI100 launched on November 15, 2020—nearly two years earlier. Both share identical physical dimensions (267 mm length, 111 mm height), dual-slot design, 300 W TDP, 700 W suggested PSU, and PCIe 4.0 x16 interface.
Head-to-Head Benchmarks
The only available head-to-head benchmark is Geekbench OpenCL, where the NVIDIA L40S delivers 330,727 points against the AMD Instinct MI100's 139,035 points. The deltaPct of 137.9% means the L40S is more than 2.3 times faster in this compute test. This is a massive margin, driven by the L40S's 4x FP32 throughput advantage and 3x shading unit count.
Contextualizing against their respective rival groups makes the gap clearer. The L40S's average score of 295,763 places it between the AMD Instinct MI300X (317,994, which is 7% faster) and the NVIDIA RTX 6000 Ada Generation (287,237, which is 3% slower). The MI100's average of 139,035 sits just above the NVIDIA Tesla V100 PCIe 16 GB (138,063, only 0.7% slower) and the Tesla V100 SXM2 32 GB (137,731, 0.9% slower). Notably, the MI100 is competitive with the V100 family—GPUs from 2017-2018—while the L40S trades blows with the MI300X and H200, which are current-generation accelerators.
The wins tally is 1-0 in favor of the L40S. This single benchmark outcome aligns with the architectural differences: the L40S's higher transistor count, newer process node, and dedicated compute cores all contribute to superior OpenCL performance. The MI100's higher memory bandwidth (1.23 TB/s) does not translate into a benchmark win here, suggesting that the workload is compute-bound rather than memory-bound. For a hypothetical bandwidth-bound kernel, the MI100's HBM2 advantage could narrow the gap, but no such benchmark data exists in the fact pack. The L40S's 99th percentile ranking versus the MI100's 96th percentile further underscores its overall superiority in the tracked benchmark universe. Given the 137.9% delta, any deployment considering the MI100 for OpenCL-based compute would face a significant performance penalty relative to the L40S.