AMD Radeon Pro W6900X vs NVIDIA L40S Comparison

AMD
RADEON

AMD Radeon Pro W6900X

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2171 MHz
TDP 300 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_metal
226,821
N/A
geekbench_opencl
130,035
330,727
geekbench_vulkan
148,865
260,799

Analysis: AMD Radeon Pro W6900X vs NVIDIA L40S

Where Each One Wins

The benchmark data splits cleanly between these two workstation-class accelerators, and the pattern is unambiguous. The NVIDIA L40S takes both head-to-head tests it shares with the AMD Radeon Pro W6900X, giving it a 2–0 win record in direct comparison. In Geekbench OpenCL, the L40S scores 330,727 against the W6900X’s 130,035—a 154.3% advantage. In Geekbench Vulkan, the L40S posts 260,799 versus 148,865, a 75.2% lead. There is no benchmark category in the data where the AMD card wins outright.

However, the W6900X holds one exclusive advantage: it has a Geekbench Metal score of 226,821, a test the L40S does not have any recorded result for. This matters because Metal is Apple’s graphics API, and the W6900X is explicitly designed for Mac systems (its generation is listed as "Radeon Pro Mac"). So while the L40S dominates in cross-platform OpenCL and Vulkan workloads, the AMD card occupies a niche where the NVIDIA part cannot even compete on paper.

Looking at broader positioning, the L40S sits in the 99th percentile of all GPUs, while the W6900X sits in the 97th percentile. That two-point percentile gap reflects a substantial average benchmark gap: the L40S averages 295,763 across its tests, versus 168,574 for the W6900X. The L40S outperforms its closest rival (NVIDIA RTX 6000 Ada Generation) by 3%, while the W6900X beats its closest rival (NVIDIA RTX 4500 Ada Generation) by a slimmer 1.5%. In short, the L40S is not just ahead of the W6900X—it occupies a higher performance tier entirely.

Architecture Differences

The architectural divide between these two is generational and foundational. The L40S uses NVIDIA’s AD102 chip built on Ada Lovelace architecture, fabricated on a 5 nm process at TSMC. It packs 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3 million per mm². The W6900X uses AMD’s Navi 21 chip based on RDNA 2.0, also made by TSMC but on a larger 7 nm node. It contains 26,800 million transistors across a 520 mm² die, with a density of just 51.5 million per mm². The L40S has roughly 2.85 times the transistor count on a die that is only 17% larger, a direct consequence of the denser manufacturing process.

Core counts diverge even more sharply. The L40S has 18,176 shading units, 568 texture mapping units, and 192 ROPs. It also includes 142 ray tracing cores and 568 tensor cores. The W6900X has 5,120 shading units, 320 TMUs, and 128 ROPs, plus 80 ray tracing cores—but no tensor core equivalent is listed. The absence of tensor cores is critical: it means the AMD card lacks dedicated hardware for the matrix math that underpins modern AI and deep learning workloads, while the L40S is explicitly equipped for them.

Memory configuration reinforces the gap. The L40S has 48 GB of GDDR6 on a 384-bit bus, delivering 864.0 GB/s of bandwidth. The W6900X has 32 GB of GDDR6 on a 256-bit bus, delivering 512.0 GB/s. That is 50% more memory and 68.75% more bandwidth for the NVIDIA card. Clock speeds tell a different story: the W6900X runs at a higher base clock (1825 MHz vs 1110 MHz) and a higher boost clock (2171 MHz vs 2520 MHz is actually lower—wait, the L40S boost is 2520 MHz, so the L40S boosts higher). Correcting that: the L40S boosts to 2520 MHz, while the W6900X boosts to 2171 MHz. The AMD card’s base clock is higher, but the NVIDIA part boosts further. Memory clocks favor the L40S at 2250 MHz (18 Gbps effective) versus 2000 MHz (16 Gbps effective) for the W6900X.

Power and interface also differ. Both cards are rated at 300 W TDP and recommend a 700 W PSU, but the L40S uses a single 16-pin power connector and a standard PCIe 4.0 x16 bus. The W6900X uses Apple MPX as its bus interface—a proprietary Apple connector—and lists no standard power connector or slot width. This is a hardware analyst’s red flag: the W6900X is not a drop-in card for standard PC systems; it is tied to Apple’s Mac Pro ecosystem.

Head-to-Head Benchmarks

The Geekbench OpenCL result is the single biggest gap in the entire comparison. The L40S scores 330,727, which is more than double the W6900X’s 130,035. The 154.3% delta means the NVIDIA card delivers roughly 2.5 times the raw compute throughput in this test. To put that in context, the L40S’s nearest rival, the AMD Instinct MI300X, scores 317,994—still 7% lower than the L40S. The W6900X, by contrast, is within 3.7% of the NVIDIA A100 PCIe 40 GB (162,504) but far below the L40S.

The Vulkan result is closer but still lopsided. The L40S scores 260,799 against 148,865 for the W6900X, a 75.2% advantage. This narrower margin suggests the W6900X is comparatively better at graphics-oriented workloads than at pure compute, but it still loses decisively. The L40S’s Vulkan score is 12.2% higher than its own OpenCL score relative to the W6900X, but that is cold comfort for AMD.

The W6900X’s Metal score of 226,821 is its only bright spot. Since the L40S has no Metal result, we cannot directly compare them, but the fact that the W6900X scores 74.5% higher in Metal than in OpenCL suggests the RDNA 2 architecture is well-tuned for Apple’s API. Still, in every test where both cards appear, the L40S wins by a wide margin. The data does not support any interpretation where the W6900X is competitive in cross-platform workloads.

FAQ

Q: Which card has more memory bandwidth?

A: The NVIDIA L40S has 864.0 GB/s of memory bandwidth, while the AMD Radeon Pro W6900X has 512.0 GB/s. The L40S also has more memory (48 GB vs 32 GB).

Q: Does the AMD card have tensor cores?

A: No. The Radeon Pro W6900X lists no tensor cores in its specifications. The NVIDIA L40S includes 568 tensor cores, which are dedicated hardware for AI and machine learning workloads.

Q: What is the performance gap in OpenCL?

A: The L40S scores 330,727 in Geekbench OpenCL, which is 154.3% higher than the W6900X’s 130,035. This is the largest margin in any shared benchmark.

Q: Why does the W6900X have a Metal benchmark but the L40S does not?

A: The W6900X is part of the Radeon Pro Mac series and uses the Apple MPX bus interface, so it is designed for Apple systems where Metal is the primary graphics API. The L40S is a server-oriented card with no recorded Metal benchmark, likely because it targets OpenCL and Vulkan environments.

Q: Which card has a higher boost clock?

A: The NVIDIA L40S boosts to 2520 MHz, while the AMD Radeon Pro W6900X boosts to 2171 MHz. However, the AMD card has a higher base clock at 1825 MHz versus 1110 MHz for the L40S.

Q: How do these cards compare to their nearest rivals?

A: The L40S is 3% ahead of the NVIDIA RTX 6000 Ada Generation and 4.1% ahead of the NVIDIA L40, but 7% behind the AMD Instinct MI300X and 11.7% behind the NVIDIA H200 NVL. The W6900X is 1.5% ahead of the NVIDIA RTX 4500 Ada Generation, 2% ahead of the RTX A5500, 2.2% ahead of the AMD Radeon PRO W7800, and 3.7% ahead of the A100 PCIe 40 GB.

The Verdict

The data points to a single conclusion: the NVIDIA L40S is the superior compute platform. It wins both shared benchmarks by margins of 154.3% and 75.2%, has 50% more memory, 68.75% more bandwidth, roughly 3.5 times the shading units, and a newer 5 nm process node. Its 99th percentile standing versus the W6900X’s 97th confirms the tier gap. For any workload that uses OpenCL or Vulkan—which covers most professional rendering, simulation, and compute tasks—the L40S is the clear choice.

The AMD Radeon Pro W6900X is not without merit, but its appeal is narrow. Its Metal score of 226,821 gives it a distinct advantage in Apple-specific environments, and its 300 W TDP matches the L40S while offering a higher base clock. But its Apple MPX bus interface locks it to Mac Pro systems, and its 32 GB memory and 512.0 GB/s bandwidth are far below the L40S. If you are building or upgrading a Mac workstation and need a powerful GPU, the W6900X is a legitimate option—its launch MSRP was 5,999 USD. For anyone running a standard PCIe-based server or workstation, the L40S is the only rational pick from these two.

The verdict is straightforward: choose the L40S for compute density, memory capacity, and cross-platform performance. Choose the W6900X only if you are constrained to Apple’s MPX ecosystem and prioritize Metal compatibility over raw compute.

Specification Differences

| Specification | NVIDIA L40S | AMD Radeon Pro W6900X |

|---|---|---|

| Architecture | Ada Lovelace | RDNA 2.0 |

| Process Node | 5 nm | 7 nm |

| Transistors | 76,300 million | 26,800 million |

| Die Size | 609 mm² | 520 mm² |

| Transistor Density | 125.3M / mm² | 51.5M / mm² |

| Base Clock | 1110 MHz | 1825 MHz |

| Boost Clock | 2520 MHz | 2171 MHz |

| Memory Clock | 2250 MHz (18 Gbps effective) | 2000 MHz (16 Gbps effective) |

| Memory Size | 48 GB | 32 GB |

| Memory Bus Width | 384 bit | 256 bit |

| Memory Bandwidth | 864.0 GB/s | 512.0 GB/s |

| Shading Units | 18,176 | 5,120 |

| TMUs | 568 | 320 |

| ROPs | 192 | 128 |

| Ray Tracing Cores | 142 | 80 |

| Tensor Cores | 568 | — |

| Pixel Rate | 483.8 GPixel/s | 277.9 GPixel/s |

| Texture Rate | 1,431.4 GTexel/s | 694.7 GTexel/s |

| FP32 Performance | 91.61 TFLOPS | 22.23 TFLOPS |

| FP16 Performance | 91.61 TFLOPS (1:1) | 44.46 TFLOPS (2:1) |

| TDP | 300 W | 300 W |

| Power Connectors | 1x 16-pin | — |

| Bus Interface | PCIe 4.0 x16 | Apple MPX |

| Display Outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a | 1x HDMI 2.1, 4x Thunderbolt |

| Release Date | 2022-10-12 | 2021-08-02 |

DETAILED SPECIFICATIONS

SPECIFICATION
Pro W6900X
L40S
Core Specs
Shading Units
5,120
18,176 +255.0%
Shaders
5,120
18,176 +255.0%
TMUs
320
568 +77.5%
ROPs
128
192 +50.0%
Compute Units
80
SM Count
142
Clocks
Base Clock
1825 MHz
1110 MHz
Boost Clock
2171 MHz
2520 MHz
Memory Clock
2000 MHz 16 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
32 GB
48 GB
VRAM (MB)
32,768
49,152 +50.0%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
384 bit
Bandwidth
512.0 GB/s
864.0 GB/s
Cache
L1 Cache
128 KB per Array
128 KB (per SM)
L2 Cache
4 MB
48 MB
L3 Cache
128 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
277.9 GPixel/s
483.8 GPixel/s
Texture Rate
694.7 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
22.23 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
1,389.4 GFLOPS (1:16)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
44.46 TFLOPS (2:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
80
142 +77.5%
Tensor Cores
568
Power
TDP
300 W
300 W
TDP (W)
300
300 0.0%
Suggested PSU
700 W
700 W
Power Connectors
1x 16-pin
Architecture
Architecture
RDNA 2.0
Ada Lovelace
GPU Name
Navi 21
AD102
Generation
Radeon Pro Mac (Navi II Series)
Server Ada (Lxx)
Process Size
7 nm
5 nm
Transistors
26,800 million
76,300 million
Die Size
520 mm²
609 mm²
Foundry
TSMC
TSMC
Density
51.5M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.1
3.0
CUDA
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
120 mm 4.7 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.14x Thunderbolt
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
Apple MPX
PCIe 4.0 x16
Other
Launch Price
5,999 USD
Production
End-of-life
End-of-life
Predecessor
Server Ampere
Successor
Server Hopper
View Radeon Pro W6900X Details View L40S Details