GPU Comparison
AMD Radeon Pro W6900X
L40
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Pro W6900X vs NVIDIA L40
# The Verdict
The NVIDIA L40 is the clear performance leader in this comparison, dominating the AMD Radeon Pro W6900X in every shared benchmark. The L40 achieves an average benchmark score of 284,111, placing it in the 99th percentile of all GPUs, while the W6900X manages 168,574 on average, sitting in the 97th percentile. The L40 wins both head-to-head tests: a 154.5% advantage in Geekbench OpenCL and a 59.4% lead in Geekbench Vulkan. This is not a close contest; the L40 is in a different performance tier entirely.
The AMD Radeon Pro W6900X is the right choice only for a very specific use case: Mac Pro systems using the Apple MPX bus interface. Its 32 GB of GDDR6 memory and 512.0 GB/s bandwidth are substantial, but the L40 counters with 48 GB and 864.0 GB/s. The L40 also leads in raw compute metrics, with 90.52 TFLOPS FP32 versus 22.23 TFLOPS, and 90.52 TFLOPS FP16 versus 44.46 TFLOPS. If your workload is built around Apple's ecosystem and requires Thunderbolt outputs, the W6900X is the only option that fits. Otherwise, the L40 is the superior GPU by every measurable performance metric in the data.
# FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA L40 scores 284,111 on average, which is 115,537 points higher than the AMD Radeon Pro W6900X's 168,574. The L40 sits in the 99th percentile of all GPUs, while the W6900X sits in the 97th percentile.
Q: How do the two compare in Geekbench OpenCL performance?
A: The L40 scores 330,926 in Geekbench OpenCL, while the W6900X scores 130,035. That is a 154.5% advantage for the L40, making it the decisive winner in this test.
Q: Is the AMD card competitive in Vulkan workloads?
A: The L40 leads in Geekbench Vulkan with 237,295 versus the W6900X's 148,865, a 59.4% margin. The AMD card is closer in Vulkan than in OpenCL, but it still trails by a significant amount.
Q: What memory configuration does each card use?
A: The L40 has 48 GB of GDDR6 memory on a 384-bit bus with 864.0 GB/s bandwidth. The W6900X has 32 GB of GDDR6 on a 256-bit bus with 512.0 GB/s bandwidth. The L40 offers 50% more capacity and 68.75% more bandwidth.
Q: Are these cards still in production?
A: Both are marked as end-of-life. The L40 was released on 2022-10-12, and the W6900X was released earlier on 2021-08-02.
Q: Which card supports which display outputs?
A: The L40 provides 4x DisplayPort 1.4a outputs, while the W6900X provides 1x HDMI 2.1 and 4x Thunderbolt. This reflects the AMD card's Mac-centric design.
# Architecture Differences
The NVIDIA L40 is built on the AD102 chip using the Ada Lovelace architecture, fabricated on a 5 nm process at TSMC. It packs 76,300 million transistors into a 609 mm² die, achieving a transistor density of 125.3M per mm². The AMD Radeon Pro W6900X uses the Navi 21 chip with RDNA 2.0 architecture, also from TSMC but on a 7 nm process. It contains 26,800 million transistors on a 520 mm² die, with a density of 51.5M per mm². The L40's newer process node and larger transistor budget give it a substantial architectural advantage.
The L40 features 18,176 shading units, 568 TMUs, and 192 ROPs, along with 142 ray tracing cores and 568 tensor cores. The W6900X has 5,120 shading units, 320 TMUs, and 128 ROPs, with 80 ray tracing cores and no tensor cores listed. The L40 also supports 1:1 FP16 throughput, matching its FP32 rate at 90.52 TFLOPS, while the W6900X uses a 2:1 ratio, delivering 44.46 TFLOPS FP16 against 22.23 TFLOPS FP32.
Both cards support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, so API coverage is identical. The L40's tensor cores are a key differentiator for AI workloads, a feature the W6900X lacks entirely. The L40 also uses a PCIe 4.0 x16 bus interface, while the W6900X uses Apple MPX, which limits its compatibility to Apple systems.
# Specification Differences
The two cards differ across nearly every major specification. The L40 uses a 5 nm process, while the W6900X uses 7 nm. Transistor counts are 76,300 million versus 26,800 million, and die sizes are 609 mm² versus 520 mm². The L40 has a base clock of 735 MHz and a boost clock of 2490 MHz, while the W6900X runs at 1825 MHz base and 2171 MHz boost. Memory speed differs as well: the L40 uses 2250 MHz (18 Gbps effective) versus 2000 MHz (16 Gbps effective) on the W6900X.
Memory capacity is 48 GB versus 32 GB, bus width is 384-bit versus 256-bit, and bandwidth is 864.0 GB/s versus 512.0 GB/s. Shading units are 18,176 versus 5,120, TMUs are 568 versus 320, and ROPs are 192 versus 128. Ray tracing cores number 142 versus 80, and tensor cores exist only on the L40 (568). Pixel rate is 478.1 GPixel/s versus 277.9 GPixel/s, and texture rate is 1,414.3 GTexel/s versus 694.7 GTexel/s.
FP32 performance is 90.52 TFLOPS versus 22.23 TFLOPS, and FP16 is 90.52 TFLOPS versus 44.46 TFLOPS. Both have a 300 W TDP and suggest a 700 W PSU. The L40 is dual-slot with a 1x 16-pin power connector, while the W6900X has no slot width or power connector data. The L40 uses PCIe 4.0 x16, the W6900X uses Apple MPX. Display outputs are 4x DisplayPort 1.4a versus 1x HDMI 2.1 plus 4x Thunderbolt. Length is identical at 267 mm, but height differs: 111 mm for the L40 versus 120 mm for the W6900X.
# Head-to-Head Benchmarks
The data shows a decisive NVIDIA L40 victory in both shared benchmarks. In Geekbench OpenCL, the L40 scores 330,926 against the W6900X's 130,035. That is a 154.5% delta, meaning the L40 is more than two and a half times faster in this workload. This is the largest gap between the two cards in any test.
In Geekbench Vulkan, the L40 scores 237,295 versus 148,865 for the W6900X, a 59.4% advantage. The gap narrows considerably compared to OpenCL, but the L40 still wins by a wide margin. The W6900X has an additional benchmark result in Geekbench Metal, scoring 226,821, but the L40 has no corresponding Metal score, so no direct comparison is possible there.
The average benchmark scores reinforce this picture. The L40 averages 284,111 across its two tests, while the W6900X averages 168,574 across three. The L40's nearest rival is the NVIDIA RTX 6000 Ada Generation at 287,237 (1.1% higher), while the W6900X's nearest rival is the NVIDIA RTX 4500 Ada Generation at 166,094 (1.5% lower). The L40 also leads its other rivals: the L40S at 295,763 (3.9% higher), the L20 at 251,147 (13.1% lower), and the AMD Instinct MI300X at 317,994 (10.7% higher). The W6900X trails the RTX A5500 at 165,217 (2% higher), the AMD Radeon PRO W7800 at 164,894 (2.2% higher), and the NVIDIA A100 PCIe 40 GB at 162,504 (3.7% higher).
# Where Each One Wins
The NVIDIA L40 wins every shared benchmark category. Its 154.5% OpenCL lead suggests a massive advantage in compute-heavy, general-purpose GPU workloads that rely on OpenCL. The 59.4% Vulkan lead indicates strong performance in graphics and compute tasks using that API. With 48 GB of memory and 864.0 GB/s of bandwidth, the L40 is better suited for large datasets and memory-intensive rendering, AI inference, or scientific computing tasks. Its 568 tensor cores provide dedicated hardware for AI and deep learning workloads, which the W6900X cannot match.
The AMD Radeon Pro W6900X wins only in the context of Apple ecosystem compatibility. Its Apple MPX bus interface and Thunderbolt display outputs make it the only viable choice for Mac Pro systems that require those connections. The W6900X also has a Geekbench Metal score of 226,821, which is higher than its OpenCL and Vulkan scores, suggesting it performs best in Apple's Metal API environment. However, the L40 has no Metal benchmark data, so a direct comparison in that API is unavailable.
For users on standard PC platforms, the L40 is the obvious pick. It offers more memory, more bandwidth, higher compute throughput, and better benchmark results across the board. The W6900X is a niche product for Mac Pro owners who need a high-end GPU with Thunderbolt support. If you are building or upgrading a PC workstation, the data leaves no question: the L40 is the stronger card.