AMD Radeon Pro W6800X vs NVIDIA L40S Comparison
AMD Radeon Pro W6800X
L40S
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Pro W6800X vs NVIDIA L40S
NVIDIA L40S vs AMD Radeon Pro W6800X is a matchup between two very different workstation-class accelerators. The L40S, built on NVIDIA’s Ada Lovelace architecture, is a server-focused compute card with 48 GB of memory, while the W6800X is an AMD RDNA 2.0 part designed specifically for Apple Mac Pro systems. Benchmark data shows a stark performance gap: the L40S delivers an average benchmark score of 295,763, placing it in the 99th percentile of all GPUs, while the W6800X scores 160,671 on average, sitting in the 97th percentile. The head-to-head results are one-sided, but the W6800X’s unique platform requirements and lower power draw make it a relevant option in its niche.
Head-to-Head Benchmarks
The only direct benchmark comparison available is Geekbench OpenCL, and the result is decisive. The NVIDIA L40S scores 330,727 points, while the AMD Radeon Pro W6800X manages 124,498 points. That is a delta of 165.6% in favor of the L40S, meaning the NVIDIA card delivers more than 2.6 times the OpenCL performance of the AMD card. This is not a close contest; the L40S absolutely dominates in raw compute throughput.
Looking at the broader benchmark context, the L40S’s average score of 295,763 puts it 3% ahead of the NVIDIA RTX 6000 Ada Generation (287,237) and 4.1% ahead of the NVIDIA L40 (284,111). It trails the AMD Instinct MI300X by 7% (317,994) and the NVIDIA H200 NVL by 11.7% (334,891). These numbers show the L40S sits comfortably in the upper echelon of compute accelerators, trading blows with massive data-center parts. The W6800X, by contrast, has an average score of 160,671, which is only 1.1% behind the NVIDIA A100 PCIe 40 GB (162,504) and 2.6% behind the AMD Radeon PRO W7800 (164,894). It is essentially at parity with those cards, but that puts it roughly 84% behind the L40S in average score.
The Geekbench Vulkan result for the L40S (260,799) is not directly compared to the W6800X, but it reinforces that the NVIDIA card is strong across different APIs. The W6800X’s best result comes from Geekbench Metal, where it scores 196,844, a test the L40S does not appear in. Still, for OpenCL workloads, the gap is enormous and consistent with the architectural differences between the two cards.
Architecture Differences
The NVIDIA L40S is built on the AD102 chip using the Ada Lovelace architecture, manufactured on a 5 nm process at TSMC. It packs 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3 million per square millimeter. The AMD Radeon Pro W6800X uses the Navi 21 chip with RDNA 2.0 architecture, also fabricated by TSMC but on a larger 7 nm node. It contains 26,800 million transistors on a 520 mm² die, giving a much lower density of 51.5 million per square millimeter.
The compute resources tell the story of their respective design goals. The L40S has 18,176 shading units, 568 texture mapping units (TMUs), and 192 raster operation units (ROPs). It also includes 142 ray tracing cores and 568 tensor cores, the latter being essential for AI and deep learning workloads. The W6800X has 3,840 shading units, 240 TMUs, and 96 ROPs, with 60 ray tracing cores and no tensor cores at all. The lack of tensor cores is a critical differentiator—the L40S is equipped for matrix math and neural network inference, while the W6800X is purely a graphics and general compute part.
Clock speeds are also notably different. The L40S runs at a base clock of 1110 MHz and boosts to 2520 MHz, while the W6800X has a higher base of 1800 MHz but a lower boost of 2087 MHz. The L40S compensates with far more parallel hardware, achieving 91.61 TFLOPS of FP32 performance and 91.61 TFLOPS of FP16 (1:1 ratio). The W6800X delivers 16.03 TFLOPS FP32 and 32.06 TFLOPS FP16 (2:1 ratio). In raw FP32, the L40S is roughly 5.7 times faster, and even in FP16, the L40S’s 1:1 rate gives it nearly three times the throughput.
Memory architecture is another major split. The L40S uses 48 GB of GDDR6 on a 384-bit bus, achieving 864.0 GB/s of bandwidth. The W6800X has 32 GB of GDDR6 on a 256-bit bus, with 512.0 GB/s bandwidth. The L40S also runs faster memory at 18 Gbps effective versus 16 Gbps on the AMD card. Pixel and texture rates follow suit: the L40S hits 483.8 GPixel/s and 1,431.4 GTexel/s, while the W6800X manages 200.4 GPixel/s and 500.9 GTexel/s.
Where Each One Wins
The NVIDIA L40S wins in every measurable compute category. For OpenCL, it is 165.6% ahead of the W6800X. For FP32 and FP16 compute, it is multiple times faster. It has more memory, more bandwidth, higher pixel and texture rates, and a significantly larger transistor budget. The L40S is clearly the choice for any workload that demands raw throughput: GPU rendering, scientific simulation, AI training, or large-scale data processing. Its 48 GB frame buffer is particularly valuable for models and datasets that exceed 32 GB, which would spill over or require partitioning on the W6800X.
The AMD Radeon Pro W6800X wins in a different sense: platform fit and power efficiency. It is designed for the Apple MPX interface, meaning it slots into Mac Pro systems without PCIe adapters or external power cabling. Its 200 W TDP is substantially lower than the L40S’s 300 W, and it requires a 550 W suggested PSU versus 700 W for the NVIDIA card. For users locked into the Apple ecosystem, the W6800X is the only one of the two that is even compatible. It also supports Metal API natively, where it scores 196,844, a result that is not directly comparable to the L40S but shows respectable performance for macOS applications. The W6800X’s quad-slot design and Apple MPX power connector are unusual, but they are exactly what a Mac Pro expects.
Specification Differences
The two cards differ in nearly every specification that matters. The L40S uses a 5 nm process; the W6800X uses 7 nm. The L40S has 76,300 million transistors versus 26,800 million, and a die size of 609 mm² versus 520 mm². Transistor density is 125.3M per mm² on the L40S versus 51.5M per mm² on the W6800X. Base clocks are 1110 MHz versus 1800 MHz, with boosts of 2520 MHz versus 2087 MHz. Memory capacity is 48 GB versus 32 GB, with bus widths of 384-bit versus 256-bit and bandwidths of 864.0 GB/s versus 512.0 GB/s. Shading units are 18,176 versus 3,840, TMUs are 568 versus 240, and ROPs are 192 versus 96. Ray tracing cores are 142 versus 60, and tensor cores are 568 versus none. FP32 is 91.61 TFLOPS versus 16.03 TFLOPS; FP16 is 91.61 TFLOPS (1:1) versus 32.06 TFLOPS (2:1). TDP is 300 W versus 200 W. Slot width is dual-slot versus quad-slot, power connectors are 1x 16-pin versus Apple MPX, and bus interface is PCIe 4.0 x16 versus Apple MPX. Display outputs are 1x HDMI 2.1 and 3x DisplayPort 1.4a versus 1x HDMI 2.1 and 4x Thunderbolt. Dimensions are identical in length (267 mm) but differ in height: 111 mm for the L40S, 120 mm for the W6800X. The L40S released on 2022-10-12, while the W6800X came earlier on 2021-08-02. Both are end-of-life products. The W6800X has a launch MSRP of 2,799 USD; the L40S has no listed launch MSRP.
FAQ
Q: Which GPU is faster in OpenCL?
A: The NVIDIA L40S is significantly faster, scoring 330,727 in Geekbench OpenCL compared to 124,498 for the AMD Radeon Pro W6800X, a 165.6% advantage.
Q: Can the AMD Radeon Pro W6800X be used in a standard PC?
A: No. The W6800X uses an Apple MPX bus interface and power connector, and it is a quad-slot card, so it is designed exclusively for Apple Mac Pro systems.
Q: Does the W6800X support AI workloads like the L40S?
A: The W6800X has no tensor cores, so it is not optimized for matrix operations that AI models require. The L40S has 568 tensor cores, giving it a clear edge in that area.
Q: How much memory do these cards have?
A: The L40S has 48 GB of GDDR6, while the W6800X has 32 GB of GDDR6. The L40S also has higher bandwidth at 864.0 GB/s versus 512.0 GB/s.
Q: Which card has lower power requirements?
A: The W6800X has a 200 W TDP and a suggested PSU of 550 W, while the L40S draws 300 W and requires a 700 W PSU.
Q: What is the average benchmark score for each?
A: The L40S has an average score of 295,763, placing it in the 99th percentile of all GPUs. The W6800X averages 160,671, in the 97th percentile.
The Verdict
The data is unambiguous: the NVIDIA L40S is the superior performer in every benchmark where both are tested. It is 165.6% faster in OpenCL, has 5.7 times the FP32 throughput, 2.7 times the memory bandwidth, and 50% more memory capacity. For anyone building a server or workstation for compute-heavy tasks—rendering, simulation, machine learning—the L40S is the obvious choice, provided the platform can accommodate its PCIe 4.0 x16 interface and dual-slot footprint. Its 99th percentile standing and proximity to cards like the H200 NVL (within 11.7%) confirm it is a top-tier accelerator.
The AMD Radeon Pro W6800X is not a competitive alternative in raw performance. Its 97th percentile score places it alongside the A100 PCIe 40 GB and RTX A5500, which are respectable but not in the L40S’s class. The W6800X’s only advantages are its lower 200 W power draw and its compatibility with Apple MPX systems. If you own a Mac Pro and need a GPU that slots in natively, the W6800X is the only one of these two that will work. Its Metal score of 196,844 shows it is capable in macOS-optimized applications. But if you have a choice of platform, the L40S is overwhelmingly the better investment of the two, even without a listed launch MSRP. The W6800X’s 2,799 USD launch price does not change the fact that it trails by over 80% in average score. Choose the L40S for performance, the W6800X only for Apple-specific builds.