AMD Radeon Pro W6800X vs NVIDIA L40S Comparison

AMD
RADEON

AMD Radeon Pro W6800X

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2087 MHz
TDP 200 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_metal
196,844
N/A
geekbench_opencl
124,498
330,727
geekbench_vulkan
N/A
260,799

Analysis: AMD Radeon Pro W6800X vs NVIDIA L40S

NVIDIA L40S vs AMD Radeon Pro W6800X is a matchup between two very different workstation-class accelerators. The L40S, built on NVIDIA’s Ada Lovelace architecture, is a server-focused compute card with 48 GB of memory, while the W6800X is an AMD RDNA 2.0 part designed specifically for Apple Mac Pro systems. Benchmark data shows a stark performance gap: the L40S delivers an average benchmark score of 295,763, placing it in the 99th percentile of all GPUs, while the W6800X scores 160,671 on average, sitting in the 97th percentile. The head-to-head results are one-sided, but the W6800X’s unique platform requirements and lower power draw make it a relevant option in its niche.

Head-to-Head Benchmarks

The only direct benchmark comparison available is Geekbench OpenCL, and the result is decisive. The NVIDIA L40S scores 330,727 points, while the AMD Radeon Pro W6800X manages 124,498 points. That is a delta of 165.6% in favor of the L40S, meaning the NVIDIA card delivers more than 2.6 times the OpenCL performance of the AMD card. This is not a close contest; the L40S absolutely dominates in raw compute throughput.

Looking at the broader benchmark context, the L40S’s average score of 295,763 puts it 3% ahead of the NVIDIA RTX 6000 Ada Generation (287,237) and 4.1% ahead of the NVIDIA L40 (284,111). It trails the AMD Instinct MI300X by 7% (317,994) and the NVIDIA H200 NVL by 11.7% (334,891). These numbers show the L40S sits comfortably in the upper echelon of compute accelerators, trading blows with massive data-center parts. The W6800X, by contrast, has an average score of 160,671, which is only 1.1% behind the NVIDIA A100 PCIe 40 GB (162,504) and 2.6% behind the AMD Radeon PRO W7800 (164,894). It is essentially at parity with those cards, but that puts it roughly 84% behind the L40S in average score.

The Geekbench Vulkan result for the L40S (260,799) is not directly compared to the W6800X, but it reinforces that the NVIDIA card is strong across different APIs. The W6800X’s best result comes from Geekbench Metal, where it scores 196,844, a test the L40S does not appear in. Still, for OpenCL workloads, the gap is enormous and consistent with the architectural differences between the two cards.

Architecture Differences

The NVIDIA L40S is built on the AD102 chip using the Ada Lovelace architecture, manufactured on a 5 nm process at TSMC. It packs 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3 million per square millimeter. The AMD Radeon Pro W6800X uses the Navi 21 chip with RDNA 2.0 architecture, also fabricated by TSMC but on a larger 7 nm node. It contains 26,800 million transistors on a 520 mm² die, giving a much lower density of 51.5 million per square millimeter.

The compute resources tell the story of their respective design goals. The L40S has 18,176 shading units, 568 texture mapping units (TMUs), and 192 raster operation units (ROPs). It also includes 142 ray tracing cores and 568 tensor cores, the latter being essential for AI and deep learning workloads. The W6800X has 3,840 shading units, 240 TMUs, and 96 ROPs, with 60 ray tracing cores and no tensor cores at all. The lack of tensor cores is a critical differentiator—the L40S is equipped for matrix math and neural network inference, while the W6800X is purely a graphics and general compute part.

Clock speeds are also notably different. The L40S runs at a base clock of 1110 MHz and boosts to 2520 MHz, while the W6800X has a higher base of 1800 MHz but a lower boost of 2087 MHz. The L40S compensates with far more parallel hardware, achieving 91.61 TFLOPS of FP32 performance and 91.61 TFLOPS of FP16 (1:1 ratio). The W6800X delivers 16.03 TFLOPS FP32 and 32.06 TFLOPS FP16 (2:1 ratio). In raw FP32, the L40S is roughly 5.7 times faster, and even in FP16, the L40S’s 1:1 rate gives it nearly three times the throughput.

Memory architecture is another major split. The L40S uses 48 GB of GDDR6 on a 384-bit bus, achieving 864.0 GB/s of bandwidth. The W6800X has 32 GB of GDDR6 on a 256-bit bus, with 512.0 GB/s bandwidth. The L40S also runs faster memory at 18 Gbps effective versus 16 Gbps on the AMD card. Pixel and texture rates follow suit: the L40S hits 483.8 GPixel/s and 1,431.4 GTexel/s, while the W6800X manages 200.4 GPixel/s and 500.9 GTexel/s.

Where Each One Wins

The NVIDIA L40S wins in every measurable compute category. For OpenCL, it is 165.6% ahead of the W6800X. For FP32 and FP16 compute, it is multiple times faster. It has more memory, more bandwidth, higher pixel and texture rates, and a significantly larger transistor budget. The L40S is clearly the choice for any workload that demands raw throughput: GPU rendering, scientific simulation, AI training, or large-scale data processing. Its 48 GB frame buffer is particularly valuable for models and datasets that exceed 32 GB, which would spill over or require partitioning on the W6800X.

The AMD Radeon Pro W6800X wins in a different sense: platform fit and power efficiency. It is designed for the Apple MPX interface, meaning it slots into Mac Pro systems without PCIe adapters or external power cabling. Its 200 W TDP is substantially lower than the L40S’s 300 W, and it requires a 550 W suggested PSU versus 700 W for the NVIDIA card. For users locked into the Apple ecosystem, the W6800X is the only one of the two that is even compatible. It also supports Metal API natively, where it scores 196,844, a result that is not directly comparable to the L40S but shows respectable performance for macOS applications. The W6800X’s quad-slot design and Apple MPX power connector are unusual, but they are exactly what a Mac Pro expects.

Specification Differences

The two cards differ in nearly every specification that matters. The L40S uses a 5 nm process; the W6800X uses 7 nm. The L40S has 76,300 million transistors versus 26,800 million, and a die size of 609 mm² versus 520 mm². Transistor density is 125.3M per mm² on the L40S versus 51.5M per mm² on the W6800X. Base clocks are 1110 MHz versus 1800 MHz, with boosts of 2520 MHz versus 2087 MHz. Memory capacity is 48 GB versus 32 GB, with bus widths of 384-bit versus 256-bit and bandwidths of 864.0 GB/s versus 512.0 GB/s. Shading units are 18,176 versus 3,840, TMUs are 568 versus 240, and ROPs are 192 versus 96. Ray tracing cores are 142 versus 60, and tensor cores are 568 versus none. FP32 is 91.61 TFLOPS versus 16.03 TFLOPS; FP16 is 91.61 TFLOPS (1:1) versus 32.06 TFLOPS (2:1). TDP is 300 W versus 200 W. Slot width is dual-slot versus quad-slot, power connectors are 1x 16-pin versus Apple MPX, and bus interface is PCIe 4.0 x16 versus Apple MPX. Display outputs are 1x HDMI 2.1 and 3x DisplayPort 1.4a versus 1x HDMI 2.1 and 4x Thunderbolt. Dimensions are identical in length (267 mm) but differ in height: 111 mm for the L40S, 120 mm for the W6800X. The L40S released on 2022-10-12, while the W6800X came earlier on 2021-08-02. Both are end-of-life products. The W6800X has a launch MSRP of 2,799 USD; the L40S has no listed launch MSRP.

FAQ

Q: Which GPU is faster in OpenCL?

A: The NVIDIA L40S is significantly faster, scoring 330,727 in Geekbench OpenCL compared to 124,498 for the AMD Radeon Pro W6800X, a 165.6% advantage.

Q: Can the AMD Radeon Pro W6800X be used in a standard PC?

A: No. The W6800X uses an Apple MPX bus interface and power connector, and it is a quad-slot card, so it is designed exclusively for Apple Mac Pro systems.

Q: Does the W6800X support AI workloads like the L40S?

A: The W6800X has no tensor cores, so it is not optimized for matrix operations that AI models require. The L40S has 568 tensor cores, giving it a clear edge in that area.

Q: How much memory do these cards have?

A: The L40S has 48 GB of GDDR6, while the W6800X has 32 GB of GDDR6. The L40S also has higher bandwidth at 864.0 GB/s versus 512.0 GB/s.

Q: Which card has lower power requirements?

A: The W6800X has a 200 W TDP and a suggested PSU of 550 W, while the L40S draws 300 W and requires a 700 W PSU.

Q: What is the average benchmark score for each?

A: The L40S has an average score of 295,763, placing it in the 99th percentile of all GPUs. The W6800X averages 160,671, in the 97th percentile.

The Verdict

The data is unambiguous: the NVIDIA L40S is the superior performer in every benchmark where both are tested. It is 165.6% faster in OpenCL, has 5.7 times the FP32 throughput, 2.7 times the memory bandwidth, and 50% more memory capacity. For anyone building a server or workstation for compute-heavy tasks—rendering, simulation, machine learning—the L40S is the obvious choice, provided the platform can accommodate its PCIe 4.0 x16 interface and dual-slot footprint. Its 99th percentile standing and proximity to cards like the H200 NVL (within 11.7%) confirm it is a top-tier accelerator.

The AMD Radeon Pro W6800X is not a competitive alternative in raw performance. Its 97th percentile score places it alongside the A100 PCIe 40 GB and RTX A5500, which are respectable but not in the L40S’s class. The W6800X’s only advantages are its lower 200 W power draw and its compatibility with Apple MPX systems. If you own a Mac Pro and need a GPU that slots in natively, the W6800X is the only one of these two that will work. Its Metal score of 196,844 shows it is capable in macOS-optimized applications. But if you have a choice of platform, the L40S is overwhelmingly the better investment of the two, even without a listed launch MSRP. The W6800X’s 2,799 USD launch price does not change the fact that it trails by over 80% in average score. Choose the L40S for performance, the W6800X only for Apple-specific builds.

DETAILED SPECIFICATIONS

SPECIFICATION
Pro W6800X
L40S
Core Specs
Shading Units
3,840
18,176 +373.3%
Shaders
3,840
18,176 +373.3%
TMUs
240
568 +136.7%
ROPs
96
192 +100.0%
Compute Units
60
SM Count
142
Clocks
Base Clock
1800 MHz
1110 MHz
Boost Clock
2087 MHz
2520 MHz
Memory Clock
2000 MHz 16 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
32 GB
48 GB
VRAM (MB)
32,768
49,152 +50.0%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
384 bit
Bandwidth
512.0 GB/s
864.0 GB/s
Cache
L1 Cache
128 KB per Array
128 KB (per SM)
L2 Cache
4 MB
48 MB
L3 Cache
128 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
200.4 GPixel/s
483.8 GPixel/s
Texture Rate
500.9 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
16.03 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
1,001.8 GFLOPS (1:16)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
32.06 TFLOPS (2:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
60
142 +136.7%
Tensor Cores
568
Power
TDP
200 W
300 W
TDP (W)
200
300 +50.0%
Suggested PSU
550 W
700 W
Power Connectors
Apple MPX
1x 16-pin
Architecture
Architecture
RDNA 2.0
Ada Lovelace
GPU Name
Navi 21
AD102
Generation
Radeon Pro Mac (Navi II Series)
Server Ada (Lxx)
Process Size
7 nm
5 nm
Transistors
26,800 million
76,300 million
Die Size
520 mm²
609 mm²
Foundry
TSMC
TSMC
Density
51.5M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.1
3.0
CUDA
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Quad-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
120 mm 4.7 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.14x Thunderbolt
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
Apple MPX
PCIe 4.0 x16
Other
Launch Price
2,799 USD
Production
End-of-life
End-of-life
Predecessor
Server Ampere
Successor
Server Hopper
View Radeon Pro W6800X Details View L40S Details