GPU Comparison

AMD
RADEON

AMD Radeon PRO W6800

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2322 MHz
TDP 250 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_metal
174,420
N/A
geekbench_opencl
121,808
330,727
geekbench_vulkan
109,961
260,799

Analysis: AMD Radeon PRO W6800 vs NVIDIA L40S

The benchmark data is unambiguous: the NVIDIA L40S is in a different performance class than the AMD Radeon PRO W6800, winning both head-to-head tests by massive margins. The L40S delivers over 2.7x the average benchmark score of the W6800, placing it in the 99th percentile of all GPUs compared to the W6800's 96th. This is not a close contest, but the specific strengths and architectural philosophies behind each card tell a detailed story.

Head-to-Head Benchmarks

The NVIDIA L40S dominates the shared test suite. In Geekbench OpenCL, the L40S scores 330,727 against the W6800's 121,808, a decisive 171.5% advantage. This delta is the single largest performance gap in the comparison, reflecting the L40S's sheer compute throughput. The Vulkan result tells a similar story: the L40S posts 260,799 versus 109,961, a 137.2% lead. Both results are consistent, showing the L40S holds its advantage across different API workloads.

The AMD Radeon PRO W6800 cannot claim a single benchmark win in this head-to-head. Its strongest result is the OpenCL score, which still trails by more than a factor of two. However, the W6800's performance profile is not without merit when viewed in context. Its average benchmark score of 135,396 places it within 0.1% of the NVIDIA A10M and RTX 4000 Ada Generation, and within 0.8% of the AMD Radeon PRO V620. This indicates the W6800 is a solid mid-pack performer among its immediate peers, even if it is completely outclassed by the L40S.

The gap between the two cards is so large that it is worth contextualizing with the L40S's own rival set. The L40S's average score of 295,763 is 3% ahead of the NVIDIA RTX 6000 Ada Generation and 4.1% ahead of the NVIDIA L40. It trails the AMD Instinct MI300X by 7% and the NVIDIA H200 NVL by 11.7%. These figures show that the L40S is not merely a top-tier card; it sits in a performance tier where the W6800 is not a participant.

Architecture Differences

The architectural divide between the two GPUs is fundamental, starting with the manufacturing process. The NVIDIA L40S is built on a 5 nm process at TSMC, while the AMD Radeon PRO W6800 uses a 7 nm process, also at TSMC. This node advantage allows the L40S to pack 76,300 million transistors onto a 609 mm² die, achieving a transistor density of 125.3M / mm². In contrast, the W6800 contains 26,800 million transistors on a 520 mm² die, with a density of 51.5M / mm². The L40S has nearly three times the transistor count and more than twice the density.

The compute architectures are equally distinct. The L40S uses Ada Lovelace architecture, featuring 18,176 shading units, 568 TMUs, and 192 ROPs. It also includes 142 RT cores and 568 tensor cores, making it a fully featured accelerator for both ray tracing and AI workloads. The W6800 uses RDNA 2.0 architecture, with 3,840 shading units, 240 TMUs, and 96 ROPs. It has 60 RT cores but no tensor core equivalent, meaning it lacks dedicated AI acceleration hardware. This architectural gap explains the massive FP32 throughput difference: the L40S delivers 91.61 TFLOPS versus the W6800's 17.83 TFLOPS.

Memory configurations further separate the two. The L40S ships with 48 GB of GDDR6 on a 384-bit bus, yielding 864.0 GB/s of bandwidth. The W6800 has 32 GB of GDDR6 on a 256-bit bus, providing 512.0 GB/s. The L40S also runs its memory at a higher effective speed of 18 Gbps compared to the W6800's 16 Gbps. Both cards use PCIe 4.0 x16, but the L40S's larger memory pool and bandwidth make it better suited for large dataset workloads. The L40S also supports a broader API feature set with Vulkan 1.4 and DirectX 12 Ultimate, matching the W6800 on API versions.

Where Each One Wins

NVIDIA L40S wins in every scenario where raw compute throughput, memory bandwidth, or AI acceleration is paramount. Its 91.61 TFLOPS of FP32 performance is over five times the W6800's output, making it the clear choice for heavy simulation, scientific computing, or any workload that scales with shading units. The 48 GB memory pool and 864.0 GB/s bandwidth allow it to handle massive datasets that would overflow the W6800's 32 GB frame buffer. The presence of 568 tensor cores gives the L40S a decisive edge in machine learning training and inference, a workload category the W6800 cannot effectively address. For users needing Vulkan performance, the L40S's 260,799 score versus 109,961 is a 137.2% improvement.

AMD Radeon PRO W6800 wins in scenarios where its specific feature set is a better fit, primarily due to its display output configuration. The W6800 offers 6x mini-DisplayPort 1.4a outputs, while the L40S provides only 1x HDMI 2.1 and 3x DisplayPort 1.4a. For multi-display professional environments, such as financial trading floors or video wall setups, the W6800's six outputs are a practical advantage. The W6800 also has a lower power draw at 250 W versus the L40S's 300 W, and its smaller physical footprint with a 50 mm width versus the L40S's unspecified width suggests easier integration into space-constrained chassis. Its FP16 performance of 35.67 TFLOPS (2:1) is double its FP32 rate, offering a relative efficiency advantage in mixed-precision workloads, though still far below the L40S's absolute numbers.

FAQ

Q: Which card has the higher average benchmark score?

A: The NVIDIA L40S has an average benchmark score of 295,763, compared to the AMD Radeon PRO W6800's 135,396, a difference of roughly 2.18x.

Q: How does the L40S compare to its nearest rivals?

A: The L40S is 3% ahead of the NVIDIA RTX 6000 Ada Generation and 4.1% ahead of the NVIDIA L40. It trails the AMD Instinct MI300X by 7% and the NVIDIA H200 NVL by 11.7%.

Q: What is the memory bandwidth difference?

A: The NVIDIA L40S provides 864.0 GB/s of bandwidth over a 384-bit bus, while the AMD Radeon PRO W6800 provides 512.0 GB/s over a 256-bit bus.

Q: Does the AMD card have any unique display capabilities?

A: Yes, the AMD Radeon PRO W6800 features 6x mini-DisplayPort 1.4a outputs, whereas the NVIDIA L40S has 1x HDMI 2.1 and 3x DisplayPort 1.4a, giving the AMD card more simultaneous display connections.

Q: What is the transistor density difference?

A: The NVIDIA L40S has a transistor density of 125.3M / mm² on a 5 nm process, while the AMD Radeon PRO W6800 has 51.5M / mm² on a 7 nm process.

Q: Which card has a higher pixel rate?

A: The NVIDIA L40S achieves 483.8 GPixel/s, more than double the AMD Radeon PRO W6800's 222.9 GPixel/s.

The Verdict

The data supports a clear conclusion: the NVIDIA L40S is the superior performer for any compute-intensive task. Its 171.5% OpenCL lead and 137.2% Vulkan lead over the W6800 are not incremental improvements; they are generational leaps. The L40S's 91.61 TFLOPS FP32 output, 48 GB memory, and tensor core support make it the only choice for AI research, large-scale rendering, or high-performance computing. The W6800's 17.83 TFLOPS and 32 GB memory are sufficient for many professional workloads, but they are bottlenecked by comparison.

The AMD Radeon PRO W6800 is not without its niche. Its 6x mini-DisplayPort outputs make it a superior option for multi-monitor professional setups, and its 250 W power draw is more modest than the L40S's 300 W. For users whose primary need is display connectivity rather than raw compute, the W6800 offers a practical advantage. However, any user prioritizing compute, AI, or memory capacity should select the NVIDIA L40S without hesitation. The benchmark data shows no scenario where the W6800 outperforms the L40S in raw performance, and the L40S's percentile ranking of 99 versus the W6800's 96 confirms its higher standing among all GPUs.

Specification Differences

| Specification | NVIDIA L40S | AMD Radeon PRO W6800 |

|---|---|---|

| Architecture | Ada Lovelace | RDNA 2.0 |

| Process Node | 5 nm | 7 nm |

| Transistors | 76,300 million | 26,800 million |

| Die Size | 609 mm² | 520 mm² |

| Transistor Density | 125.3M / mm² | 51.5M / mm² |

| Base Clock | 1110 MHz | 1575 MHz |

| Boost Clock | 2520 MHz | 2322 MHz |

| Memory Size | 48 GB | 32 GB |

| Memory Bus Width | 384 bit | 256 bit |

| Memory Bandwidth | 864.0 GB/s | 512.0 GB/s |

| Memory Effective Speed | 18 Gbps | 16 Gbps |

| Shading Units | 18,176 | 3,840 |

| TMUs | 568 | 240 |

| ROPs | 192 | 96 |

| RT Cores | 142 | 60 |

| Tensor Cores | 568 | N/A |

| Pixel Rate | 483.8 GPixel/s | 222.9 GPixel/s |

| Texture Rate | 1,431.4 GTexel/s | 557.3 GTexel/s |

| FP32 Performance | 91.61 TFLOPS | 17.83 TFLOPS |

| FP16 Performance | 91.61 TFLOPS (1:1) | 35.67 TFLOPS (2:1) |

| TDP | 300 W | 250 W |

| Power Connectors | 1x 16-pin | 1x 6-pin + 1x 8-pin |

| Suggested PSU | 700 W | 600 W |

| Display Outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a | 6x mini-DisplayPort 1.4a |

| Dimensions (Height) | 111 mm | 120 mm |

| Dimensions (Width) | N/A | 50 mm |

| Release Date | 2022-10-12 | 2021-06-07 |

| Launch MSRP | N/A | 2,249 USD |

DETAILED SPECIFICATIONS

SPECIFICATION
PRO W6800
L40S
Core Specs
Shading Units
3,840
18,176 +373.3%
Shaders
3,840
18,176 +373.3%
TMUs
240
568 +136.7%
ROPs
96
192 +100.0%
Compute Units
60
SM Count
142
Clocks
Base Clock
1575 MHz
1110 MHz
Boost Clock
2322 MHz
2520 MHz
Memory Clock
2000 MHz 16 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
32 GB
48 GB
VRAM (MB)
32,768
49,152 +50.0%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
384 bit
Bandwidth
512.0 GB/s
864.0 GB/s
Cache
L1 Cache
128 KB per Array
128 KB (per SM)
L2 Cache
4 MB
48 MB
L3 Cache
128 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
222.9 GPixel/s
483.8 GPixel/s
Texture Rate
557.3 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
17.83 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
1,114.6 GFLOPS (1:16)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
35.67 TFLOPS (2:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
60
142 +136.7%
Tensor Cores
568
Power
TDP
250 W
300 W
TDP (W)
250
300 +20.0%
Suggested PSU
600 W
700 W
Power Connectors
1x 6-pin + 1x 8-pin
1x 16-pin
Architecture
Architecture
RDNA 2.0
Ada Lovelace
GPU Name
Navi 21
AD102
Generation
Radeon Pro Navi (Navi II Series)
Server Ada (Lxx)
Process Size
7 nm
5 nm
Transistors
26,800 million
76,300 million
Die Size
520 mm²
609 mm²
Foundry
TSMC
TSMC
Density
51.5M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.1
3.0
CUDA
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
120 mm 4.7 inches
111 mm 4.4 inches
Outputs
6x mini-DisplayPort 1.4a
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
2,249 USD
Production
End-of-life
End-of-life
Predecessor
Radeon Pro Vega
Server Ampere
Successor
Server Hopper
View Radeon PRO W6800 Details View L40S Details