GPU Comparison
AMD Radeon Pro W6900X
L20
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Pro W6900X vs NVIDIA L20
The Verdict
The benchmark data splits these two workstation cards cleanly by workload and platform. The NVIDIA L20 is the decisive compute winner in the two shared tests, leading by 110.9% in Geekbench OpenCL and by 53.2% in Geekbench Vulkan. Its average benchmark score of 251,147 places it in the 99th percentile of all GPUs, while the AMD Radeon Pro W6900X averages 168,574 and sits in the 97th percentile. If your work is OpenCL or Vulkan-heavy rendering, simulation, or AI inference, the L20 is the clear choice from these numbers alone.
However, the W6900X is not without a niche. It was built for Apple MPX systems, uses RDNA 2.0, and has a Geekbench Metal score of 226,821, a test the L20 does not appear in. For Mac Pro users locked into Apple's ecosystem, the W6900X is the only one of these two that fits that bus interface. The L20 uses PCIe 4.0 x16, which is standard in PC servers but not in Apple MPX slots. So the verdict is practical: pick the L20 for raw compute and PC/server flexibility; pick the W6900X only if you need an Apple MPX card with strong Metal performance and can accept its 32 GB memory ceiling and lower OpenCL/Vulkan scores.
FAQ
Q: Which card has the higher average benchmark score?
A: The NVIDIA L20 averages 251,147 across its benchmarks, which is 49.0% higher than the AMD Radeon Pro W6900X's 168,574 average. The L20 also ranks in the 99th percentile of all GPUs, versus the W6900X's 97th percentile.
Q: How large is the performance gap in the shared OpenCL test?
A: In Geekbench OpenCL, the L20 scores 274,276 versus the W6900X's 130,035. That is a 110.9% delta, the L20 more than doubles the AMD card's result.
Q: Does the W6900X win any head-to-head test?
A: No. In the two common benchmarks (OpenCL and Vulkan), the L20 wins both. The W6900X's only unique benchmark is Geekbench Metal, where it scores 226,821, but there is no L20 Metal result to compare against.
Q: What is the memory capacity difference?
A: The L20 has 48 GB of GDDR6 on a 384-bit bus, delivering 864.0 GB/s bandwidth. The W6900X has 32 GB of GDDR6 on a 256-bit bus, with 512.0 GB/s bandwidth. That is 16 GB more capacity and 68.8% more bandwidth for the L20.
Q: Which card is more power-hungry according to the data?
A: The W6900X has a 300 W TDP and a suggested PSU of 700 W. The L20 has a 275 W TDP and a suggested PSU of 600 W. So the AMD card draws more power, despite delivering far lower OpenCL and Vulkan scores.
Q: What is the production status of each card?
A: The L20 is listed as "Active" production, released 2023-11-15, with a predecessor of Server Ampere and a successor of Server Hopper. The W6900X is "End-of-life", released 2021-08-02, with no predecessor or successor listed.
Architecture Differences
The L20 uses the AD102 chip built on TSMC's 5 nm process, with 76,300 million transistors on a 609 mm² die. That yields a transistor density of 125.3M per mm². The architecture is Ada Lovelace, part of the Server Ada (Lxx) generation. The W6900X uses the Navi 21 chip on TSMC's 7 nm process, with 26,800 million transistors on a 520 mm² die, a density of 51.5M per mm². Its architecture is RDNA 2.0, from the Radeon Pro Mac (Navi II Series) generation.
The transistor count difference is stark: the L20 has roughly 2.85 times more transistors than the W6900X. The newer 5 nm process allows that density, while the 7 nm process limits the AMD chip. In terms of compute resources, the L20 has 11,776 shading units, 368 TMUs, and 128 ROPs. It also includes 92 ray tracing cores and 368 tensor cores. The W6900X has 5,120 shading units, 320 TMUs, and 128 ROPs, the same ROP count but less than half the shading units. It has 80 ray tracing cores and no tensor cores listed.
The L20 supports FP32 at 59.35 TFLOPS and FP16 at 59.35 TFLOPS with a 1:1 ratio. The W6900X delivers 22.23 TFLOPS FP32 and 44.46 TFLOPS FP16 with a 2:1 ratio. That means the L20 is 2.67 times faster in FP32, but the W6900X has a higher FP16-to-FP32 ratio, suggesting it can double throughput on half-precision workloads. The L20's tensor cores are the key architectural addition for AI workloads, which the AMD card lacks entirely.
Specification Differences
| Field | NVIDIA L20 | AMD Radeon Pro W6900X |
|---|---|---|
| Process node | 5 nm | 7 nm |
| Transistors | 76,300 million | 26,800 million |
| Die size | 609 mm² | 520 mm² |
| Base clock | 1440 MHz | 1825 MHz |
| Boost clock | 2520 MHz | 2171 MHz |
| Memory clock | 2250 MHz (18 Gbps effective) | 2000 MHz (16 Gbps effective) |
| Memory size | 48 GB GDDR6 | 32 GB GDDR6 |
| Memory bus | 384 bit | 256 bit |
| Memory bandwidth | 864.0 GB/s | 512.0 GB/s |
| Shading units | 11776 | 5120 |
| TMUs | 368 | 320 |
| ROPs | 128 | 128 |
| Ray tracing cores | 92 | 80 |
| Tensor cores | 368 | None |
| Pixel rate | 322.6 GPixel/s | 277.9 GPixel/s |
| Texture rate | 927.4 GTexel/s | 694.7 GTexel/s |
| FP32 | 59.35 TFLOPS | 22.23 TFLOPS |
| FP16 | 59.35 TFLOPS (1:1) | 44.46 TFLOPS (2:1) |
| TDP | 275 W | 300 W |
| Suggested PSU | 600 W | 700 W |
| Bus interface | PCIe 4.0 x16 | Apple MPX |
| Display outputs | 4x DisplayPort 1.4a | 1x HDMI 2.1, 4x Thunderbolt |
| Dimensions | 267 mm x 111 mm | 267 mm x 120 mm |
| Production status | Active | End-of-life |
| Release date | 2023-11-15 | 2021-08-02 |
| Launch MSRP | None listed | 5,999 USD |
The L20 has a higher boost clock (2520 MHz vs 2171 MHz) but a lower base clock (1440 MHz vs 1825 MHz). The W6900X actually has a higher base clock, which may help in sustained workloads that don't hit boost. The L20's memory subsystem is clearly superior: 48 GB vs 32 GB, 384-bit vs 256-bit bus, and 864.0 GB/s vs 512.0 GB/s bandwidth. The L20 is also shorter in height (111 mm vs 120 mm), making it easier to fit in dense server chassis.
Head-to-Head Benchmarks
The two cards share two benchmark tests: Geekbench OpenCL and Geekbench Vulkan. The L20 wins both decisively.
In Geekbench OpenCL, the L20 scores 274,276 against the W6900X's 130,035. The delta is 110.9%, meaning the L20 is more than twice as fast. This is the largest gap in any shared test. The L20's nearest rival in this metric is the NVIDIA L40 (284,111, -11.6% delta), which is actually faster, and the RTX 6000 Ada Generation (287,237, -12.6%). The W6900X's OpenCL score of 130,035 is far below its own average of 168,574, suggesting OpenCL is not a strong workload for the RDNA 2.0 architecture.
In Geekbench Vulkan, the L20 scores 228,018 versus the W6900X's 148,865. The delta is 53.2%. While still a clear win for the L20, the gap narrows compared to OpenCL. The L20's Vulkan score is closer to its average, while the W6900X's Vulkan result is also below its average but less dramatically so.
The W6900X has one additional benchmark: Geekbench Metal, scoring 226,821. There is no L20 Metal score in the data, so no direct comparison is possible. However, this Metal score is higher than the W6900X's OpenCL (130,035) and Vulkan (148,865) results, indicating the card is optimized for Apple's Metal API.
The overall win count is 2-0 in favor of the L20. The average benchmark score difference is 82,573 points (251,147 vs 168,574), which is a 49.0% advantage for the L20.
Where Each One Wins
NVIDIA L20 wins in: OpenCL compute, Vulkan rendering, AI inference (due to 368 tensor cores), high-bandwidth memory workloads (864.0 GB/s vs 512.0 GB/s), large model fitting (48 GB vs 32 GB), and any PCIe 4.0 server environment. The L20's 59.35 TFLOPS FP32 is 2.67 times the W6900X's 22.23 TFLOPS, so any FP32-heavy simulation or rendering task will favor it strongly. Its texture rate of 927.4 GTexel/s versus 694.7 GTexel/s also helps in texture-bound scenes. The L20 is in active production, so it will be available for new builds, and its 275 W TDP with a 600 W PSU suggestion makes it easier to power than the AMD card.
AMD Radeon Pro W6900X wins in: Apple MPX systems exclusively. The bus interface is Apple MPX, which the L20 cannot use. For Mac Pro owners, the W6900X is the only option of the two. Its Geekbench Metal score of 226,821 is strong, and Metal is the primary graphics API on macOS. The card also has a higher FP16 throughput relative to its FP32 (44.46 TFLOPS vs 22.23 TFLOPS, a 2:1 ratio), which could benefit half-precision workflows that use Metal. The W6900X's base clock of 1825 MHz is higher than the L20's 1440 MHz, which may help in latency-sensitive tasks that don't scale with core count. Its display outputs (1x HDMI 2.1 and 4x Thunderbolt) are tailored for Apple monitors and peripherals, whereas the L20 offers only DisplayPort 1.4a.
For anyone building a PC or server, the L20 is the data-backed choice: it wins every shared benchmark, has more memory, more bandwidth, more compute units, and a lower power draw. The W6900X is a niche product for Apple MPX slots, where its Metal performance and Thunderbolt outputs make it the only viable pick, but its 32 GB memory and 512.0 GB/s bandwidth are substantial limitations compared to the L20. The W6900X is also end-of-life, so new supply may be limited, while the L20 remains in active production.