AMD Radeon Pro W6600X vs NVIDIA L40S Comparison
AMD Radeon Pro W6600X
L40S
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Pro W6600X vs NVIDIA L40S
Head-to-Head Benchmarks
The recorded data shows no direct head-to-head benchmark runs between the NVIDIA L40S and the AMD Radeon Pro W6600X. Instead, the database places each card against its own nearest rivals, and the two GPUs occupy very different performance tiers. The L40S achieves an average benchmark score of 295,763, placing it in the 99th percentile of all GPUs. The W6600X averages 107,342, which sits in the 94th percentile. That is a gap of roughly 175%, but the more useful comparison comes from looking at each card's specific benchmark results and how they stack up within their respective peer groups.
For the L40S, the Geekbench OpenCL score is 330,727, and the Vulkan score is 260,799. These are massive numbers, consistent with a GPU designed for compute-heavy server workloads. Against its nearest rivals, the L40S leads the NVIDIA RTX 6000 Ada Generation by 3%, with the RTX 6000 averaging 287,237. It also beats the NVIDIA L40 by 4.1%, as the L40 averages 284,111. However, the L40S trails the AMD Instinct MI300X by 7% (that card scores 317,994) and falls 11.7% behind the NVIDIA H200 NVL (334,891). So within its own tier, the L40S is competitive but not the absolute top; it is faster than the RTX 6000 Ada and the L40, but slower than the MI300X and H200 NVL.
The W6600X has only one recorded benchmark result: a Geekbench Metal score of 107,342. That is its entire average, since no other tests appear in the database. Among its nearest rivals, the W6600X edges out the AMD Radeon Pro Vega II Duo by 0.6% (that card averages 106,750) and beats the NVIDIA Quadro RTX 6000 by 5.4% (101,872). It falls behind the AMD Radeon Pro Vega II by 2.1% (109,617) and the AMD Radeon PRO W7900 by 3.1% (110,725). The pattern here is clear: the W6600X is a mid-pack performer among older pro GPUs, trading blows with Vega-era cards and the Quadro RTX 6000.
The biggest single-score difference between the two cards is in OpenCL: the L40S posts 330,727 versus the W6600X's sole Metal result of 107,342. Even accounting for different APIs, the L40S is in a different performance class. The Vulkan score of 260,799 for the L40S further reinforces that the NVIDIA card is built for raw throughput, while the W6600X is a much smaller, lower-power part aimed at macOS workstation use.
Architecture Differences
The L40S uses the AD102 chip on NVIDIA's Ada Lovelace architecture, built on a 5 nm process at TSMC. The die is 609 mm² and contains 76,300 million transistors, giving a transistor density of 125.3 million per square millimeter. The W6600X uses the Navi 23 chip on AMD's RDNA 2.0 architecture, also fabricated by TSMC but on a 7 nm node. Its die is 237 mm² with 11,060 million transistors, for a density of 46.7 million per square millimeter. The L40S is the larger, denser, and more complex chip by a wide margin.
The L40S ships with 18,176 shading units, 568 texture mapping units, and 192 ROPs. It also includes 142 RT cores and 568 tensor cores. The W6600X has 2,048 shading units, 128 TMUs, and 64 ROPs, plus 32 RT cores and no tensor cores at all, as the field is recorded as null. That absence of tensor cores is a major architectural difference: the L40S is built for AI and deep learning workloads that rely on tensor operations, while the W6600X has no such hardware.
Clock speeds tell a different story. The L40S has a base clock of 1110 MHz and a boost of 2520 MHz. The W6600X has a higher base clock at 2068 MHz and a boost of 2479 MHz. So the AMD chip starts at nearly double the L40S's base frequency, though the boost clocks are close. The memory clocks differ as well: the L40S runs at 2250 MHz with 18 Gbps effective, while the W6600X runs at 2000 MHz with 16 Gbps effective. The L40S has 48 GB of GDDR6 on a 384-bit bus, yielding 864.0 GB/s of bandwidth. The W6600X has 8 GB of GDDR6 on a 128-bit bus, for 256.0 GB/s. That is a 3.4x bandwidth advantage for the L40S, which matters enormously in compute scenarios.
The L40S supports PCIe 4.0 x16, while the W6600X uses Apple MPX as its bus interface. The L40S has display outputs (1x HDMI 2.1 and 3x DisplayPort 1.4a), while the W6600X has no outputs at all, meaning it is purely an internal compute or render card for Mac systems. The L40S draws 300 W with a 700 W suggested PSU and a 16-pin power connector; the W6600X draws 120 W with a 300 W suggested PSU and no power connector listed. The W6600X is a much lighter power load.
The Verdict
From the data, the L40S is the clear choice for anyone needing high compute throughput. Its average score of 295,763 versus 107,342 for the W6600X is not a close contest. The L40S is in the 99th percentile of all GPUs, while the W6600X sits in the 94th. If your workload involves OpenCL or Vulkan compute, the L40S delivers 330,727 and 260,799 respectively, numbers the W6600X cannot approach with its single Metal score of 107,342.
The W6600X has its own place, however. It is a low-power card at 120 W, requires no external power connector, and uses the Apple MPX interface. The L40S needs a 16-pin connector and a 700 W PSU. For a Mac Pro user who needs a modest GPU for Metal-based tasks and cannot accommodate a high-power card, the W6600X fits. Its 8 GB of memory is small compared to 48 GB, but for lighter workloads that may be sufficient. The L40S is end-of-life, as is the W6600X, so neither is a future-proof purchase.
Choose the L40S if you need raw compute, tensor cores, high bandwidth, and large memory. Choose the W6600X if you are constrained by power, slot space, or the Apple MPX bus, and if your software runs on Metal. The data does not support any other conclusion: the L40S is the performance king, the W6600X is the niche workstation part.
Specification Differences
The two cards differ in nearly every specification. The L40S uses a 5 nm process; the W6600X uses 7 nm. The L40S has 76,300 million transistors on a 609 mm² die; the W6600X has 11,060 million on 237 mm². Transistor density is 125.3 million per mm² for the L40S versus 46.7 million for the W6600X.
Memory is a major split: the L40S has 48 GB of GDDR6 on a 384-bit bus with 864.0 GB/s bandwidth; the W6600X has 8 GB on a 128-bit bus with 256.0 GB/s. The L40S's memory clock is 2250 MHz (18 Gbps effective), while the W6600X runs at 2000 MHz (16 Gbps effective).
Compute units differ sharply. The L40S has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The W6600X has 2,048 shading units, 128 TMUs, 64 ROPs, and 32 RT cores, with no tensor cores. Pixel rate is 483.8 GPixel/s for the L40S versus 158.7 GPixel/s for the W6600X. Texture rate is 1,431.4 GTexel/s versus 317.3 GTexel/s. FP32 performance is 91.61 TFLOPS for the L40S versus 10.15 TFLOPS for the W6600X. FP16 is 91.61 TFLOPS (1:1) for the L40S versus 20.31 TFLOPS (2:1) for the W6600X.
Power and connectivity also differ. The L40S has a 300 W TDP and a 700 W suggested PSU; the W6600X has a 120 W TDP and a 300 W suggested PSU. The L40S uses a 16-pin power connector; the W6600X lists none. The L40S uses PCIe 4.0 x16; the W6600X uses Apple MPX. Display outputs: the L40S has 1x HDMI 2.1 and 3x DisplayPort 1.4a; the W6600X has no outputs. The L40S is 267 mm long and 111 mm tall; the W6600X has no recorded dimensions.
Both cards support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. Both are end-of-life. The L40S launched on 2022-10-12; the W6600X on 2021-08-02. The W6600X has a launch MSRP of 699 USD; the L40S has no recorded launch MSRP.
FAQ
Q: Which card has higher FP32 performance?
A: The L40S delivers 91.61 TFLOPS of FP32, while the W6600X delivers 10.15 TFLOPS. That is roughly a 9x advantage for the L40S.
Q: Can the W6600X handle ray tracing?
A: Yes, the W6600X has 32 RT cores and supports DirectX 12 Ultimate, which includes ray tracing features. The L40S has 142 RT cores, so it has more ray tracing hardware, but the W6600X is not without RT capability.
Q: Why does the W6600X have a higher base clock than the L40S?
A: The W6600X has a base clock of 2068 MHz versus 1110 MHz for the L40S. This is partly because the W6600X is a smaller, lower-power chip (120 W TDP) that can run at higher frequencies at its scale, while the L40S, at 300 W with far more compute units, starts at a lower base clock but boosts to 2520 MHz, close to the W6600X's 2479 MHz boost.
Q: Which card has more memory bandwidth?
A: The L40S has 864.0 GB/s of bandwidth thanks to 48 GB on a 384-bit bus. The W6600X has 256.0 GB/s from 8 GB on a 128-bit bus. The L40S offers over 3x the bandwidth.
Q: Does the W6600X work in a standard PCIe slot?
A: No, the W6600X uses the Apple MPX bus interface, not PCIe. The L40S uses PCIe 4.0 x16. The W6600X is designed for Apple Mac systems.
Q: Which card is better for Metal-based workloads?
A: The W6600X has a recorded Geekbench Metal score of 107,342, while the L40S has no Metal benchmark in the database. For Metal-specific tasks, the W6600X is the only one with direct data, and its 94th percentile ranking suggests it is competent within that API.
Where Each One Wins
The L40S wins in every category where raw numbers are available. It leads in shading units (18,176 versus 2,048), TMUs (568 versus 128), ROPs (192 versus 64), RT cores (142 versus 32), and it has tensor cores that the W6600X lacks. Its FP32 rate is 91.61 TFLOPS versus 10.15 TFLOPS, and its FP16 rate is 91.61 TFLOPS versus 20.31 TFLOPS. Memory capacity is 48 GB versus 8 GB, and bandwidth is 864.0 GB/s versus 256.0 GB/s. The L40S also has display outputs, while the W6600X has none. For any compute-heavy task, whether OpenCL, Vulkan, or AI inference using tensor cores, the L40S is the superior choice.
The W6600X wins on power efficiency and integration. It draws 120 W versus 300 W, and its suggested PSU is 300 W versus 700 W. It needs no power connector, whereas the L40S requires a 16-pin connector. The W6600X uses the Apple MPX bus, which makes it a direct fit for Mac Pro systems, while the L40S needs a PCIe 4.0 slot. The W6600X also has a higher base clock (2068 MHz versus 1110 MHz), which can help in lightly threaded or latency-sensitive tasks, though its boost clock is slightly lower (2479 MHz versus 2520 MHz). For a Mac user with limited power budget and no need for display outputs, the W6600X is the practical pick. For anyone else, the L40S is the data-backed winner.