AMD Radeon PRO W7600 vs NVIDIA L40S Comparison

AMD
RADEON

AMD Radeon PRO W7600

CORE STATE Navi 33
VRAM 8 GB
CLOCK SPEED 2440 MHz
TDP 130 W
BUS WIDTH 128 bit
ARCHITECTURE RDNA 3.0
nm
PROCESS 6 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
81,528
330,727
geekbench_vulkan
92,688
260,799

Analysis: AMD Radeon PRO W7600 vs NVIDIA L40S

Head-to-Head Benchmarks

The recorded data shows a decisive advantage for the NVIDIA L40S in every measured workload. In Geekbench OpenCL, the L40S scores 330,727 against the AMD Radeon PRO W7600's 81,528, a delta of 305.7%. That is not a marginal gap; it is a four-fold difference in raw compute throughput. The Vulkan results tell a similar story, with the L40S posting 260,799 versus 92,688, a 181.4% lead.

What stands out here is the consistency of the margin. OpenCL and Vulkan stress different parts of the GPU stack, yet the L40S wins both by a wide margin. The Radeon PRO W7600 does not take a single benchmark in this comparison. The data shows two wins for the L40S and zero for the AMD card. Looking at averages, the L40S sits at 295,763, while the W7600 averages 87,108, a spread that places them in entirely different performance tiers.

The percentile rankings reinforce this gap. The L40S sits at the 99th percentile among all GPUs, meaning it outperforms 99% of the database's entries. The W7600 holds the 93rd percentile, which is still strong, but it is a full six percentile points lower. For context, the L40S's nearest rivals include the AMD Instinct MI300X, which scores 317,994, making the L40S 7% slower than that accelerator. The NVIDIA H200 NVL sits 11.7% ahead of the L40S. The W7600's nearest rivals are far closer, with the NVIDIA Quadro GP100 only 0.4% ahead and the NVIDIA CMP 40HX just 1.7% behind. The W7600 is fighting in a crowded field of mid-range workstation cards, while the L40S is competing in the top 1% of all accelerators.

Architecture Differences

The underlying silicon explains the performance chasm. The L40S uses the AD102 chip on the Ada Lovelace architecture, built by TSMC at a 5 nm process. The W7600 uses the Navi 33 chip on the RDNA 3.0 architecture, also from TSMC but at a 6 nm node. The transistor counts are staggering in comparison: the L40S houses 76,300 million transistors on a 609 mm² die, yielding a density of 125.3 million transistors per mm². The W7600 has 13,300 million transistors on a 204 mm² die, with a density of 65.2 million per mm². That is a 5.7x difference in raw transistor count and nearly double the density, showcasing how much newer the 5 nm process is versus the older 6 nm node.

The memory subsystems diverge dramatically. The L40S has 48 GB of GDDR6 across a 384-bit bus, delivering 864.0 GB/s of bandwidth. The W7600 has 8 GB of GDDR6 on a 128-bit bus, achieving 288.0 GB/s. That is a 6x memory capacity advantage and a 3x bandwidth advantage for the NVIDIA card. For workloads that are memory-bound, such as large model inference or high-resolution rendering, the L40S's bandwidth is a critical differentiator.

Core counts are equally lopsided. The L40S has 18,176 shading units, 568 TMUs, and 192 ROPs. It also carries 142 RT cores and 568 tensor cores. The W7600 has 2,048 shading units, 128 TMUs, and 64 ROPs, with 32 RT cores and no tensor cores at all. The FP32 performance is 91.61 TFLOPS for the L40S versus 19.99 TFLOPS for the W7600. The FP16 ratio is interesting: the L40S runs at 1:1, giving 91.61 TFLOPS, while the W7600 runs at 2:1, reaching 39.98 TFLOPS. The 1:1 ratio on the L40S suggests it is designed for workloads that need uniform precision, whereas the W7600's 2:1 ratio indicates a focus on rendering where FP16 is acceptable.

The bus interface differs too: the L40S uses PCIe 4.0 x16, while the W7600 uses PCIe 4.0 x8. The power targets reflect the performance gap: the L40S is rated at 300 W, requiring a single 16-pin connector and a 700 W suggested PSU, while the W7600 is rated at 130 W with a 6-pin connector and a 300 W PSU. The L40S is dual-slot, the W7600 is single-slot. Display outputs also differ: the L40S has one HDMI 2.1 and three DisplayPort 1.4a, while the W7600 has four DisplayPort 2.1, which is a newer standard.

Where Each One Wins

The L40S wins everywhere the data has a measurement. In OpenCL, it is 305.7% ahead. In Vulkan, it is 181.4% ahead. This suggests it is the better choice for any compute-heavy workload that uses these APIs, including general-purpose GPU compute, machine learning inference, physics simulation, and video processing. The presence of tensor cores and a 1:1 FP16 ratio make the L40S particularly suited for AI workloads that rely on FP16 or FP32 precision. The 48 GB memory is a practical advantage for datasets that exceed 8 GB, which is the W7600's entire memory capacity.

The W7600's wins are not in raw performance but in the attributes that surround it. It is a single-slot card, which is unusual and valuable for dense server environments. It draws 130 W, which is less than half the L40S's 300 W, making it easier to cool and power. It has four DisplayPort 2.1 outputs, which is a modern connectivity set, whereas the L40S's DisplayPort 1.4a is older, though it does have an HDMI 2.1 port. The W7600 is also a smaller card: 241 mm length versus 267 mm for the L40S.

For use cases that do not require massive memory or extreme compute, the W7600 is the more practical card. Think of it as a professional visual workstation card that can drive multiple 4K monitors or handle light to moderate rendering tasks. The L40S is a server-grade accelerator, with a production status of end-of-life, whereas the W7600 is active. The L40S was released on October 12, 2022, while the W7600 came later on August 2, 2023.

FAQ

Q: Which card is faster in OpenCL compute?

A: The NVIDIA L40S scores 330,727 in Geekbench Open, versus the AMD Radeon PRO W7600's 81,528, making the L40S 305.7% faster.

Q: Does the AMD Radeon PRO W7600 win any benchmark?

A: No. The recorded data shows zero wins for the W7600, with the L40S taking both the OpenCL and Vulkan tests.

Q: What is the difference in memory capacity?

A: The L40S has 48 GB of GDDR6, while the W7600 has 8 GB of GDDR6. The L40S also has a wider 384-bit bus versus a 128-bit bus, resulting in 864.0 GB/s versus 288.0 GB/s.

Q: Does the AMD card have tensor cores?

A: No. The W7600 has no tensor cores, while the L40S has 568 tensor cores. This is a major differentiator for AI and deep learning workloads.

Q: What is the power consumption of each?

A: The L40S has a TDP of 300 W and requires a 700 W PSU, while the W7600 has a TDP of 130 W and requires a 300 W PSU.

Q: How do they compare to their nearest rivals?

A: The L40S is 3% faster than the NVIDIA RTX 6000 Ada Generation and 4.1% faster than the NVIDIA L40. The W7600 is 0.4% slower than the NVIDIA Quadro GP100 and 1.7% faster than the NVIDIA CMP 40HX.

The Verdict

The data is unambiguous: the NVIDIA L40S is the superior accelerator by every measured metric. It outclasses the AMD Radeon PRO W7600 by over 300% in OpenCL and 180% in Vulkan. It has 6x the memory capacity, 4.5x the FP32 throughput, and a 99th percentile rating versus 93rd. The L40S is a top-tier server accelerator, competing with the NVIDIA H200 NVL and AMD Instinct MI300X, though it trails both by 11.7% and 7% respectively. There is no workload in the database where the W7600 comes out ahead.

However, the L40S is not for everyone. It is end-of-life, dual-slot, and requires a 700 W PSU. The AMD W7600 is a modern, 130 W, single-slot card with DisplayPort 2.1. If the use case involves driving multiple 8K displays, power efficiency, or a compact chassis, the W7600 is the rational choice, provided the workload fits within 8 GB memory and less than 20 TFLOPS of FP32 compute. For any serious compute, AI, or rendering task, the L40S is the only choice, and the benchmark scores show it dominates.

Specification Differences

The two cards differ in nearly every specification. The L40S uses an AD1020 chip on Ada Lovelace at 5 nm, while the W7600 uses Navi 33 on RDNA 3.0 at 6 nm. Transistor count is 76,300 million versus 13,300 million, and die size is 609 mm² versus 204 mm². The L40S has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores; the W7600 has 2,048 shading units, 128 TMUs, 64 ROPs, and 32 RT cores, with no tensor cores. Memory is 48 GB GDDR6 on a 384-bit bus (864.0 GB/s) versus 8 GB GDDR6 on a 128-bit bus (288.0 GB/s). FP32 is 91.61 TFLOPS vs 19.99 TFLOPS. FP16 is 91.61 TFLOPS (1:1) vs 39.98 TFLOPS (2:1). TDP is 300 W vs 130 W. The L40S is dual-slot with a 16-pin connector and 700 W PSU, while the W7600 is single-slot with a 6-pin connector and 300 W PSU. Bus interface is PCIe 4.0 x16 vs PCIe 4.0 x8. Display outputs are 1x HDMI 2.1 and 3x DisplayPort 1.4a vs 4x DisplayPort 2.1. The L40S measures 267 mm x 111 mm, the W7600 measures 241 mm x 115 mm. Release dates are October 2022 for the L40S and August 2023 for the W7600. The L40S is end-of-life, the W7600 is active.

DETAILED SPECIFICATIONS

SPECIFICATION
PRO W7600
L40S
Core Specs
Shading Units
2,048
18,176 +787.5%
Shaders
2,048
18,176 +787.5%
TMUs
128
568 +343.8%
ROPs
64
192 +200.0%
Compute Units
32
SM Count
142
Clocks
Base Clock
1720 MHz
1110 MHz
Boost Clock
2440 MHz
2520 MHz
Memory Clock
2250 MHz 18 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
8 GB
48 GB
VRAM (MB)
8,192
49,152 +500.0%
Memory Type
GDDR6
GDDR6
Memory Bus
128 bit
384 bit
Bandwidth
288.0 GB/s
864.0 GB/s
Cache
L1 Cache
128 KB per Array
128 KB (per SM)
L2 Cache
2 MB
48 MB
L3 Cache
32 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
156.2 GPixel/s
483.8 GPixel/s
Texture Rate
312.3 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
19.99 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
624.6 GFLOPS (1:32)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
39.98 TFLOPS (2:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
32
142 +343.8%
Tensor Cores
568
Matrix Cores
64
Power
TDP
130 W
300 W
TDP (W)
130
300 +130.8%
Suggested PSU
300 W
700 W
Power Connectors
1x 6-pin
1x 16-pin
Architecture
Architecture
RDNA 3.0
Ada Lovelace
GPU Name
Navi 33
AD102
Codename
Hotpink Bonefish
Generation
Radeon Pro Navi (Navi III Series)
Server Ada (Lxx)
Process Size
6 nm
5 nm
Transistors
13,300 million
76,300 million
Die Size
204 mm²
609 mm²
Foundry
TSMC
TSMC
Density
65.2M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.2
3.0
CUDA
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Dual-slot
Length
241 mm 9.5 inches
267 mm 10.5 inches
Height
115 mm 4.5 inches
111 mm 4.4 inches
Outputs
4x DisplayPort 2.1
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x8
PCIe 4.0 x16
Other
Launch Price
599 USD
Production
Active
End-of-life
Predecessor
Radeon Pro Vega
Server Ampere
Successor
Server Hopper
View Radeon PRO W7600 Details View L40S Details