NVIDIA L40S vs NVIDIA PG506-232 Comparison

NVIDIA
GEFORCE

NVIDIA L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022
VS
NVIDIA
GEFORCE

PG506-232

CORE STATE GA100
VRAM 24 GB
CLOCK SPEED 1440 MHz
TDP 165 W
BUS WIDTH 3072 bit
ARCHITECTURE Ampere
nm
PROCESS 7 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
330,727
225,124
geekbench_vulkan
260,799
N/A

Analysis: NVIDIA L40S vs NVIDIA PG506-232

NVIDIA’s L40S and PG506-232 are both end-of-life server accelerators, but they occupy very different corners of the product stack. The L40S is a modern Ada Lovelace part aimed at graphics and AI inference, while the PG506-232 is an older Ampere compute card with a focus on memory bandwidth. The data shows a decisive performance gap between them, but each card has a distinct technical identity worth unpacking.

Head-to-Head Benchmarks

The single available head-to-head benchmark, Geekbench OpenCL, paints a stark picture. The NVIDIA L40S scores 330,727 points, while the NVIDIA PG506-232 manages 225,124 points. That is a 46.9% advantage for the L40S, a massive margin that underscores the generational leap between the two architectures. In practical terms, the L40S delivers nearly half again as much raw compute throughput in this OpenCL workload.

Looking at the broader benchmark context, the L40S’s average benchmark score is 295,763, which places it in the 99th percentile of all GPUs. Its nearest rival, the NVIDIA RTX 6000 Ada Generation, scores 287,237 (3% lower), and the NVIDIA L40 scores 284,111 (4.1% lower). The L40S even beats the AMD Instinct MI300X, which scores 317,994 (7% lower), though it trails the NVIDIA H200 NVL’s 334,891 by 11.7%. These numbers confirm the L40S is not merely faster than the PG506-232; it is competitive with the top-tier accelerators of its era.

The PG506-232, by contrast, has an average benchmark score of 225,124, also in the 99th percentile. Its nearest rivals include the AMD Radeon PRO W7900D at 219,827 (2.4% lower), the NVIDIA A100 PCIe 80 GB at 207,124 (8.7% lower), and the NVIDIA L20 at 251,147 (10.4% higher). Interestingly, the PG506-232 beats the RTX 6000D by 14.9%, showing it is no slouch in its own generation. Still, the raw delta to the L40S is enormous: the L40S outperforms the PG506-232 by nearly 47% in the only direct comparison available.

Where Each One Wins

The L40S wins the only head-to-head benchmark outright, and its architecture suggests it dominates in most general-purpose and graphics-heavy workloads. With 18,176 shading units and 568 tensor cores, the L40S is built for parallel throughput. Its FP32 compute is rated at 91.61 TFLOPS, a figure that dwarfs the PG506-232’s 10.32 TFLOPS. For FP16, the L40S also delivers 91.61 TFLOPS (1:1), while the PG506-232 offers 10.32 TFLOPS (1:1) — an 8.9x gap in both precision formats. Any workload that relies on raw shader or tensor math will heavily favor the L40S.

The PG506-232, however, has one clear advantage: memory bandwidth. It packs 24 GB of HBM2 across a 3072-bit bus, delivering 933.1 GB/s of bandwidth. The L40S, with 48 GB of GDDR6 on a 384-bit bus, reaches 864.0 GB/s. That means the PG506-232 moves data 8% faster, despite having half the memory capacity. For memory-bound tasks like large matrix operations or certain HPC simulations, the PG506-232’s bandwidth could make it competitive despite its lower compute throughput. Also, the PG506-232 has a much lower TDP of 165 W versus the L40S’s 300 W, which could be relevant in power-constrained environments.

The L40S also wins on display outputs, offering 1x HDMI 2.1 and 3x DisplayPort 1.4a, whereas the PG506-232 has no display outputs at all. The L40S is clearly intended for visual computing, including rendering and possibly workstation use, while the PG506-232 is a headless compute card.

Architecture Differences

The two cards are separated by a full generation. The L40S uses the AD102 chip built on TSMC’s 5 nm process, part of the Ada Lovelace architecture. The PG506-232 uses the GA100 chip on TSMC’s 7 nm process, from the Ampere architecture. This process shrink is fundamental: the L40S packs 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3M per mm². The PG506-232 has 54,200 million transistors on a larger 826 mm² die, with a density of just 65.6M per mm². The L40S achieves nearly double the transistor density, which explains its massive compute advantage.

The L40S includes 142 ray tracing cores, a feature entirely absent from the PG506-232 (its RT core count is null). This makes the L40S suitable for ray-traced rendering workloads, while the PG506-232 is purely a compute device. Both cards have 568 tensor cores on the L40S versus 224 on the PG506-232, but the L40S’s newer tensor core design likely contributes to its higher FP16 throughput. The L40S also has 568 texture mapping units and 192 ROPs, compared to 224 TMUs and 96 ROPs on the PG506-232, further widening the gap in texture and pixel processing rates.

The memory types differ fundamentally: the L40S uses GDDR6, while the PG506-232 uses HBM2. The PG506-232’s HBM2 provides higher bandwidth per pin, but the L40S’s larger 384-bit bus and faster effective speed (18 Gbps versus 2.4 Gbps) allow it to nearly match bandwidth despite the older memory technology. The L40S’s boost clock of 2520 MHz is also substantially higher than the PG506-232’s 1440 MHz, reflecting the efficiency of the 5 nm process.

Specification Differences

The specifications diverge sharply in almost every measurable category. The L40S has 48 GB of GDDR6 memory, while the PG506-232 has 24 GB of HBM2. The memory bus is 384-bit on the L40S versus 3072-bit on the PG506-232, though the L40S’s higher effective memory speed (18 Gbps vs 2.4 Gbps) brings bandwidth to 864.0 GB/s versus 933.1 GB/s — a 7.4% bandwidth advantage for the PG506-232. Shading units are 18,176 on the L40S versus 3,584 on the PG506-232. TMUs are 568 versus 224, and ROPs are 192 versus 96. Tensor cores are 568 versus 224, and the L40S adds 142 RT cores where the PG506-232 has none.

Clocks are also very different: the L40S runs at a base of 1110 MHz and boosts to 2520 MHz, while the PG506-232 runs at 930 MHz base and 1440 MHz boost. The L40S’s pixel rate is 483.8 GPixel/s versus 138.2 GPixel/s, and its texture rate is 1,431.4 GTexel/s versus 322.6 GTexel/s. FP32 and FP16 are both 91.61 TFLOPS on the L40S versus 10.32 TFLOPS on the PG506-232. The L40S draws 300 W with a 1x 16-pin connector and a suggested 700 W PSU, while the PG506-232 draws 165 W with an 8-pin EPS connector and a suggested 450 W PSU. Both are dual-slot cards with identical dimensions (267 mm length, 111-112 mm height) and PCIe 4.0 x16 interfaces. The L40S supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4; the PG506-232 has no listed API support. The L40S was released in October 2022, while the PG506-232 came out in April 2021. The L40S’s predecessor is Server Ampere (which includes the PG506-232’s generation), and its successor is Server Hopper; the PG506-232’s predecessor is Tesla Turing and its successor is Server Ada.

FAQ

Q: Which GPU is faster in Geekbench OpenCL?

A: The NVIDIA L40S scores 330,727 versus the PG506-232’s 225,124, a 46.9% advantage for the L40S.

Q: Does the PG506-232 have any advantage over the L40S?

A: Yes, in memory bandwidth: the PG506-232 delivers 933.1 GB/s versus 864.0 GB/s on the L40S, an 8% lead. It also draws less power (165 W vs 300 W).

Q: How do the two cards compare in FP32 compute?

A: The L40S delivers 91.61 TFLOPS, while the PG506-232 manages 10.32 TFLOPS — an 8.9x gap in favor of the L40S.

Q: Can the PG506-232 handle ray tracing workloads?

A: No, the PG506-232 has no ray tracing cores, while the L40S includes 142 RT cores.

Q: What are the memory capacities and types?

A: The L40S has 48 GB of GDDR6 on a 384-bit bus, while the PG506-232 has 24 GB of HBM2 on a 3072-bit bus.

Q: Which card supports display outputs?

A: Only the L40S, which offers 1x HDMI 2.1 and 3x DisplayPort 1.4a. The PG506-232 has no display outputs.

DETAILED SPECIFICATIONS

SPECIFICATION
L40S
PG506-232
Core Specs
Shading Units
18,176
3,584 -80.3%
Shaders
18,176
3,584 -80.3%
TMUs
568
224 -60.6%
ROPs
192
96 -50.0%
SM Count
142
56 -60.6%
Clocks
Base Clock
1110 MHz
930 MHz
Boost Clock
2520 MHz
1440 MHz
Memory Clock
2250 MHz 18 Gbps effective
1215 MHz 2.4 Gbps effective
Memory
Memory Size
48 GB
24 GB
VRAM (MB)
49,152
24,576 -50.0%
Memory Type
GDDR6
HBM2
Memory Bus
384 bit
3072 bit
Bandwidth
864.0 GB/s
933.1 GB/s
Cache
L1 Cache
128 KB (per SM)
192 KB (per SM)
L2 Cache
48 MB
24 MB
Performance
Pixel Rate
483.8 GPixel/s
138.2 GPixel/s
Texture Rate
1,431.4 GTexel/s
322.6 GTexel/s
FP32 (TFLOPS)
91.61 TFLOPS
10.32 TFLOPS
FP64 (TFLOPS)
1,431.4 GFLOPS (1:64)
5.161 TFLOPS (1:2)
FP16 (TFLOPS)
91.61 TFLOPS (1:1)
10.32 TFLOPS (1:1)
AI/RT
RT Cores
142
Tensor Cores
568
224 -60.6%
Power
TDP
300 W
165 W
TDP (W)
300
165 -45.0%
Suggested PSU
700 W
450 W
Power Connectors
1x 16-pin
8-pin EPS
Architecture
Architecture
Ada Lovelace
Ampere
GPU Name
AD102
GA100
Generation
Server Ada (Lxx)
Server Ampere (Axx)
Process Size
5 nm
7 nm
Transistors
76,300 million
54,200 million
Die Size
609 mm²
826 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
65.6M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.9
8.0
Shader Model
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
112 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Server Ampere
Tesla Turing
Successor
Server Hopper
Server Ada
View L40S Details View PG506-232 Details