NVIDIA B300 SXM6 AC vs NVIDIA L40S Comparison

NVIDIA
GEFORCE

NVIDIA B300 SXM6 AC

CORE STATE GB110
VRAM 288 GB
CLOCK SPEED 2032 MHz
TDP 1100 W
BUS WIDTH 8192 bit
ARCHITECTURE Blackwell Ultra
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
369,831
330,727
geekbench_vulkan
N/A
260,799

Analysis: NVIDIA B300 SXM6 AC vs NVIDIA L40S

The Verdict

The data presents a clear split between two very different NVIDIA server accelerators. The NVIDIA B300 SXM6 AC is the overwhelming performance leader, posting an average benchmark score of 369,831 across its OpenCL result, which places it in the 100th percentile of all GPUs. The NVIDIA L40S, by contrast, averages 295,763 across OpenCL and Vulkan tests, sitting in the 99th percentile. The B300 leads the L40S by 25% in average score, and in the direct OpenCL head-to-head, the B300 wins by 11.8%, scoring 369,831 against 330,727.

For buyers who need maximum raw compute throughput, the B300 SXM6 AC is the only choice from this data. It outpaces not just the L40S but also the NVIDIA B200 by 7%, the H200 NVL by 10.4%, and the AMD Instinct MI300X by 16.3%. Its 288 GB of HBM3e memory with 8.19 TB/s bandwidth dwarfs the L40S’s 48 GB GDDR6 at 864 GB/s. This is a system designed for the largest models and most demanding data-center workloads.

The L40S, however, is not without merit. It is a dual-slot PCIe card with display outputs, a 300 W TDP, and a 700 W suggested PSU, making it far more adaptable to standard server chassis. It supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, whereas the B300 has no API support listed. The L40S also has a higher boost clock (2520 MHz vs 2032 MHz) and dramatically higher pixel rate (483.8 GPixel/s vs 48.77 GPixel/s). For workloads that need graphics output, rasterization, or lower power envelopes, the L40S is the practical option.

Strictly from the data, the verdict is: choose the B300 SXM6 AC for pure compute density and massive memory capacity in a data-center rack; choose the L40S for flexible deployment, display output, and graphics-adjacent tasks where its 300 W power draw and standard PCIe form factor are decisive advantages.

Architecture Differences

The two chips share a foundry and process node — TSMC at 5 nm — but diverge completely in design philosophy. The B300 uses the GB110 chip under the Blackwell Ultra architecture, while the L40S uses the AD102 chip under Ada Lovelace. The B300’s die is 1628 mm² with 208,000 million transistors, giving a density of 127.8M transistors per mm². The L40S’s die is 609 mm² with 76,300 million transistors, at 125.3M per mm². The B300 is not just bigger; it is marginally denser per square millimeter.

Memory architecture is the starkest difference. The B300 uses 288 GB of HBM3e across an 8192-bit bus, achieving 8.19 TB/s bandwidth. The L40S uses 48 GB of GDDR6 on a 384-bit bus, delivering 864 GB/s. That is an approximate 9.5x bandwidth advantage for the B300 and a 6x capacity advantage. The memory clock also differs: the B300 runs at 2000 MHz (8 Gbps effective), while the L40S runs at 2250 MHz (18 Gbps effective). The L40S’s faster memory clock per pin does not compensate for its narrower bus.

Compute resources tell a similar story. The B300 has 18,944 shading units, 592 TMUs, and 592 tensor cores, but only 24 ROPs. The L40S has 18,176 shading units, 568 TMUs, 568 tensor cores, and 142 RT cores, with 192 ROPs. The shading unit counts are close, but the ROP disparity is enormous — the L40S has 8x the ROPs of the B300. This explains the pixel rate gap: 483.8 GPixel/s for the L40S versus 48.77 GPixel/s for the B300.

The B300 is a compute monster with FP32 and FP16 both at 76.99 TFLOPS (1:1 ratio). The L40S pushes higher raw FP32 and FP16 at 91.61 TFLOPS each. The B300’s texture rate is 1,202.9 GTexel/s, while the L40S achieves 1,431.4 GTexel/s. Interestingly, the L40S has higher raw FP32, texture, and pixel throughput, yet loses decisively in the OpenCL benchmark. This suggests the B300’s advantage lies in memory bandwidth, capacity, and architecture efficiency rather than raw shader throughput.

Interfaces and physical specs differ as well. The B300 is an SXM module with PCIe 6.0 x16 and no display outputs. The L40S is dual-slot with PCIe 4.0 x16, one HDMI 2.1, and three DisplayPort 1.4a outputs. The B300’s TDP is 1100 W with a 1500 W suggested PSU; the L40S draws 300 W with a 700 W PSU. The L40S has physical dimensions listed (267 mm length, 111 mm height); the B300 has none listed.

FAQ

Q: Which GPU has higher raw FP32 compute?

A: The L40S has higher FP32 at 91.61 TFLOPS, versus 76.99 TFLOPS for the B300. However, the B300 wins the OpenCL benchmark by 11.8%, indicating that raw FP32 throughput does not determine overall compute performance in this test.

Q: How much memory do these GPUs have, and why does it matter?

A: The B300 has 288 GB of HBM3e with 8.19 TB/s bandwidth, while the L40S has 48 GB of GDDR6 with 864 GB/s bandwidth. The B300’s 6x capacity and roughly 9.5x bandwidth advantage are critical for large model inference and training datasets that cannot fit in 48 GB.

Q: Does the L40S support graphics APIs?

A: Yes. The L40S supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, and has display outputs including 1x HDMI 2.1 and 3x DisplayPort 1.4a. The B300 lists no API support and no display outputs.

Q: What is the power requirement difference?

A: The B300 has a TDP of 1100 W and a suggested PSU of 1500 W. The L40S has a TDP of 300 W and a suggested PSU of 700 W. This makes the L40S far easier to integrate into existing infrastructure.

Q: How does the B300 compare to the NVIDIA B200?

A: The B300’s average benchmark score of 369,831 is 7% higher than the B200’s 345,482. Both are in the same Blackwell family, but the B300 leads in this data.

Q: Is the L40S still a competitive accelerator despite being end-of-life?

A: Yes. Despite its end-of-life production status, the L40S scores 295,763 on average, which is 3% ahead of the RTX 6000 Ada Generation and 4.1% ahead of the L40. It trails the MI300X by 7% and the H200 NVL by 11.7%, but remains a strong mid-tier option.

Specification Differences

The two GPUs differ across nearly every measurable specification. The B300 uses the GB110 chip on Blackwell Ultra architecture; the L40S uses AD102 on Ada Lovelace. The B300 has 208,000 million transistors on a 1628 mm² die; the L40S has 76,300 million on 609 mm². Transistor density is nearly identical (127.8M vs 125.3M per mm²), but the B300 packs nearly 2.7x the transistors.

Clock speeds differ significantly. The B300 has a base clock of 1665 MHz and a boost of 2032 MHz. The L40S has a base of 1110 MHz and a boost of 2520 MHz. The L40S boosts 488 MHz higher, but the B300 has a 555 MHz higher base clock. Memory clocks: the B300 runs at 2000 MHz (8 Gbps effective), while the L40S runs at 2250 MHz (18 Gbps effective).

Memory capacity and type are the biggest differentiators: 288 GB HBM3e for the B300 versus 48 GB GDDR6 for the L40S. Bus width is 8192-bit versus 384-bit, and bandwidth is 8.19 TB/s versus 864 GB/s. The B300 has 18,944 shading units, 592 TMUs, 24 ROPs, and 592 tensor cores; the L40S has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores.

Pixel rate is 48.77 GPixel/s for the B300 versus 483.8 GPixel/s for the L40S — a 10x difference favoring the L40S. Texture rate is 1,202.9 GTexel/s versus 1,431.4 GTexel/s, favoring the L40S by about 19%. FP32 and FP16 are both 76.99 TFLOPS for the B300 and 91.61 TFLOPS for the L40S.

Power and physical specs diverge: the B300 is an SXM module with 1100 W TDP and 1500 W suggested PSU; the L40S is dual-slot with 300 W TDP and 700 W suggested PSU, plus a 1x 16-pin power connector. The B300 uses PCIe 6.0 x16; the L40S uses PCIe 4.0 x16. The B300 has no display outputs; the L40S has HDMI 2.1 and DisplayPort 1.4a. The B300 is listed as Active production; the L40S is End-of-life. Release dates: the B300 launched on 2025-09-10; the L40S launched on 2022-10-12.

Head-to-Head Benchmarks

The single direct comparison available is Geekbench OpenCL. The B300 scores 369,831, and the L40S scores 330,727. The B300 wins by 11.8%. This is the only head-to-head result, but it is consistent with the average score gap: the B300’s average is 369,831 (from its one benchmark), while the L40S’s average is 295,763 (averaging its OpenCL and Vulkan results). The L40S’s Vulkan score of 260,799 is notably lower than its OpenCL score, dragging its average down.

Looking at the nearest rivals provides context. The B300’s 25% lead over the L40S in average score is its largest margin among the listed rivals. The B300 leads the B200 by 7%, the H200 NVL by 10.4%, and the MI300X by 16.3%. The L40S, from its side, trails the MI300X by 7% and the H200 NVL by 11.7%, while leading the RTX 6000 Ada Generation by 3% and the L40 by 4.1%. The data shows the B300 operating in a higher performance tier entirely.

The B300’s OpenCL score of 369,831 is 11.8% above the L40S’s OpenCL score of 330,727. This is a substantial gap, but not overwhelming. The L40S’s higher FP32 (91.61 vs 76.99 TFLOPS) and higher texture rate (1,431.4 vs 1,202.9 GTexel/s) suggest that the B300’s win comes from other factors — most plausibly its 8.19 TB/s memory bandwidth versus 864 GB/s, which allows the B300 to feed its compute units far more efficiently in memory-bound OpenCL workloads.

Where Each One Wins

The B300 SXM6 AC wins decisively in compute-heavy, memory-intensive data-center workloads. Its 288 GB HBM3e capacity and 8.19 TB/s bandwidth are unmatched by the L40S’s 48 GB GDDR6. For large language model inference, training datasets exceeding 48 GB, or any workload that thrashs memory, the B300’s 25% average score lead and 11.8% OpenCL head-to-head win make it the clear choice. It also leads every listed rival — B200, H200 NVL, MI300X, and L40S — in average score, cementing its position at the 100th percentile.

The L40S wins in graphics and rasterization-adjacent tasks. Its 192 ROPs versus the B300’s 24 ROPs, and its 483.8 GPixel/s pixel rate versus 48.77 GPixel/s, make it the only option for pixel-heavy work. It has display outputs (HDMI 2.1 and three DisplayPort 1.4a), supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, and fits in a dual-slot PCIe 4.0 x16 form factor. The B300 has no display outputs and no API support. The L40S also wins on power efficiency: 300 W TDP versus 1100 W, and a 700 W suggested PSU versus 1500 W.

For FP32 compute, the L40S’s 91.61 TFLOPS exceeds the B300’s 76.99 TFLOPS, and its texture rate of 1,431.4 GTexel/s beats 1,202.9 GTexel/s. In scenarios where raw shader throughput matters more than memory bandwidth, the L40S has a theoretical edge — though the benchmark data shows the B300 winning overall. The L40S also has a higher boost clock (2520 vs 2032 MHz), which can help in latency-sensitive, low-occupancy workloads.

The use-case split is clean: the B300 for maximum compute density, massive memory, and top-tier data-center performance; the L40S for graphics output, lower power, flexible deployment, and pixel-rate-heavy tasks. The L40S’s end-of-life status and the B300’s active production status further reinforce that the B300 is the forward-looking choice for compute, while the L40S remains a viable option for specific graphics or power-constrained roles.

DETAILED SPECIFICATIONS

SPECIFICATION
B300 SXM6 AC
L40S
Core Specs
Shading Units
18,944
18,176 -4.1%
Shaders
18,944
18,176 -4.1%
TMUs
592
568 -4.1%
ROPs
24
192 +700.0%
SM Count
148
142 -4.1%
Clocks
Base Clock
1665 MHz
1110 MHz
Boost Clock
2032 MHz
2520 MHz
Memory Clock
2000 MHz 8 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
288 GB
48 GB
VRAM (MB)
294,912
49,152 -83.3%
Memory Type
HBM3e
GDDR6
Memory Bus
8192 bit
384 bit
Bandwidth
8.19 TB/s
864.0 GB/s
Cache
L1 Cache
256 KB (per SM)
128 KB (per SM)
L2 Cache
126 MB
48 MB
Performance
Pixel Rate
48.77 GPixel/s
483.8 GPixel/s
Texture Rate
1,202.9 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
76.99 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
1,202.9 GFLOPS (1:64)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
76.99 TFLOPS (1:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
142
Tensor Cores
592
568 -4.1%
Power
TDP
1100 W
300 W
TDP (W)
1,100
300 -72.7%
Suggested PSU
1500 W
700 W
Power Connectors
1x 16-pin
Architecture
Architecture
Blackwell Ultra
Ada Lovelace
GPU Name
GB110
AD102
Generation
Server Blackwell (Bxx)
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
208,000 million
76,300 million
Die Size
1628 mm²
609 mm²
Foundry
TSMC
TSMC
Density
127.8M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
10.3
8.9
Shader Model
6.8
Physical
Slot Width
SXM Module
Dual-slot
Length
267 mm 10.5 inches
Height
111 mm 4.4 inches
Outputs
No outputs
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 6.0 x16
PCIe 4.0 x16
Other
Production
Active
End-of-life
Predecessor
Server Hopper
Server Ampere
Successor
Server Rubin
Server Hopper
View B300 SXM6 AC Details View L40S Details