NVIDIA L40S vs NVIDIA RTX 4000 SFF Ada Generation Comparison

NVIDIA
GEFORCE

NVIDIA L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022
VS
NVIDIA
GEFORCE

RTX 4000 SFF Ada Generation

CORE STATE AD104
VRAM 20 GB
CLOCK SPEED 1560 MHz
TDP 70 W
BUS WIDTH 160 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

geekbench_opencl
330,727
124,812
geekbench_vulkan
260,799
109,364

Analysis: NVIDIA L40S vs NVIDIA RTX 4000 SFF Ada Generation

Head-to-Head Benchmarks

The benchmark data leaves no ambiguity about the performance hierarchy between these two Ada Lovelace cards. In the Geekbench OpenCL test, the NVIDIA L40S scores 330,727 points, while the NVIDIA RTX 4000 SFF Ada Generation scores 124,812 points. That represents a delta of 165%, meaning the L40S more than doubles the SFF card's compute throughput in this workload. The Vulkan results tell a similar story, though with a slightly narrower margin: the L40S posts 260,799 points against the RTX 4000 SFF's 109,364, a 138.5% advantage.

The record shows the L40S wins both head-to-head benchmarks, with a 2-0 overall tally. But the shape of those wins matters. The OpenCL gap is larger than the Vulkan gap, which suggests the L40S's advantage scales differently depending on the API and workload characteristics. OpenCL tends to expose raw compute resources more directly, and the L40S has vastly more of those resources available. Vulkan, being a lower-level API, can sometimes mask architectural differences, yet even there the L40S holds a commanding lead.

Contextualizing these scores against the nearest rivals in the database adds depth. The L40S's average benchmark score is 295,763, placing it in the 99th percentile of all GPUs tracked. Its closest competitor, the NVIDIA RTX 6000 Ada Generation, averages 287,237, a 3% gap in the L40S's favor. The NVIDIA L40 sits at 284,111, 4.1% behind. Looking the other direction, the AMD Instinct MI300X averages 317,994, which is 7% ahead of the L40S, and the NVIDIA H200 NVL reaches 334,891, 11.7% ahead. So the L40S is not the absolute fastest accelerator in the database, but it sits comfortably near the top, within striking distance of much larger and more power-hungry parts.

The RTX 4000 SFF, by contrast, averages 117,088 across its benchmark suite, landing in the 95th percentile. Its nearest rivals cluster tightly around it: the NVIDIA GB10 averages 117,393 (0.3% ahead), the AMD Radeon PRO W7700 averages 118,976 (1.6% ahead), the NVIDIA Tesla V100 SXM2 16 GB averages 114,395 (2.4% behind), and the NVIDIA RTX A5500 Mobile averages 113,944 (2.8% behind). This is a dense pack of mid-range performers, and the RTX 4000 SFF holds its own within it. The data shows a clear stratification: the L40S operates in a performance tier roughly 2.5 times higher than the RTX 4000 SFF, while the SFF card competes in a much more crowded and closely matched segment.

FAQ

Q: Which GPU has the higher raw compute throughput?

A: The NVIDIA L40S delivers 91.61 TFLOPS of FP32 performance, while the NVIDIA RTX 4000 SFF Ada Generation delivers 19.17 TFLOPS. The L40S is approximately 4.8 times faster in this metric.

Q: How do their memory subsystems compare?

A: The L40S has 48 GB of GDDR6 memory on a 384-bit bus, yielding 864.0 GB/s of bandwidth. The RTX 4000 SFF has 20 GB of GDDR6 on a 160-bit bus, yielding 280.0 GB/s. The L40S offers 2.4 times the capacity and roughly 3.1 times the bandwidth.

Q: Are these GPUs from the same architecture generation?

A: Yes, both use the Ada Lovelace architecture and are fabricated on TSMC's 5 nm process. However, they use different chips: the L40S uses the AD102 die (609 mm², 76,300 million transistors), while the RTX 4000 SFF uses the AD104 die (294 mm², 35,800 million transistors).

Q: What is the power consumption difference?

A: The L40S has a TDP of 300 W and requires a 1x 16-pin power connector with a suggested 700 W power supply. The RTX 4000 SFF has a TDP of 70 W, requires no power connectors, and suggests a 250 W power supply. The SFF card consumes about 23% of the L40S's power budget.

Q: Which GPU has better API support?

A: Both cards support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. There is no difference in API feature levels between the two.

Q: What are the physical size differences?

A: The L40S is 267 mm long and 111 mm tall. The RTX 4000 SFF is 168 mm long and 69 mm tall. Both are dual-slot cards, but the SFF is dramatically shorter and lower-profile.

Architecture Differences

Both GPUs share the Ada Lovelace architecture and the 5 nm TSMC process, but they diverge significantly in die size and resource allocation. The L40S uses the AD102 chip, which is NVIDIA's largest Ada die at 609 mm² and houses 76,300 million transistors. The RTX 4000 SFF uses AD104, a much smaller die at 294 mm² with 35,800 million transistors. Transistor density is similar, 125.3M per mm² for the L40S versus 121.8M per mm² for the RTX 4000 SFF, indicating the process node is being utilized at comparable efficiency on both parts.

The compute resource disparity is stark. The L40S fields 18,176 shading units, 568 texture mapping units, and 192 ROPs. The RTX 4000 SFF has 6,144 shading units, 192 TMUs, and 64 ROPs. That is roughly 3 times more shading units, 3 times more TMUs, and 3 times more ROPs on the L40S. The ray tracing and tensor core counts follow the same ratio: 142 RT cores and 568 tensor cores on the L40S versus 48 RT cores and 192 tensor cores on the SFF card.

These resource ratios translate directly into throughput metrics. The L40S achieves a pixel rate of 483.8 GPixel/s and a texture rate of 1,431.4 GTexel/s. The RTX 4000 SFF manages 99.84 GPixel/s and 299.5 GTexel/s. The L40S is roughly 4.8 times faster in pixel throughput and 4.8 times faster in texture throughput, which aligns closely with the shading unit ratio. FP16 performance mirrors FP32 on both cards at a 1:1 ratio, so neither has a dedicated half-precision advantage over the other.

Clock speeds tell an interesting counter-narrative. The L40S runs at a base clock of 1110 MHz and boosts to 2520 MHz. The RTX 4000 SFF runs at a much more conservative 720 MHz base and 1560 MHz boost. Despite the SFF card's lower clocks, the L40S's massive resource advantage overwhelms any clock-rate differences. The L40S compensates for its higher clocks with a 300 W TDP, while the SFF card's modest 70 W TDP explains its restrained clock profile. The SFF card is clearly engineered for thermal and power constraints, not peak performance.

Specification Differences

The two cards differ across nearly every measurable specification, reflecting their distinct design intents. Memory capacity is 48 GB on the L40S versus 20 GB on the RTX 4000 SFF. Memory type is GDDR6 for both, but the bus width differs substantially: 384-bit versus 160-bit. Effective memory speed is 18 Gbps on the L40S versus 14 Gbps on the SFF card. The resulting bandwidth gap is 864.0 GB/s versus 280.0 GB/s.

Clock specifications diverge as discussed: the L40S has a 1110 MHz base and 2520 MHz boost, while the RTX 4000 SFF has a 720 MHz base and 1560 MHz boost. The L40S's memory runs at 2250 MHz (18 Gbps effective), while the SFF card's memory runs at 1750 MHz (14 Gbps effective).

Power delivery is another major differentiator. The L40S draws 300 W and requires a 16-pin power connector, with a suggested power supply of 700 W. The RTX 4000 SFF draws only 70 W, requires no external power connector, and suggests a 250 W power supply. The physical dimensions reflect this: the L40S is 267 mm long and 111 mm tall, while the SFF card is 168 mm long and 69 mm tall.

Display outputs also differ. The L40S offers 1x HDMI 2.1 and 3x DisplayPort 1.4a. The RTX 4000 SFF offers 4x mini-DisplayPort 1.4a. Neither card includes a VGA or DVI output. Bus interface is identical: PCIe 4.0 x16 for both. Production status separates them further: the L40S is marked end-of-life, while the RTX 4000 SFF is active. Release dates are October 12, 2022 for the L40S and March 20, 2023 for the SFF card.

The Verdict

The data supports a clear division of purpose. The NVIDIA L40S is a high-throughput compute accelerator designed for workloads that demand maximum resources: 48 GB of memory, 864.0 GB/s of bandwidth, and 91.61 TFLOPS of FP32 compute. Its 99th percentile standing among all GPUs and its position near the top of its nearest rival group, within 11.7% of the much larger H200 NVL, confirm that it belongs in the upper echelon of server accelerators. The 165% and 138.5% benchmark leads over the RTX 4000 SFF are not marginal improvements; they represent a completely different performance class.

The NVIDIA RTX 4000 SFF Ada Generation is engineered for the opposite constraint set. Its 70 W TDP, lack of power connectors, and compact dimensions (168 mm length, 69 mm height) make it suitable for space-constrained and power-limited environments. Its 95th percentile ranking shows it is still a capable performer in absolute terms, and its nearest rivals cluster within a 2.8% band, indicating it is competitively positioned in its segment. But the benchmark results show it cannot approach the L40S's raw throughput.

For a buyer choosing between these two, the decision hinges on workload scale and physical constraints. If the task requires large model residency, massive memory bandwidth, or sustained compute throughput, the L40S is the only logical choice from this pairing. If the installation demands a low-profile, low-power card that can operate without external power cabling, the RTX 4000 SFF is the appropriate pick. There is no middle ground in the data: the L40S wins every benchmark by a wide margin, and the SFF card wins on power efficiency, physical footprint, and installation flexibility.

Where Each One Wins

The L40S wins decisively in raw compute performance. Its FP32 throughput of 91.61 TFLOPS is 4.8 times the SFF card's 19.17 TFLOPS. Its texture rate of 1,431.4 GTexel/s is 4.8 times higher, and its pixel rate of 483.8 GPixel/s is 4.8 times higher. Memory bandwidth of 864.0 GB/s is 3.1 times the SFF card's 280.0 GB/s. The L40S also offers more than double the memory capacity, 48 GB versus 20 GB, which is critical for large datasets and models that exceed the SFF card's capacity. In the head-to-head benchmark suite, the L40S wins both tests with deltas of 165% (OpenCL) and 138.5% (Vulkan).

The RTX 4000 SFF wins on power and physical integration. Its 70 W TDP is less than a quarter of the L40S's 300 W, and it requires no external power connector, meaning it can be installed in systems without dedicated GPU power cabling. Its suggested power supply of 250 W versus the L40S's 700 W makes it viable in pre-existing workstations with modest PSUs. Its dimensions, 168 mm by 69 mm, are dramatically smaller than the L40S's 267 mm by 111 mm, enabling installation in compact chassis where the L40S would not physically fit. The SFF card also offers four mini-DisplayPort outputs versus the L40S's single HDMI and three DisplayPort connections, which may be advantageous for multi-display setups in constrained spaces.

The RTX 4000 SFF's production status is active, while the L40S is marked end-of-life. For organizations planning long-term deployments, the SFF card has a clearer availability path. However, the L40S's successor is listed as Server Hopper, indicating NVIDIA's roadmap moves upward rather than toward the SFF segment. The SFF card's successor is Blackwell PRO W, which suggests the workstation-oriented lineage continues. The data ultimately points to complementary roles: the L40S for maximum compute density and the RTX 4000 SFF for maximum deployment flexibility.

DETAILED SPECIFICATIONS

SPECIFICATION
L40S
RTX 4000 SFF Ada Generation
Core Specs
Shading Units
18,176
6,144 -66.2%
Shaders
18,176
6,144 -66.2%
TMUs
568
192 -66.2%
ROPs
192
64 -66.7%
SM Count
142
48 -66.2%
Clocks
Base Clock
1110 MHz
720 MHz
Boost Clock
2520 MHz
1560 MHz
Memory Clock
2250 MHz 18 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
48 GB
20 GB
VRAM (MB)
49,152
20,480 -58.3%
Memory Type
GDDR6
GDDR6
Memory Bus
384 bit
160 bit
Bandwidth
864.0 GB/s
280.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
48 MB
48 MB
Performance
Pixel Rate
483.8 GPixel/s
99.84 GPixel/s
Texture Rate
1,431.4 GTexel/s
299.5 GTexel/s
FP32 (TFLOPS)
91.61 TFLOPS
19.17 TFLOPS
FP64 (TFLOPS)
1,431.4 GFLOPS (1:64)
299.5 GFLOPS (1:64)
FP16 (TFLOPS)
91.61 TFLOPS (1:1)
19.17 TFLOPS (1:1)
AI/RT
RT Cores
142
48 -66.2%
Tensor Cores
568
192 -66.2%
Power
TDP
300 W
70 W
TDP (W)
300
70 -76.7%
Suggested PSU
700 W
250 W
Power Connectors
1x 16-pin
None
Architecture
Architecture
Ada Lovelace
Ada Lovelace
GPU Name
AD102
AD104
Generation
Server Ada (Lxx)
Workstation Ada (x000A)
Process Size
5 nm
5 nm
Transistors
76,300 million
35,800 million
Die Size
609 mm²
294 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
121.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
168 mm 6.6 inches
Height
111 mm 4.4 inches
69 mm 2.7 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
4x mini-DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
Active
Predecessor
Server Ampere
Workstation Ampere
Successor
Server Hopper
Blackwell PRO W
View L40S Details View RTX 4000 SFF Ada Generation Details