NVIDIA L40S
NVIDIA graphics card specifications and benchmark scores
At a Glance
NVIDIANVIDIA L40S Specifications
GPU Core
Shader units and compute resources
The NVIDIA L40S GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.
L40S Clock Speeds
GPU and memory frequencies
Clock speeds directly impact the L40S's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The L40S by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.
NVIDIA's L40S Memory
VRAM capacity and bandwidth
VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The L40S's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.
L40S by NVIDIA Cache
On-chip cache hierarchy
On-chip cache provides ultra-fast data access for the L40S, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.
L40S Theoretical Performance
Compute and fill rates
Theoretical performance metrics provide a baseline for comparing the NVIDIA L40S against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.
L40S Ray Tracing & AI
Hardware acceleration features
The NVIDIA L40S includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the L40S capable of delivering both stunning graphics and smooth frame rates in modern titles.
Ada Lovelace Architecture & Process
Manufacturing and design details
The NVIDIA L40S is built on NVIDIA's Ada Lovelace architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the L40S will perform in GPU benchmarks compared to previous generations.
Power & Thermal
TDP and power requirements
Power specifications for the NVIDIA L40S determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the L40S to maintain boost clocks without throttling.
L40S by NVIDIA Physical & Connectivity
Dimensions and outputs
Physical dimensions of the NVIDIA L40S are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.
NVIDIA API Support
Graphics and compute APIs
API support determines which games and applications can fully utilize the NVIDIA L40S. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.
L40S Product Information
Release and pricing details
The NVIDIA L40S is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the L40S by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.
About NVIDIA L40S
The NVIDIA L40S is a server GPU in NVIDIA's Ada Lovelace product line, built on the AD102 chip. TSMC manufactures the die on a 5 nm process, packing 76,300 million transistors into 609 mm² for a density of 125.3M per mm². The board carries 48 GB of GDDR6 memory across a 384-bit bus, producing 864.0 GB/s of bandwidth. The base clock is 1110 MHz with a 2520 MHz boost; the memory clock is 2250 MHz, or 18 Gbps effective. Benchmark aggregation places the L40S at 292603 average score, which is the 100th percentile among all GPUs in the database. The L40S is end-of-life, and its product predecessors and successors are Server Ampere and Server Hopper.
Power and Cooling
The L40S is specified with a TDP of 300 W. The suggested power supply is 700 W, and the card uses a single 16-pin power connector. Mechanically, the card is dual-slot, 267 mm long (10.5 inches), and 111 mm tall (4.4 inches). It connects via PCIe 4.0 x16. The power and mechanical envelope is consistent with a server accelerator that can be installed in a dual-slot server chassis. The display outputs are 1x HDMI 2.1 and 3x DisplayPort 1.4a, which gives the card optional display connectivity despite its server orientation.
Ray Tracing and Feature Set
The architecture is Ada Lovelace, with 142 RT cores, 568 tensor cores, 18,176 shading units, 568 texture mapping units, and 192 ROPs. The pixel rate is 483.8 GPixel/s and the texture rate is 1,431.4 GTexel/s. FP32 and FP16 compute are both rated at 91.61 TFLOPS, at a 1:1 ratio, so there is no FP16 rate penalty in this data. API support includes DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The hardware and API list together cover ray-traced workloads and general-purpose compute paths.
Who Should Consider It
The L40S is a candidate for compute-heavy server environments. Its 48 GB frame buffer and 864.0 GB/s bandwidth allow large data sets to stay local to the GPU. The Geekbench OpenCL score of 334437 and the Geekbench Vulkan score of 250769 are both high, but OpenCL is the stronger result. In the aggregate, average score 292603 and 100th percentile rank put the card at the top of the database, with only the H200 NVL in the nearest rival list scoring higher. Workloads that can use 91.61 TFLOPS of FP32 or FP16 throughput will extract the most from the hardware. The 1:1 FP16 ratio means FP16-heavy models do not lose throughput relative to FP32. For system builders, the 300 W TDP, 700 W PSU recommendation, dual-slot width, and 267 mm length define the integration constraints. The 16-pin connector is also a system design consideration.
How It Compares
Compared to the NVIDIA RTX 6000 Ada Generation, the L40S posts an average score of 292603 against 281932. The delta is 3.8% in favor of the L40S. The two cards are close, and the L40S leads by a small margin. This puts them in the same performance band, with the L40S at the upper edge.
Against the NVIDIA L40, the L40S scores 292603 versus 281655. The delta is 3.9%, a shade larger than the gap to the RTX 6000 Ada Generation. Both are L-series server cards, and the L40S is the faster of the two in aggregate.
Relative to the NVIDIA H200 NVL, the L40S average score of 292603 trails 305608. The deltaPct is -4.3%, meaning the H200 NVL is ahead. This is the only nearest rival with a higher average score. Despite the 100th percentile all-GPU rank, the L40S does not lead this particular comparison.
Against the NVIDIA L20, the L40S has an average score of 292603 versus 266428. The delta is 9.8%, the largest margin among the four nearest rivals. The L20 is clearly behind in aggregate benchmark performance.
Benchmark Performance
The L40S aggregate benchmark score of 292603 is built from a Geekbench OpenCL result of 334437 and a Geekbench Vulkan result of 250769. The OpenCL number is meaningfully higher than the Vulkan number, which is useful when selecting an API for compute workloads. In the nearest rival set, the L40S is 3.8% above the RTX 6000 Ada Generation's 281932 and 3.9% above the L40's 281655. It is 9.8% above the L20's 266428. The lone negative delta is against the H200 NVL, whose 305608 average score is 4.3% higher than the L40S. The distribution of scores shows a tight cluster at the top: the deltas of 3.8% and 3.9% to the L40S imply that the RTX 6000 Ada Generation and the L40 are close to each other as well, while the L20 sits further down. The H200 NVL separates upward, making it the performance reference in this group. At the architecture level, the 91.61 TFLOPS FP32 and FP16 rates, 483.8 GPixel/s pixel rate, and 1,431.4 GTexel/s texture rate are consistent with a high-end compute accelerator. Memory bandwidth of 864.0 GB/s and capacity of 48 GB support large working sets. The combined data points to a GPU that is at the top of the database, with one nearest rival ahead in average score.
Detailed benchmark scores and charts for the NVIDIA L40S are below.
Benchmark Scores
geekbench_openclSource
Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA L40S handles parallel computing tasks like video encoding and scientific simulations.
geekbench_vulkanSource
Geekbench Vulkan tests GPU compute using the modern low-overhead Vulkan API. This shows how NVIDIA L40S performs with next-generation graphics and compute workloads. Vulkan offers better CPU efficiency than older APIs like OpenGL.
Popular NVIDIA L40S Comparisons
See how the L40S stacks up against similar graphics cards from the same generation and competing brands.
Compare with Other GPUs
Select another GPU to compare specifications and benchmarks side-by-side.
Browse GPUs