NVIDIA L40
NVIDIA graphics card specifications and benchmark scores
At a Glance
NVIDIANVIDIA L40 Specifications
L40 GPU Core
Shader units and compute resources
The NVIDIA L40 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.
L40 Clock Speeds
GPU and memory frequencies
Clock speeds directly impact the L40's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The L40 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.
NVIDIA's L40 Memory
VRAM capacity and bandwidth
VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The L40's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.
L40 by NVIDIA Cache
On-chip cache hierarchy
On-chip cache provides ultra-fast data access for the L40, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.
L40 Theoretical Performance
Compute and fill rates
Theoretical performance metrics provide a baseline for comparing the NVIDIA L40 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.
L40 Ray Tracing & AI
Hardware acceleration features
The NVIDIA L40 includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the L40 capable of delivering both stunning graphics and smooth frame rates in modern titles.
Ada Lovelace Architecture & Process
Manufacturing and design details
The NVIDIA L40 is built on NVIDIA's Ada Lovelace architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the L40 will perform in GPU benchmarks compared to previous generations.
NVIDIA's L40 Power & Thermal
TDP and power requirements
Power specifications for the NVIDIA L40 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the L40 to maintain boost clocks without throttling.
L40 by NVIDIA Physical & Connectivity
Dimensions and outputs
Physical dimensions of the NVIDIA L40 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.
NVIDIA API Support
Graphics and compute APIs
API support determines which games and applications can fully utilize the NVIDIA L40. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.
L40 Product Information
Release and pricing details
The NVIDIA L40 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the L40 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.
L40 Benchmark Scores
geekbench_openclSource
Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA L40 handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms.
geekbench_vulkanSource
Geekbench Vulkan tests GPU compute using the modern low-overhead Vulkan API. This shows how NVIDIA L40 performs with next-generation graphics and compute workloads. Vulkan offers better CPU efficiency than older APIs like OpenGL. Modern games and applications increasingly use Vulkan for cross-platform GPU acceleration.
About NVIDIA L40
The NVIDIA L40 is a server-class GPU built on the Ada Lovelace architecture, using the AD102 chip fabricated at TSMC's 5 nm process. The die packs 76,300 million transistors across 609 mm², yielding a transistor density of 125.3 million per square millimeter. The card ships with 48 GB of GDDR6 memory on a 384-bit bus, delivering 864.0 GB/s of bandwidth. Benchmark results place the L40 in the 99th percentile of all GPUs, with an average score of 281,655 across Geekbench OpenCL and Vulkan workloads. The card connects via PCIe 4.0 x16 and outputs video through four DisplayPort 1.4a ports.
Who Should Consider It
The L40 targets server and workstation deployments, and the data reflects a card built for heavy compute. Its 99th-percentile standing means it outperforms nearly every other GPU in the database. The 48 GB frame buffer is the standout feature: with 864.0 GB/s of bandwidth, the card can handle texture-heavy workloads and large datasets without running out of memory.
Comparing to its nearest rivals, the L40 sits just 0.1% behind the RTX 6000 Ada Generation in average score (281,655 vs 281,932). That is a statistical tie. Anyone choosing between these two cards should weigh other factors — the benchmark data will not separate them.
Against the L40S, the L40 trails by 3.7% (292,603 vs 281,655). The L40S is the faster sibling. Against the L20, the L40 leads by 5.7% (281,655 vs 266,428). Against the H200 NVL, the L40 is 7.8% behind (281,655 vs 305,608).
The Geekbench OpenCL score of 330,683 versus the Vulkan score of 232,627 shows a significant gap between compute APIs — OpenCL workloads extract substantially more performance from this architecture.
For users running high-resolution rendering or large model inference, the 48 GB capacity and the 99th-percentile compute performance make the L40 a viable option. The card's dual-slot form factor and 267 mm length (10.5 inches) mean it fits in standard server chassis. The 111 mm height (4.4 inches) is a practical dimension for dense deployments.
Ray Tracing and Feature Set
The L40 includes 142 RT cores and 568 tensor cores. These are the dedicated hardware units for ray tracing and AI acceleration. The tensor core count is substantial — 568 tensor cores paired with 90.52 TFLOPS of FP16 compute (at a 1:1 ratio with FP32) suggests strong AI inference throughput.
On the API front, the L40 supports DirectX 12 Ultimate (feature level 12_2), OpenGL 4.6, and Vulkan 1.4. DirectX 12 Ultimate support means hardware ray tracing and mesh shaders are available in applications that use those features. Vulkan 1.4 is the latest iteration of that API, providing low-level access for compute-heavy workloads.
The pixel rate of 478.1 GPixel/s and texture rate of 1,414.3 GTexel/s are strong figures. The texture rate in particular — over 1.4 trillion texels per second — indicates the card can feed its 568 texture mapping units efficiently.
Benchmark Performance
The average benchmark score for the L40 is 281,655. This is derived from two Geekbench tests: OpenCL at 330,683 and Vulkan at 232,627. The OpenCL result is significantly higher than the Vulkan score, which suggests the card's compute architecture responds better to OpenCL's programming model.
The L40 runs at a base clock of 735 MHz with a boost clock of 2490 MHz. The 18,176 shading units, 568 TMUs, and 192 ROPs deliver a pixel rate of 478.1 GPixel/s and a texture rate of 1,414.3 GTexel/s. FP32 compute is rated at 90.52 TFLOPS, with FP16 matching at a 1:1 ratio.
Against the nearest rivals:
- RTX 6000 Ada Generation: 281,932 average score, 0.1% ahead of the L40. This is effectively a tie.
- L40S: 292,603 average score, 3.7% ahead of the L40. The L40S is the clear winner in this pairing.
- L20: 266,428 average score, 5.7% behind the L40. The L40 holds a comfortable lead.
- H200 NVL: 305,608 average score, 7.8% ahead of the L40. The H200 NVL is the strongest of the group.
The 99th-percentile ranking means the L40 outperforms 99% of all GPUs in the database. Even where it trails its closest rivals, it does so by single-digit percentages. The L40 is not the fastest card in its class, but it is firmly at the top end of the overall GPU landscape.
Power and Cooling
The L40 carries a 300 W TDP. NVIDIA recommends a 700 W power supply for systems using this card. Power is delivered through a single 16-pin connector.
The card is dual-slot, which is modest for a GPU with this compute capability. The dimensions are 267 mm in length (10.5 inches) and 111 mm in height (4.4 inches).
The 300 W TDP is worth noting relative to the performance. The L40 delivers 90.52 TFLOPS of FP32 compute within a 300 W envelope, which speaks to the efficiency of the 5 nm TSMC process.
FAQ
Q: How does the NVIDIA L40 compare to the RTX 6000 Ada Generation?
A: The two cards are effectively tied. The RTX 6000 Ada scores 281,932 on average, just 0.1% higher than the L40's 281,655. Benchmark results will not separate them.
Q: What is the difference between the L40 and the L40S?
A: The L40S is 3.7% faster on average, scoring 292,603 versus the L40's 281,655. The L40S is the higher-performing variant in this pairing.
Q: How much memory does the L40 have and what type is it?
A: The L40 has 48 GB of GDDR6 memory on a 384-bit bus, with 864.0 GB/s of bandwidth and an effective memory speed of 18 Gbps.
Q: What APIs does the L40 support?
A: The L40 supports DirectX 12 Ultimate (feature level 12_2), OpenGL 4.6, and Vulkan 1.4.
Q: What power supply is recommended for the L40?
A: NVIDIA suggests a 700 W power supply. The card itself has a 300 W TDP and uses a single 16-pin power connector.
Q: Is the L40 still in production?
A: No, the production status is listed as end-of-life. It was released on 2022-10-12, with the Server Ampere generation as its predecessor and Server Hopper as its successor.
Memory Subsystem
The L40's memory subsystem is one of its defining features. It ships with 48 GB of GDDR6 memory, which is a substantial capacity for server workloads. The memory clock runs at 2250 MHz, with an effective data rate of 18 Gbps.
The 384-bit memory bus is wide, and combined with the 18 Gbps effective speed, it yields 864.0 GB/s of bandwidth. This is a high-bandwidth configuration that supports high-resolution rendering and large dataset processing.
For high-resolution workloads, the combination of 48 GB capacity and 864 GB/s bandwidth means the card can hold entire scenes in VRAM without spilling to system memory. The 99th-percentile overall performance ranking is supported by this memory configuration.
The pixel rate of 478.1 GPixel/s and texture rate of 1,414.3 GTexel/s are also relevant here — they indicate how fast the card can fill frames and sample textures, which matters at high resolutions. The 192 ROPs handle pixel output, while the 568 TMUs handle texture sampling.
The FP16 performance of 90.52 TFLOPS (at a 1:1 ratio with FP32) is notable for mixed-precision workloads, and the 568 tensor cores provide the AI acceleration infrastructure that pairs with the memory capacity for large model inference.
The AMD Equivalent of L40
Looking for a similar graphics card from AMD? The AMD Radeon RX 7900 XTX offers comparable performance and features in the AMD lineup.
Popular NVIDIA L40 Comparisons
See how the L40 stacks up against similar graphics cards from the same generation and competing brands.
Compare L40 with Other GPUs
Select another GPU to compare specifications and benchmarks side-by-side.
Browse GPUs