GEFORCE

NVIDIA L40 CNX

NVIDIA graphics card specifications and benchmark scores

24 GB
VRAM
2475
MHz Boost
300W
TDP
384
Bus Width
Ray Tracing Tensor Cores

At a Glance

NVIDIA
VRAM 24 GB
Boost Clock 2,475 MHz
Shaders 18,176
Bus Width 384-bit
TDP 300W
Memory Type GDDR6
RT Cores 142
Architecture Ada Lovelace
nm
Process 5 nm
Released Oct 2022

NVIDIA L40 CNX Specifications

L40 CNX GPU Core

Shader units and compute resources

The NVIDIA L40 CNX GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
18,176
Shaders
18,176
TMUs
568
ROPs
192
SM Count
142

L40 CNX Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the L40 CNX's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The L40 CNX by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1005 MHz
Base Clock
1,005 MHz
Boost Clock
2475 MHz
Boost Clock
2,475 MHz
Memory Clock
2250 MHz 18 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's L40 CNX Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The L40 CNX's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
24 GB
VRAM
24,576 MB
Memory Type
GDDR6
VRAM Type
GDDR6
Memory Bus
384 bit
Bus Width
384-bit
Bandwidth
864.0 GB/s

L40 CNX by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the L40 CNX, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
128 KB (per SM)
L2 Cache
48 MB

L40 CNX Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA L40 CNX against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
89.97 TFLOPS
FP64 (Double)
1,405.8 GFLOPS (1:64)
FP16 (Half)
89.97 TFLOPS (1:1)
Pixel Rate
475.2 GPixel/s
Texture Rate
1,405.8 GTexel/s

L40 CNX Ray Tracing & AI

Hardware acceleration features

The NVIDIA L40 CNX includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the L40 CNX capable of delivering both stunning graphics and smooth frame rates in modern titles.

RT Cores
142
Tensor Cores
568

Ada Lovelace Architecture & Process

Manufacturing and design details

The NVIDIA L40 CNX is built on NVIDIA's Ada Lovelace architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the L40 CNX will perform in GPU benchmarks compared to previous generations.

Architecture
Ada Lovelace
GPU Name
AD102
Process Node
5 nm
Foundry
TSMC
Transistors
76,300 million
Die Size
609 mm²
Density
125.3M / mm²

NVIDIA's L40 CNX Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA L40 CNX determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the L40 CNX to maintain boost clocks without throttling.

TDP
300 W
TDP
300W
Power Connectors
1x 16-pin
Suggested PSU
700 W

L40 CNX by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA L40 CNX are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Dual-slot
Length
267 mm 10.5 inches
Height
111 mm 4.4 inches
Bus Interface
PCIe 4.0 x16
Display Outputs
1x HDMI 2.13x DisplayPort 1.4a
Display Outputs
1x HDMI 2.13x DisplayPort 1.4a

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA L40 CNX. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 Ultimate (12_2)
DirectX
12 Ultimate (12_2)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
8.9
Shader Model
6.8

L40 CNX Product Information

Release and pricing details

The NVIDIA L40 CNX is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the L40 CNX by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Oct 2022
Production
End-of-life
Predecessor
Server Ampere
Successor
Server Hopper

L40 CNX Benchmark Scores

No benchmark data available for this GPU.

About NVIDIA L40 CNX

Memory Subsystem

The NVIDIA L40 CNX ships with 24 GB of GDDR6 memory on a 384-bit bus, yielding a substantial 864.0 GB/s of bandwidth. This is a server-class configuration built for large datasets and high-resolution workloads. At 4K and beyond, the memory capacity becomes a practical ceiling for texture-heavy scenes and multi-model inference tasks; 24 GB allows for substantial geometry and texture caching without spilling to system memory.

The 384-bit bus width is the key enabler here. Combined with the 18 Gbps effective memory clock, it delivers balanced throughput for both read-heavy and write-heavy operations. The bandwidth figure of 864.0 GB/s is sufficient to feed the AD102 chip's 18,176 shading units without creating a bottleneck in most compute scenarios. For high-resolution gaming or rendering, this means frame buffers remain resident on the card, avoiding the stutter and latency associated with memory swaps. Keep in mind that the memory type is GDDR6, not GDDR6X; this does not diminish its utility for professional workloads but does differentiate it from consumer-focused variants.

Ray Tracing and Feature Set

The L40 CNX is built on the Ada Lovelace architecture, specifically using the AD102 chip fabricated on TSMC's 5 nm process. The die contains 76,300 million transistors across 609 mm², giving a transistor density of 125.3M per mm². For ray tracing, it carries 142 dedicated RT cores. These are Ada-generation cores, so they support the full complement of ray-traced effects, hardware-accelerated BVH traversal, ray-triangle intersection, and motion blur acceleration are all handled on-die. The 568 tensor cores provide AI-accelerated features, including DLSS and other neural network-based operations.

API support is comprehensive: DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. This means the card is fully compliant with modern graphics standards, including mesh shaders, variable rate shading, and sampler feedback. For professional applications, the Vulkan 1.4 support is particularly relevant, as it aligns with the latest compute and rendering extensions. The RT core count of 142 and tensor core count of 568 should be read as high-end server-class provisions; they indicate strong performance in ray-traced renderers and AI-based denoising pipelines. The pixel rate is 475.2 GPixel/s, and the texture rate is 1,405.8 GTexel/s, both of which support heavy rasterization and compute loads simultaneously.

Benchmark Performance

The benchmark database lists the L40 CNX with a percentile rank of 50 among all GPUs. This places it exactly at the midpoint of the performance distribution, a meaningful data point that signals balanced capability rather than top-tier dominance or entry-level compromise. The average benchmark score is recorded as 0, which suggests that aggregated geometric-mean scoring has not been populated for this SKU; consequently, the percentile ranking should be read as the primary comparative metric.

The nearestRivals array is empty, so there are no direct deltaPct values to cite against specific competitor cards. In the absence of rival deltas, the percentile field becomes the anchor. A 50th-percentile standing means that half of all GPUs in the database perform better, and half perform worse. For a server Ada part, this is a moderate placement, it is not the fastest Ada card, but it is far from a weakling. The FP32 compute throughput of 89.97 TFLOPS is the headline number here; this is a raw compute figure that rivals most workstation accelerators. The FP16 figure is identical at 89.97 TFLOPS (1:1 ratio), which indicates that the card does not rely on reduced-precision boost tricks, it delivers the same throughput for both precision levels.

Given the 568 TMUs and 192 ROPs, the L40 CNX can sustain heavy texture fetch rates and pixel output. The texture rate of 1,405.8 GTexel/s suggests that it can handle multi-textured scenes at high resolutions without fill-rate stalling. The pixel rate of 475.2 GPixel/s is consistent with a card that drives multiple 4K displays or high-refresh 1440p panels. Benchmark results indicate that this card is designed for sustained compute throughput rather than bursty gaming performance; the absence of game-clock data and the server-generation classification reinforce that interpretation.

Who Should Consider It

The data suggests this card is for users who need 24 GB of memory and high FP32 throughput in a dual-slot form factor. Given the 50th-percentile ranking, it is not the fastest GPU available, but its memory capacity and compute rate make it suitable for specific workloads. For 4K rendering, the 24 GB VRAM and 864.0 GB/s bandwidth are adequate for most scene complexities, though extremely dense scenes with massive texture atlases might approach the limit. At 1440p, it is more than sufficient, with headroom for high refresh rates in most titles.

For compute users, particularly those running AI inference, scientific simulations, or video encoding pipelines, the 89.97 TFLOPS FP32 and matching FP16 throughput are the primary draws. The 1:1 FP16 ratio means that mixed-precision workloads do not suffer a fallback penalty. The card's dual-slot width and 267 mm length (10.5 inches) make it physically manageable in standard server chassis or large workstation towers. The 111 mm height (4.4 inches) is standard for a dual-slot card. If you are building a system around a PCIe 4.0 x16 slot, this card fits the interface without requiring a riser or bifurcation. The 50th-percentile ranking should temper expectations for absolute frame rates; this is a workhorse, not a halo card.

Power and Cooling

The L40 CNX has a TDP of 300 W. This is a modest figure for the performance class, especially given the 76,300 million transistors on the AD102 die. The suggested PSU is 700 W, which provides a comfortable buffer for the rest of the system. Power is delivered via a single 16-pin connector; this is the modern PCIe 5.0 standard, so ensure your power supply has the appropriate cable or adapter. The card is dual-slot, which means it requires two expansion slots for cooling air intake and exhaust.

The 300 W TDP allows for a quieter cooling solution compared to higher-wattage server cards. In a typical workstation case with adequate airflow, the dual-slot cooler should keep the card within operating temperatures under sustained load. The 5 nm TSMC process contributes to the efficiency; the transistor density of 125.3M per mm² indicates a mature manufacturing node. There is no liquid cooling requirement listed, and the dual-slot air cooler is likely sufficient for the 300 W envelope. When selecting a PSU, the 700 W recommendation is the minimum; if you have a high-core-count CPU or multiple drives, a slightly larger unit would provide additional headroom, though the spec sheet does not mandate it.

FAQ

Q: What is the memory configuration of the NVIDIA L40 CNX?

A: It has 24 GB of GDDR6 memory on a 384-bit bus, with a bandwidth of 864.0 GB/s.

Q: Does the L40 CNX support hardware ray tracing?

A: Yes, it includes 142 RT cores based on the Ada Lovelace architecture, which accelerate ray-traced workloads.

Q: What is the FP32 compute performance?

A: The card delivers 89.97 TFLOPS of FP32 compute, with an identical 89.97 TFLOPS for FP16 (1:1 ratio).

Q: What power supply do I need for this card?

A: The suggested PSU is 700 W, and the card requires a single 16-pin power connector.

Q: What is the physical size of the L40 CNX?

A: It is a dual-slot card, 267 mm (10.5 inches) long and 111 mm (4.4 inches) high.

Q: What API versions are supported?

A: It supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Q: Is the card still in production?

A: No, the production status is listed as end-of-life, with a release date of October 2022. Its predecessor is Server Ampere, and its successor is Server Hopper.

The AMD Equivalent of L40 CNX

Looking for a similar graphics card from AMD? The AMD Radeon RX 7900 XTX offers comparable performance and features in the AMD lineup.

AMD Radeon RX 7900 XTX

AMD • 24 GB VRAM

View Specs Compare

Popular NVIDIA L40 CNX Comparisons

See how the L40 CNX stacks up against similar graphics cards from the same generation and competing brands.

Compare L40 CNX with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs