GEFORCE

NVIDIA L4

NVIDIA graphics card specifications and benchmark scores

24 GB
VRAM
2040
MHz Boost
72W
TDP
192
Bus Width
Ray Tracing Tensor Cores

At a Glance

NVIDIA
VRAM 24 GB
Boost Clock 2,040 MHz
Shaders 7,424
Bus Width 192-bit
TDP 72W
Memory Type GDDR6
RT Cores 60
Architecture Ada Lovelace
nm
Process 5 nm
Released Mar 2023

NVIDIA L4 Specifications

L4 GPU Core

Shader units and compute resources

The NVIDIA L4 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
7,424
Shaders
7,424
TMUs
240
ROPs
80
SM Count
60

L4 Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the L4's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The L4 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
795 MHz
Base Clock
795 MHz
Boost Clock
2040 MHz
Boost Clock
2,040 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's L4 Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The L4's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
24 GB
VRAM
24,576 MB
Memory Type
GDDR6
VRAM Type
GDDR6
Memory Bus
192 bit
Bus Width
192-bit
Bandwidth
300.1 GB/s

L4 by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the L4, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
128 KB (per SM)
L2 Cache
48 MB

L4 Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA L4 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
30.29 TFLOPS
FP64 (Double)
473.3 GFLOPS (1:64)
FP16 (Half)
30.29 TFLOPS (1:1)
Pixel Rate
163.2 GPixel/s
Texture Rate
489.6 GTexel/s

L4 Ray Tracing & AI

Hardware acceleration features

The NVIDIA L4 includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the L4 capable of delivering both stunning graphics and smooth frame rates in modern titles.

RT Cores
60
Tensor Cores
240

Ada Lovelace Architecture & Process

Manufacturing and design details

The NVIDIA L4 is built on NVIDIA's Ada Lovelace architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the L4 will perform in GPU benchmarks compared to previous generations.

Architecture
Ada Lovelace
GPU Name
AD104
Process Node
5 nm
Foundry
TSMC
Transistors
35,800 million
Die Size
294 mm²
Density
121.8M / mm²

NVIDIA's L4 Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA L4 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the L4 to maintain boost clocks without throttling.

TDP
72 W
TDP
72W
Power Connectors
None
Suggested PSU
250 W

L4 by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA L4 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Single-slot
Length
169 mm 6.7 inches
Height
56 mm 2.2 inches
Bus Interface
PCIe 4.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA L4. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 Ultimate (12_2)
DirectX
12 Ultimate (12_2)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
8.9
Shader Model
6.8

L4 Product Information

Release and pricing details

The NVIDIA L4 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the L4 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Mar 2023
Production
Active
Predecessor
Server Ampere
Successor
Server Hopper

L4 Benchmark Scores

geekbench_openclSource

Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA L4 handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms.

geekbench_opencl #54 of 643
140,838
36%
Max: 388,405
Compare with other GPUs

geekbench_vulkanSource

Geekbench Vulkan tests GPU compute using the modern low-overhead Vulkan API. This shows how NVIDIA L4 performs with next-generation graphics and compute workloads. Vulkan offers better CPU efficiency than older APIs like OpenGL. Modern games and applications increasingly use Vulkan for cross-platform GPU acceleration.

geekbench_vulkan #60 of 444
121,306
32%
Max: 376,915

About NVIDIA L4

The NVIDIA L4 is a single-slot server accelerator built on the Ada Lovelace architecture, featuring 24 GB of GDDR6 memory and a 72 W TDP. Its average benchmark score of 128,665 places it in the 97th percentile of all GPUs, a strong showing for a card that requires no auxiliary power connectors. Against its nearest rivals, the L4 lands in a tight cluster: it trails the GeForce RTX 3090 Ti by 2.5%, the Radeon PRO W6800 by 3.7%, and the Radeon RX 9070 GRE by 4.3%, while leading the Radeon PRO W7700 by 4.2%. These margins are narrow, making the L4 a competitive option in its performance tier.

How It Compares

The L4 sits 2.5% behind the NVIDIA GeForce RTX 3090 Ti, which posts an average score of 131,911. That is a remarkably small gap given the L4’s 72 W TDP and single-slot design, and it suggests the Ada Lovelace architecture extracts a high level of performance from a much lower power envelope. In practice, the L4 will deliver near-parity with the 3090 Ti in many compute workloads, though the 3090 Ti retains a slight edge in aggregate benchmarks.

Against the AMD Radeon PRO W6800, the L4 is 3.7% slower, with the PRO W6800 averaging 133,588. This is another close result, and the L4’s 24 GB memory capacity matches the PRO W6800’s professional positioning. The data shows a meaningful but not decisive advantage for the AMD card, and the L4’s power efficiency may offset that gap in dense server deployments.

The L4 is 4.2% faster than the AMD Radeon PRO W7700, which averages 123,434. This is the L4’s clearest win among its nearest rivals. The PRO W7700 is a capable workstation card, but the L4’s higher average score and larger memory capacity give it an edge in memory-intensive tasks. The 4.2% delta is consistent across the benchmark suite, indicating a stable performance advantage.

The AMD Radeon RX 9070 GRE leads the L4 by 4.3%, with an average score of 134,417. This is the largest deficit in the rival group, but it is still a narrow margin. The RX 9070 GRE is a consumer-oriented card, while the L4 is designed for server use, so the comparison highlights how close the L4 comes to a high-end consumer GPU despite its headless design and low power draw.

Ray Tracing and Feature Set

The L4 includes 60 RT cores and 240 tensor cores, both built on the Ada Lovelace architecture. This hardware enables hardware-accelerated ray tracing and tensor-based compute, making the card suitable for rendering, AI inference, and similar workloads. The API support is comprehensive: DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4 are all present. DirectX 12 Ultimate support indicates feature level 12_2, which includes advanced ray tracing and mesh shader capabilities.

The card has no display outputs, which confirms its role as a server accelerator rather than a desktop graphics card. It belongs to the Server Ada (Lxx) generation, with the predecessor listed as Server Ampere and the successor as Server Hopper. The tensor cores are a key differentiator for AI workloads, while the RT cores provide dedicated ray tracing throughput. The FP32 and FP16 performance are both rated at 30.29 TFLOPS, with a 1:1 ratio, meaning the card does not lose half-rate FP16 throughput, which is advantageous for mixed-precision compute.

Power and Cooling

The L4 has a TDP of 72 W, an exceptionally low figure for a card with this level of performance. It requires no power connectors, and the suggested PSU is 250 W. This makes the L4 easy to integrate into existing server chassis without additional power cabling. The single-slot form factor is 169 mm in length and 56 mm in height, allowing for dense multi-card configurations.

The low power draw also means cooling requirements are modest. A single-slot cooler is sufficient to handle the 72 W TDP, and the compact dimensions fit standard server racks. The lack of power connectors simplifies installation, and the 250 W PSU recommendation leaves ample headroom for the rest of the system. This combination of low power and compact size is a major advantage for data center environments where space and thermal budgets are tight.

FAQ

Q: What is the NVIDIA L4’s average benchmark score?

A: The average benchmark score is 128,665, placing it in the 97th percentile of all GPUs.

Q: How does the L4 compare to the GeForce RTX 3090 Ti?

A: The L4 trails the RTX 3090 Ti by 2.5%, with the 3090 Ti averaging 131,911.

Q: Does the L4 require power connectors?

A: No. The L4 has a 72 W TDP and no power connectors, and the suggested PSU is 250 W.

Q: What memory does the L4 have?

A: It has 24 GB of GDDR6 memory on a 192-bit bus, with 300.1 GB/s of bandwidth.

Q: What APIs does the L4 support?

A: It supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Q: What is the L4’s form factor?

A: It is a single-slot card measuring 169 mm in length and 56 mm in height, with a PCIe 4.0 x16 interface.

Benchmark Performance

The L4’s benchmark results show a clear performance profile. In Geekbench OpenCL, it scores 140,838, while in Geekbench Vulkan it scores 116,491. The OpenCL result is substantially higher, indicating that the card’s compute performance is better expressed through OpenCL workloads. The average of these two tests is 128,665, which is the figure used for comparisons against rivals.

Relative to the RTX 3090 Ti, the L4 is 2.5% slower. That is a narrow margin, and the L4’s 24 GB memory capacity may close the gap in memory-bound tasks. Against the Radeon PRO W6800, the L4 is 3.7% slower, with the PRO W6800 averaging 133,588. The Radeon RX 9070 GRE is the strongest rival, leading by 4.3% with an average of 134,417. The L4’s best showing is against the Radeon PRO W7700, where it is 4.2% faster.

The 97th percentile ranking puts the L4 above the vast majority of GPUs, even though it is not the fastest in its immediate rival group. The deltas are all within a 4.3% range, meaning the L4 is highly competitive with these cards. The Vulkan score is notably lower than the OpenCL score, which may be relevant for workloads that rely on Vulkan compute. Overall, the benchmark data indicates a well-rounded accelerator that trades a small amount of peak performance for drastically lower power consumption.

Who Should Consider It

The L4 is best suited for server and data center deployments where power efficiency and density are priorities. Its 72 W TDP and single-slot design make it ideal for multi-GPU systems, and the absence of power connectors simplifies cabling. The 24 GB memory capacity is substantial for large models and datasets, and the 300.1 GB/s bandwidth provides enough throughput for many compute tasks.

For high-resolution compute workloads, the L4’s 24 GB GDDR6 buffer is a key asset. The card’s 97th percentile ranking ensures it can handle demanding tasks, while the 60 RT cores and 240 tensor cores add specialized acceleration for ray tracing and AI. However, because the L4 has no display outputs, it is not suitable for direct desktop use. It is a headless accelerator designed for servers, and its performance is competitive with high-end workstation cards despite its low power draw.

Users who prioritize raw benchmark scores above all else may prefer the RX 9070 GRE or RTX 3090 Ti, which lead by 4.3% and 2.5%, respectively. But for environments where power, space, and thermal limits are critical, the L4’s combination of performance and efficiency is compelling. The single-slot form factor and 169 mm length allow for high-density configurations, and the 250 W PSU recommendation means even modest server power supplies can support it.

Memory Subsystem

The L4 is equipped with 24 GB of GDDR6 memory, a substantial amount for a server accelerator. The memory bus is 192-bit, and the total bandwidth is 300.1 GB/s. The memory clock is 1563 MHz, with an effective data rate of 12.5 Gbps. This configuration provides a balance between capacity and bandwidth, though the 192-bit bus limits raw throughput compared to wider-memory cards.

For high-resolution workloads, the 24 GB capacity is the standout feature. Large textures, big batch sizes, and complex models can reside entirely in memory, reducing the need for frequent data transfers. The 300.1 GB/s bandwidth is sufficient for many compute tasks, but memory-bound workloads may see a bottleneck when accessing large datasets. Still, the capacity far exceeds what most consumer GPUs offer, making the L4 a strong choice for inference, rendering, and other memory-intensive applications.

The 1:1 FP32/FP16 ratio (30.29 TFLOPS each) further enhances compute flexibility, and the 240 tensor cores can accelerate AI workloads. The memory subsystem, combined with the Ada Lovelace architecture, positions the L4 as a versatile accelerator for server environments. Its 24 GB buffer and 300.1 GB/s bandwidth are well matched to the card’s compute capabilities, and the low 72 W TDP means that memory capacity does not come at the cost of power efficiency.

The AMD Equivalent of L4

Looking for a similar graphics card from AMD? The AMD Radeon RX 7600 offers comparable performance and features in the AMD lineup.

AMD Radeon RX 7600

AMD • 8 GB VRAM

View Specs Compare

Popular NVIDIA L4 Comparisons

See how the L4 stacks up against similar graphics cards from the same generation and competing brands.

Compare L4 with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs