NVIDIA H20 NVL16 vs NVIDIA L20 Comparison

NVIDIA
GEFORCE

NVIDIA H20 NVL16

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 400 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

L20

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 275 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

geekbench_opencl
N/A
274,276
geekbench_vulkan
N/A
228,018

Analysis: NVIDIA H20 NVL16 vs NVIDIA L20

Head-to-Head Benchmarks

The recorded database contains no direct head-to-head benchmark comparisons between the NVIDIA H20 NVL16 and the NVIDIA L20. The head-to-head benchmark array is empty, and neither part lists a wins count. The H20 NVL16 also has no benchmark entries, no average score, and no nearest rivals in the database. The L20, by contrast, has two recorded benchmark scores: 274276 in Geekbench OpenCL and 228018 in Geekbench Vulkan. Its average benchmark score is 251147.

The absence of direct comparisons means the only numerical basis for relative standing comes from the L20's percentile and rival data. The L20 sits in the 99th percentile of all GPUs in the database. Its nearest rivals are the NVIDIA PG506-232, which scores 225124 on average, and the AMD Radeon PRO W7900D, which scores 219827. The L20 leads the PG506-232 by 11.6 percent and the Radeon PRO W7900D by 14.2 percent. It trails the NVIDIA L40, which averages 284111, by 11.6 percent, and the NVIDIA RTX 6000 Ada Generation, which averages 287237, by 12.6 percent.

The H20 NVL16 has no comparable data points. Its percentile versus all GPUs is 50, meaning the database places it at the median of recorded GPUs, but no average score accompanies that figure. The L20's 99th percentile ranking, paired with its concrete scores, indicates that the L20 is the only one of the two with measurable, verified performance data in the database. The H20 NVL16's 50th percentile is a placeholder position without supporting benchmark results.

For the H20 NVL16, the theoretical compute figures are the only available performance indicators. Its FP32 throughput is 39.54 TFLOPS, and its FP16 throughput is 79.07 TFLOPS with a 2:1 ratio. The L20 delivers 59.35 TFLOPS in FP32 and 59.35 TFLOPS in FP16 with a 1:1 ratio. In FP32, the L20 is roughly 50 percent ahead based on these raw figures. In FP16, the H20 NVL16 is roughly 33 percent ahead. These are specification-derived estimates, not measured benchmark outcomes, so they carry less weight than recorded scores.

The L20's Geekbench OpenCL score of 274276 is its stronger result. The Vulkan score of 228018 is notably lower, a gap of over 46 thousand points. That spread suggests the L20's compute advantage expresses itself more fully in OpenCL workloads than in Vulkan workloads, likely due to driver and API scheduling differences that favor the former.

FAQ

Q: Which GPU has a higher recorded average benchmark score?

A: The NVIDIA L20 has a recorded average benchmark score of 251147. The NVIDIA H20 NVL16 has no recorded benchmark scores and an average score of 0 in the database.

Q: How does the L20 compare to its nearest rivals in the database?

A: The L20 leads the NVIDIA PG506-232, which averages 225124, by 11.6 percent. It also leads the AMD Radeon PRO W7900D, which averages 219827, by 14.2 percent. The L20 trails the NVIDIA L40, averaging 284111, by 11.6 percent, and the NVIDIA RTX 6000 Ada Generation, averaging 287237, by 12.6 percent.

Q: What is the memory configuration difference between the two GPUs?

A: The H20 NVL16 uses 96 GB of HBM3 memory on a 6144-bit bus, delivering 4.03 TB/s of bandwidth. The L20 uses 48 GB of GDDR6 memory on a 384-bit bus, delivering 864.0 GB/s of bandwidth.

Q: Which GPU has higher FP32 throughput?

A: The L20 has higher FP32 throughput at 59.35 TFLOPS. The H20 NVL16 delivers 39.54 TFLOPS in FP32.

Q: Which GPU has higher FP16 throughput?

A: The H20 NVL16 has higher FP16 throughput at 79.07 TFLOPS with a 2:1 ratio. The L20 delivers 59.35 TFLOPS in FP16 with a 1:1 ratio.

Q: What are the thermal design power figures for each GPU?

A: The H20 NVL16 has a TDP of 400 W and a suggested PSU of 800 W. The L20 has a TDP of 275 W and a suggested PSU of 600 W.

Where Each One Wins

The L20 wins on measured performance. The database contains actual benchmark results for it, and those results place it in the 99th percentile of all GPUs. The Geekbench OpenCL score of 274276 and the Vulkan score of 228018 give it a verifiable performance footprint. The H20 NVL16 has no measured results, so any performance claim for it must rely on theoretical specifications.

The L20 also wins on FP32 compute. Its 59.35 TFLOPS exceeds the H20 NVL16's 39.54 TFLOPS by a wide margin. This makes the L20 the stronger choice for workloads that depend on single-precision floating-point math, such as traditional graphics rendering and many simulation tasks. The L20 also carries a full API stack, with DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4 support, plus four DisplayPort 1.4a outputs. The H20 NVL16 has no display outputs and no recorded API support, which makes it unsuitable for any interactive or graphics-output role.

The H20 NVL16 wins on memory capacity and bandwidth. Its 96 GB of HBM3 on a 6144-bit bus delivers 4.03 TB/s, which is over 4.6 times the L20's bandwidth of 864.0 GB/s on its 384-bit GDDR6 bus. That bandwidth advantage is decisive for large-model inference and training workloads where data movement dominates. The H20 NVL16 also wins on FP16 throughput, delivering 79.07 TFLOPS versus the L20's 59.35 TFLOPS, which matters for AI inference and mixed-precision training.

The H20 NVL16's slot form factor, an SXM module, indicates it is designed for dense server integration, while the L20's dual-slot PCIe form factor with a 16-pin connector is more flexible for standard server chassis. The L20 uses PCIe 4.0 x16, while the H20 NVL16 uses PCIe 5.0 x16, giving the latter a newer bus interface.

Specification Differences

The core specification sheets diverge on nearly every measurable parameter. The H20 NVL16 uses the GH100 chip on the Hopper architecture, while the L20 uses the AD102 chip on Ada Lovelace. The H20 NVL16 has 80,000 million transistors on an 814 mm² die, with a transistor density of 98.3M per mm². The L20 has 76,300 million transistors on a 609 mm² die, with a higher transistor density of 125.3M per mm².

Clock speeds differ significantly. The H20 NVL16 runs at a base clock of 1830 MHz and a boost clock of 1980 MHz. The L20 runs at a lower base clock of 1440 MHz but a much higher boost clock of 2520 MHz. Memory clocks also differ: the H20 NVL16's HBM3 operates at 1313 MHz with 5.3 Gbps effective, while the L20's GDDR6 operates at 2250 MHz with 18 Gbps effective.

The compute unit counts are different. The H20 NVL16 has 9984 shading units, 312 TMUs, 24 ROPs, and 312 tensor cores. The L20 has 11776 shading units, 368 TMUs, 128 ROPs, 92 RT cores, and 368 tensor cores. The L20 has RT cores, which the H20 NVL16 lacks entirely, reflecting the latter's compute-focused design.

Pixel and texture rates reflect the ROP and TMU differences. The H20 NVL16 delivers 47.52 GPixel/s and 617.8 GTexel/s. The L20 delivers 322.6 GPixel/s and 927.4 GTexel/s. The L20's pixel rate is nearly 7 times higher, and its texture rate is about 50 percent higher.

Power and physical specifications differ. The H20 NVL16 has a TDP of 400 W, a suggested PSU of 800 W, and uses an SXM module slot. The L20 has a TDP of 275 W, a suggested PSU of 600 W, uses a dual-slot form factor, a single 16-pin power connector, and measures 267 mm in length and 111 mm in height. The H20 NVL16 has no display outputs; the L20 has 4x DisplayPort 1.4a.

Release dates and generations differ. The H20 NVL16 belongs to the Server Hopper (Hxx) generation and was released on 2025-09-01. The L20 belongs to the Server Ada (Lxx) generation and was released on 2023-11-15. The H20 NVL16's predecessor is Server Ada and its successor is Server Blackwell. The L20's predecessor is Server Ampere and its successor is Server Hopper.

Architecture Differences

The two GPUs represent different NVIDIA server architectures. The H20 NVL16 is built on Hopper, the L20 on Ada Lovelace. Both use a 5 nm process from TSMC, but the design philosophies diverge sharply.

Hopper, as implemented in the H20 NVL16, prioritizes memory bandwidth and mixed-precision throughput for AI and high-performance computing. The HBM3 memory subsystem with a 6144-bit bus and 4.03 TB/s bandwidth is the defining feature. The FP16 throughput of 79.07 TFLOPS at a 2:1 ratio indicates a design tuned for tensor-heavy workloads where reduced precision is acceptable. The absence of RT cores and display outputs confirms that the H20 NVL16 is not intended for graphics or rendering tasks.

Ada Lovelace, as implemented in the L20, retains a more general-purpose profile. The 92 RT cores and full DirectX 12 Ultimate support, along with OpenGL 4.6 and Vulkan 1.4, make it usable for graphics, rendering, and compute. The L20's FP16 throughput equals its FP32 throughput at 59.35 TFLOPS with a 1:1 ratio, meaning it does not accelerate half-precision workloads beyond single-precision rates. The GDDR6 memory on a 384-bit bus is conventional and far lower bandwidth than HBM3.

Transistor density favors the L20: 125.3M per mm² versus 98.3M per mm² for the H20 NVL16. This reflects the Ada Lovelace design's more compact layout on a smaller die, 609 mm² versus 814 mm². The H20 NVL16's larger die with lower density suggests a design with more memory controllers and a wider memory interface, which is consistent with its HBM3 configuration.

The L20's boost clock of 2520 MHz versus the H20 NVL16's 1980 MHz indicates that Ada Lovelace operates at higher frequencies. The L20's base clock of 1440 MHz is lower than the H20 NVL16's 1830 MHz, but the boost delta is substantial. The H20 NVL16's higher base clock with lower boost suggests a more stable, sustained operating profile, while the L20 relies on aggressive boosting.

The bus interfaces differ. The H20 NVL16 uses PCIe 5.0 x16, the L20 uses PCIe 4.0 x16. This matters for host data transfer, but the H20 NVL16's HBM3 bandwidth dominates any interconnect consideration.

The L20 has display outputs, the H20 NVL16 has none. The L20 supports a conventional graphics API stack, the H20 NVL16 supports none. These differences define the use cases: the L20 is a general-purpose server GPU that can handle graphics, the H20 NVL16 is a specialized compute accelerator for memory-bound AI workloads.

DETAILED SPECIFICATIONS

SPECIFICATION
H20 NVL16
L20
Core Specs
Shading Units
9,984
11,776 +17.9%
Shaders
9,984
11,776 +17.9%
TMUs
312
368 +17.9%
ROPs
24
128 +433.3%
SM Count
78
92 +17.9%
Clocks
Base Clock
1830 MHz
1440 MHz
Boost Clock
1980 MHz
2520 MHz
Memory Clock
1313 MHz 5.3 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
96 GB
48 GB
VRAM (MB)
98,304
49,152 -50.0%
Memory Type
HBM3
GDDR6
Memory Bus
6144 bit
384 bit
Bandwidth
4.03 TB/s
864.0 GB/s
Cache
L1 Cache
256 KB (per SM)
128 KB (per SM)
L2 Cache
60 MB
96 MB
Performance
Pixel Rate
47.52 GPixel/s
322.6 GPixel/s
Texture Rate
617.8 GTexel/s
927.4 GTexel/s
FP32 (TFLOPS)
39.54 TFLOPS
59.35 TFLOPS
FP64 (TFLOPS)
19.77 TFLOPS (1:2)
927.4 GFLOPS (1:64)
FP16 (TFLOPS)
79.07 TFLOPS (2:1)
59.35 TFLOPS (1:1)
AI/RT
RT Cores
92
Tensor Cores
312
368 +17.9%
Power
TDP
400 W
275 W
TDP (W)
400
275 -31.3%
Suggested PSU
800 W
600 W
Power Connectors
1x 16-pin
Architecture
Architecture
Hopper
Ada Lovelace
GPU Name
GH100
AD102
Generation
Server Hopper (Hxx)
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
80,000 million
76,300 million
Die Size
814 mm²
609 mm²
Foundry
TSMC
TSMC
Density
98.3M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
9.0
8.9
Shader Model
6.8
Physical
Slot Width
SXM Module
Dual-slot
Length
267 mm 10.5 inches
Height
111 mm 4.4 inches
Outputs
No outputs
4x DisplayPort 1.4a
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x16
Other
Production
Active
Active
Predecessor
Server Ada
Server Ampere
Successor
Server Blackwell
Server Hopper
View H20 NVL16 Details View L20 Details