GEFORCE

NVIDIA H20

NVIDIA graphics card specifications and benchmark scores

96 GB
VRAM
1980
MHz Boost
500W
TDP
6144
Bus Width
Tensor Cores

At a Glance

NVIDIA
VRAM 96 GB
Boost Clock 1,980 MHz
Shaders 9,984
Bus Width 6144-bit
TDP 500W
Memory Type HBM3
Architecture Hopper
nm
Process 5 nm
Released Feb 2024

NVIDIA H20 Specifications

H20 GPU Core

Shader units and compute resources

The NVIDIA H20 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
9,984
Shaders
9,984
TMUs
312
ROPs
24
SM Count
78

H20 Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the H20's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The H20 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1830 MHz
Base Clock
1,830 MHz
Boost Clock
1980 MHz
Boost Clock
1,980 MHz
Memory Clock
1313 MHz 5.3 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's H20 Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The H20's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
96 GB
VRAM
98,304 MB
Memory Type
HBM3
VRAM Type
HBM3
Memory Bus
6144 bit
Bus Width
6144-bit
Bandwidth
4.03 TB/s

H20 by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the H20, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
256 KB (per SM)
L2 Cache
60 MB

H20 Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA H20 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
39.54 TFLOPS
FP64 (Double)
19.77 TFLOPS (1:2)
FP16 (Half)
79.07 TFLOPS (2:1)
Pixel Rate
47.52 GPixel/s
Texture Rate
617.8 GTexel/s

H20 Ray Tracing & AI

Hardware acceleration features

The NVIDIA H20 includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the H20 capable of delivering both stunning graphics and smooth frame rates in modern titles.

Tensor Cores
312

Hopper Architecture & Process

Manufacturing and design details

The NVIDIA H20 is built on NVIDIA's Hopper architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the H20 will perform in GPU benchmarks compared to previous generations.

Architecture
Hopper
GPU Name
GH100
Process Node
5 nm
Foundry
TSMC
Transistors
80,000 million
Die Size
814 mm²
Density
98.3M / mm²

NVIDIA's H20 Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA H20 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the H20 to maintain boost clocks without throttling.

TDP
500 W
TDP
500W
Suggested PSU
900 W

H20 by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA H20 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
SXM Module
Bus Interface
PCIe 5.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA H20. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
N/A
DirectX
N/A
OpenGL
N/A
OpenGL
N/A
Vulkan
N/A
Vulkan
N/A
OpenCL
3.0
CUDA
9.0
Shader Model
N/A

H20 Product Information

Release and pricing details

The NVIDIA H20 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the H20 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Feb 2024
Production
Active
Predecessor
Server Ada
Successor
Server Blackwell

H20 Benchmark Scores

No benchmark data available for this GPU.

About NVIDIA H20

Benchmark Performance

The NVIDIA H20 occupies a unique position in the server GPU landscape, with a benchmark profile that prioritizes compute throughput over traditional graphics workloads. The data shows a raw FP32 performance of 39.54 TFLOPS, which places this card in the mid-to-upper tier of the current server lineup. More telling is the FP16 figure of 79.07 TFLOPS (2:1), which doubles the FP32 output and signals that the architecture is heavily optimized for mixed-precision and AI inference tasks rather than pure rasterization.

The percentile ranking of 50 against all GPUs indicates that the H20 sits exactly at the median of the entire GPU database. This is a critical data point because it contextualizes the card's position: it is neither a top-tier flagship nor a low-end entry, but rather a specialized workhorse that excels in specific workloads while remaining average in overall synthetic benchmarks. The average benchmark score of 0 further reinforces that this card is not designed for conventional gaming or graphics benchmarks, as those tests produce no meaningful scores for this hardware.

Clock speeds are modest for a server part, with a base clock of 1830 MHz and a boost clock of 1980 MHz. These frequencies are lower than what one might expect from a 500 W TDP part, which suggests the design prioritizes stability and sustained throughput over peak single-thread performance. The boost clock represents a modest 8.2% uplift over the base clock, indicating a relatively flat frequency curve that maintains consistent performance under load.

The compute metrics reveal the card's true strengths. The texture rate of 617.8 GTexel/s and pixel rate of 47.52 GPixel/s are respectable figures, but they are not exceptional for a GPU of this class. What stands out is the shading unit count of 9984, which when combined with the 312 TMUs and 24 ROPs, creates a configuration that is heavily weighted toward compute shaders and tensor operations rather than traditional graphics rendering. The low ROP count relative to shading units is a clear indicator that this is not a rasterization-focused card.

Ray Tracing and Feature Set

The H20 does not include dedicated ray tracing cores, a significant omission that clearly delineates its purpose. The API support list confirms this positioning, with DirectX, OpenGL, and Vulkan all marked as "N/A." This means the card is not intended for any real-time graphics workloads, ray-traced or otherwise. Instead, the feature set is built around the 312 tensor cores, which are the primary compute engines for AI and machine learning workloads.

The Hopper architecture brings with it a specific set of capabilities that are optimized for data center operations. The tensor cores are designed to accelerate matrix operations, which are fundamental to neural network training and inference. The lack of display outputs further confirms that this is a compute-only accelerator, meant to be installed in server racks and accessed remotely rather than connected to a monitor.

The PCIe 5.0 x16 interface provides high-bandwidth connectivity to the host system, which is essential for feeding data to the compute cores efficiently. The architecture's support for FP16 operations at 2:1 ratio is particularly important for AI workloads, as many neural networks are trained using mixed-precision techniques that leverage FP16 for faster computation while maintaining accuracy through FP32 accumulation.

In terms of production status, the H20 is currently listed as "Active," meaning it is still being manufactured and sold. The release date of January 31, 2024, places it within the current generation of server hardware, positioned between the Server Ada predecessor and the Server Blackwell successor. This timing suggests the H20 is a bridging product that fills a specific niche in the market while the next generation is being rolled out.

Memory Subsystem

The memory subsystem is where the H20 demonstrates its most impressive specifications. It is equipped with 96 GB of HBM3 memory, which is a substantial capacity that allows for large models and datasets to be held entirely in GPU memory. The 6144-bit bus width is exceptionally wide, enabling the memory controller to access data across a massive number of parallel pathways.

The resulting memory bandwidth of 4.03 TB/s is a staggering figure that places the H20 in the upper echelon of memory bandwidth available in any GPU. To put this in perspective, this bandwidth allows the card to move over 4 terabytes of data per second between the compute cores and the memory, which is essential for memory-intensive workloads like large language model inference and training.

The memory clock runs at 1313 MHz with an effective data rate of 5.3 Gbps, which is relatively modest for HBM3 technology. However, the extreme bus width compensates for the lower clock speed, resulting in the exceptional overall bandwidth figure. This design choice prioritizes power efficiency and thermal management over raw memory clock speeds, as the wide bus allows for high bandwidth without requiring excessive memory frequencies.

For high-resolution workloads, the memory subsystem is more than adequate. While traditional high-resolution gaming would be constrained by the lack of display outputs, the memory capacity and bandwidth are perfectly suited for high-resolution scientific simulations, large-scale data analysis, and AI model processing. The 96 GB capacity means that even the largest models can be loaded entirely into memory without requiring data to be swapped to system RAM or disk storage.

Power and Cooling

The H20 carries a TDP of 500 W, which is a substantial power draw that requires serious cooling and power delivery infrastructure. The suggested PSU rating is 900 W, which provides a reasonable headroom above the card's maximum power consumption to account for other system components and transient power spikes.

The card is designed as an SXM Module, which means it is not a standard PCIe card that can be installed in a typical desktop chassis. Instead, it is meant to be mounted on a server motherboard or in a specialized chassis that can accommodate SXM form factors. This form factor also dictates the cooling solution, as SXM modules typically rely on server-level cooling systems such as high-pressure fans or liquid cooling solutions.

The power delivery is handled through the SXM module's connector, which provides the necessary power to the card. There are no standard PCIe power connectors listed, as the SXM form factor uses a proprietary connection that is part of the module socket. This means the H20 cannot be retrofitted into existing desktop systems without significant modification.

The 5 nm manufacturing process from TSMC, combined with the 80,000 million transistors on an 814 mm² die, presents a significant thermal management challenge. The transistor density of 98.3M per mm² is among the highest in the industry, which means the heat is concentrated in a relatively small area. The 500 W TDP must be dissipated effectively to maintain stable operation under sustained load.

How It Compares

The H20's position in the market is defined by its specialized nature. The benchmark data shows no nearest rivals, which indicates that this card does not have direct competitors in the same performance and feature segment. This is a unique position that reflects the card's specific design goals.

When compared to traditional server GPUs from the previous generation, the H20 offers a different balance of compute and memory resources. While it may not match the raw FP32 performance of larger server parts, its 96 GB memory capacity and 4.03 TB/s bandwidth are competitive with or superior to many data center accelerators. The lack of ray tracing capabilities and graphics APIs further differentiates it from consumer-oriented or even professional visualization cards.

The 50th percentile ranking against all GPUs is a double-edged sword. On one hand, it means the card is not exceptional in overall synthetic benchmarks. On the other hand, this ranking is heavily influenced by gaming-oriented benchmarks that the H20 cannot run, making the percentile less meaningful for its actual use case. In compute-specific benchmarks, the H20 would likely rank much higher.

The successor relationship with Server Blackwell suggests that the H20 is part of a transitional period in NVIDIA's server lineup. The predecessor, Server Ada, represents the older generation, while Blackwell represents the future. The H20 sits between these two, offering Hopper architecture features at a time when the industry is moving toward new technologies.

FAQ

Q: What is the memory capacity of the NVIDIA H20?

A: The H20 is equipped with 96 GB of HBM3 memory, which provides substantial capacity for large AI models and datasets.

Q: Does the H20 support DirectX or Vulkan?

A: No, the H20 has no support for DirectX, OpenGL, or Vulkan. All graphics APIs are marked as N/A, as the card is designed for compute workloads only.

Q: What is the power consumption of the H20?

A: The H20 has a TDP of 500 W, and NVIDIA suggests a 900 W power supply for systems using this card.

Q: Can the H20 be used for gaming?

A: No, the H20 has no display outputs and no graphics API support, making it unsuitable for gaming or any real-time visualization tasks.

Q: What is the form factor of the H20?

A: The H20 uses an SXM Module form factor, which requires a compatible server motherboard or chassis that supports this mounting standard.

Q: What is the memory bandwidth of the H20?

A: The H20 provides 4.03 TB/s of memory bandwidth through a 6144-bit bus interface running HBM3 memory.

Who Should Consider It

The NVIDIA H20 is a specialized compute accelerator that is appropriate for a narrow set of use cases. The data shows that this card is designed for AI and machine learning workloads, particularly those that require large memory capacity and high memory bandwidth. Organizations running large language models, complex neural networks, or scientific simulations that need to hold massive datasets in GPU memory will find the H20's 96 GB capacity and 4.03 TB/s bandwidth to be exceptional assets.

For AI inference tasks, the FP16 performance of 79.07 TFLOPS is the key metric. This level of mixed-precision throughput is well-suited for serving models in production environments where latency and throughput are critical. The 312 tensor cores are specifically designed to accelerate the matrix operations that dominate neural network inference.

The H20 is not appropriate for any graphics-related workloads. The lack of display outputs, ray tracing cores, and graphics API support makes it useless for gaming, professional visualization, or any task that requires rendering to a screen. Similarly, the 24 ROPs and 47.52 GPixel/s pixel rate are far below what would be needed for even modest graphics performance.

For high-resolution compute workloads, the memory subsystem is exceptional. The 4.03 TB/s bandwidth ensures that the compute cores are never starved for data, which is a common bottleneck in memory-intensive applications. The 6144-bit bus width allows for massive parallel data access, which is critical for workloads that require frequent memory reads and writes.

The power and cooling requirements mean that the H20 is only suitable for data center environments with proper infrastructure. The 500 W TDP and 900 W suggested PSU rating require servers with robust power delivery and cooling systems. The SXM form factor further restricts deployment to compatible server platforms.

In summary, the H20 is a purpose-built AI accelerator that excels at what it is designed to do. The 50th percentile ranking in overall benchmarks is misleading, as it does not reflect the card's capabilities in its intended workloads. For organizations that need large memory capacity, high bandwidth, and strong FP16 compute performance for AI applications, the H20 is a viable choice. For anyone else, the card's limitations in graphics and its specialized form factor make it an impractical option.

The AMD Equivalent of H20

Looking for a similar graphics card from AMD? The AMD Radeon RX 7600 XT offers comparable performance and features in the AMD lineup.

AMD Radeon RX 7600 XT

AMD • 16 GB VRAM

View Specs Compare

Popular NVIDIA H20 Comparisons

See how the H20 stacks up against similar graphics cards from the same generation and competing brands.

Compare H20 with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs