NVIDIA H20 NVL16
NVIDIA graphics card specifications and benchmark scores
At a Glance
NVIDIANVIDIA H20 NVL16 Specifications
H20 NVL16 GPU Core
Shader units and compute resources
The NVIDIA H20 NVL16 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.
H20 NVL16 Clock Speeds
GPU and memory frequencies
Clock speeds directly impact the H20 NVL16's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The H20 NVL16 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.
NVIDIA's H20 NVL16 Memory
VRAM capacity and bandwidth
VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The H20 NVL16's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.
H20 NVL16 by NVIDIA Cache
On-chip cache hierarchy
On-chip cache provides ultra-fast data access for the H20 NVL16, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.
H20 NVL16 Theoretical Performance
Compute and fill rates
Theoretical performance metrics provide a baseline for comparing the NVIDIA H20 NVL16 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.
H20 NVL16 Ray Tracing & AI
Hardware acceleration features
The NVIDIA H20 NVL16 includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the H20 NVL16 capable of delivering both stunning graphics and smooth frame rates in modern titles.
Hopper Architecture & Process
Manufacturing and design details
The NVIDIA H20 NVL16 is built on NVIDIA's Hopper architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the H20 NVL16 will perform in GPU benchmarks compared to previous generations.
NVIDIA's H20 NVL16 Power & Thermal
TDP and power requirements
Power specifications for the NVIDIA H20 NVL16 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the H20 NVL16 to maintain boost clocks without throttling.
H20 NVL16 by NVIDIA Physical & Connectivity
Dimensions and outputs
Physical dimensions of the NVIDIA H20 NVL16 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.
NVIDIA API Support
Graphics and compute APIs
API support determines which games and applications can fully utilize the NVIDIA H20 NVL16. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.
H20 NVL16 Product Information
Release and pricing details
The NVIDIA H20 NVL16 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the H20 NVL16 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.
H20 NVL16 Benchmark Scores
No benchmark data available for this GPU.
About NVIDIA H20 NVL16
How It Compares
The NVIDIA H20 NVL16 is a server-oriented accelerator built on the Hopper architecture, occupying a distinct position within NVIDIA’s data center lineup. Its percentile standing among all GPUs is 50, placing it exactly at the median of the database’s tracked graphics processors. This positioning reflects a product engineered for balanced throughput rather than peak consumer gaming performance, which aligns with its server-class design goals.
As a member of the Server Hopper generation, the H20 NVL16 succeeds the Server Ada family and precedes the Server Blackwell generation. This generational placement means the card inherits Hopper’s compute-focused design philosophy while looking ahead to Blackwell’s architectural refinements. The chip itself, designated GH100, is a massive 814 mm² die manufactured on TSMC’s 5 nm process, housing 80,000 million transistors for a transistor density of 98.3 million per square millimeter. These physical characteristics situate the H20 NVL16 as a high-complexity, high-yield part intended for dense compute workloads.
The benchmark database lists no nearest rivals for this accelerator, meaning comparative analysis relies on its internal specifications and percentile ranking rather than direct head-to-head scores. With an average benchmark score of zero and no entries in the benchmarks array, the H20 NVL16’s performance must be inferred from its raw compute metrics and memory subsystem. The 50th percentile status suggests that when measured across the entire GPU landscape, including consumer and professional parts, this card lands in the middle, but its server orientation means that comparison is not entirely apples-to-apples.
Ray Tracing and Feature Set
The H20 NVL16 does not expose dedicated ray tracing cores in its specification sheet, and its API support is listed as N/A for DirectX, OpenGL, and Vulkan. This absence is consistent with a compute-focused server accelerator, where rasterization and real-time graphics APIs are not the primary workload drivers. The card is not designed for gaming or interactive rendering, so the lack of RT core counts and graphics API compatibility should not be viewed as a deficiency but rather as a reflection of its intended deployment environment.
Instead, the H20 NVL16 emphasizes tensor processing, featuring 312 tensor cores. These tensor cores are the architectural engines for AI inference, deep learning training, and matrix math operations that dominate modern data center workloads. The Hopper architecture’s tensor core design enables mixed-precision computation, and the H20 NVL16 delivers 79.07 TFLOPS of FP16 performance at a 2:1 ratio, effectively doubling its FP32 throughput when operating in reduced precision. This positions the card as a capable workhorse for neural network inference and fine-tuning tasks that benefit from high FP16 throughput.
The card also provides 9984 shading units, 312 texture mapping units, and 24 raster operations pipelines. While the shading units could theoretically handle graphics workloads, the lack of display outputs and graphics API support confirms that this silicon is repurposed for compute. The texture rate of 617.8 GTexel/s and pixel rate of 47.52 GPixel/s are included in the specification, but these metrics are more relevant for understanding the chip’s raw processing capacity rather than any real-world graphics application. In a server rack, the H20 NVL16 operates as a headless compute accelerator, communicating via PCIe 5.0 x16 for host connectivity.
Power and Cooling
The H20 NVL16 carries a thermal design power of 400 W, a figure that reflects its high transistor count and sustained compute capability. This TDP is typical for a server accelerator module designed for dense deployment in data center chassis, where cooling infrastructure is already provisioned for high wattage. The card’s form factor is specified as an SXM Module, which means it does not use a standard PCIe slot for mounting or cooling; instead, it plugs into a proprietary SXM socket on a server motherboard or baseboard, with cooling provided by the system’s chassis fans or liquid cooling loops.
Power delivery for the H20 NVL16 is handled through the SXM connector, and the specification sheet lists no separate power connectors because the module draws its power from the socket interface. The suggested power supply rating is 800 W, which accounts for the accelerator’s 400 W draw plus overhead for the host system’s CPU, memory, storage, and networking. System integrators should ensure that the server’s power supply can sustain peak loads, especially during sustained compute bursts where the card may operate near its TDP limit for extended periods.
The production status is Active, meaning the H20 NVL16 is currently available for server manufacturers and cloud providers to integrate into their offerings. The release date is listed as 2025-09-01, which places it in the near-future product cycle for Hopper-based accelerators. Since the launch MSRP field is null, no pricing information is available, and the card’s cost is negotiated through enterprise channels rather than retail listings. The absence of display outputs reinforces that this is not a consumer product; it has no video ports because it is never intended to drive a monitor.
FAQ
Q: What architecture is the NVIDIA H20 NVL16 based on?
A: The H20 NVL16 is built on the Hopper architecture, specifically using the GH100 chip manufactured on TSMC’s 5 nm process.
Q: How much memory does the H20 NVL16 have and what type is it?
A: The card features 96 GB of HBM3 memory with a 6144-bit bus interface, providing a total bandwidth of 4.03 TB/s.
Q: Does the H20 NVL16 support DirectX or Vulkan for gaming?
A: No, the API support is listed as N/A for DirectX, OpenGL, and Vulkan, and the card has no display outputs, making it unsuitable for gaming or graphics rendering.
Q: What is the power consumption of this accelerator?
A: The thermal design power is 400 W, and the suggested power supply rating for the host system is 800 W.
Q: What is the form factor of the H20 NVL16?
A: It is an SXM Module, meaning it installs into an SXM socket on a server motherboard rather than a standard PCIe slot, and it uses the PCIe 5.0 x16 bus interface for data transfer.
Q: What is the FP32 compute performance of the H20 NVL16?
A: The card delivers 39.54 TFLOPS of FP32 performance, with FP16 performance reaching 79.07 TFLOPS at a 2:1 ratio.
Benchmark Performance
The H20 NVL16’s benchmark data is notably sparse, with an empty benchmarks array and an average benchmark score of zero. This absence of direct performance numbers means that the card’s capabilities must be evaluated through its architectural specifications and compute throughput figures. The percentile rank of 50 indicates that, when compared to all GPUs in the database, the H20 NVL16 sits at the median, half of tracked GPUs score higher, half score lower. However, this percentile is heavily influenced by consumer gaming cards that dominate the database, which are not the H20 NVL16’s intended competitors.
In terms of raw compute, the H20 NVL16’s FP32 throughput of 39.54 TFLOPS is substantial, but its FP16 performance of 79.07 TFLOPS (2:1 ratio) is where its tensor core strength lies. The 312 tensor cores are designed to accelerate matrix operations, and the 2:1 FP16 ratio indicates the card is optimized for AI workloads that rely on half-precision arithmetic. The FP32 figure, while lower than the FP16 number, still represents a significant compute resource for scientific simulation and other single-precision tasks.
The memory subsystem plays a critical role in benchmark performance, as the 4.03 TB/s bandwidth ensures that compute units are fed with data without bottlenecks. The 96 GB of HBM3 memory allows for large model weights and datasets to reside on-card, reducing the need for PCIe transfers. This is particularly important for transformer-based AI models, which require substantial memory capacity for parameters and activations. The 6144-bit bus width is among the widest in the database, enabling the high bandwidth figure.
Given the lack of nearest rivals and direct benchmark scores, the H20 NVL16’s performance profile is best understood through its architectural positioning. It is not a gaming card, nor is it a low-power inference chip; it is a full-fat Hopper accelerator with 9984 shading units and 312 texture mapping units. The pixel rate of 47.52 GPixel/s and texture rate of 617.8 GTexel/s are high numbers that would be impressive in a graphics context, but they are secondary to the tensor and memory capabilities. The 24 ROPs are low compared to consumer cards, further confirming that rasterization is not a priority. In a server context, the H20 NVL16’s benchmark relevance comes from its ability to sustain high FP16 throughput while holding massive models in its 96 GB memory pool.
Who Should Consider It
The H20 NVL16 is a server accelerator, which immediately narrows its target audience to data center operators, cloud service providers, and enterprises running AI or high-performance computing workloads. The card’s 96 GB HBM3 memory and 4.03 TB/s bandwidth make it well-suited for large language model inference, where the entire model can reside in on-card memory, eliminating the latency of host-side swapping. The 79.07 TFLOPS FP16 performance provides the compute headroom needed for token generation and batch inference, while the 39.54 TFLOPS FP32 capability supports scientific computing tasks that require higher precision.
For organizations running AI training workloads, the H20 NVL16’s tensor cores are the key attraction. The 312 tensor cores are optimized for the matrix multiplications that form the backbone of neural network backpropagation, and the FP16 2:1 ratio allows for faster training when precision requirements permit mixed-precision techniques. However, the card’s 400 W TDP and SXM form factor mean that it is not a drop-in solution; it requires a compatible server platform with SXM sockets and sufficient cooling. The suggested 800 W PSU rating indicates that system integration must account for the accelerator’s power draw alongside the rest of the server’s components.
The 50th percentile ranking suggests that this card is not an extreme outlier in the GPU landscape, but that ranking is skewed by consumer parts. For its intended server workloads, the H20 NVL16 offers a balanced combination of memory capacity, bandwidth, and compute throughput. It is not positioned for edge inference or low-power deployments, as its 400 W TDP rules out most compact or power-constrained environments. Instead, it is designed for rack-mounted servers in climate-controlled data centers, where power and cooling are abundant.
Organizations that need to serve large AI models to many concurrent users would benefit from the H20 NVL16’s memory capacity, as it reduces the need for model sharding across multiple GPUs. Similarly, researchers working with massive scientific datasets that require single-precision arithmetic would find the 39.54 TFLOPS FP32 performance adequate for many simulation tasks. The card is not suitable for gaming, workstation graphics, or any workload requiring a display output, as it has none.
Memory Subsystem
The H20 NVL16’s memory subsystem is one of its most defining features, comprising 96 GB of HBM3 memory arranged on a 6144-bit bus. This configuration yields a peak bandwidth of 4.03 TB/s, which is among the highest in the database and critical for feeding the card’s compute units. The 6144-bit bus width is exceptionally wide, allowing data to flow in parallel across thousands of memory pins. This width is a hallmark of HBM technology, which stacks memory dies vertically and connects them through an interposer, achieving much higher bandwidth than traditional GDDR memory on a narrower bus.
The 96 GB capacity is particularly significant for AI workloads, as it allows models with billions of parameters to be loaded entirely into on-card memory. For inference tasks, this means the model weights and KV cache can reside on the GPU, eliminating the performance penalty of fetching data over PCIe. The 4.03 TB/s bandwidth ensures that the tensor cores can retrieve data at a rate that keeps them busy, which is essential for sustain high utilization during batch processing. The memory clock is listed as 1313 MHz, with an effective data rate of 5.3 Gbps, which is the per-pin transfer rate that, when multiplied by the 6144-bit bus, produces the aggregate bandwidth figure.
For high-resolution input data, such as large images or long sequences in natural language processing, the memory bandwidth determines how quickly data can be moved into the compute units. The H20 NVL16’s 4.03 TB/s bandwidth is sufficient to handle such workloads without becoming a bottleneck, provided the host system can supply data at a comparable rate via the PCIe 5.0 x16 interface. The 96 GB capacity also supports multi-tenant scenarios where multiple smaller models are loaded simultaneously, increasing the card’s utilization and return on investment for server operators.
The HBM3 memory type is a notable advancement over previous HBM generations, offering higher density and bandwidth per stack. The 80,000 million transistors on the 814 mm² die are complemented by the memory stacks, which are not included in the die size but are part of the overall package. The 98.3 million transistors per square millimeter density figure reflects the chip’s complexity, and the memory subsystem is designed to match that complexity with equally high-end storage. For workloads that are memory-bound, such as graph analytics or database operations, the H20 NVL16’s memory subsystem would provide a significant advantage over cards with smaller capacities or lower bandwidth. The 4.03 TB/s figure is a standout specification that positions this accelerator for the most demanding data-centric tasks in a server environment.
The AMD Equivalent of H20 NVL16
Looking for a similar graphics card from AMD? The AMD Radeon RX 7700 offers comparable performance and features in the AMD lineup.
Popular NVIDIA H20 NVL16 Comparisons
See how the H20 NVL16 stacks up against similar graphics cards from the same generation and competing brands.
Compare H20 NVL16 with Other GPUs
Select another GPU to compare specifications and benchmarks side-by-side.
Browse GPUs