NVIDIA H200 NVL vs NVIDIA N1 20SM Comparison

NVIDIA
GEFORCE

NVIDIA H200 NVL

CORE STATE GH100
VRAM 141 GB
CLOCK SPEED 1785 MHz
TDP 600 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

N1 20SM

CORE STATE GB20B
VRAM 128 GB
CLOCK SPEED 2346 MHz
TDP unknown
BUS WIDTH 256 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2026

PERFORMANCE BENCHMARKS

geekbench_opencl
334,891
N/A

Analysis: NVIDIA H200 NVL vs NVIDIA N1 20SM

Head-to-Head Benchmarks

The database contains a single recorded OpenCL benchmark for the NVIDIA H200 NVL, scoring 334,891 points. The NVIDIA N1 20SM has no recorded benchmark scores, placing it at the 50th percentile among all GPUs with an average benchmark score of zero. This makes a direct numerical comparison impossible; the H200 NVL is the only one of the pair with measurable performance data.

Against its nearest rivals, the H200 NVL sits 3.1% behind the NVIDIA B200, which averages 345,482 points. It is 5.3% ahead of the AMD Instinct MI300X, which scores 317,994 points. The H200 NVL also trails the NVIDIA B300 SXM6 AC by 9.4%, as that part averages 369,831 points. Finally, the H200 NVL leads the NVIDIA L40S by 13.2%, with the L40S averaging 295,763 points. These deltas show the H200 NVL occupies a solid mid-to-upper tier among server accelerators, clearly outpacing the L40S and MI300X while sitting just behind the B200 and further behind the B300.

Because the N1 20SM has no recorded benchmarks and no nearest rivals listed, there is no head-to-head score to cite. The data shows the N1 20SM is an active product, but its performance profile remains unmeasured in the database. The absence of any benchmark entry means the H200 NVL holds every measurable win by default, though this reflects missing data rather than a demonstrated advantage.

Where Each One Wins

The H200 NVL wins in every category where the database has recorded information. Its OpenCL score of 334,891 places it at the 100th percentile among all GPUs, meaning it outperforms every other listed part in the database. The N1 20SM, by contrast, sits at the 50th percentile with no score, indicating it has not been tested or has not produced a result.

For compute-heavy workloads, the H200 NVL delivers 60.32 TFLOPS of FP32 performance and 120.6 TFLOPS of FP16 performance at a 2:1 ratio. The N1 20SM offers 12.01 TFLOPS in both FP32 and FP16 at a 1:1 ratio. The H200 NVL provides roughly five times the FP32 throughput, a substantial gap for tasks like scientific simulation or AI inference that rely on raw floating-point math.

Memory bandwidth tells a similar story. The H200 NVL uses 141 GB of HBM3e across a 6144-bit bus, achieving 4.89 TB/s. The N1 20SM uses 128 GB of LPDDR5X across a 256-bit bus, achieving 273.2 GB/s. The H200 NVL offers nearly 18 times the memory bandwidth, which directly benefits workloads that stream large datasets, such as training large language models or processing high-resolution imagery.

Texture and pixel rates also favor the H200 NVL. Its texture rate is 942.5 GTexel/s versus 375.4 GTexel/s for the N1 20SM. The pixel rates are closer: 42.84 GPixel/s for the H200 NVL and 56.30 GPixel/s for the N1 20SM, a case where the N1 20SM actually produces a higher fill rate despite its smaller overall compute footprint. This suggests the N1 20SM could handle rasterization-style tasks with better efficiency per pixel, though neither part targets typical graphics workloads.

The N1 20SM does include 20 RT cores and a single HDMI output, features absent from the H200 NVL, which has no display outputs. This makes the N1 20SM the only one of the two with any video-out capability, though its API support is listed as N/A for DirectX, OpenGL, and Vulkan, so its display output may serve specialized or embedded uses rather than consumer gaming.

Architecture Differences

The H200 NVL uses the GH100 chip on the Hopper architecture, belonging to the Server Hopper (Hxx) generation. The N1 20SM uses the GB20B chip on the Blackwell 2.0 architecture, belonging to the Blackwell IGP (N1x) generation. Both are fabricated on a 5 nm process at TSMC, so the manufacturing node is identical.

Transistor counts differ sharply. The H200 NVL integrates 80,000 million transistors on a 814 mm² die, yielding a density of 98.3 million transistors per square millimeter. The N1 20SM has an unknown transistor count but a much smaller 382 mm² die. The H200 NVL's die is more than double the size, reflecting its far larger compute and memory subsystems.

Core configurations diverge considerably. The H200 NVL has 16,896 shading units, 528 TMUs, 24 ROPs, and 528 tensor cores. It has no RT cores listed. The N1 20SM has 2,560 shading units, 160 TMUs, 24 ROPs, 80 tensor cores, and 20 RT cores. The H200 NVL carries 6.6 times the shading units and 6.6 times the tensor cores, while the N1 20SM brings dedicated ray tracing hardware that the H200 NVL lacks.

Clock speeds also differ. The H200 NVL runs at a 1365 MHz base and 1785 MHz boost, while the N1 20SM runs at a 741 MHz base and 2346 MHz boost. The N1 20SM has a higher boost clock by 561 MHz, which helps its smaller core count achieve respectable per-core throughput, but the H200 NVL's massive core count still dominates aggregate compute.

Memory architecture is fundamentally different. The H200 NVL uses HBM3e with a 6144-bit bus and 141 GB capacity. The N1 20SM uses LPDDR5X with a 256-bit bus and 128 GB capacity. The H200 NVL's memory clock is 1593 MHz with 6.4 Gbps effective data rate, while the N1 20SM runs at 1067 MHz with 8.5 Gbps effective. Despite the N1 20SM's higher effective data rate per pin, the H200 NVL's vastly wider bus delivers 4.89 TB/s versus 273.2 GB/s.

Power and physical design also separate the two. The H200 NVL draws 600 W, requires an 8-pin EPS connector, and a suggested 1000 W PSU. It is a dual-slot card measuring 267 mm in length and 111 mm in height. The N1 20SM has unknown TDP, no power connectors, and is listed as an IGP (integrated graphics processor), suggesting it draws power from its host system rather than a dedicated supply.

Release timing differs by about a year and a half. The H200 NVL launched on 2024-11-17, while the N1 20SM launched on 2026-05-31. The H200 NVL lists its predecessor as Server Ada and successor as Server Blackwell, while the N1 20SM has no predecessor or successor listed. Both are marked as Active in production.

The Verdict

From the recorded data, the H200 NVL is the clear choice for any compute-intensive server workload. Its 334,891 OpenCL score at the 100th percentile places it above all other GPUs in the database, and its 4.89 TB/s memory bandwidth and 60.32 TFLOPS FP32 throughput position it for large-scale AI training, scientific computing, and data center acceleration. The 141 GB HBM3e capacity provides ample room for models and datasets that exceed the 128 GB LPDDR5X of the N1 20SM.

The N1 20SM, with no benchmark score and only 12.01 TFLOPS FP32, appears suited to lighter or integrated roles. Its 20 RT cores and single HDMI output suggest it may target edge or embedded systems where display output and moderate compute coexist. The 273.2 GB/s memory bandwidth is far below the H200 NVL, so memory-bound tasks would stall quickly on the N1 20SM.

Buyers needing raw performance should select the H200 NVL. Buyers needing a compact, integrated part with display output and ray tracing support should consider the N1 20SM, but they must accept that the database contains no performance evidence for it. The H200 NVL also holds a 600 W TDP versus an unknown TDP for the N1 20SM, so power planning depends on the N1 20SM's unspecified draw.

FAQ

Q: Which GPU has a higher OpenCL benchmark score?

A: The H200 NVL has a recorded score of 334,891, while the N1 20SM has no recorded benchmark score.

Q: How does the H200 NVL compare to the AMD Instinct MI300X?

A: The H200 NVL scores 5.3% higher than the MI300X, which averages 317,994 points.

Q: Which GPU has more memory bandwidth?

A: The H200 NVL delivers 4.89 TB/s using HBM3e, while the N1 20SM delivers 273.2 GB/s using LPDDR5X.

Q: Does the N1 20SM support display output?

A: Yes, the N1 20SM has 1x HDMI output, while the H200 NVL has no display outputs.

Q: What is the transistor count for each GPU?

A: The H200 NVL has 80,000 million transistors. The N1 20SM has an unknown transistor count.

Q: Which GPU has a larger die size?

A: The H200 NVL has an 814 mm² die, while the N1 20SM has a 382 mm² die.

Specification Differences

| Specification | NVIDIA H200 NVL | NVIDIA N1 20SM |

|---------------|-----------------|----------------|

| Chip | GH100 | GB20B |

| Architecture | Hopper | Blackwell 2.0 |

| Generation | Server Hopper (Hxx) | Blackwell IGP (N1x) |

| Process Node | 5 nm | 5 nm |

| Foundry | TSMC | TSMC |

| Transistors | 80,000 million | unknown |

| Die Size | 814 mm² | 382 mm² |

| Base Clock | 1365 MHz | 741 MHz |

| Boost Clock | 1785 MHz | 2346 MHz |

| Memory Clock | 1593 MHz 6.4 Gbps effective | 1067 MHz 8.5 Gbps effective |

| Memory Size | 141 GB | 128 GB |

| Memory Type | HBM3e | LPDDR5X |

| Memory Bus Width | 6144 bit | 256 bit |

| Memory Bandwidth | 4.89 TB/s | 273.2 GB/s |

| Shading Units | 16896 | 2560 |

| TMUs | 528 | 160 |

| ROPs | 24 | 24 |

| RT Cores | null | 20 |

| Tensor Cores | 528 | 80 |

| Pixel Rate | 42.84 GPixel/s | 56.30 GPixel/s |

| Texture Rate | 942.5 GTexel/s | 375.4 GTexel/s |

| FP32 | 60.32 TFLOPS | 12.01 TFLOPS |

| FP16 | 120.6 TFLOPS (2:1) | 12.01 TFLOPS (1:1) |

| TDP | 600 W | unknown |

| Slot Width | Dual-slot | IGP |

| Power Connectors | 8-pin EPS | None |

| Suggested PSU | 1000 W | null |

| Bus Interface | PCIe 5.0 x16 | PCIe 5.0 x16 |

| Display Outputs | No outputs | 1x HDMI |

| Length | 267 mm 10.5 inches | null |

| Height | 111 mm 4.4 inches | null |

| Release Date | 2024-11-17 | 2026-05-31 |

| Predecessor | Server Ada | null |

| Successor | Server Blackwell | null |

| Percentile vs All GPUs | 100 | 50 |

| Average Benchmark Score | 334891 | 0 |

DETAILED SPECIFICATIONS

SPECIFICATION
H200 NVL
N1 20SM
Core Specs
Shading Units
16,896
2,560 -84.8%
Shaders
16,896
2,560 -84.8%
TMUs
528
160 -69.7%
ROPs
24
24 0.0%
SM Count
132
20 -84.8%
Clocks
Base Clock
1365 MHz
741 MHz
Boost Clock
1785 MHz
2346 MHz
Memory Clock
1593 MHz 6.4 Gbps effective
1067 MHz 8.5 Gbps effective
Memory
Memory Size
141 GB
128 GB
VRAM (MB)
144,384
131,072 -9.2%
Memory Type
HBM3e
LPDDR5X
Memory Bus
6144 bit
256 bit
Bandwidth
4.89 TB/s
273.2 GB/s
Cache
L1 Cache
256 KB (per SM)
128 KB (per SM)
L2 Cache
50 MB
50 MB
Performance
Pixel Rate
42.84 GPixel/s
56.30 GPixel/s
Texture Rate
942.5 GTexel/s
375.4 GTexel/s
FP32 (TFLOPS)
60.32 TFLOPS
12.01 TFLOPS
FP64 (TFLOPS)
30.16 TFLOPS (1:2)
187.7 GFLOPS (1:64)
FP16 (TFLOPS)
120.6 TFLOPS (2:1)
12.01 TFLOPS (1:1)
AI/RT
RT Cores
—
20
Tensor Cores
528
80 -84.8%
Power
TDP
600 W
unknown
TDP (W)
600
—
Suggested PSU
1000 W
—
Power Connectors
8-pin EPS
None
Architecture
Architecture
Hopper
Blackwell 2.0
GPU Name
GH100
GB20B
Generation
Server Hopper (Hxx)
Blackwell IGP (N1x)
Process Size
5 nm
5 nm
Transistors
80,000 million
unknown
Die Size
814 mm²
382 mm²
Foundry
TSMC
TSMC
Density
98.3M / mm²
—
API Support
OpenCL
3.0
3.0
CUDA
9.0
12.1
Physical
Slot Width
Dual-slot
IGP
Length
267 mm 10.5 inches
—
Height
111 mm 4.4 inches
—
Outputs
No outputs
1x HDMI
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Production
Active
Active
Predecessor
Server Ada
—
Successor
Server Blackwell
—
View H200 NVL Details View N1 20SM Details