GEFORCE

NVIDIA B200

NVIDIA graphics card specifications and benchmark scores

90 GB
VRAM
1965
MHz Boost
1000W
TDP
4096
Bus Width
Tensor Cores

At a Glance

NVIDIA
VRAM 90 GB
Boost Clock 1,965 MHz
Shaders 18,944
Bus Width 4096-bit
TDP 1000W
Memory Type HBM3e
Architecture Blackwell
nm
Process 5 nm

NVIDIA B200 Specifications

GPU Core

Shader units and compute resources

The NVIDIA B200 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
18,944
Shaders
18,944
TMUs
592
ROPs
24
SM Count
148

B200 Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the B200's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The B200 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
700 MHz
Base Clock
700 MHz
Boost Clock
1965 MHz
Boost Clock
1,965 MHz
Memory Clock
2000 MHz 8 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's B200 Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The B200's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
90 GB
VRAM
92,160 MB
Memory Type
HBM3e
VRAM Type
HBM3e
Memory Bus
4096 bit
Bus Width
4096-bit
Bandwidth
4.10 TB/s

B200 by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the B200, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
256 KB (per SM)
L2 Cache
50 MB

B200 Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA B200 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
74.45 TFLOPS
FP64 (Double)
37.22 TFLOPS (1:2)
FP16 (Half)
1,191.2 TFLOPS (16:1)
Pixel Rate
47.16 GPixel/s
Texture Rate
1,163.3 GTexel/s

B200 Ray Tracing & AI

Hardware acceleration features

The NVIDIA B200 includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the B200 capable of delivering both stunning graphics and smooth frame rates in modern titles.

Tensor Cores
592

Blackwell Architecture & Process

Manufacturing and design details

The NVIDIA B200 is built on NVIDIA's Blackwell architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the B200 will perform in GPU benchmarks compared to previous generations.

Architecture
Blackwell
GPU Name
GB100
Process Node
5 nm
Foundry
TSMC
Transistors
104,000 million

Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA B200 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the B200 to maintain boost clocks without throttling.

TDP
1000 W
TDP
1000W
Suggested PSU
1400 W

B200 by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA B200 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
SXM Module
Bus Interface
PCIe 5.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA B200. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

OpenCL
3.0
CUDA
10.0

B200 Product Information

Release and pricing details

The NVIDIA B200 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the B200 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Production
Active
Predecessor
Server Hopper
Successor
Server Rubin

About NVIDIA B200

The NVIDIA B200 is a server-grade accelerator built on the Blackwell architecture and GB100 chip, fabricated on a 5 nm process at TSMC with 104,000 million transistors. It targets high-performance computing and AI workloads, offering a distinct profile: massive FP32 throughput, enormous FP16 capability, and a 90 GB HBM3e memory pool, all within a 1000 W power envelope. The data shows a device positioned for compute density rather than consumer graphics, with no display outputs and a focus on raw processing power.

Benchmark Performance

The NVIDIA B200’s benchmark positioning is defined by its absence of traditional gaming or workstation benchmark scores. The available data lists an average benchmark score of 0 and a percentile versus all GPUs of 50, which places it at the median of the database’s tracked devices. However, this percentile is misleading without score comparisons—the nearestRivals array is empty, meaning no direct percentage deltas can be calculated against specific competitors. Instead, the performance analysis must rely on the theoretical throughput figures provided.

In raw FP32 compute, the B200 delivers 74.45 TFLOPS. This is a substantial figure for general compute tasks, though the architecture clearly prioritizes mixed-precision and tensor operations. The FP16 performance is dramatically higher at 1,191.2 TFLOPS, achieved via a 16:1 ratio—indicating that the tensor cores are the primary execution units for AI and deep learning workloads. The texture rate of 1,163.3 GTexel/s and pixel rate of 47.16 GPixel/s are less relevant for server applications but confirm the hardware’s capability to handle rasterization if ever tasked with it. The clock speeds—700 MHz base and 1965 MHz boost—show a wide dynamic range, allowing the accelerator to idle efficiently and then ramp up under load.

Because no rival scores are provided, the benchmark interpretation is absolute rather than relative. The 50th percentile implies that half of all tracked GPUs score higher and half lower, but the B200’s zero average score suggests it is not being evaluated in standard suites. This is typical for server accelerators, which are often excluded from consumer benchmarks. The conclusion is that the B200’s performance is best understood through its compute ceilings: 74.45 TFLOPS for FP32 and 1,191.2 TFLOPS for FP16, with the latter representing the dominant use case.

Who Should Consider It

Given the absence of gaming benchmarks and the lack of display outputs, the B200 is not intended for desktop use. The data indicates a target audience of data centers and research institutions running AI training, inference, or scientific simulations. The 74.45 TFLOPS FP32 performance supports traditional HPC tasks like fluid dynamics or molecular modeling, where double-precision is less critical. For AI workloads, the 1,191.2 TFLOPS FP16 throughput is the headline feature—this is where the accelerator excels, enabling large-scale model training with mixed-precision techniques.

Resolution-based recommendations are inapplicable here, as the B200 has no video outputs. Instead, consider the memory capacity: 90 GB of HBM3e allows holding very large datasets or model parameters in memory, reducing the need for frequent host-device transfers. If your workload requires processing batches of high-resolution imagery or large transformer models, the 90 GB pool is a decisive advantage. The 4096-bit bus width and 4.10 TB/s bandwidth further ensure that memory-bound operations—common in AI—do not stall on data movement. The data suggests this is for users who need massive parallel compute and memory capacity, not for those seeking a consumer graphics card.

Memory Subsystem

The memory configuration is a cornerstone of the B200’s design. It features 90 GB of HBM3e memory, which is a high-bandwidth, stackable memory type optimized for throughput over latency. The bus width is 4096 bits, an exceptionally wide interface that allows simultaneous data transfer across many channels. This yields a bandwidth of 4.10 TB/s—a figure that dwarfs typical GDDR6 or even older HBM implementations.

For high-resolution workloads, this memory subsystem matters in two ways. First, capacity: 90 GB allows entire datasets or model checkpoints to reside on-chip. For example, a large language model with billions of parameters can fit entirely in memory, eliminating the need for sharding across multiple devices. Second, bandwidth: 4.10 TB/s ensures that the 18944 shading units and 592 tensor cores are fed with data without bottlenecks. In practical terms, this means faster iteration times for training loops and higher throughput for inference batch processing. The memory clock is listed as 2000 MHz with 8 Gbps effective, which, combined with the 4096-bit bus, produces the stated bandwidth. This is a balanced design: the memory is not overclocked but relies on width and efficiency.

Power and Cooling

The B200 has a TDP of 1000 W, which is a very high power draw, reflecting its compute density. The suggested PSU rating is 1400 W, meaning that any system integrating this accelerator must have a power supply capable of delivering at least that amount, likely with headroom for other components. The slot width is listed as "SXM Module," indicating it is designed for NVIDIA’s proprietary SXM form factor, which typically requires a baseboard or chassis with integrated cooling—not a standard PCIe slot with a heatsink.

The power connector details are not provided, but the SXM form factor implies a board-level power delivery system rather than the 8-pin or 12VHPWR connectors seen on consumer cards. Cooling is not specified with a cooler type, but the 1000 W TDP necessitates a robust solution—likely liquid cooling or high-flow server airflow. The bus interface is PCIe 5.0 x16, which provides ample bandwidth for host communication, though the primary data path may be through NVLink or other proprietary interconnects in a multi-GPU setup. The 1000 W TDP is the highest among typical server accelerators, so power delivery and thermal management are critical design considerations for any deployment.

How It Compares

The nearestRivals array is empty in the FACT PACK, so no direct competitor comparisons are available. However, the predecessor and successor are identified: "Server Hopper" and "Server Rubin," respectively. This positional data indicates that the B200 sits between two generations of NVIDIA server accelerators. The predecessor, Server Hopper, would logically be based on the Hopper architecture, which is known for its strong FP16 and tensor performance, but the B200’s Blackwell architecture and 1,191.2 TFLOPS FP16 output suggest a significant generational leap. The successor, Server Rubin, is presumably the next iteration, but no specs are provided for either.

Without rival scores, the comparison is qualitative. The B200’s 90 GB memory and 4.10 TB/s bandwidth exceed typical HBM2e or HBM3 configurations seen in prior server parts, indicating a focus on memory capacity as a differentiator. The 592 tensor cores are a key feature, likely enabling the high FP16 throughput. The 1000 W TDP is higher than many data center GPUs, suggesting that performance is prioritized over efficiency. In the absence of numerical rivals, the verdict is that the B200 represents a top-tier, power-hungry choice for AI and HPC, positioned between Hopper and Rubin in the product stack.

FAQ

Q: What is the memory size and type of the NVIDIA B200?

A: The B200 features 90 GB of HBM3e memory, which is a high-bandwidth memory type optimized for throughput.

Q: What is the FP16 performance of the B200?

A: The FP16 throughput is 1,191.2 TFLOPS, achieved via a 16:1 ratio, indicating heavy reliance on tensor cores.

Q: Does the B200 have any display outputs?

A: No, the B200 has "No outputs" listed, confirming it is a compute-only accelerator with no video connectivity.

Q: What is the suggested power supply rating for a system with the B200?

A: The suggested PSU rating is 1400 W, which is higher than the B200’s 1000 W TDP to provide headroom.

Q: What is the bus interface of the B200?

A: The bus interface is PCIe 5.0 x16, though the physical form factor is an SXM Module, not a standard PCIe card.

Q: What is the process node and transistor count?

A: The B200 is fabricated on a 5 nm process at TSMC and contains 104,000 million transistors.

Ray Tracing and Feature Set

The B200’s ray tracing capabilities are not directly specified—the rtCores field is null in the FACT PACK. This absence suggests that ray tracing is not a primary feature for this server accelerator, as the focus is on compute and AI workloads. The tensor cores, however, are explicitly listed at 592, which is a substantial count for accelerating matrix operations, deep learning inference, and training. The FP16 performance of 1,191.2 TFLOPS is directly tied to these tensor cores, which use mixed-precision to achieve high throughput.

The API support is also not provided—directx, opengl, and vulkan are all null. This is consistent with a server product that does not target consumer graphics APIs. Instead, the B200 likely relies on CUDA or other compute frameworks, though those are not listed in the FACT PACK. The architecture is Blackwell, which is the successor to Hopper and predecessor to Rubin, indicating a focus on AI and HPC rather than rasterization. The lack of RT cores and display outputs, combined with the high tensor core count, positions the B200 as a pure compute engine. The 592 TMUs and 24 ROPs are present but secondary; the 1,163.3 GTexel/s texture rate and 47.16 GPixel/s pixel rate are nominal for a device of this class. In summary, the feature set is dominated by tensor processing, with ray tracing and conventional graphics features absent or unreported.

Detailed benchmark scores and charts for the NVIDIA B200 are below.

Benchmark Scores

geekbench_openclSource

Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA B200 handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms.

geekbench_opencl #3 of 650
345,482
89%
Max: 388,405
Compare with other GPUs

Top 5 Performers

#1 NVIDIA RTX 6000D
388,405
#2 NVIDIA B300 SXM6 AC
369,831
#3 NVIDIA B200
345,482
#4 NVIDIA H200 NVL
334,891

Popular NVIDIA B200 Comparisons

See how the B200 stacks up against similar graphics cards from the same generation and competing brands.

Compare with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs