GEFORCE

NVIDIA B100

NVIDIA graphics card specifications and benchmark scores

96 GB
VRAM
975
MHz Boost
1000W
TDP
4096
Bus Width
Tensor Cores

At a Glance

NVIDIA
VRAM 96 GB
Boost Clock 975 MHz
Shaders 16,896
Bus Width 4096-bit
TDP 1000W
Memory Type HBM3e
Architecture Blackwell
nm
Process 5 nm

NVIDIA B100 Specifications

GPU Core

Shader units and compute resources

The NVIDIA B100 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
16,896
Shaders
16,896
TMUs
528
ROPs
24
SM Count
132

B100 Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the B100's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The B100 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
700 MHz
Base Clock
700 MHz
Boost Clock
975 MHz
Boost Clock
975 MHz
Memory Clock
2000 MHz 8 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's B100 Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The B100's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
96 GB
VRAM
98,304 MB
Memory Type
HBM3e
VRAM Type
HBM3e
Memory Bus
4096 bit
Bus Width
4096-bit
Bandwidth
4.10 TB/s

B100 by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the B100, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
256 KB (per SM)
L2 Cache
50 MB

B100 Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA B100 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
32.95 TFLOPS
FP64 (Double)
16.47 TFLOPS (1:2)
FP16 (Half)
131.8 TFLOPS (4:1)
Pixel Rate
23.40 GPixel/s
Texture Rate
514.8 GTexel/s

B100 Ray Tracing & AI

Hardware acceleration features

The NVIDIA B100 includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the B100 capable of delivering both stunning graphics and smooth frame rates in modern titles.

Tensor Cores
528

Blackwell Architecture & Process

Manufacturing and design details

The NVIDIA B100 is built on NVIDIA's Blackwell architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the B100 will perform in GPU benchmarks compared to previous generations.

Architecture
Blackwell
GPU Name
GB102
Process Node
5 nm
Foundry
TSMC
Transistors
104,000 million

Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA B100 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the B100 to maintain boost clocks without throttling.

TDP
1000 W
TDP
1000W
Suggested PSU
1400 W

B100 by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA B100 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
SXM Module
Bus Interface
PCIe 5.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA B100. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

OpenCL
3.0
CUDA
10.1

B100 Product Information

Release and pricing details

The NVIDIA B100 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the B100 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Production
Active
Predecessor
Server Hopper
Successor
Server Rubin

About NVIDIA B100

NVIDIA B100 is a server-focused Blackwell architecture GPU built on TSMC’s 5 nm process, packing 104,000 million transistors into a GB102 chip, and the data shows it is positioned as a compute-oriented accelerator rather than a consumer graphics card, with a 50th percentile ranking among all GPUs and no benchmark scores or nearest rivals on record, making its performance profile defined almost entirely by its raw specification sheet.

Benchmark Performance

The FACT PACK lists no benchmark scores, no average benchmark score, and no nearest rivals for the NVIDIA B100. The `avgBenchmarkScore` field is 0, and `nearestRivals` is an empty array. The `percentileVsAllGpus` stands at 50, which places it exactly at the midpoint of all GPUs in the database — but without any measured scores, this percentile likely reflects its specification-based tier rather than direct testing. The FP32 compute is 32.95 TFLOPS, while FP16 reaches 131.8 TFLOPS using a 4:1 ratio. Texture rate is 514.8 GTexel/s, and pixel rate is 23.40 GPixel/s. These figures indicate a design where FP32 throughput is substantial but FP16 is the dominant mode, suggesting the B100 is optimized for mixed-precision workloads. The 24 ROPs are notably low relative to the 16,896 shading units and 528 TMUs, which implies the chip prioritizes compute throughput over rasterization efficiency. In the absence of rival comparisons, the data cannot show deltas, but the raw numbers alone suggest a card that would likely lead in FP16-heavy tasks while lagging in traditional pixel-pushing scenarios.

Ray Tracing and Feature Set

The FACT PACK contains no `rtCores` field value — it is listed as `null`. Similarly, there are no DirectX, OpenGL, or Vulkan API version numbers provided. The tensor core count is 528, which is identical to the TMU count, indicating a heavy investment in matrix math acceleration. The architecture is Blackwell, and the generation is listed as "Server Blackwell (Bxx)," confirming this is not a consumer gaming product. Display outputs are explicitly "No outputs," meaning the B100 is not designed to drive a monitor. The bus interface is PCIe 5.0 x16, which provides modern host connectivity, but the absence of any display or API data means ray tracing capabilities cannot be assessed from this pack. The tensor cores are the only specialized acceleration units confirmed, and with 528 of them, the data implies a focus on AI inference and training rather than real-time graphics. Without RT core counts or API support, any statement about ray tracing performance would be speculation, so the dataset restricts analysis to tensor core presence and the compute-oriented feature set.

Power and Cooling

The thermal design power is 1000 W, which is a high figure that demands serious cooling infrastructure. The suggested PSU is 1400 W, indicating that a power supply rated for that output is recommended for systems hosting this module. The slot width is listed as "SXM Module," which means the B100 is designed for server chassis with SXM sockets rather than standard PCIe slots. No power connector details are provided — the field is `null` — so the exact pin configuration is unknown, but the SXM form factor typically implies a board-level power delivery system. The 1000 W TDP combined with the 1400 W PSU recommendation suggests that the B100 draws close to its TDP under load, and system designers must account for that draw alongside other components. The absence of a length, height, or width measurement reinforces that this is a module, not a plug-in card. Cooling solutions are not specified, but a 1000 W TDP in a server context generally requires active cooling, often with high-static-pressure fans or liquid cooling loops, though no such specifics appear in the FACT PACK.

Who Should Consider It

Given the lack of benchmark scores, the recommendation must be based on specification data. The 96 GB HBM3e memory with a 4096-bit bus and 4.10 TB/s bandwidth makes the B100 suitable for workloads that require massive datasets resident on the GPU — training large language models, running scientific simulations, or processing high-resolution volumetric data. The FP16 throughput of 131.8 TFLOPS suggests that half-precision compute tasks, such as neural network training with mixed precision, would see strong performance. The FP32 figure of 32.95 TFLOPS is also robust for single-precision scientific computing. The 1000 W TDP and SXM form factor mean this is not for desktop users; it targets data center racks with adequate power and cooling. The absence of display outputs confirms it is not for gaming or workstation visualization. Users who need to render frames to a screen should look elsewhere. The data implies that the B100 suits researchers and cloud providers who can leverage its memory capacity and tensor cores for AI workloads, not consumers seeking high-resolution gaming performance.

How It Compares

No nearest rivals are listed in the FACT PACK, so a direct comparison against specific competing products cannot be made. The `nearestRivals` array is empty, and no rival names, scores, or deltaPct values are provided. The predecessor is "Server Hopper" and the successor is "Server Rubin," which places the B100 in a generational timeline, but no performance data for those predecessors or successors exists in this pack. Without rival benchmarks, the B100’s 50th percentile ranking cannot be contextualized against other GPUs. The absence of comparisons is itself informative: it suggests the database has not yet accumulated enough testing data for this SKU, or that its server-oriented nature places it outside the typical consumer GPU comparison pool. The data cannot show whether the B100 outperforms or trails any specific rival, so any assertion of relative performance would be unsupported.

Memory Subsystem

The B100 is equipped with 96 GB of HBM3e memory, which is a high-capacity configuration aimed at large-scale compute. The bus width is 4096 bits, a wide interface that enables the memory to deliver 4.10 TB/s of bandwidth. This bandwidth figure is substantial, meaning the GPU can feed its 16,896 shading units and 528 tensor cores with data at very high rates. For high-resolution workloads, such as 4K or 8K rendering, the memory capacity would allow entire scenes or datasets to reside on the GPU without spilling to system RAM. However, the 24 ROPs are a limiting factor for rasterization: even with massive bandwidth, the pixel throughput of 23.40 GPixel/s caps how fast the GPU can write final pixels to a framebuffer. The memory clock is listed as 2000 MHz with 8 Gbps effective, which is a modest clock speed but compensated by the extremely wide bus. HBM3e is a high-bandwidth memory type designed for server AI accelerators, not consumer graphics, so this subsystem is optimized for sustained data movement in compute kernels rather than frame-buffer access patterns. The 4.10 TB/s bandwidth is likely the key differentiator for memory-bound tasks, as it allows rapid access to the 96 GB pool.

FAQ

Q: What is the FP32 compute performance of the NVIDIA B100?

A: The FACT PACK lists FP32 throughput as 32.95 TFLOPS.

Q: How much memory does the B100 have, and what type is it?

A: It has 96 GB of HBM3e memory with a 4096-bit bus width and 4.10 TB/s bandwidth.

Q: Does the B100 support any display outputs?

A: No, the FACT PACK lists "No outputs" for display outputs, meaning it is not designed for monitor connection.

Q: What is the recommended power supply for a system with the B100?

A: The suggested PSU is 1400 W, while the GPU’s TDP is 1000 W.

Q: What is the tensor core count on the B100?

A: The B100 has 528 tensor cores, according to the FACT PACK.

Q: What bus interface does the B100 use?

A: It uses PCIe 5.0 x16, though the slot width is listed as SXM Module, indicating a server form factor.

Detailed benchmark scores and charts for the NVIDIA B100 are below.

Benchmark Scores

No benchmark data available for this GPU.

Compare with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs