NVIDIA B200 vs NVIDIA Tesla V100 PCIe 32 GB Comparison

NVIDIA
GEFORCE

NVIDIA B200

CORE STATE GB100
VRAM 90 GB
CLOCK SPEED 1965 MHz
TDP 1000 W
BUS WIDTH 4096 bit
ARCHITECTURE Blackwell
nm
PROCESS 5 nm
LAUNCH DATE —
VS
NVIDIA
GEFORCE

Tesla V100 PCIe 32 GB

CORE STATE GV100
VRAM 32 GB
CLOCK SPEED 1380 MHz
TDP 250 W
BUS WIDTH 4096 bit
ARCHITECTURE Volta
nm
PROCESS 12 nm
LAUNCH DATE 2018

PERFORMANCE BENCHMARKS

geekbench_opencl
345,482
168,763
geekbench_vulkan
N/A
131,847

Analysis: NVIDIA B200 vs NVIDIA Tesla V100 PCIe 32 GB

The NVIDIA B200 and NVIDIA Tesla V100 PCIe 32 GB represent two distinct eras of NVIDIA’s server acceleration, separated by a massive generational gap in both architecture and performance. The data shows that the B200, built on the Blackwell architecture, holds a decisive advantage in the only shared benchmark, while the V100 remains a legacy product with a different set of trade-offs. This analysis breaks down the numbers to help you decide which card fits your workload, strictly based on the available facts.

Head-to-Head Benchmarks

The only direct comparison available is the Geekbench OpenCL test, and the results are not close. The NVIDIA B200 scores 345,482, while the Tesla V100 PCIe 32 GB scores 168,763. That is a delta of 104.7% — the B200 is more than twice as fast in this compute-heavy workload. The B200 wins the head-to-head with 1 win to 0.

To put that in context, the B200’s score places it at the 100th percentile of all GPUs, while the V100 sits at the 96th percentile. The B200’s nearest rivals include the NVIDIA B300 SXM6 AC (369,831, 6.6% ahead), the NVIDIA H200 NVL (334,891, 3.2% behind), and the AMD Instinct MI300X (317,994, 8.6% behind). The V100, by contrast, is in a completely different tier; its nearest rival is the NVIDIA A10G (151,963, 1.1% behind), and it trails the AMD Radeon Pro W6800X (160,671) by 6.5%. The B200’s OpenCL result is 104.7% higher than the V100’s, meaning the older card would need roughly two of its own units to match a single B200 in this test.

The delta between the two is larger than the gap between the B200 and any of its listed rivals. For example, the B200 is only 8.6% ahead of the AMD Instinct MI300X, but it is 104.7% ahead of the V100. This illustrates how the V100 has been left behind by the current generation of accelerators. The V100’s Vulkan score of 131,847 is not directly comparable to the B200, as the B200 has no listed Vulkan result, but it does indicate the V100’s performance ceiling in graphics-oriented APIs.

Where Each One Wins

Based on the data, the B200 is the clear winner in raw compute throughput. Its OpenCL score of 345,482 dwarfs the V100’s 168,763, making it the obvious choice for any workload that relies on general-purpose GPU compute, such as large-scale AI training or scientific simulations. The B200’s 74.45 TFLOPS of FP32 performance and 1,191.2 TFLOPS of FP16 (16:1) performance are far beyond the V100’s 14.13 TFLOPS FP32 and 28.26 TFLOPS FP16 (2:1). For tasks that need massive floating-point throughput, the B200 is the only rational option.

The V100, however, retains a niche. Its 32 GB of HBM2 memory, while smaller than the B200’s 90 GB of HBM3e, is still sufficient for many legacy workloads. The V100’s 897.0 GB/s of memory bandwidth is respectable, though it is less than a quarter of the B200’s 4.10 TB/s. The V100’s advantage lies in its compatibility and lower power footprint — it is a dual-slot card with a 250 W TDP, whereas the B200 is an SXM module with a 1000 W TDP. If your system is built around PCIe 3.0 and cannot accommodate an SXM module, the V100 is the only one of the two that fits. The V100 also has API support for DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4, while the B200 lists no such APIs, making the V100 the more flexible choice for any compute tasks that touch graphics APIs.

The data does not show any workload where the V100 wins on performance. It wins only on form factor, power consumption, and legacy software compatibility. For modern high-performance computing, the B200’s benchmark lead is insurmountable.

Architecture Differences

The two cards are built on fundamentally different architectures. The B200 uses the GB100 chip on the Blackwell architecture, fabricated on a 5 nm process at TSMC. It packs 104,000 million transistors. The V100 uses the GV100 chip on the Volta architecture, built on a 12 nm process, also at TSMC, with 21,100 million transistors and a die size of 815 mm². The B200’s process node is more advanced, allowing for a significantly higher transistor count.

The memory subsystems are also vastly different. The B200 has 90 GB of HBM3e memory on a 4096-bit bus, yielding 4.10 TB/s of bandwidth. The V100 has 32 GB of HBM2 on the same 4096-bit bus width, but only achieves 897.0 GB/s. The B200’s memory clock is 2000 MHz (8 Gbps effective), while the V100’s is 876 MHz (1752 Mbps effective). The B200’s memory bandwidth advantage is roughly 4.6x.

Core configurations diverge sharply. The B200 has 18,944 shading units, 592 TMUs, and 24 ROPs, with 592 tensor cores. The V100 has 5,120 shading units, 320 TMUs, and 128 ROPs, with 640 tensor cores. Notably, the V100 has more ROPs (128 vs. 24) and more tensor cores (640 vs. 592), but the B200’s clock speeds and architecture efficiency more than compensate. The B200’s base clock is 700 MHz with a boost of 1965 MHz; the V100’s base is 1230 MHz with a boost of 1380 MHz. The B200’s pixel rate is 47.16 GPixel/s, but its texture rate is 1,163.3 GTexel/s; the V100’s pixel rate is 176.6 GPixel/s and its texture rate is 441.6 GTexel/s. The V100’s higher ROP count gives it a pixel rate advantage, but the B200 dominates texture fill.

The B200 is a PCIe 5.0 x16 device, while the V100 uses PCIe 3.0 x16. The B200 is an SXM module, requiring a different physical slot than the V100’s dual-slot PCIe form factor. The B200’s TDP is 1000 W, with a suggested PSU of 1400 W; the V100’s TDP is 250 W, with a suggested PSU of 600 W. The production status also differs: the B200 is Active, while the V100 is End-of-life.

FAQ

Q: Which card has the higher OpenCL benchmark score?

A: The NVIDIA B200 scores 345,482, which is 104.7% higher than the Tesla V100’s 168,763.

Q: How does the B200 compare to its closest rivals?

A: The B200 is 3.2% ahead of the NVIDIA H200 NVL (334,891), 8.6% ahead of the AMD Instinct MI300X (317,994), and 16.8% ahead of the NVIDIA L40S (295,763). It is 6.6% behind the NVIDIA B300 SXM6 AC (369,831).

Q: Is the V100 still competitive with modern accelerators?

A: No. The V100’s nearest rival is the NVIDIA A10G (151,963), which is 1.1% behind it. It also trails the AMD Radeon Pro W6800X (160,671) by 6.5% and the NVIDIA A100 PCIe 40 GB (162,504) by 7.5%. Its 96th percentile ranking is far below the B200’s 100th.

Q: What are the memory size and bandwidth differences?

A: The B200 has 90 GB of HBM3e with 4.10 TB/s bandwidth, while the V100 has 32 GB of HBM2 with 897.0 GB/s bandwidth. Both use a 4096-bit bus.

Q: Can the V100 run graphical workloads?

A: Yes, the V100 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The B200 lists no API support, making the V100 the only one with documented graphics API compatibility.

Q: What is the power requirement for each card?

A: The B200 has a TDP of 1000 W and a suggested PSU of 1400 W. The V100 has a TDP of 250 W and a suggested PSU of 600 W.

Specification Differences

  • Process Node: B200 is 5 nm; V100 is 12 nm.
  • Transistors: B200 has 104,000 million; V100 has 21,100 million.
  • Die Size: V100 is 815 mm²; B200 has no listed die size.
  • Base Clock: B200 is 700 MHz; V100 is 1230 MHz.
  • Boost Clock: B200 is 1965 MHz; V100 is 1380 MHz.
  • Memory Size: B200 is 90 GB; V100 is 32 GB.
  • Memory Type: B200 is HBM3e; V100 is HBM2.
  • Memory Bandwidth: B200 is 4.10 TB/s; V100 is 897.0 GB/s.
  • Shading Units: B200 has 18,944; V100 has 5,120.
  • TMUs: B200 has 592; V100 has 320.
  • ROPs: B200 has 24; V100 has 128.
  • Tensor Cores: B200 has 592; V100 has 640.
  • FP32 Performance: B200 is 74.45 TFLOPS; V100 is 14.13 TFLOPS.
  • FP16 Performance: B200 is 1,191.2 TFLOPS (16:1); V100 is 28.26 TFLOPS (2:1).
  • TDP: B200 is 1000 W; V100 is 250 W.
  • Slot Width: B200 is SXM Module; V100 is Dual-slot.
  • Power Connectors: V100 uses 2x 8-pin; B200 has none listed.
  • Suggested PSU: B200 is 1400 W; V100 is 600 W.
  • Bus Interface: B200 is PCIe 5.0 x16; V100 is PCIe 3.0 x16.
  • APIs: V100 supports DirectX 12, OpenGL 4.6, Vulkan 1.4; B200 lists none.
  • Production Status: B200 is Active; V100 is End-of-life.

The Verdict

The data is unambiguous. The NVIDIA B200 is the superior accelerator for any compute-intensive task, offering 104.7% higher OpenCL performance than the Tesla V100. Its 90 GB of HBM3e memory, 4.10 TB/s bandwidth, and 74.45 TFLOPS FP32 throughput make it a top-tier choice, ranking in the 100th percentile of all GPUs. The B200 is the only option here if you need maximum performance for modern AI or scientific workloads.

The Tesla V100 PCIe 32 GB, however, is not without purpose. Its 250 W TDP and dual-slot form factor make it far easier to integrate into existing servers, and its API support for Vulkan and DirectX extends its utility beyond pure compute. It is an end-of-life product, but it remains a capable option for legacy systems or workloads that do not require the B200’s immense power. If your system cannot support an SXM module or a 1000 W TDP, the V100 is the practical choice. Otherwise, the B200’s benchmark results make it the clear winner.

DETAILED SPECIFICATIONS

SPECIFICATION
B200
Tesla V100 PCIe 32 GB
Core Specs
Shading Units
18,944
5,120 -73.0%
Shaders
18,944
5,120 -73.0%
TMUs
592
320 -45.9%
ROPs
24
128 +433.3%
SM Count
148
80 -45.9%
Clocks
Base Clock
700 MHz
1230 MHz
Boost Clock
1965 MHz
1380 MHz
Memory Clock
2000 MHz 8 Gbps effective
876 MHz 1752 Mbps effective
Memory
Memory Size
90 GB
32 GB
VRAM (MB)
92,160
32,768 -64.4%
Memory Type
HBM3e
HBM2
Memory Bus
4096 bit
4096 bit
Bandwidth
4.10 TB/s
897.0 GB/s
Cache
L1 Cache
256 KB (per SM)
128 KB (per SM)
L2 Cache
50 MB
6 MB
Performance
Pixel Rate
47.16 GPixel/s
176.6 GPixel/s
Texture Rate
1,163.3 GTexel/s
441.6 GTexel/s
FP32 (TFLOPS)
74.45 TFLOPS
14.13 TFLOPS
FP64 (TFLOPS)
37.22 TFLOPS (1:2)
7.066 TFLOPS (1:2)
FP16 (TFLOPS)
1,191.2 TFLOPS (16:1)
28.26 TFLOPS (2:1)
AI/RT
Tensor Cores
592
640 +8.1%
Power
TDP
1000 W
250 W
TDP (W)
1,000
250 -75.0%
Suggested PSU
1400 W
600 W
Power Connectors
—
2x 8-pin
Architecture
Architecture
Blackwell
Volta
GPU Name
GB100
GV100
Generation
Server Blackwell (Bxx)
Tesla Volta (Vxx)
Process Size
5 nm
12 nm
Transistors
104,000 million
21,100 million
Die Size
—
815 mm²
Foundry
TSMC
TSMC
Density
—
25.9M / mm²
API Support
DirectX
—
12 (12_1)
OpenGL
—
4.6
Vulkan
—
1.4
OpenCL
3.0
3.0
CUDA
10.0
7.0
Shader Model
—
6.8
Physical
Slot Width
SXM Module
Dual-slot
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 3.0 x16
Other
Production
Active
End-of-life
Predecessor
Server Hopper
Tesla Pascal
Successor
Server Rubin
Tesla Turing
View B200 Details View Tesla V100 PCIe 32 GB Details