NVIDIA B200 SXM6 vs NVIDIA GeForce RTX 4090 D Comparison

NVIDIA
GEFORCE

NVIDIA B200 SXM6

CORE STATE GB100
VRAM 180 GB
CLOCK SPEED 1830 MHz
TDP 1000 W
BUS WIDTH 8192 bit
ARCHITECTURE Blackwell
nm
PROCESS 5 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

GeForce RTX 4090 D

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 425 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
N/A
8,587
geekbench_opencl
N/A
278,621
geekbench_vulkan
N/A
246,941

Analysis: NVIDIA B200 SXM6 vs NVIDIA GeForce RTX 4090 D

The Verdict

The database records two NVIDIA accelerators with fundamentally different design targets. The NVIDIA B200 SXM6 is a server-oriented Blackwell part with no display outputs, while the NVIDIA GeForce RTX 4090 D is a consumer Ada Lovelace graphics card. The recorded benchmark data only covers the RTX 4090 D, which holds a 98th percentile ranking across all GPUs. The B200 SXM6 has no benchmark scores in the database, placing it at the 50th percentile with an average score of zero. The RTX 4090 D shows an average benchmark score of 178,050 across three tests. The B200 SXM6 targets compute-heavy server workloads, while the RTX 4090 D addresses desktop rendering and general graphics tasks. The RTX 4090 D is marked end-of-life, whereas the B200 SXM6 remains active in production. The B200 SXM6 uses a 1000 W TDP with a 1400 W suggested power supply, while the RTX 4090 D draws 425 W with an 800 W suggested PSU. The B200 SXM6 carries a launch MSRP of 34,999 USD; the RTX 4090 D lists at 1,599 USD.

Architecture Differences

The B200 SXM6 uses the GB100 chip on the Blackwell architecture, built on a 5 nm process at TSMC. The RTX 4090 D uses the AD102 chip on Ada Lovelace, also fabricated on a 5 nm node at the same foundry. Transistor counts differ sharply: the B200 SXM6 integrates 208,000 million transistors across a 1628 mm² die, yielding a transistor density of 127.8M per mm². The RTX 4090 D packs 76,300 million transistors on a 609 mm² die, with a density of 125.3M per mm². The B200 SXM6 belongs to the Server Blackwell (Bxx) generation and succeeds Server Hopper, while the RTX 4090 D is part of the GeForce 40-series, succeeding GeForce 30.

Memory architecture separates the two decisively. The B200 SXM6 uses 180 GB of HBM3e on an 8192-bit bus, delivering 8.19 TB/s of bandwidth. The RTX 4090 D uses 24 GB of GDDR6X on a 384-bit bus, with 1.01 TB/s bandwidth. The B200 SXM6 lists a memory clock of 2000 MHz at 8 Gbps effective; the RTX 4090 D runs at 1313 MHz at 21 Gbps effective. Clock behavior also diverges. The B200 SXM6 has a base clock of 120 MHz and a boost of 1830 MHz. The RTX 4090 D starts at 2280 MHz base and boosts to 2520 MHz.

Shader resources differ in count and arrangement. The B200 SXM6 has 18,944 shading units, 592 TMUs, and only 24 ROPs. The RTX 4090 D has 14,592 shading units, 456 TMUs, and 176 ROPs. Tensor core counts show 592 for the B200 SXM6 versus 456 for the RTX 4090 D. The RTX 4090 D includes 114 ray tracing cores; the B200 SXM6 lists none. The B200 SXM6 has no API support for DirectX, OpenGL, or Vulkan, while the RTX 4090 D supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The B200 SXM6 connects via PCIe 6.0 x16; the RTX 4090 D uses PCIe 4.0 x16. The B200 SXM6 is an SXM module with no display outputs; the RTX 4090 D is a triple-slot card with 1x HDMI 2.1 and 3x DisplayPort 1.4a, measuring 304 mm in length, 137 mm in height, and 61 mm in width.

FAQ

Q: Which card has the higher FP32 compute throughput?

A: The RTX 4090 D delivers 73.54 TFLOPS FP32, while the B200 SXM6 provides 69.34 TFLOPS. The difference is 4.2 TFLOPS in favor of the RTX 4090 D.

Q: What are the memory capacities and bandwidths?

A: The B200 SXM6 has 180 GB of HBM3e with 8.19 TB/s bandwidth. The RTX 4090 D has 24 GB of GDDR6X with 1.01 TB/s bandwidth. The B200 SXM6 provides over eight times the memory bandwidth.

Q: Which product is still in production?

A: The B200 SXM6 is listed as Active in production status. The RTX 4090 D is marked End-of-life.

Q: Does the B200 SXM6 support graphics APIs?

A: No. The B200 SXM6 lists DirectX, OpenGL, and Vulkan as N/A. The RTX 4090 D supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4.

Q: What is the power requirement difference?

A: The B200 SXM6 has a TDP of 1000 W with a 1400 W suggested power supply. The RTX 4090 D has a TDP of 425 W with an 800 W suggested PSU.

Q: How do the pixel and texture rates compare?

A: The RTX 4090 D achieves 443.5 GPixel/s pixel rate and 1,149.1 GTexel/s texture rate. The B200 SXM6 achieves 43.92 GPixel/s and 1,083.4 GTexel/s. The RTX 4090 D is over ten times faster in pixel throughput.

Specification Differences

The two GPUs differ across nearly every measurable field. The B200 SXM6 uses the GB100 chip; the RTX 4090 D uses AD102. Architecture: Blackwell versus Ada Lovelace. Generation: Server Blackwell (Bxx) versus GeForce 40. Transistors: 208,000 million versus 76,300 million. Die size: 1628 mm² versus 609 mm². Transistor density: 127.8M per mm² versus 125.3M per mm².

Base clock: 120 MHz versus 2280 MHz. Boost clock: 1830 MHz versus 2520 MHz. Memory clock: 2000 MHz 8 Gbps effective versus 1313 MHz 21 Gbps effective. Memory size: 180 GB versus 24 GB. Memory type: HBM3e versus GDDR6X. Bus width: 8192 bit versus 384 bit. Bandwidth: 8.19 TB/s versus 1.01 TB/s.

Shading units: 18,944 versus 14,592. TMUs: 592 versus 456. ROPs: 24 versus 176. Tensor cores: 592 versus 456. RT cores: none listed versus 114. Pixel rate: 43.92 GPixel/s versus 443.5 GPixel/s. Texture rate: 1,083.4 GTexel/s versus 1,149.1 GTexel/s. FP32: 69.34 TFLOPS versus 73.54 TFLOPS. FP16: 69.34 TFLOPS versus 73.54 TFLOPS, both at 1:1 ratio.

TDP: 1000 W versus 425 W. Slot width: SXM Module versus Triple-slot. Power connectors: none listed versus 1x 16-pin. Suggested PSU: 1400 W versus 800 W. Bus interface: PCIe 6.0 x16 versus PCIe 4.0 x16. Display outputs: none versus 1x HDMI 2.1 and 3x DisplayPort 1.4a. APIs: N/A versus DirectX 12 Ultimate, OpenGL 4.6, Vulkan 1.4. Dimensions: not listed versus 304 mm length, 137 mm height, 61 mm width. Production status: Active versus End-of-life. Release date: 2024-10-31 versus 2023-12-27. Predecessor: Server Hopper versus GeForce 30. Successor: Server Rubin versus GeForce 50. Launch MSRP: 34,999 USD versus 1,599 USD.

Head-to-Head Benchmarks

The database contains no direct head-to-head benchmark entries between the B200 SXM6 and the RTX 4090 D. The wins counter shows zero for both sides. However, the RTX 4090 D has three recorded benchmark scores. In 3dmark_3dmark_steel_nomad_dx12, it scores 8,587. In geekbench_opencl, it scores 278,621. In geekbench_vulkan, it scores 246,941. The average benchmark score across these tests is 178,050.

The B200 SXM6 has no benchmark entries and thus an average score of zero. The percentile ranking confirms the gap: the RTX 4090 D sits at the 98th percentile among all GPUs, while the B200 SXM6 sits at the 50th percentile. The nearest rivals for the RTX 4090 D provide context for its standing. The NVIDIA RTX PRO 5000 Blackwell averages 182,109, which is 2.2% ahead of the RTX 4090 D. The NVIDIA A100 SXM4 80 GB averages 183,725, 3.1% ahead. The NVIDIA RTX 5000 Ada Generation averages 184,664, 3.6% ahead. The NVIDIA A100 SXM4 40 GB averages 187,147, 4.9% ahead.

The FP32 compute numbers show the RTX 4090 D ahead at 73.54 TFLOPS versus 69.34 TFLOPS for the B200 SXM6, a 4.2 TFLOPS margin. Texture rate favors the RTX 4090 D slightly at 1,149.1 GTexel/s versus 1,083.4 GTexel/s. Pixel rate heavily favors the RTX 4090 D at 443.5 GPixel/s versus 43.92 GPixel/s, a tenfold difference. The B200 SXM6 counters with memory bandwidth: 8.19 TB/s versus 1.01 TB/s, an eightfold advantage. The B200 SXM6 also holds a capacity advantage at 180 GB versus 24 GB.

Where Each One Wins

The RTX 4090 D wins in rasterization-oriented metrics. Its 176 ROPs versus 24 ROPs drives the pixel rate advantage. Its 73.54 TFLOPS FP32 exceeds the B200 SXM6's 69.34 TFLOPS. The texture rate margin, while smaller, still favors the RTX 4090 D. The presence of 114 ray tracing cores, graphics API support, and display outputs makes the RTX 4090 D the only one of the two suited for rendering workloads, gaming, or any task requiring a video output. Its triple-slot form factor and 425 W TDP fit standard desktop configurations.

The B200 SXM6 wins in memory-bound and capacity-driven workloads. The 180 GB HBM3e pool with 8.19 TB/s bandwidth dwarfs the RTX 4090 D's 24 GB at 1.01 TB/s. The 8192-bit bus width enables data movement at a scale the RTX 4090 D cannot approach. The B200 SXM6 also carries more tensor cores at 592 versus 456, suggesting an advantage in matrix-heavy operations. Its SXM module form factor and PCIe 6.0 x16 interface target server platforms. The lack of display outputs and graphics APIs confirms a compute-only role.

The recorded benchmark scores apply exclusively to the RTX 4090 D, which holds a 98th percentile position. Its nearest rivals all score within 4.9% of its average, indicating a tight competitive cluster. The B200 SXM6 has no comparable scores in the database, so its relative standing cannot be quantified from benchmark results. The production status differs: the B200 SXM6 remains active while the RTX 4090 D is end-of-life. The release dates show the B200 SXM6 arrived on 2024-10-31, roughly ten months after the RTX 4090 D's 2023-12-27 launch. The database records the B200 SXM6's successor as Server Rubin and the RTX 4090 D's successor as GeForce 50.

DETAILED SPECIFICATIONS

SPECIFICATION
B200 SXM6
RTX 4090 D
Core Specs
Shading Units
18,944
14,592 -23.0%
Shaders
18,944
14,592 -23.0%
TMUs
592
456 -23.0%
ROPs
24
176 +633.3%
SM Count
148
114 -23.0%
Clocks
Base Clock
120 MHz
2280 MHz
Boost Clock
1830 MHz
2520 MHz
Memory Clock
2000 MHz 8 Gbps effective
1313 MHz 21 Gbps effective
Memory
Memory Size
180 GB
24 GB
VRAM (MB)
184,320
24,576 -86.7%
Memory Type
HBM3e
GDDR6X
Memory Bus
8192 bit
384 bit
Bandwidth
8.19 TB/s
1.01 TB/s
Cache
L1 Cache
256 KB (per SM)
128 KB (per SM)
L2 Cache
126 MB
72 MB
Performance
Pixel Rate
43.92 GPixel/s
443.5 GPixel/s
Texture Rate
1,083.4 GTexel/s
1,149.1 GTexel/s
FP32 (TFLOPS)
69.34 TFLOPS
73.54 TFLOPS
FP64 (TFLOPS)
34.67 TFLOPS (1:2)
1,149.1 GFLOPS (1:64)
FP16 (TFLOPS)
69.34 TFLOPS (1:1)
73.54 TFLOPS (1:1)
AI/RT
RT Cores
114
Tensor Cores
592
456 -23.0%
Power
TDP
1000 W
425 W
TDP (W)
1,000
425 -57.5%
Suggested PSU
1400 W
800 W
Power Connectors
1x 16-pin
Architecture
Architecture
Blackwell
Ada Lovelace
GPU Name
GB100
AD102
Generation
Server Blackwell (Bxx)
GeForce 40
Process Size
5 nm
5 nm
Transistors
208,000 million
76,300 million
Die Size
1628 mm²
609 mm²
Foundry
TSMC
TSMC
Density
127.8M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
10.0
8.9
Shader Model
6.8
Physical
Slot Width
SXM Module
Triple-slot
Length
304 mm 12 inches
Height
137 mm 5.4 inches
Outputs
No outputs
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 6.0 x16
PCIe 4.0 x16
Other
Launch Price
34,999 USD
1,599 USD
Production
Active
End-of-life
Predecessor
Server Hopper
GeForce 30
Successor
Server Rubin
GeForce 50
View B200 SXM6 Details View GeForce RTX 4090 D Details