NVIDIA GB10 vs NVIDIA GeForce RTX 4090 D Comparison

NVIDIA
GEFORCE

NVIDIA GB10

CORE STATE GB20B
VRAM 128 GB
CLOCK SPEED 2418 MHz
TDP 140 W
BUS WIDTH 256 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

GeForce RTX 4090 D

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 425 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

geekbench_opencl
120,137
278,621
geekbench_vulkan
114,648
246,941
3dmark_3dmark_steel_nomad_dx12
N/A
8,587

Analysis: NVIDIA GB10 vs NVIDIA GeForce RTX 4090 D

The data is unambiguous: the NVIDIA GeForce RTX 4090 D dominates the NVIDIA GB10 in the head-to-head metrics available, delivering more than double the performance in both OpenCL and Vulkan workloads. However, the GB10 counters with a vastly larger memory pool, a newer architecture, and a significantly lower power envelope, positioning it as a specialized server compute device rather than a direct competitor to a consumer flagship graphics card.

Head-to-Head Benchmarks

The benchmark results are decisively in favor of the RTX 4090 D. In the Geekbench OpenCL test, the RTX 4090 D scores 278,621, while the GB10 manages 120,137. This represents a 131.9% advantage for the RTX 4090 D, meaning it is roughly 2.3 times faster in raw compute throughput as measured by this test. The gap is nearly identical in the Geekbench Vulkan test, where the RTX 4090 D scores 246,941 against the GB10's 114,648, a 115.4% lead. This consistent doubling of performance across both API workloads suggests a fundamental difference in compute capability rather than an optimization quirk.

Looking at the broader context, the RTX 4090 D's average benchmark score of 178,050 places it in the 98th percentile of all GPUs. Its nearest rivals are all professional or data-center oriented cards: it trails the NVIDIA RTX PRO 5000 Blackwell by 2.2%, the NVIDIA A100 SXM4 80 GB by 3.1%, the NVIDIA RTX 5000 Ada Generation by 3.6%, and the NVIDIA A100 SXM4 40 GB by 4.9%. This shows that despite its consumer branding, the 4090 D is performing at a level comparable to top-tier workstation accelerators.

The GB10, by contrast, has an average benchmark score of 117,393, placing it in the 95th percentile. Its closest rival is the NVIDIA RTX 4000 SFF Ada Generation, which it edges out by a slim 0.3%. It also sits 1.3% ahead of the AMD Radeon PRO W7700, 2.6% ahead of the NVIDIA Tesla V100 SXM2 16 GB, and 3.0% ahead of the NVIDIA RTX A5500 Mobile. While the 95th percentile is still high, the absolute scores show that the GB10 operates in a different performance class, roughly 34% below the 4090 D's average.

The deltaPct values against rivals tell a clear story. The 4090 D's closest competitor is only 2.2% faster, while the GB10's closest competitor is 0.3% slower. This indicates the 4090 D is near the top of its performance tier, whereas the GB10 is in the middle of a lower tier. The head-to-head numbers confirm that in pure compute performance, the 4090 D is the unequivocal winner, with its wins totaling 2 out of 2 benchmarks.

The Verdict

The verdict hinges on what the user prioritizes: raw compute throughput or memory capacity within a power-constrained environment. If the primary workload is graphics rendering, gaming, or general-purpose compute where FP32 performance is king, the RTX 4090 D is the clear choice. Its 73.54 TFLOPS FP32 output is more than double the GB10's 29.71 TFLOPS, which directly translates to the 115-132% benchmark lead observed.

For tasks that require massive memory capacity, the GB10 is the only option here. Its 128 GB of LPDDR5X memory is over five times the 24 GB available on the 4090 D. This makes the GB10 suitable for large language model inference or datasets that exceed the 4090 D's VRAM. However, the GB10's memory bandwidth is significantly lower at 273.2 GB/s versus 1.01 TB/s, meaning that even with more capacity, data throughput is substantially slower.

The production status also matters. The 4090 D is marked as "End-of-life," while the GB10 is "Active." This suggests the 4090 D is not a forward-looking investment, whereas the GB10 is a current and active product line. The GB10 also has a much lower TDP of 140 W versus 425 W, making it suitable for dense server deployments where power and cooling are at a premium. The 4090 D’s triple-slot design and 800 W suggested PSU requirement further cement its role as a desktop or workstation component, not a server blade.

Architecture Differences

The two GPUs are built on fundamentally different architectures. The RTX 4090 D uses the Ada Lovelace architecture with the AD102 chip, while the GB10 uses the newer Blackwell 2.0 architecture with the GB20B chip. Both are fabricated on a 5 nm process at TSMC, but the similarities end there. The 4090 D's die is much larger at 609 mm², housing 76,300 million transistors, while the GB10's die is 382 mm² with a transistor count listed as unknown.

In terms of compute resources, the 4090 D has 14,592 shading units, 456 TMUs, and 176 ROPs. The GB10 has 6,144 shading units, 384 TMUs, and only 48 ROPs. This disparity in ROPs is particularly stark and explains the massive difference in pixel rate: 443.5 GPixel/s for the 4090 D versus 116.1 GPixel/s for the GB10. The 4090 D also fields 114 RT cores and 456 tensor cores, compared to 48 RT cores and 384 tensor cores on the GB10.

Clock speeds differ as well. The 4090 D has a base clock of 2280 MHz and a boost clock of 2520 MHz. The GB10 has a lower base clock of 1665 MHz but a boost clock of 2418 MHz, which is close to the 4090 D's boost. The memory configurations are entirely different: the 4090 D uses 24 GB of GDDR6X on a 384-bit bus, while the GB10 uses 128 GB of LPDDR5X on a 256-bit bus. This explains the bandwidth disparity: 1.01 TB/s versus 273.2 GB/s.

The bus interface also differs: the 4090 D uses PCIe 4.0 x16, while the GB10 uses the newer PCIe 5.0 x16. The display outputs are minimal on both, with the 4090 D offering 1x HDMI 2.1 and 3x DisplayPort 1.4a, while the GB10 has a single HDMI port. The GB10 has no dedicated power connectors, relying on the PCIe slot, while the 4090 D requires a 1x 16-pin connector.

FAQ

Q: Which GPU is faster in raw compute performance?

A: The NVIDIA GeForce RTX 4090 D is significantly faster. In the Geekbench OpenCL test, it scores 278,621 versus 120,137 for the GB10, a 131.9% difference. In the Vulkan test, it scores 246,941 versus 114,648, a 115.4% lead.

Q: Does the GB10 have any advantage in memory capacity?

A: Yes, the GB10 has 128 GB of LPDDR5X memory, which is substantially more than the 24 GB of GDDR6X on the 4090 D. However, the 4090 D has much higher memory bandwidth at 1.01 TB/s versus 273.2 GB/s.

Q: What is the power consumption difference?

A: The 4090 D has a TDP of 425 W and requires a suggested PSU of 800 W. The GB10 has a TDP of 140 W and a suggested PSU of 300 W.

Q: Which GPU is better for a server environment?

A: The GB10 is better suited for servers. It is an IGP (Integrated Graphics Processor) form factor, has no power connectors, and uses PCIe 5.0 x16. The 4090 D is a triple-slot card with a 1x 16-pin power connector and is marked as end-of-life.

Q: How do these GPUs rank against all other GPUs?

A: The 4090 D is in the 98th percentile of all GPUs, with an average benchmark score of 178,050. The GB10 is in the 95th percentile, with an average benchmark score of 117,393.

Q: What are the production statuses of these cards?

A: The RTX 4090 D is listed as "End-of-life," while the GB10 is listed as "Active."

Where Each One Wins

The RTX 4090 D wins decisively in every benchmark category where both have data. Its 2-0 record in head-to-head tests reflects a fundamental compute advantage. The 4090 D is the winner for any application that relies on FP32 throughput, pixel fill rate, or texture processing. Its 443.5 GPixel/s pixel rate and 1,149.1 GTexel/s texture rate dwarf the GB10's 116.1 GPixel/s and 928.5 GTexel/s. For gaming, real-time ray tracing, or GPU-accelerated rendering, the 4090 D is the superior part.

The GB10 wins in the unbenchmarked categories of memory capacity and power efficiency. Its 128 GB memory pool is unmatched by the 4090 D, making it the choice for workloads that require loading massive datasets into VRAM. Its 140 W TDP is less than a third of the 4090 D's 425 W, and its IGP form factor with no power connectors allows for high-density server deployments. The GB10's 384 tensor cores, while fewer than the 4090 D's 456, still provide substantial AI acceleration capability within a low-power envelope. The GB10's active production status also gives it a lifecycle advantage over the end-of-life 4090 D.

Specification Differences

| Specification | NVIDIA GeForce RTX 4090 D | NVIDIA GB10 |

|---|---|---|

| Architecture | Ada Lovelace | Blackwell 2.0 |

| Process Node | 5 nm | 5 nm |

| Transistors | 76,300 million | unknown |

| Die Size | 609 mm² | 382 mm² |

| Base Clock | 2280 MHz | 1665 MHz |

| Boost Clock | 2520 MHz | 2418 MHz |

| Memory Size | 24 GB | 128 GB |

| Memory Type | GDDR6X | LPDDR5X |

| Memory Bus | 384 bit | 256 bit |

| Memory Bandwidth | 1.01 TB/s | 273.2 GB/s |

| Shading Units | 14592 | 6144 |

| TMUs | 456 | 384 |

| ROPs | 176 | 48 |

| RT Cores | 114 | 48 |

| Tensor Cores | 456 | 384 |

| Pixel Rate | 443.5 GPixel/s | 116.1 GPixel/s |

| Texture Rate | 1,149.1 GTexel/s | 928.5 GTexel/s |

| FP32 Performance | 73.54 TFLOPS | 29.71 TFLOPS |

| TDP | 425 W | 140 W |

| Slot Width | Triple-slot | IGP |

| Power Connectors | 1x 16-pin | None |

| Suggested PSU | 800 W | 300 W |

| Bus Interface | PCIe 4.0 x16 | PCIe 5.0 x16 |

| Production Status | End-of-life | Active |

DETAILED SPECIFICATIONS

SPECIFICATION
GB10
RTX 4090 D
Core Specs
Shading Units
6,144
14,592 +137.5%
Shaders
6,144
14,592 +137.5%
TMUs
384
456 +18.8%
ROPs
48
176 +266.7%
SM Count
48
114 +137.5%
Clocks
Base Clock
1665 MHz
2280 MHz
Boost Clock
2418 MHz
2520 MHz
Memory Clock
1067 MHz 8.5 Gbps effective
1313 MHz 21 Gbps effective
Memory
Memory Size
128 GB
24 GB
VRAM (MB)
131,072
24,576 -81.3%
Memory Type
LPDDR5X
GDDR6X
Memory Bus
256 bit
384 bit
Bandwidth
273.2 GB/s
1.01 TB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
50 MB
72 MB
Performance
Pixel Rate
116.1 GPixel/s
443.5 GPixel/s
Texture Rate
928.5 GTexel/s
1,149.1 GTexel/s
FP32 (TFLOPS)
29.71 TFLOPS
73.54 TFLOPS
FP64 (TFLOPS)
464.3 GFLOPS (1:64)
1,149.1 GFLOPS (1:64)
FP16 (TFLOPS)
29.71 TFLOPS (1:1)
73.54 TFLOPS (1:1)
AI/RT
RT Cores
48
114 +137.5%
Tensor Cores
384
456 +18.8%
Power
TDP
140 W
425 W
TDP (W)
140
425 +203.6%
Suggested PSU
300 W
800 W
Power Connectors
None
1x 16-pin
Architecture
Architecture
Blackwell 2.0
Ada Lovelace
GPU Name
GB20B
AD102
Generation
Server Blackwell (Bxx)
GeForce 40
Process Size
5 nm
5 nm
Transistors
unknown
76,300 million
Die Size
382 mm²
609 mm²
Foundry
TSMC
TSMC
Density
—
125.3M / mm²
API Support
DirectX
—
12 Ultimate (12_2)
OpenGL
—
4.6
Vulkan
—
1.4
OpenCL
3.0
3.0
CUDA
12.1
8.9
Shader Model
—
6.8
Physical
Slot Width
IGP
Triple-slot
Length
150 mm 5.9 inches
304 mm 12 inches
Height
51 mm 2 inches
137 mm 5.4 inches
Outputs
1x HDMI
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x16
Other
Launch Price
3,999 USD
1,599 USD
Production
Active
End-of-life
Predecessor
Server Hopper
GeForce 30
Successor
Server Rubin
GeForce 50
View GB10 Details View GeForce RTX 4090 D Details