GPU Comparison

NVIDIA
GEFORCE

NVIDIA GB10

CORE STATE GB20B
VRAM 128 GB
CLOCK SPEED 2418 MHz
TDP 140 W
BUS WIDTH 256 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

GeForce RTX 3090 Ti

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1860 MHz
TDP 450 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
120,137
174,441
geekbench_vulkan
114,648
215,633
3dmark_3dmark_steel_nomad_dx12
N/A
5,741

Analysis: NVIDIA GB10 vs NVIDIA GeForce RTX 3090 Ti

The NVIDIA GeForce RTX 3090 Ti and the NVIDIA GB10 represent two radically different interpretations of high-end computing. The RTX 3090 Ti is a massive, triple-slot Ampere-based graphics card designed for peak client performance, while the GB10 is a compact, 140 W Blackwell server component aimed at dense, power-efficient deployments. The benchmark data shows a clear performance hierarchy, but the choice between them hinges on workload type, physical constraints, and architectural priorities rather than raw speed alone. The RTX 3090 Ti wins both shared benchmark comparisons decisively, yet the GB10 counters with double the memory capacity, a newer process node, and a fraction of the power draw.

Where Each One Wins

The RTX 3090 Ti dominates in raw compute and graphics-accelerated tasks. In the two head-to-head benchmarks available, it wins both: Geekbench OpenCL and Geekbench Vulkan. The Vulkan advantage is particularly stark, with the 3090 Ti posting a score of 215,633 against the GB10’s 114,648, a 88.1% lead. This indicates a massive advantage in GPU-compute and rendering APIs that rely on low-level hardware access. The 3090 Ti also holds a significant edge in FP32 throughput, delivering 40.00 TFLOPS versus the GB10’s 29.71 TFLOPS, and its pixel rate of 208.3 GPixel/s is nearly double the GB10’s 116.1 GPixel/s. For any workload that stresses shading units, rasterization, or raw FP32 math, the 3090 Ti is the clear winner.

The GB10 wins in capacity, efficiency, and form-factor flexibility. Its 128 GB of LPDDR5X memory is over five times the 24 GB of GDDR6X on the 3090 Ti, making it the superior choice for large model inference, massive datasets, or in-memory databases. Its 140 W TDP is less than one-third of the 3090 Ti’s 450 W, and it requires no external power connectors, fitting into an IGP form factor at 150 mm by 150 mm. The GB10 also runs on the newer Blackwell 2.0 architecture, built on a 5 nm TSMC process versus the 8 nm Samsung node of the 3090 Ti. While its texture rate is higher at 928.5 GTexel/s versus 625.0 GTexel/s, the GB10’s overall compute ceiling is lower, so its wins are contextual: memory-heavy, power-constrained, or space-constrained environments.

Architecture Differences

The architectural gap between these two is generational. The RTX 3090 Ti uses the GA102 chip on the Ampere architecture, fabricated on Samsung’s 8 nm process. This chip packs 28,300 million transistors into a 628 mm² die, yielding a transistor density of 45.1M per mm². In contrast, the GB10 uses the GB20B chip on the Blackwell 2.0 architecture, built on TSMC’s 5 nm process, with a die size of 382 mm². Transistor count for the GB10 is listed as unknown, but the smaller die on a denser process suggests a fundamentally different design philosophy focused on efficiency rather than absolute transistor count.

Core configurations differ substantially. The 3090 Ti carries 10,752 shading units, 336 TMUs, 112 ROPs, 84 RT cores, and 336 tensor cores. The GB10 has 6,144 shading units, 384 TMUs, 48 ROPs, 48 RT cores, and 384 tensor cores. Notably, the GB10 has more TMUs and tensor cores than the 3090 Ti, which explains its higher texture rate (928.5 GTexel/s vs. 625.0 GTexel/s) despite fewer shading units. The memory subsystems are also divergent: the 3090 Ti uses a 384-bit bus with GDDR6X at 21 Gbps effective, achieving 1.01 TB/s of bandwidth, while the GB10 uses a 256-bit bus with LPDDR5X at 8.5 Gbps effective, delivering 273.2 GB/s. The GB10’s bandwidth is far lower, but its capacity is unmatched.

Clock speeds favor the GB10, with a base of 1665 MHz and boost of 2418 MHz, versus the 3090 Ti’s 1560 MHz base and 1860 MHz boost. The GB10 also supports PCIe 5.0 x16, while the 3090 Ti is limited to PCIe 4.0 x16. However, the 3090 Ti supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the GB10 lists N/A for all three APIs, indicating it is not a client graphics card but a compute or server-focused part. Power delivery reflects this: the 3090 Ti requires a 1x 16-pin connector and an 850 W suggested PSU, while the GB10 has no power connectors and a 300 W suggested PSU.

Head-to-Head Benchmarks

The shared benchmark suite reveals a lopsided contest. In Geekbench OpenCL, the RTX 3090 Ti scores 174,441 against the GB10’s 120,137, a 45.2% advantage. This margin is substantial and reflects the 3090 Ti’s higher FP32 throughput and memory bandwidth. In Geekbench Vulkan, the gap widens dramatically: the 3090 Ti hits 215,633, while the GB10 manages only 114,648, a 88.1% delta. The Vulkan result suggests that the 3090 Ti’s architecture is far more efficient at translating driver-level commands into hardware execution, or that the GB10’s lack of Vulkan API support (listed as N/A) handicaps its performance in this test.

Looking at the nearest rivals provides context for each card’s standing. The 3090 Ti’s average benchmark score is 131,938, which places it 0.7% ahead of the NVIDIA L4 (131,072) and 2.4% behind both the NVIDIA RTX 4000 Ada Generation (135,218) and the NVIDIA A10M (135,230). It also trails the AMD Radeon PRO W6800 (135,396) by 2.6%. This indicates that while the 3090 Ti is a top-tier consumer card, it sits just below the leading workstation offerings in average performance. The GB10’s average score is 117,393, which is 0.3% ahead of the NVIDIA RTX 4000 SFF Ada Generation (117,088) but 1.3% behind the AMD Radeon PRO W7700 (118,976). It holds a 2.6% lead over the NVIDIA Tesla V100 SXM2 16 GB (114,395) and a 3% lead over the NVIDIA RTX A5500 Mobile (113,944). Both cards sit at the 95th percentile versus all GPUs, meaning they are both exceptionally high-performing parts in the broader market.

The wins tally is 2-0 in favor of the 3090 Ti, but the GB10’s advantages in memory capacity and power efficiency do not show up in these compute benchmarks. The 3090 Ti’s 1.01 TB/s bandwidth is a decisive factor in memory-bound workloads, while the GB10’s 273.2 GB/s is a bottleneck for high-throughput compute but sufficient for its intended server tasks.

The Verdict

The data supports a clear split: the RTX 3090 Ti is the pick for raw performance, client graphics, and any workload where compute throughput is king. Its 45.2% lead in OpenCL and 88.1% lead in Vulkan over the GB10 are decisive. It also boasts a higher pixel rate (208.3 GPixel/s vs. 116.1 GPixel/s) and more shading units (10,752 vs. 6,144), making it the superior choice for rendering, gaming, and general-purpose GPU compute. The 3090 Ti’s 1.01 TB/s memory bandwidth is nearly four times the GB10’s, which is critical for high-resolution textures and large data transfers.

The GB10 is the pick for memory capacity, power efficiency, and server density. Its 128 GB of LPDDR5X memory is unmatched by the 3090 Ti’s 24 GB, making it ideal for large language model inference, scientific simulations with massive state spaces, or memory-cached databases. Its 140 W TDP and lack of power connectors allow for deployment in environments where the 3090 Ti’s 450 W requirement and triple-slot footprint would be prohibitive. The GB10’s newer 5 nm process and 384 tensor cores (versus 336 on the 3090 Ti) suggest it is optimized for AI workloads, despite its lower FP32 ceiling.

FAQ

Q: Which GPU has a higher average benchmark score?

A: The NVIDIA GeForce RTX 3090 Ti has an average benchmark score of 131,938, while the NVIDIA GB10 scores 117,393. The 3090 Ti is ahead by 14,545 points.

Q: How much faster is the RTX 3090 Ti in Vulkan compared to the GB10?

A: In Geekbench Vulkan, the RTX 3090 Ti scores 215,633 versus the GB10’s 114,648, a 88.1% advantage for the 3090 Ti.

Q: What is the memory capacity difference between the two?

A: The GB10 has 128 GB of LPDDR5X memory, while the RTX 3090 Ti has 24 GB of GDDR6X. The GB10 offers over five times the capacity.

Q: Which GPU has a higher boost clock?

A: The NVIDIA GB10 has a boost clock of 2418 MHz, compared to the RTX 3090 Ti’s 1860 MHz.

Q: What are the power requirements for each?

A: The RTX 3090 Ti has a TDP of 450 W and requires an 850 W suggested PSU with a 1x 16-pin connector. The GB10 has a TDP of 140 W, requires no power connectors, and has a 300 W suggested PSU.

Q: Which GPU has more tensor cores?

A: The NVIDIA GB10 has 384 tensor cores, while the RTX 3090 Ti has 336 tensor cores. The GB10 also has more TMUs (384 vs. 336).

DETAILED SPECIFICATIONS

SPECIFICATION
GB10
RTX 3090 Ti
Core Specs
Shading Units
6,144
10,752 +75.0%
Shaders
6,144
10,752 +75.0%
TMUs
384
336 -12.5%
ROPs
48
112 +133.3%
SM Count
48
84 +75.0%
Clocks
Base Clock
1665 MHz
1560 MHz
Boost Clock
2418 MHz
1860 MHz
Memory Clock
1067 MHz 8.5 Gbps effective
1313 MHz 21 Gbps effective
Memory
Memory Size
128 GB
24 GB
VRAM (MB)
131,072
24,576 -81.3%
Memory Type
LPDDR5X
GDDR6X
Memory Bus
256 bit
384 bit
Bandwidth
273.2 GB/s
1.01 TB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
50 MB
6 MB
Performance
Pixel Rate
116.1 GPixel/s
208.3 GPixel/s
Texture Rate
928.5 GTexel/s
625.0 GTexel/s
FP32 (TFLOPS)
29.71 TFLOPS
40.00 TFLOPS
FP64 (TFLOPS)
464.3 GFLOPS (1:64)
625.0 GFLOPS (1:64)
FP16 (TFLOPS)
29.71 TFLOPS (1:1)
40.00 TFLOPS (1:1)
AI/RT
RT Cores
48
84 +75.0%
Tensor Cores
384
336 -12.5%
Power
TDP
140 W
450 W
TDP (W)
140
450 +221.4%
Suggested PSU
300 W
850 W
Power Connectors
None
1x 16-pin
Architecture
Architecture
Blackwell 2.0
Ampere
GPU Name
GB20B
GA102
Generation
Server Blackwell (Bxx)
GeForce 30
Process Size
5 nm
8 nm
Transistors
unknown
28,300 million
Die Size
382 mm²
628 mm²
Foundry
TSMC
Samsung
Density
45.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
12.1
8.6
Shader Model
6.8
Physical
Slot Width
IGP
Triple-slot
Length
150 mm 5.9 inches
336 mm 13.2 inches
Height
51 mm 2 inches
140 mm 5.5 inches
Outputs
1x HDMI
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x16
Other
Launch Price
3,999 USD
1,999 USD
Production
Active
End-of-life
Predecessor
Server Hopper
GeForce 20
Successor
Server Rubin
GeForce 40
View GB10 Details View GeForce RTX 3090 Ti Details