NVIDIA GeForce RTX 4090 D vs NVIDIA Rubin GPU Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4090 D

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 425 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

Rubin GPU

CORE STATE GR100
VRAM 288 GB
CLOCK SPEED 2267 MHz
TDP 2300 W
BUS WIDTH 16384 bit
ARCHITECTURE Rubin
nm
PROCESS 3 nm
LAUNCH DATE 2026

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
8,587
N/A
geekbench_opencl
278,621
N/A
geekbench_vulkan
246,941
N/A

Analysis: NVIDIA GeForce RTX 4090 D vs NVIDIA Rubin GPU

Head-to-Head Benchmarks

The recorded database contains no direct head-to-head benchmark comparisons between the NVIDIA GeForce RTX 4090 D and the NVIDIA Rubin GPU. The RTX 4090 D has three benchmark entries: a 3DMark Steel Nomad DX12 score of 8587, a Geekbench OpenCL score of 278621, and a Geekbench Vulkan score of 246941. The Rubin GPU has no benchmark entries in the database, with an average benchmark score of 0 and a percentile rank of 50 against all GPUs.

The RTX 4090 D holds an average benchmark score of 178050 across all recorded tests, placing it in the 98th percentile of all GPUs. Its nearest rivals in the database show how tightly clustered this performance tier is: the NVIDIA RTX PRO 5000 Blackwell averages 182109 (2.2% higher), the NVIDIA A100 SXM4 80 GB averages 183725 (3.1% higher), the NVIDIA RTX 5000 Ada Generation averages 184664 (3.6% higher), and the NVIDIA A100 SXM4 40 GB averages 187147 (4.9% higher). These deltas indicate the RTX 4090 D sits just below a group of professional-grade accelerators, trailing the closest rival by a narrow margin.

The Rubin GPU's lack of recorded benchmarks means no comparative score analysis is possible from the database. Its percentile rank of 50 is the default baseline for an unmeasured device, not an indicator of real-world performance. The data shows a complete absence of measurable results for this part, which contrasts sharply with the RTX 4090 D's substantial benchmark presence.

Architecture Differences

The two GPUs come from entirely different architectural lineages. The RTX 4090 D uses the AD102 chip based on Ada Lovelace architecture, built on a 5 nm process at TSMC. The Rubin GPU uses the GR100 chip based on Rubin architecture, also fabricated by TSMC but on a 3 nm node. This process difference represents a significant jump in manufacturing technology between the two designs.

Transistor counts reveal the scale gap: the RTX 4090 D packs 76,300 million transistors on a 609 mm² die, yielding a transistor density of 125.3 million per square millimeter. The Rubin GPU contains 336,000 million transistors on a 1456 mm² die, with a density of 230.8 million per square millimeter. The Rubin GPU carries more than four times the transistor count and more than double the die area, while also achieving nearly double the density per square millimeter.

Memory configurations diverge dramatically. The RTX 4090 D uses 24 GB of GDDR6X across a 384-bit bus, delivering 1.01 TB/s of bandwidth. The Rubin GPU uses 288 GB of HBM4 across a 16384-bit bus, delivering 22.1 TB/s. That is roughly 22 times the bandwidth and 12 times the capacity. The memory clock rates also differ: the RTX 4090 D runs at 1313 MHz (21 Gbps effective), while the Rubin GPU runs at 2695 MHz (10.8 Gbps effective), reflecting the different memory technologies in play.

Compute resources show a clear hierarchy. The RTX 4090 D has 14592 shading units, 456 texture mapping units, 176 render output units, 114 ray tracing cores, and 456 tensor cores. The Rubin GPU has 28672 shading units, 896 texture mapping units, 24 render output units, no recorded ray tracing cores, and 896 tensor cores. The Rubin GPU doubles the shading units, TMUs, and tensor cores, but its ROPS count is dramatically lower at 24 versus 176.

Clock speeds tell a different story. The RTX 4090 D runs at a base clock of 2280 MHz and boosts to 2520 MHz. The Rubin GPU has a much lower base clock of 700 MHz but boosts to 2267 MHz. The RTX 4090 D operates at a higher sustained frequency, while the Rubin GPU relies on its massive core count and memory bandwidth rather than raw clock speed.

The power envelope separates these parts into different categories entirely. The RTX 4090 D has a TDP of 425 W with a suggested PSU of 800 W. The Rubin GPU has a TDP of 2300 W with a suggested PSU of 2700 W. The Rubin GPU consumes more than five times the power of the RTX 4090 D.

Physical formats also differ. The RTX 4090 D is a triple-slot card measuring 304 mm in length, 137 mm in height, and 61 mm in width, using a single 16-pin power connector. The Rubin GPU is an SXM module with no recorded dimensions, no display outputs, and no power connector listed. The bus interface moves from PCIe 4.0 x16 on the RTX 4090 D to PCIe 6.0 x16 on the Rubin GPU.

API support is another dividing line. The RTX 4090 D supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The Rubin GPU lists N/A for all graphics APIs, confirming its server-oriented design with no consumer graphics path.

Where Each One Wins

The RTX 4090 D wins in every measured category because it is the only one with recorded data. Its 98th percentile ranking against all GPUs indicates it sits at the top tier of measured hardware. The OpenCL score of 278621 and Vulkan score of 246941 show strong compute and graphics API performance, while the 3DMark Steel Nomad DX12 score of 8587 demonstrates modern DirectX 12 workload capability.

The RTX 4090 D also wins on practical graphics features. It offers display outputs (1x HDMI 2.1 and 3x DisplayPort 1.4a), making it usable in consumer and workstation display environments. Its triple-slot form factor fits standard PC chassis, and its 425 W TDP is manageable with an 800 W suggested PSU. The 24 GB of GDDR6X memory with 1.01 TB/s bandwidth serves high-resolution textures and large datasets.

The Rubin GPU wins on raw theoretical compute and memory specifications. Its FP32 performance of 130.0 TFLOPS nearly doubles the RTX 4090 D's 73.54 TFLOPS. Its FP16 performance of 260.0 TFLOPS (2:1 ratio) is more than triple the RTX 4090 D's 73.54 TFLOPS (1:1 ratio). The 22.1 TB/s memory bandwidth and 288 GB capacity dwarf the RTX 4090 D's resources, indicating a design aimed at massive AI training and inference workloads rather than rendering.

The Rubin GPU also leads on node technology, using a 3 nm process versus 5 nm, and on interconnect, with PCIe 6.0 versus PCIe 4.0. Its SXM module form factor and lack of display outputs confirm its role as a server accelerator, not a consumer graphics card.

The Verdict

The database shows two products with fundamentally different purposes. The RTX 4090 D is a consumer and workstation graphics card with measured benchmark results, display outputs, and a conventional power profile. The Rubin GPU is a server accelerator with no recorded benchmarks, no display outputs, and extreme power and memory specifications.

For any application requiring measured performance data, the RTX 4090 D is the only choice supported by evidence. Its 98th percentile ranking and average score of 178050 place it among the top measured GPUs. The narrow delta percentages against its nearest rivals (2.2% to 4.9%) show it competes closely with professional Blackwell and Ada generation accelerators.

For applications demanding maximum memory capacity, bandwidth, and raw compute throughput, the Rubin GPU's specifications indicate a superior theoretical platform. The 288 GB HBM4 memory with 22.1 TB/s bandwidth and 130.0 TFLOPS FP32 performance represent a different performance class entirely. However, the absence of any benchmark results means the database cannot confirm how these specifications translate into real-world performance.

The RTX 4090 D supports standard graphics APIs including DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, making it suitable for rendering, gaming, and general GPU compute. The Rubin GPU lists no API support, reinforcing its specialized server role. Users requiring display output, standard software compatibility, or measured performance should choose the RTX 4090 D. Users building massive-scale AI systems with extreme memory requirements should consider the Rubin GPU based on its specifications alone.

FAQ

Q: How does the RTX 4090 D compare to its nearest rivals in average benchmark score?

A: The RTX 4090 D averages 178050 across all benchmark tests. The RTX PRO 5000 Blackwell averages 182109 (2.2% higher), the A100 SXM4 80 GB averages 183725 (3.1% higher), the RTX 5000 Ada Generation averages 184664 (3.6% higher), and the A100 SXM4 40 GB averages 187147 (4.9% higher). The RTX 4090 D trails each rival by a small margin.

Q: What benchmark results exist for the Rubin GPU?

A: The database contains no benchmark entries for the Rubin GPU. Its average benchmark score is 0 and its percentile rank against all GPUs is 50, which is the baseline value for unmeasured devices.

Q: What are the memory specifications for each GPU?

A: The RTX 4090 D has 24 GB of GDDR6X memory on a 384-bit bus with 1.01 TB/s bandwidth. The Rubin GPU has 288 GB of HBM4 memory on a 16384-bit bus with 22.1 TB/s bandwidth. The Rubin GPU offers 12 times the capacity and roughly 22 times the bandwidth.

Q: Which GPU has higher FP32 compute performance?

A: The Rubin GPU delivers 130.0 TFLOPS FP32 performance, while the RTX 4090 D delivers 73.54 TFLOPS. The Rubin GPU also delivers 260.0 TFLOPS FP16 performance (2:1 ratio), compared to the RTX 4090 D's 73.54 TFLOPS FP16 (1:1 ratio).

Q: What are the power requirements for each GPU?

A: The RTX 4090 D has a TDP of 425 W and a suggested PSU of 800 W. The Rubin GPU has a TDP of 2300 W and a suggested PSU of 2700 W. The Rubin GPU consumes over five times the power of the RTX 4090 D.

Q: Does the Rubin GPU support display outputs or standard graphics APIs?

A: No. The Rubin GPU lists no display outputs and no support for DirectX, OpenGL, or Vulkan. The RTX 4090 D supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, and offers 1x HDMI 2.1 plus 3x DisplayPort 1.4a outputs.

Specification Differences

| Specification | NVIDIA GeForce RTX 4090 D | NVIDIA Rubin GPU |

|---|---|---|

| Architecture | Ada Lovelace | Rubin |

| Process Node | 5 nm | 3 nm |

| Transistors | 76,300 million | 336,000 million |

| Die Size | 609 mm² | 1456 mm² |

| Transistor Density | 125.3M / mm² | 230.8M / mm² |

| Base Clock | 2280 MHz | 700 MHz |

| Boost Clock | 2520 MHz | 2267 MHz |

| Memory Size | 24 GB | 288 GB |

| Memory Type | GDDR6X | HBM4 |

| Memory Bus Width | 384 bit | 16384 bit |

| Memory Bandwidth | 1.01 TB/s | 22.1 TB/s |

| Memory Clock | 1313 MHz (21 Gbps effective) | 2695 MHz (10.8 Gbps effective) |

| Shading Units | 14592 | 28672 |

| TMUs | 456 | 896 |

| ROPs | 176 | 24 |

| RT Cores | 114 | None recorded |

| Tensor Cores | 456 | 896 |

| Pixel Rate | 443.5 GPixel/s | 54.41 GPixel/s |

| Texture Rate | 1,149.1 GTexel/s | 2,031.2 GTexel/s |

| FP32 | 73.54 TFLOPS | 130.0 TFLOPS |

| FP16 | 73.54 TFLOPS (1:1) | 260.0 TFLOPS (2:1) |

| TDP | 425 W | 2300 W |

| Slot Width | Triple-slot | SXM Module |

| Power Connectors | 1x 16-pin | None recorded |

| Suggested PSU | 800 W | 2700 W |

| Bus Interface | PCIe 4.0 x16 | PCIe 6.0 x16 |

| Display Outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a | No outputs |

| DirectX Support | 12 Ultimate (12_2) | N/A |

| OpenGL Support | 4.6 | N/A |

| Vulkan Support | 1.4 | N/A |

| Dimensions | 304 mm x 137 mm x 61 mm | Not recorded |

| Production Status | End-of-life | Active |

| Release Date | 2023-12-27 | 2025-12-31 |

| Predecessor | GeForce 30 | Server Blackwell |

| Successor | GeForce 50 | None recorded |

| Launch MSRP | 1,599 USD | None recorded |

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4090 D
Rubin GPU
Core Specs
Shading Units
14,592
28,672 +96.5%
Shaders
14,592
28,672 +96.5%
TMUs
456
896 +96.5%
ROPs
176
24 -86.4%
SM Count
114
224 +96.5%
Clocks
Base Clock
2280 MHz
700 MHz
Boost Clock
2520 MHz
2267 MHz
Memory Clock
1313 MHz 21 Gbps effective
2695 MHz 10.8 Gbps effective
Memory
Memory Size
24 GB
288 GB
VRAM (MB)
24,576
294,912 +1100.0%
Memory Type
GDDR6X
HBM4
Memory Bus
384 bit
16384 bit
Bandwidth
1.01 TB/s
22.1 TB/s
Cache
L1 Cache
128 KB (per SM)
256 KB (per SM)
L2 Cache
72 MB
128 MB
Performance
Pixel Rate
443.5 GPixel/s
54.41 GPixel/s
Texture Rate
1,149.1 GTexel/s
2,031.2 GTexel/s
FP32 (TFLOPS)
73.54 TFLOPS
130.0 TFLOPS
FP64 (TFLOPS)
1,149.1 GFLOPS (1:64)
32.50 TFLOPS (1:4)
FP16 (TFLOPS)
73.54 TFLOPS (1:1)
260.0 TFLOPS (2:1)
AI/RT
RT Cores
114
Tensor Cores
456
896 +96.5%
Power
TDP
425 W
2300 W
TDP (W)
425
2,300 +441.2%
Suggested PSU
800 W
2700 W
Power Connectors
1x 16-pin
Architecture
Architecture
Ada Lovelace
Rubin
GPU Name
AD102
GR100
Generation
GeForce 40
Server Rubin (Rxx)
Process Size
5 nm
3 nm
Transistors
76,300 million
336,000 million
Die Size
609 mm²
1456 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
230.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.9
10.7
Shader Model
6.8
Physical
Slot Width
Triple-slot
SXM Module
Length
304 mm 12 inches
Height
137 mm 5.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 6.0 x16
Other
Launch Price
1,599 USD
Production
End-of-life
Active
Predecessor
GeForce 30
Server Blackwell
Successor
GeForce 50
View GeForce RTX 4090 D Details View Rubin GPU Details