GEFORCE

NVIDIA Rubin GPU

NVIDIA graphics card specifications and benchmark scores

288 GB
VRAM
2267
MHz Boost
2300W
TDP
16384
Bus Width
Tensor Cores

At a Glance

NVIDIA
VRAM 288 GB
Boost Clock 2,267 MHz
Shaders 28,672
Bus Width 16384-bit
TDP 2300W
Memory Type HBM4
Architecture Rubin
nm
Process 3 nm
Released Jan 2026

NVIDIA Rubin GPU Specifications

Rubin GPU GPU Core

Shader units and compute resources

The NVIDIA Rubin GPU GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
28,672
Shaders
28,672
TMUs
896
ROPs
24
SM Count
224

Rubin GPU Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Rubin GPU's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Rubin GPU by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
700 MHz
Base Clock
700 MHz
Boost Clock
2267 MHz
Boost Clock
2,267 MHz
Memory Clock
2695 MHz 10.8 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's Rubin GPU Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Rubin GPU's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
288 GB
VRAM
294,912 MB
Memory Type
HBM4
VRAM Type
HBM4
Memory Bus
16384 bit
Bus Width
16384-bit
Bandwidth
22.1 TB/s

Rubin GPU by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Rubin GPU, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
256 KB (per SM)
L2 Cache
128 MB

Rubin GPU Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Rubin GPU against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
130.0 TFLOPS
FP64 (Double)
32.50 TFLOPS (1:4)
FP16 (Half)
260.0 TFLOPS (2:1)
Pixel Rate
54.41 GPixel/s
Texture Rate
2,031.2 GTexel/s

Rubin GPU Ray Tracing & AI

Hardware acceleration features

The NVIDIA Rubin GPU includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the Rubin GPU capable of delivering both stunning graphics and smooth frame rates in modern titles.

Tensor Cores
896

Rubin Architecture & Process

Manufacturing and design details

The NVIDIA Rubin GPU is built on NVIDIA's Rubin architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Rubin GPU will perform in GPU benchmarks compared to previous generations.

Architecture
Rubin
GPU Name
GR100
Process Node
3 nm
Foundry
TSMC
Transistors
336,000 million
Die Size
1456 mm²
Density
230.8M / mm²

NVIDIA's Rubin GPU Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Rubin GPU determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Rubin GPU to maintain boost clocks without throttling.

TDP
2300 W
TDP
2300W
Suggested PSU
2700 W

Rubin GPU by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Rubin GPU are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
SXM Module
Bus Interface
PCIe 6.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Rubin GPU. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
N/A
DirectX
N/A
OpenGL
N/A
OpenGL
N/A
Vulkan
N/A
Vulkan
N/A
OpenCL
3.0
CUDA
10.7
Shader Model
N/A

Rubin GPU Product Information

Release and pricing details

The NVIDIA Rubin GPU is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Rubin GPU by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Jan 2026
Production
Active
Predecessor
Server Blackwell

Rubin GPU Benchmark Scores

No benchmark data available for this GPU.

About NVIDIA Rubin GPU

Memory Subsystem

The NVIDIA Rubin GPU is equipped with a massive 288 GB of HBM4 memory, operating across a 16,384-bit bus. This configuration yields a peak bandwidth of 22.1 TB/s, placing it in an entirely different class from consumer or even most workstation graphics cards. The memory clock is listed at 2,695 MHz, translating to 10.8 Gbps effective data rate.

For high-resolution workloads, the implications are straightforward. At 4K and beyond, frame buffers grow exponentially, and texture-heavy scenes demand constant, high-throughput access to VRAM. A 288 GB pool means that even the most demanding scientific visualizations, large-scale AI inference batches, or 8K video processing tasks can reside entirely in memory without spillover to system RAM. The 22.1 TB/s bandwidth ensures that the shading units and tensor cores are never starved for data, a critical factor given the GPU's compute throughput.

The 16,384-bit bus width is the enabling factor here. While HBM4 itself offers per-stack bandwidth improvements, the sheer width of the interface allows the Rubin GPU to sustain such an extreme data rate. This is not a card designed for 1080p gaming; it is a memory subsystem built for data-center-scale problems where capacity and bandwidth are the primary constraints. The pixel rate of 54.41 GPixel/s, while seemingly modest, is irrelevant in this context, the memory system is optimized for streaming large datasets, not rasterizing pixels.

How It Compares

The FACT PACK lists no nearest rivals for the NVIDIA Rubin GPU. This absence is itself telling. The benchmark database shows a percentile rank of 50 among all GPUs, but with zero benchmark scores recorded, this is a default positioning rather than a measured outcome. The lack of rival data means that any comparative analysis must be grounded in the absolute specifications provided rather than relative performance metrics.

Without nearestRivals entries, there are no deltaPct values or rival names to reference. The Rubin GPU exists in a vacuum of comparative data, its 130.0 TFLOPS FP32 and 260.0 TFLOPS FP16 (2:1) compute figures stand alone. The predecessor is listed as "Server Blackwell," which suggests a generational leap, but no specific scores are provided for that prior product. The production status is "Active," and the release date falls on December 31, 2025, indicating a current-generation server part with no direct competition yet cataloged in this database.

Benchmark Performance

The benchmark data for the NVIDIA Rubin GPU is sparse: the benchmarks array is empty, and the average benchmark score is 0. The percentileVsAllGpus field shows 50, which would typically indicate a median performer, but without actual scores, this is a placeholder value rather than a measured result. The absence of benchmark entries means that the raw compute specifications are the only quantitative performance indicators available.

The FP32 throughput of 130.0 TFLOPS is a headline figure. This represents the GPU's peak single-precision floating-point capability, a metric that scales with shading unit count (28,672) and boost clock (2,267 MHz). The FP16 performance of 260.0 TFLOPS (2:1) doubles this throughput, reflecting the tensor core workload where reduced precision is acceptable. The texture rate of 2,031.2 GTexel/s and pixel rate of 54.41 GPixel/s provide additional context: the former is substantial, the latter is minimal, confirming that this is not a rasterization-focused part.

Since there are no rival scores to compare against, the analysis must focus on internal consistency. The 896 tensor cores and 896 TMUs are balanced for compute-heavy workloads, while the 24 ROPs are a token presence, likely included for basic display output rather than rendering capability. The FP32-to-FP16 ratio of exactly 2:1 is typical for NVIDIA architectures, indicating that the FP16 path is not a separate hardware unit but rather a doubled-throughput mode on the same data path.

Who Should Consider It

Given the specifications, the NVIDIA Rubin GPU is squarely aimed at data-center operators and high-performance computing facilities. The 288 GB HBM4 memory, combined with 22.1 TB/s bandwidth, makes it suitable for large language model training, massive scientific simulations, and real-time inference across multiple concurrent users. The absence of display outputs confirms this, it is not a card that will ever drive a monitor.

For resolution-based recommendations, the data suggests this is not a 4K or 8K gaming card. The pixel rate of 54.41 GPixel/s is far lower than what a high-end consumer GPU would deliver, and the 24 ROPs are a fraction of what even a midrange desktop card offers. The FP32 compute of 130.0 TFLOPS is the relevant metric: tasks that require massive parallel floating-point operations, such as molecular dynamics, climate modeling, or financial risk simulation, would see direct benefit.

The FP16 throughput of 260.0 TFLOPS is particularly relevant for AI workloads. Training a transformer model, for example, relies heavily on mixed-precision operations, and the 2:1 FP16 ratio means the Rubin GPU can halve its training time compared to FP32-bound processes. The 896 tensor cores are the dedicated hardware for this, and their presence alongside 28,672 shading units indicates a balanced design for both traditional HPC and AI/ML tasks. Users with workloads that fit within 288 GB of memory and can leverage 22.1 TB/s of bandwidth are the target audience.

Ray Tracing and Feature Set

The FACT PACK lists rtCores as null, which means no ray tracing core count is specified. This absence is notable. The APIs section shows DirectX as N/A, OpenGL as N/A, and Vulkan as N/A, which is consistent with a server part that has no graphics output and no consumer-facing rendering API support. The display outputs field confirms "No outputs."

The tensor cores, however, are present at 896 units. This is the key feature for this GPU. Tensor cores accelerate matrix multiplication and convolution operations, which are foundational to deep learning inference and training. The FP16 throughput of 260.0 TFLOPS is directly tied to these tensor cores, and the 2:1 ratio against FP32 indicates that the hardware can double its throughput when operating in reduced precision.

Without ray tracing cores and without graphics API support, the Rubin GPU is not a ray tracing accelerator in any conventional sense. It does not render frames, produce images, or interact with game engines. Its feature set is entirely compute-oriented. The PCIe 6.0 x16 interface is the sole connectivity method, providing a high-bandwidth link to the host system for data transfer. The absence of display outputs and graphics APIs positions this as a pure accelerator, not a GPU in the traditional consumer sense.

Power and Cooling

The thermal design power for the NVIDIA Rubin GPU is 2,300 W. This is an extremely high figure, reflecting the massive compute density on a 3 nm process node. The suggested PSU rating is 2,700 W, which indicates that the system power supply must be sized with a significant margin above the GPU's TDP to account for transient spikes and other system components.

The slot width is listed as "SXM Module," which means this is not a PCIe card that installs into a standard server chassis slot. Instead, it is designed for NVIDIA's proprietary SXM form factor, which provides direct power delivery and cooling through a baseboard. The power connectors field is null, which is consistent with the SXM design, power is delivered through the module's edge connector rather than external 8-pin or 12VHPWR cables.

Cooling requirements are implicit in the 2,300 W TDP. The SXM form factor typically uses liquid cooling or high-flow air cooling in a server chassis designed to handle such thermal loads. The 3 nm process node from TSMC helps mitigate some of the thermal density, but 2,300 W is a data-center-scale power draw that requires industrial-grade cooling infrastructure. The transistor count of 336,000 million (336 billion) on a 1,456 mm² die results in a density of 230.8 million transistors per square millimeter, which is a staggering figure that directly correlates with the power draw.

FAQ

Q: What is the memory size of the NVIDIA Rubin GPU?

A: The GPU is equipped with 288 GB of HBM4 memory.

Q: What is the peak FP32 compute performance?

A: The FP32 throughput is 130.0 TFLOPS, with FP16 performance at 260.0 TFLOPS (2:1).

Q: Does this GPU have display outputs?

A: No, the display outputs field is listed as "No outputs."

Q: What is the suggested PSU rating for this GPU?

A: The suggested PSU is 2,700 W, while the GPU's TDP is 2,300 W.

Q: What is the bus interface and memory bandwidth?

A: The bus interface is PCIe 6.0 x16, and the memory bandwidth is 22.1 TB/s across a 16,384-bit bus.

Q: What is the production status and release date?

A: The production status is "Active," and the release date is December 31, 2025.

Architecture and Design

The NVIDIA Rubin GPU is built on the GR100 chip, which is part of the Rubin architecture. This is a server-class design, categorized under the "Server Rubin (Rxx)" generation. The manufacturing process is 3 nm at TSMC, which is a leading-edge node that allows for the integration of 336,000 million transistors on a die size of 1,456 mm². The resulting transistor density of 230.8 million per square millimeter is among the highest of any GPU produced.

The core configuration consists of 28,672 shading units, 896 TMUs, and 24 ROPs. The shading units are the primary compute elements, handling FP32 and FP16 arithmetic. The TMUs are responsible for texture filtering, which is relevant in compute workloads that involve sampling, though with 24 ROPs, the pixel output stage is minimal. The 896 tensor cores are dedicated to matrix operations, and their count matches the TMU count, suggesting a design where each tensor core is paired with a texture unit for certain operations.

The clock speeds are set at a 700 MHz base and 2,267 MHz boost. This is a wide boost range, indicating that the GPU can dynamically scale its power draw based on workload. The memory clock is 2,695 MHz, which translates to 10.8 Gbps effective when accounting for the HBM4 double-data-rate signaling. The FP32 peak of 130.0 TFLOPS is calculated from the shading units multiplied by the boost clock and two operations per clock, and the FP16 figure of 260.0 TFLOPS represents a 2:1 ratio, which is achieved by running the same data path at half precision.

The design is a departure from consumer GPUs in every respect. The 24 ROPs are a token presence, the 1,456 mm² die is enormous, and the 2,300 W TDP is unrivaled in any non-server context. The bus interface is PCIe 6.0 x16, which offers substantial bandwidth for host communication, but the primary data path is through the HBM4 memory. The architecture is optimized for throughput in compute-heavy, memory-bound workloads, with no consideration for rasterization efficiency or video output. The production status is "Active," and the release date is set for the end of 2025, positioning this as a current-generation product for data-center deployment.

The AMD Equivalent of Rubin GPU

Looking for a similar graphics card from AMD? The AMD Radeon RX 9060 XT LP offers comparable performance and features in the AMD lineup.

AMD Radeon RX 9060 XT LP

AMD • 16 GB VRAM

View Specs Compare

Popular NVIDIA Rubin GPU Comparisons

See how the Rubin GPU stacks up against similar graphics cards from the same generation and competing brands.

Compare Rubin GPU with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs