GEFORCE

NVIDIA RTX PRO 4000 Blackwell SFF

NVIDIA graphics card specifications and benchmark scores

24 GB
VRAM
1432
MHz Boost
70W
TDP
192
Bus Width
Ray Tracing Tensor Cores

At a Glance

NVIDIA
VRAM 24 GB
Boost Clock 1,432 MHz
Shaders 8,960
Bus Width 192-bit
TDP 70W
Memory Type GDDR7
RT Cores 70
Architecture Blackwell 2.0
nm
Process 5 nm
Released Aug 2025

NVIDIA RTX PRO 4000 Blackwell SFF Specifications

RTX PRO 4000 Blackwell SFF GPU Core

Shader units and compute resources

The NVIDIA RTX PRO 4000 Blackwell SFF GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
8,960
Shaders
8,960
TMUs
280
ROPs
96
SM Count
70

RTX PRO 4000 Blackwell SFF Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the RTX PRO 4000 Blackwell SFF's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The RTX PRO 4000 Blackwell SFF by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
577 MHz
Base Clock
577 MHz
Boost Clock
1432 MHz
Boost Clock
1,432 MHz
Memory Clock
1125 MHz 18 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's RTX PRO 4000 Blackwell SFF Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The RTX PRO 4000 Blackwell SFF's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
24 GB
VRAM
24,576 MB
Memory Type
GDDR7
VRAM Type
GDDR7
Memory Bus
192 bit
Bus Width
192-bit
Bandwidth
432.0 GB/s

RTX PRO 4000 Blackwell SFF by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the RTX PRO 4000 Blackwell SFF, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
128 KB (per SM)
L2 Cache
48 MB

RTX PRO 4000 Blackwell SFF Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA RTX PRO 4000 Blackwell SFF against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
25.66 TFLOPS
FP64 (Double)
401.0 GFLOPS (1:64)
FP16 (Half)
25.66 TFLOPS (1:1)
Pixel Rate
137.5 GPixel/s
Texture Rate
401.0 GTexel/s

RTX PRO 4000 Blackwell SFF Ray Tracing & AI

Hardware acceleration features

The NVIDIA RTX PRO 4000 Blackwell SFF includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the RTX PRO 4000 Blackwell SFF capable of delivering both stunning graphics and smooth frame rates in modern titles.

RT Cores
70
Tensor Cores
280

Blackwell 2.0 Architecture & Process

Manufacturing and design details

The NVIDIA RTX PRO 4000 Blackwell SFF is built on NVIDIA's Blackwell 2.0 architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the RTX PRO 4000 Blackwell SFF will perform in GPU benchmarks compared to previous generations.

Architecture
Blackwell 2.0
GPU Name
GB203
Process Node
5 nm
Foundry
TSMC
Transistors
45,600 million
Die Size
378 mm²
Density
120.6M / mm²

NVIDIA's RTX PRO 4000 Blackwell SFF Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA RTX PRO 4000 Blackwell SFF determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the RTX PRO 4000 Blackwell SFF to maintain boost clocks without throttling.

TDP
70 W
TDP
70W
Power Connectors
None
Suggested PSU
250 W

RTX PRO 4000 Blackwell SFF by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA RTX PRO 4000 Blackwell SFF are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Dual-slot
Length
167 mm 6.6 inches
Height
69 mm 2.7 inches
Bus Interface
PCIe 5.0 x8
Display Outputs
4x mini-DisplayPort 2.1b
Display Outputs
4x mini-DisplayPort 2.1b

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA RTX PRO 4000 Blackwell SFF. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 Ultimate (12_2)
DirectX
12 Ultimate (12_2)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
12.0
Shader Model
6.8

RTX PRO 4000 Blackwell SFF Product Information

Release and pricing details

The NVIDIA RTX PRO 4000 Blackwell SFF is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the RTX PRO 4000 Blackwell SFF by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Aug 2025
Production
Active
Predecessor
Workstation Ada

RTX PRO 4000 Blackwell SFF Benchmark Scores

3dmark_3dmark_steel_nomad_dx12Source

3DMark Steel Nomad is the latest GPU benchmark running at native 4K with DirectX 12. It's roughly 3x more demanding than Time Spy, testing NVIDIA RTX PRO 4000 Blackwell SFF with cutting-edge rendering techniques. The benchmark uses state-of-the-art graphics technologies to stress modern hardware.

About NVIDIA RTX PRO 4000 Blackwell SFF

The NVIDIA RTX PRO 4000 Blackwell SFF is a dual-slot workstation GPU built around the GB203 chip and Blackwell 2.0 architecture. It is manufactured by NVIDIA on a 5 nm TSMC process, integrating 45,600 million transistors on a 378 mm² die for a transistor density of 120.6M per mm². The board is listed as Active production, with a release date of 2025-08-10, and the predecessor generation is Workstation Ada. This data profile contains no benchmark entries and no nearest rival entries, so the following analysis is based entirely on the specification fields: clocks, memory, compute rates, thermal limits, and API support.

Power and Cooling

The thermal design point of the RTX PRO 4000 Blackwell SFF is 70 W. That is a low power envelope for a GPU with this compute configuration, and the suggested power supply rating is 250 W. The power connector field is listed as None, meaning no auxiliary power connectors are required. The board draws its specified 70 W without additional cabling, which simplifies installation in systems with modest power delivery.

The card occupies a dual-slot width. Its physical dimensions are 167 mm in length, 69 mm in height, and 40 mm in width. The compact length and the 70 W thermal limit are consistent with the SFF label in the product name. Cooling requirements remain modest because of the 70 W envelope, although the dual-slot design provides a larger thermal surface than a single-slot board. The bus interface is PCIe 5.0 x8, and the system power recommendation of 250 W reflects a build that does not need a high-capacity power supply. With no power connectors present, the primary constraint for integration is physical clearance for a dual-slot, 167 mm card rather than power connector availability.

The low TDP is notable when paired with the compute resources on the board. The chip contains 8960 shading units, 280 texture mapping units, and 96 ROPs, all running within the 70 W limit. The boost clock is 1432 MHz, while the base clock is 577 MHz. The resulting pixel rate is 137.5 GPixel/s, and the texture rate is 401.0 GTexel/s. These throughput numbers are achieved under a power budget that a 250 W PSU can support. For system design, the absence of a power connector means there is no 12VHPWR or 8-pin requirement to plan around, and the dual-slot width is the only form-factor consideration beyond the 167 mm board length.

Ray Tracing and Feature Set

The RTX PRO 4000 Blackwell SFF implements 70 RT cores and 280 tensor cores. The architecture is Blackwell 2.0, which provides the base hardware support for ray-traced workloads and tensor-based compute operations. The RT core count represents a dedicated hardware path for ray intersection and traversal work, while the 280 tensor cores handle matrix operations. The FP16 compute rate is 25.66 TFLOPS, and it is listed as 1:1 with FP32, which is also 25.66 TFLOPS. This means half-precision throughput does not drop relative to single-precision throughput on this specification sheet; both rates are identical.

API support includes DirectX 12 Ultimate at feature level 12_2, OpenGL 4.6, and Vulkan 1.4. DirectX 12 Ultimate support in the API list indicates the board exposes the feature level associated with modern DirectX 12 workloads. OpenGL 4.6 covers legacy and cross-platform compute and rendering contexts, while Vulkan 1.4 provides a current low-level graphics and compute API. The combination of 70 RT cores, 280 tensor cores, and these API levels defines the feature set for ray tracing and accelerated compute. The tensor cores are not tied to a listed tensor TFLOPS value in this data, but their count of 280 is the relevant hardware resource. Similarly, the RT core count of 70 is the dedicated ray tracing resource.

Because the architecture is Blackwell 2.0 and the FP16 rate equals the FP32 rate at 25.66 TFLOPS, the compute profile is balanced between precision formats. The 8960 shading units provide the integer and FP32 execution foundation, while the 280 tensor cores extend the accelerator beyond graphics into neural or matrix-based workloads. For a workstation user, the feature set is defined by RT cores for ray tracing, tensor cores for AI-style compute, and the API surface that includes DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4.

Memory Subsystem

The card ships with 24 GB of GDDR7 memory on a 192-bit bus. The memory clock is 1125 MHz, with an effective data rate of 18 Gbps. This configuration yields a memory bandwidth of 432.0 GB/s. For high-resolution workloads, the 24 GB capacity is a substantial resource; a pool of this size can hold large geometry buffers, high-resolution textures, and multi-frame render targets without spilling to system memory. The 432.0 GB/s bandwidth is the sustained data movement ceiling for the frame buffer, and it is delivered over a 192-bit interface paired with GDDR7’s effective transfer rate.

The combination of 24 GB and 432.0 GB/s is significant for scenes that are both large and bandwidth-sensitive. At high resolutions, resolution scaling increases both the capacity needed for framebuffers and the bandwidth required to move pixels and texture data. The 24 GB capacity addresses the former, while the 432.0 GB/s figure addresses the latter. The memory interface is 192-bit, which is narrower than a 256-bit or wider bus, but the GDDR7 data rate offsets the narrower bus through a higher effective transfer rate. The memory subsystem is one of the more distinctive aspects of this product: 24 GB of VRAM is a large allocation, and 432.0 GB/s of bandwidth is the rate at which that memory can be accessed.

The memory type GDDR7 is newer than older memory types, and the effective 18 Gbps data rate is listed explicitly in the specification. The 1125 MHz memory clock is the base transfer clock, while the effective rate of 18 Gbps is the data rate after signaling factors. The product’s memory bandwidth works with the 70 W TDP to keep power consumption limited while still providing a 24 GB frame buffer. For high-resolution scenarios, the limiting factor is most likely to be the 432.0 GB/s bandwidth; capacity at 24 GB is less likely to be the bottleneck, unless the workload demands more than 24 GB of resident data. The 192-bit bus width is a defining parameter: with 24 GB of GDDR7, the per-module density must be high, but the aggregate bandwidth is set by the bus width and the effective data rate.

How It Compares

The nearestRivals field in the data is empty, and the benchmark array is empty as well. The only positional metric present is percentileVsAllGpus, which is 50. That value places the RTX PRO 4000 Blackwell SFF at the midpoint of the tracked GPU population in the database, though avgBenchmarkScore is 0 and no benchmark entries are listed. Because no neighbor entries are supplied, there are no rival names, scores, or deltaPct values to cite. The comparison section is therefore limited to the percentile value and the specification-level position relative to the rest of the database.

The predecessor generation, Workstation Ada, is listed, but no predecessor benchmark score is included. The product’s own production status is Active, and it has no successor listed. Within this data profile, there is no direct score-based comparison to other GPUs. The percentileVsAllGpus of 50 indicates a median standing, but without concrete benchmark numbers or rival deltas, the exact performance gap cannot be quantified. The RTX PRO 4000 Blackwell SFF’s position in the database is thus defined by its specification sheet rather than by measured performance. Its 70 RT cores, 280 tensor cores, and 25.66 TFLOPS FP32/FP16 rates are the numerical anchors available for position, but they are not rival-relative figures. The absence of rival data means any claim about being faster or slower than a specific card would be unsupported.

Who Should Consider It

The RTX PRO 4000 Blackwell SFF is aimed at workstation builds where power draw and physical space are constrained. The 70 W TDP, dual-slot width, 167 mm length, and lack of power connectors make it a plausible fit for compact systems. The SFF designation in the product name reinforces that intent. The 24 GB GDDR7 frame buffer is the standout capacity figure, and it is suitable for workloads that need large resident datasets: high-resolution textures, large scenes, or multi-buffer rendering. The 432.0 GB/s bandwidth is the related performance boundary for moving that data.

Compute features also guide system fit. The 8960 shading units deliver a known FP32 throughput of 25.66 TFLOPS, and the identical FP16 rate means mixed-precision work does not suffer a half-rate penalty. The 280 tensor cores open up tensor-accelerated workflows, while the 70 RT cores provide dedicated ray tracing resources. The API list includes DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, making the card compatible with a broad range of rendering stacks. For users working at high resolutions, the 24 GB capacity is the primary asset, and the 192-bit bus with 432.0 GB/s is the bandwidth that sustains frame-buffer traffic. The 137.5 GPixel/s pixel rate and 401.0 GTexel/s texture rate quantify the rasterization throughput available within the 70 W envelope.

This is not a card defined by benchmark scores, because none are published in the data. It is a card defined by a specific thermal and mechanical profile: 70 W, dual-slot, no power connectors, 167 mm length. Anyone considering it should weigh the 24 GB memory capacity and balanced FP32/FP16 compute against the 432.0 GB/s bandwidth and the 192-bit bus. For a compact workstation that requires 24 GB of GDDR7 and Blackwell 2.0 architecture, the RTX PRO 4000 Blackwell SFF is the hardware in question. Those who prioritize pure bandwidth may find the 432.0 GB/s figure modest compared with wider-bus parts, but no rival comparisons are available to make that case within this data. The placement at the 50th percentile in the database is a neutral signal, and the empty benchmark list means no measured performance ranking can be assigned.

The AMD Equivalent of RTX PRO 4000 Blackwell SFF

Looking for a similar graphics card from AMD? The AMD Radeon RX 7400 offers comparable performance and features in the AMD lineup.

AMD Radeon RX 7400

AMD • 8 GB VRAM

View Specs Compare

Popular NVIDIA RTX PRO 4000 Blackwell SFF Comparisons

See how the RTX PRO 4000 Blackwell SFF stacks up against similar graphics cards from the same generation and competing brands.

Compare RTX PRO 4000 Blackwell SFF with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs