GEFORCE

NVIDIA Quadro RTX 4000

NVIDIA graphics card specifications and benchmark scores

8 GB
VRAM
1545
MHz Boost
160W
TDP
256
Bus Width
Ray Tracing Tensor Cores

At a Glance

NVIDIA
VRAM 8 GB
Boost Clock 1,545 MHz
Shaders 2,304
Bus Width 256-bit
TDP 160W
Memory Type GDDR6
RT Cores 36
Architecture Turing
nm
Process 12 nm
Released Nov 2018

NVIDIA Quadro RTX 4000 Specifications

GPU Core

Shader units and compute resources

The NVIDIA Quadro RTX 4000 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
2,304
Shaders
2,304
TMUs
144
ROPs
64
SM Count
36

Quadro RTX 4000 Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Quadro RTX 4000's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Quadro RTX 4000 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1005 MHz
Base Clock
1,005 MHz
Boost Clock
1545 MHz
Boost Clock
1,545 MHz
Memory Clock
1625 MHz 13 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's Quadro RTX 4000 Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Quadro RTX 4000's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
8 GB
VRAM
8,192 MB
Memory Type
GDDR6
VRAM Type
GDDR6
Memory Bus
256 bit
Bus Width
256-bit
Bandwidth
416.0 GB/s

Quadro RTX 4000 by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Quadro RTX 4000, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
64 KB (per SM)
L2 Cache
4 MB

Quadro RTX 4000 Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Quadro RTX 4000 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
7.119 TFLOPS
FP64 (Double)
222.5 GFLOPS (1:32)
FP16 (Half)
14.24 TFLOPS (2:1)
Pixel Rate
98.88 GPixel/s
Texture Rate
222.5 GTexel/s

Quadro RTX 4000 Ray Tracing & AI

Hardware acceleration features

The NVIDIA Quadro RTX 4000 includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the Quadro RTX 4000 capable of delivering both stunning graphics and smooth frame rates in modern titles.

RT Cores
36
Tensor Cores
288

Turing Architecture & Process

Manufacturing and design details

The NVIDIA Quadro RTX 4000 is built on NVIDIA's Turing architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Quadro RTX 4000 will perform in GPU benchmarks compared to previous generations.

Architecture
Turing
GPU Name
TU104
Process Node
12 nm
Foundry
TSMC
Transistors
13,600 million
Die Size
545 mm²
Density
25.0M / mm²

Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Quadro RTX 4000 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Quadro RTX 4000 to maintain boost clocks without throttling.

TDP
160 W
TDP
160W
Power Connectors
1x 8-pin
Suggested PSU
450 W

Quadro RTX 4000 by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Quadro RTX 4000 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Single-slot
Length
241 mm 9.5 inches
Height
111 mm 4.4 inches
Bus Interface
PCIe 3.0 x16
Display Outputs
3x DisplayPort 1.4a1x USB Type-C
Display Outputs
3x DisplayPort 1.4a1x USB Type-C

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Quadro RTX 4000. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 Ultimate (12_2)
DirectX
12 Ultimate (12_2)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
7.5
Shader Model
6.8

Quadro RTX 4000 Product Information

Release and pricing details

The NVIDIA Quadro RTX 4000 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Quadro RTX 4000 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Nov 2018
Launch Price
899 USD
Production
End-of-life
Predecessor
Quadro Volta
Successor
Workstation Ampere

About NVIDIA Quadro RTX 4000

The NVIDIA Quadro RTX 4000 is a Turing-architecture GPU built around the TU104 chip, fabricated by TSMC on a 12 nm process with 13,600 million transistors on a 545 mm² die. It belongs to the Quadro Turing (Tx000) generation, was released on 2018-11-12, and sits between the Quadro Volta and Workstation Ampere product lines. In the benchmark database it records an average benchmark score of 18852 and a 60th percentile ranking across all GPUs, with a launch MSRP of 899 USD. The card carries a base clock of 1005 MHz, a boost clock of 1545 MHz, and uses a PCIe 3.0 x16 bus interface.

Memory Subsystem

The NVIDIA Quadro RTX 4000 is equipped with 8 GB of GDDR6 memory on a 256-bit bus. The memory clock is 1625 MHz, quoted as 13 Gbps effective, and the resulting memory bandwidth is 416.0 GB/s. This combination defines the card’s ability to feed a large, continuously accessed frame buffer, which is especially relevant for high-resolution rendering where the amount of pixel data being written and read grows quickly.

The memory path serves a GPU with 2304 shading units, 144 texture mapping units, and 64 ROPs. The card also reaches a pixel rate of 98.88 GPixel/s and a texture rate of 222.5 GTexel/s. Those rates are heavily dependent on bandwidth, and the 256-bit bus width is the hardware mechanism that allows the 416.0 GB/s figure to be sustained. In practical terms, the 8 GB VRAM capacity is the ceiling for how much frame data and geometry can be held locally, while the bandwidth determines how quickly the compute and raster units can read from and write to that memory. For high-resolution work, 416.0 GB/s is a meaningful data throughput figure, and the 8 GB frame buffer is the amount of immediate video memory available before data must be swapped.

How It Compares

Against the NVIDIA Tesla K80, the Quadro RTX 4000 is effectively tied. The K80’s average score is 18866, and the Quadro’s average score is 18852, producing a delta of -0.1%. This is the closest rival in the database by a large margin.

Against the NVIDIA Tesla K20m, the Quadro trails by 0.8%. The K20m holds an average score of 19011, so the Quadro’s 18852 aggregate result is slightly lower, but the gap is still under one percent.

Against the NVIDIA GeForce RTX 4050 Mobile, the Quadro sits 1.0% behind. The RTX 4050 Mobile posts an average score of 19049, while the Quadro RTX 4000 posts 18852. The delta of -1.0% shows a consistent, if modest, deficit in the aggregate benchmark index.

Against the AMD Radeon 780M, the Quadro is 1.1% behind. The Radeon 780M has the highest average score among the four nearest rivals at 19057, and the Quadro’s delta of -1.1% is the largest negative gap in this group. Even so, the entire rival cluster is separated by a narrow range of average scores.

Benchmark Performance

The overall benchmark position can be summarized by the average benchmark score of 18852 and the 60th percentile figure. The Quadro RTX 4000 is not a leading GPU in the database; it sits in the middle of the distribution, and all four nearest rival deltas are negative. The deltas run from -0.1% against the Tesla K80 to -1.1% against the Radeon 780M, meaning the aggregate scores of the nearest competitors are tightly grouped.

Individual benchmark results vary more widely than the aggregate score might suggest. In the 3DMark Steel Nomad DX12 test, the Quadro scores 1873. Geekbench OpenCL returns 85166, while Geekbench Vulkan returns 78844, showing substantially higher results in compute-oriented tests than in the DirectX 12 3D test. PassMark results are split by API: DirectX 9 scores 205, DirectX 10 scores 108, DirectX 11 scores 128, and DirectX 12 scores 52. The general-purpose PassMark G2D score is 846, the G3D score is 15117, and the GPU compute score is 6176.

Raw compute figures from the fact pack align with that mixed profile. FP32 performance is 7.119 TFLOPS, while FP16 performance is 14.24 TFLOPS with a 2:1 ratio. These are the theoretical compute ceilings, while the PassMark GPU compute result of 6176 is a measured workload data point. The 3DMark Steel Nomad DX12 score of 1873 and PassMark DirectX 12 score of 52 are the least remarkable results in the suite, which suggests modern DirectX 12 rendering is not the primary strength of this configuration.

FAQ

Q: What memory configuration does the NVIDIA Quadro RTX 4000 use?

A: It uses 8 GB of GDDR6 memory on a 256-bit bus, with 416.0 GB/s memory bandwidth, a 1625 MHz memory clock, and 13 Gbps effective signaling.

Q: Which graphics APIs are supported?

A: The card supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Q: What are the power requirements?

A: The TDP is 160 W, the power connector is 1x 8-pin, and the suggested PSU is 450 W.

Q: How many RT and tensor cores are included?

A: The GPU includes 36 RT cores and 288 tensor cores.

Q: What was the launch MSRP?

A: The launch MSRP was 899 USD.

Q: How does the Quadro RTX 4000 compare with the GeForce RTX 4050 Mobile?

A: The Quadro’s average benchmark score is 18852, while the RTX 4050 Mobile averages 19049, giving a delta of -1.0%. The Quadro trails by 1.0% in the aggregate score.

Ray Tracing and Feature Set

The Quadro RTX 4000 is built on the Turing architecture with the TU104 chip. It includes 36 RT cores for ray tracing work and 288 tensor cores for tensor-accelerated operations. The underlying chip is fabricated on TSMC’s 12 nm process with 13,600 million transistors, a die size of 545 mm², and a transistor density of 25.0M per mm².

The API feature set includes DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. This gives the card support for modern graphics APIs at the levels listed in the data. The display outputs are 3x DisplayPort 1.4a and 1x USB Type-C, providing the physical output options for multi-display workstation setups. The ray tracing and tensor core counts are the specific hardware facts relevant to workloads that use those paths; the 36 RT cores are the ray tracing block, and the 288 tensor cores are the compute-oriented block for tensor operations.

Power and Cooling

The Quadro RTX 4000 is a single-slot design. Its physical dimensions are 241 mm in length, which is 9.5 inches, and 111 mm in height, which is 4.4 inches. The TDP is 160 W, and power is supplied through a single 1x 8-pin power connector. The suggested PSU rating is 450 W, which gives a defined power-supply target for system builders. The base clock is 1005 MHz and the boost clock is 1545 MHz, both running within that 160 W power envelope. The production status is end-of-life, so this is a card that has already completed its production cycle rather than a newly available product.

Who Should Consider It

The Quadro RTX 4000 fits a middle-of-the-database GPU profile, with a 60th percentile ranking and an average score of 18852. Its nearest rivals are all within a narrow range: the Tesla K80 is 0.1% ahead, the Tesla K20m is 0.8% ahead, the GeForce RTX 4050 Mobile is 1.0% ahead, and the Radeon 780M is 1.1% ahead. Users comparing these cards on the aggregate benchmark index should expect performance parity rather than a decisive winner.

The strongest data points are compute-oriented: Geekbench OpenCL scores 85166, Geekbench Vulkan scores 78844, and PassMark GPU compute scores 6176. Those numbers make the card more compelling for workloads that exercise compute throughput. The 8 GB GDDR6 frame buffer and 416.0 GB/s memory bandwidth are the relevant hardware facts for high-resolution frame buffering, and the single-slot width plus 160 W TDP define the physical and power requirements. Modern DirectX 12 3D workloads are not the strongest area, as the 3DMark Steel Nomad DX12 score of 1873 and PassMark DirectX 12 score of 52 show. Users who prioritize compute results, need a single-slot card, and can work with the 8-pin power connector and 450 W suggested PSU have a clearly documented hardware profile to evaluate.

Detailed benchmark scores and charts for the NVIDIA Quadro RTX 4000 are below.

Benchmark Scores

3dmark_3dmark_steel_nomad_dx12Source

3DMark Steel Nomad is the latest GPU benchmark running at native 4K with DirectX 12. It's roughly 3x more demanding than Time Spy, testing NVIDIA Quadro RTX 4000 with cutting-edge rendering techniques. The benchmark uses state-of-the-art graphics technologies to stress modern hardware. Scores accurately predict NVIDIA Quadro RTX 4000 performance in demanding AAA games at 4K resolution.

3dmark_3dmark_steel_nomad_dx12 #109 of 188
1,873
10%
Max: 18,355

geekbench_openclSource

Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA Quadro RTX 4000 handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms. Higher scores benefit applications that leverage GPU acceleration for non-graphics workloads.

geekbench_opencl #146 of 650
74,540
19%
Max: 388,405
Compare with other GPUs

geekbench_vulkanSource

Geekbench Vulkan tests GPU compute using the modern low-overhead Vulkan API. This shows how NVIDIA Quadro RTX 4000 performs with next-generation graphics and compute workloads.

geekbench_vulkan #110 of 446
78,844
21%
Max: 376,915

passmark_directx_10Source

DirectX 10 tests NVIDIA Quadro RTX 4000 with the graphics API introduced with Windows Vista. This shows performance in games from the 2007-2009 era that targeted this feature level. DX10 introduced geometry shaders and other features still used today.

passmark_directx_11Source

DirectX 11 tests NVIDIA Quadro RTX 4000 with the widely-used graphics API powering most current games. This shows mainstream gaming performance across the majority of today's titles. DX11 remains the most common rendering path even in newer games. Tessellation and compute shaders introduced in DX11 are heavily used in modern game engines.

passmark_directx_12Source

DirectX 12 tests NVIDIA Quadro RTX 4000 with the modern low-overhead graphics API. This shows performance in next-gen games that leverage DX12 features like ray tracing and mesh shaders.

passmark_directx_9Source

DirectX 9 tests NVIDIA Quadro RTX 4000 performance with the legacy graphics API still used by older games. This shows compatibility and performance with classic titles from the 2000s era.

passmark_g2dSource

PassMark G2D tests 2D graphics performance for desktop rendering, UI elements, and productivity applications. This shows how NVIDIA Quadro RTX 4000 handles everyday visual tasks.

passmark_g3dSource

PassMark G3D measures overall 3D graphics performance of NVIDIA Quadro RTX 4000 across DirectX 9 through 12 tests. This provides a comprehensive gaming capability score. The combined result predicts performance across various game engines and API versions.

passmark_g3d #90 of 186
15,117
34%
Max: 44,065

passmark_gpu_computeSource

GPU compute tests parallel processing capability of NVIDIA Quadro RTX 4000 using OpenCL. This shows performance in video encoding, scientific computing, and AI workloads.

passmark_gpu_compute #98 of 184
6,176
22%
Max: 28,396

Popular NVIDIA Quadro RTX 4000 Comparisons

See how the Quadro RTX 4000 stacks up against similar graphics cards from the same generation and competing brands.

Compare with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs