NVIDIA GeForce RTX 4090 D vs NVIDIA PG506-232 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4090 D

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 425 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

PG506-232

CORE STATE GA100
VRAM 24 GB
CLOCK SPEED 1440 MHz
TDP 165 W
BUS WIDTH 3072 bit
ARCHITECTURE Ampere
nm
PROCESS 7 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
8,587
N/A
geekbench_opencl
278,621
225,124
geekbench_vulkan
246,941
N/A

Analysis: NVIDIA GeForce RTX 4090 D vs NVIDIA PG506-232

The NVIDIA PG506-232 and the NVIDIA GeForce RTX 4090 D are both end-of-life, 24 GB NVIDIA boards, but they target fundamentally different workloads and performance tiers. The data shows a clear overall winner: the RTX 4090 D, which outperforms the PG506-232 by 19.2% in the sole shared OpenCL benchmark, while also offering a dramatically higher compute ceiling. However, the PG506-232 is not without merit; it is a far more power-efficient, compact, and server-oriented card that holds its own in specific professional contexts where the RTX 4090 D's consumer-focused design and higher power draw are liabilities. The verdict depends entirely on whether raw performance or specialized form factor and efficiency take priority.

The Verdict

For buyers seeking maximum raw compute performance in a single slot, the NVIDIA GeForce RTX 4090 D is the definitive choice. It wins the only head-to-head benchmark available, scoring 278,621 in Geekbench OpenCL against the PG506-232's 225,124, a 19.2% lead. This advantage is compounded by its massively higher peak specifications, including 73.54 TFLOPS FP32 versus 10.32 TFLOPS, and a 1.01 TB/s memory bandwidth versus 933.1 GB/s. The RTX 4090 D is the superior card for any workload that can leverage its raw shader throughput and modern Ada Lovelace features like 114 RT cores.

Conversely, the NVIDIA PG506-232 is the pick for specific server environments where power and physical footprint are critical constraints. Its 165 W TDP is less than half of the RTX 4090 D's 425 W, and its dual-slot, 267 mm length profile is significantly more accommodating than the triple-slot, 304 mm RTX 4090 D. The data shows the PG506-232 sits in the 99th percentile of all GPUs, just one point below the RTX 4090 D's 98th percentile, indicating it is still a highly capable compute card. For a dense server chassis with strict power budgets, the PG506-232 is the logical, purpose-built option.

The RTX 4090 D is the better overall product for most users, but the PG506-232 wins on efficiency and server integration. The choice is not about which is faster—the RTX 4090 D is unequivocally faster—but about which set of trade-offs fits the deployment.

Architecture Differences

The two cards are built on different architectures and process nodes, leading to distinct design philosophies. The PG506-232 uses the Ampere architecture (GA100 chip) on a 7 nm TSMC process, while the RTX 4090 D uses the Ada Lovelace architecture (AD102 chip) on a 5 nm TSMC process. This node difference is stark: the RTX 4090 D packs 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3M / mm², compared to the PG506-232's 54,200 million transistors on a larger 826 mm² die, for a density of 65.6M / mm². The Ada Lovelace chip is more than twice as dense.

Core configurations diverge significantly. The PG506-232 has 3,584 shading units, 224 TMUs, and 96 ROPs, while the RTX 4090 D has 14,592 shading units, 456 TMUs, and 176 ROPs. The RTX 4090 D also features 114 dedicated RT cores and 456 tensor cores, whereas the PG506-232 lists no RT cores but has 224 tensor cores. The memory subsystems also differ fundamentally: the PG506-232 uses HBM2 with a 3072-bit bus, while the RTX 4090 D uses GDDR6X with a 384-bit bus. The HBM2 setup provides higher bandwidth per watt, but the GDDR6X solution offers higher absolute bandwidth at 1.01 TB/s.

Clock speeds heavily favor the RTX 4090 D. It boosts to 2,520 MHz from a 2,280 MHz base, while the PG506-232 boosts to only 1,440 MHz from a 930 MHz base. This clock advantage, combined with more cores, explains the massive FP32 throughput gap. The PG506-232 has no display outputs, confirming its server-only role, while the RTX 4090 D includes 1x HDMI 2.1 and 3x DisplayPort 1.4a, making it suitable for workstation or consumer use. The PG506-232 also lacks listed API support for DirectX, OpenGL, and Vulkan, while the RTX 4090 D supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4.

Head-to-Head Benchmarks

The only directly comparable benchmark in the data is Geekbench OpenCL, and the results are decisive. The RTX 4090 D scores 278,621, while the PG506-232 scores 225,124. This 19.2% delta is the single data point for a direct comparison and it favors the RTX 4090 D in a general compute workload.

Looking at the broader context from the rival sets, the PG506-232's score places it 2.4% ahead of the AMD Radeon PRO W7900D (219,827) and 8.7% ahead of the NVIDIA A100 PCIe 80 GB (207,124). It trails the NVIDIA L20 (251,147) by 10.4% but beats the NVIDIA RTX 6000D (195,964) by 14.9%. This shows the PG506-232 is competitive within its professional server class, sitting near the middle of the pack.

The RTX 4090 D's average benchmark score of 178,050 is lower than its OpenCL score due to the inclusion of other tests, but its nearest rivals are all within a tight band. It is 2.2% behind the NVIDIA RTX PRO 5000 Blackwell (182,109), 3.1% behind the NVIDIA A100 SXM4 80 GB (183,725), 3.6% behind the NVIDIA RTX 5000 Ada Generation (184,664), and 4.9% behind the NVIDIA A100 SXM4 40 GB (187,147). The RTX 4090 D's 278,621 OpenCL score, however, is substantially higher than any of these rivals' average scores, suggesting its compute peak is among the highest in its class.

FAQ

Q: Which card has higher raw compute performance?

A: The NVIDIA GeForce RTX 4090 D is clearly superior. It scores 278,621 in Geekbench OpenCL versus 225,124 for the PG506-232, a 19.2% advantage, and offers 73.54 TFLOPS FP32 compared to 10.32 TFLOPS.

Q: Is the PG506-232 more power-efficient?

A: Yes, significantly. The PG506-232 has a 165 W TDP, while the RTX 4090 D has a 425 W TDP. The PG506-232 also suggests a 450 W PSU versus an 800 W PSU for the RTX 4090 D.

Q: Can either card be used for display output?

A: No, the PG506-232 has no display outputs, making it a pure compute card. The RTX 4090 D has 1x HDMI 2.1 and 3x DisplayPort 1.4a outputs.

Q: How do their memory types differ?

A: The PG506-232 uses 24 GB of HBM2 on a 3072-bit bus with 933.1 GB/s bandwidth. The RTX 4090 D uses 24 GB of GDDR6X on a 384-bit bus with 1.01 TB/s bandwidth.

Q: Which card has a smaller physical footprint?

A: The PG506-232 is dual-slot and 267 mm long, while the RTX 4090 D is triple-slot and 304 mm long. The PG506-232 is also shorter in height at 112 mm versus 137 mm.

Q: What is the launch MSRP of the RTX 4090 D?

A: The launch MSRP of the NVIDIA GeForce RTX 4090 D is 1,599 USD.

Where Each One Wins

The RTX 4090 D wins decisively in all measured performance categories. It is 19.2% faster in OpenCL, has a 7.1x higher FP32 throughput (73.54 vs 10.32 TFLOPS), a 3.2x higher pixel rate (443.5 vs 138.2 GPixel/s), and a 3.6x higher texture rate (1,149.1 vs 322.6 GTexel/s). It also has a higher memory bandwidth (1.01 vs 933.1 GB/s) and 4x the shading units (14,592 vs 3,584). This card is for workloads that demand maximum shader compute, ray tracing (114 RT cores), and modern graphics API support (DirectX 12 Ultimate, Vulkan 1.4). Its display outputs also make it viable for local rendering or visualization tasks.

The PG506-232 wins on physical and power efficiency. Its 165 W TDP is 61% lower than the RTX 4090 D's 425 W, and its dual-slot, 267 mm length is far more server-friendly than the RTX 4090 D's triple-slot, 304 mm design. It also uses HBM2 memory, which is often preferred in HPC environments for its density and reliability characteristics. While its 225,124 OpenCL score is lower, it still ranks in the 99th percentile of all GPUs, outperforming the AMD Radeon PRO W7900D by 2.4% and the NVIDIA A100 PCIe 80 GB by 8.7%. This card is for dense server deployments where power limits, chassis space, and memory type are more important than raw speed.

Specification Differences

The two cards differ across nearly every core specification. The PG506-232 uses the GA100 chip on 7 nm, while the RTX 4090 D uses the AD102 chip on 5 nm. Transistor counts are 54,200 million versus 76,300 million, with die sizes of 826 mm² versus 609 mm². Clocks are starkly different: base 930 MHz vs 2,280 MHz, boost 1,440 MHz vs 2,520 MHz. Memory is HBM2 24 GB with a 3072-bit bus versus GDDR6X 24 GB with a 384-bit bus.

Core counts differ: 3,584 vs 14,592 shading units, 224 vs 456 TMUs, and 96 vs 176 ROPs. The RTX 4090 D has 114 RT cores; the PG506-232 has none listed. Tensor cores are 224 vs 456. The PG506-232's TDP is 165 W versus 425 W, with a dual-slot versus triple-slot design. Power connectors are 8-pin EPS versus 1x 16-pin. Display outputs are absent on the PG506-232, while the RTX 4090 D has 1x HDMI 2.1 and 3x DisplayPort 1.4a. API support is absent for the PG506-232, but the RTX 4090 D lists DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4. Release dates are April 2021 for the PG506-232 and December 2023 for the RTX 4090 D. The RTX 4090 D has a launch MSRP of 1,599 USD; the PG506-232 has no listed MSRP.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4090 D
PG506-232
Core Specs
Shading Units
14,592
3,584 -75.4%
Shaders
14,592
3,584 -75.4%
TMUs
456
224 -50.9%
ROPs
176
96 -45.5%
SM Count
114
56 -50.9%
Clocks
Base Clock
2280 MHz
930 MHz
Boost Clock
2520 MHz
1440 MHz
Memory Clock
1313 MHz 21 Gbps effective
1215 MHz 2.4 Gbps effective
Memory
Memory Size
24 GB
24 GB
VRAM (MB)
24,576
24,576 0.0%
Memory Type
GDDR6X
HBM2
Memory Bus
384 bit
3072 bit
Bandwidth
1.01 TB/s
933.1 GB/s
Cache
L1 Cache
128 KB (per SM)
192 KB (per SM)
L2 Cache
72 MB
24 MB
Performance
Pixel Rate
443.5 GPixel/s
138.2 GPixel/s
Texture Rate
1,149.1 GTexel/s
322.6 GTexel/s
FP32 (TFLOPS)
73.54 TFLOPS
10.32 TFLOPS
FP64 (TFLOPS)
1,149.1 GFLOPS (1:64)
5.161 TFLOPS (1:2)
FP16 (TFLOPS)
73.54 TFLOPS (1:1)
10.32 TFLOPS (1:1)
AI/RT
RT Cores
114
—
Tensor Cores
456
224 -50.9%
Power
TDP
425 W
165 W
TDP (W)
425
165 -61.2%
Suggested PSU
800 W
450 W
Power Connectors
1x 16-pin
8-pin EPS
Architecture
Architecture
Ada Lovelace
Ampere
GPU Name
AD102
GA100
Generation
GeForce 40
Server Ampere (Axx)
Process Size
5 nm
7 nm
Transistors
76,300 million
54,200 million
Die Size
609 mm²
826 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
65.6M / mm²
API Support
DirectX
12 Ultimate (12_2)
—
OpenGL
4.6
—
Vulkan
1.4
—
OpenCL
3.0
3.0
CUDA
8.9
8.0
Shader Model
6.8
—
Physical
Slot Width
Triple-slot
Dual-slot
Length
304 mm 12 inches
267 mm 10.5 inches
Height
137 mm 5.4 inches
112 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
1,599 USD
—
Production
End-of-life
End-of-life
Predecessor
GeForce 30
Tesla Turing
Successor
GeForce 50
Server Ada
View GeForce RTX 4090 D Details View PG506-232 Details