NVIDIA PG506-232 vs NVIDIA Quadro GP100 Comparison

NVIDIA
GEFORCE

NVIDIA PG506-232

CORE STATE GA100
VRAM 24 GB
CLOCK SPEED 1440 MHz
TDP 165 W
BUS WIDTH 3072 bit
ARCHITECTURE Ampere
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Quadro GP100

CORE STATE GP100
VRAM 16 GB
CLOCK SPEED 1443 MHz
TDP 235 W
BUS WIDTH 4096 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_opencl
225,124
87,445

Analysis: NVIDIA PG506-232 vs NVIDIA Quadro GP100

Head-to-Head Benchmarks

The recorded database contains only one common benchmark between these two accelerators: Geekbench OpenCL. The result is decisive. The NVIDIA PG506-232 scores 225,124, while the NVIDIA Quadro GP100 scores 87,445. That is a delta of 157.4% in favor of the PG506-232. In practical terms, the Ampere-based card delivers more than two and a half times the raw OpenCL throughput of the Pascal-era Quadro.

The PG506-232 sits at the 99th percentile of all GPUs in the database. Its nearest rivals include the AMD Radeon PRO W7900D at 219,827 (2.4% behind), the NVIDIA A100 PCIe 80 GB at 207,124 (8.7% behind), and the NVIDIA RTX 6000D at 195,964 (14.9% behind). Only the NVIDIA L20, at 251,147, places ahead of it among the listed rivals, and that card leads by 10.4%. The PG506-232 is clearly positioned near the top of the professional compute stack.

The Quadro GP100, by contrast, sits at the 93rd percentile. Its nearest rivals are much closer in score. The AMD Radeon PRO W7600 scores 87,108, just 0.4% behind. The NVIDIA CMP 40HX scores 85,637, 2.1% behind. The NVIDIA RTX A4500 Mobile leads it by 4% with a score of 91,134, and the NVIDIA RTX A4500 leads by 4.6% with 91,671. The GP100 is competitive with mid-range workstation cards of a later generation, but it is not in the same performance class as the PG506-232.

The head-to-head result shows a single win for the PG506-232 and zero wins for the Quadro GP100. There is no benchmark in the database where the older card comes out ahead. The magnitude of the OpenCL gap, 157.4%, is far larger than the typical generational step seen between adjacent product tiers. This is not a marginal improvement; it is a wholesale leap in compute capability.

Architecture Differences

The two cards come from different architectural eras. The PG506-232 is built on the Ampere architecture with the GA100 chip, fabricated on TSMC's 7 nm process. The Quadro GP100 uses the Pascal architecture with the GP100 chip, on TSMC's 16 nm process. The node shrink is substantial, and the transistor counts reflect it. The GA100 packs 54,200 million transistors on an 826 mm² die, yielding a transistor density of 65.6 million per square millimeter. The GP100 contains 15,300 million transistors on a 610 mm² die, for a density of 25.1 million per square millimeter. The Ampere chip more than triples the transistor count while increasing die size by only about a third.

Both cards share the same count of shading units (3,584), TMUs (224), and ROPs (96). The pixel rates are nearly identical: 138.2 GPixel/s for the PG506-232 and 138.5 GPixel/s for the Quadro GP100. Texture rates are similarly close: 322.6 GTexel/s versus 323.2 GTexel/s. In terms of pure rasterization throughput, the two are effectively matched. The FP32 figures are also almost equal, 10.32 TFLOPS for the PG506-232 and 10.34 TFLOPS for the Quadro GP100.

The major architectural split appears in FP16 and in tensor hardware. The Quadro GP100 reaches 20.69 TFLOPS FP16, at a 2:1 ratio relative to FP32. The PG506-232 lists 10.32 TFLOPS FP16 at a 1:1 ratio. On paper, the older card has a higher FP16 throughput. However, the PG506-232 includes 224 tensor cores, while the Quadro GP100 has none. The database does not provide a tensor-specific benchmark, but the presence of tensor cores is a defining feature of the Ampere generation and is absent from Pascal. For workloads that use tensor operations, the PG506-232 has dedicated hardware that the GP100 simply lacks.

Memory configurations also diverge. The PG506-232 uses 24 GB of HBM2 on a 3072-bit bus, with a bandwidth of 933.1 GB/s. The Quadro GP100 uses 16 GB of HBM2 on a wider 4096-bit bus, but its bandwidth is lower at 732.2 GB/s. The effective memory clock is 2.4 Gbps for the PG506-232 versus 1430 Mbps for the GP100. The newer card has more capacity and higher bandwidth despite a narrower bus, thanks to faster memory signaling.

The PG506-232 is a server part with no display outputs. The Quadro GP100 is a workstation card with 1x DVI and 4x DisplayPort 1.4a outputs. The PG506-232 uses PCIe 4.0 x16, while the Quadro GP100 uses PCIe 3.0 x16. The newer card also has a lower TDP of 165 W versus 235 W for the older one. The power connector differs as well: 8-pin EPS for the PG506-232 versus 1x 8-pin for the Quadro GP100.

FAQ

Q: Which card has the higher Geekbench OpenCL score?

A: The NVIDIA PG506-232 scores 225,124, which is 157.4% higher than the Quadro GP100's 87,445.

Q: Does the Quadro GP100 have any advantage in FP16 performance?

A: Yes. The Quadro GP100 lists FP16 at 20.69 TFLOPS with a 2:1 ratio, while the PG506-232 lists 10.32 TFLOPS with a 1:1 ratio. However, the PG506-232 includes 224 tensor cores, which the GP100 lacks.

Q: Which card has more memory?

A: The PG506-232 has 24 GB of HBM2, while the Quadro GP100 has 16 GB of HBM2. The PG506-232 also has higher bandwidth at 933.1 GB/s versus 732.2 GB/s.

Q: Are the shading unit counts the same?

A: Yes. Both cards have 3,584 shading units, 224 TMUs, and 96 ROPs. Their FP32 performance is nearly identical at 10.32 TFLOPS and 10.34 TFLOPS respectively.

Q: Which card supports display outputs?

A: The Quadro GP100 has 1x DVI and 4x DisplayPort 1.4a outputs. The PG506-232 has no display outputs, as it is a server accelerator.

Q: How do the cards compare in power consumption?

A: The PG506-232 has a TDP of 165 W, while the Quadro GP100 has a TDP of 235 W.

Specification Differences

The two cards differ in nearly every major specification category. The process node is 7 nm for the PG506-232 versus 16 nm for the Quadro GP100. Transistor count is 54,200 million versus 15,300 million. Die size is 826 mm² versus 610 mm². Transistor density is 65.6M per mm² versus 25.1M per mm².

Clock speeds differ in base and memory. The PG506-232 has a base clock of 930 MHz and a boost clock of 1440 MHz. The Quadro GP100 has a base clock of 1304 MHz and a boost clock of 1443 MHz. Memory clocks are 1215 MHz (2.4 Gbps effective) for the PG506-232 and 715 MHz (1430 Mbps effective) for the Quadro GP100.

Memory capacity is 24 GB versus 16 GB. Bus width is 3072 bit versus 4096 bit. Bandwidth is 933.1 GB/s versus 732.2 GB/s. FP16 performance is 10.32 TFLOPS versus 20.69 TFLOPS. The PG506-232 has 224 tensor cores; the Quadro GP100 has none.

TDP is 165 W versus 235 W. Power connectors are 8-pin EPS versus 1x 8-pin. Suggested PSU is 450 W versus 550 W. Bus interface is PCIe 4.0 x16 versus PCIe 3.0 x16. Display outputs are none versus 1x DVI and 4x DisplayPort 1.4a.

API support also differs. The Quadro GP100 lists DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.3. The PG506-232 lists no API values in the database, consistent with its compute-focused server role. Physical dimensions are nearly the same: both are 267 mm long, with heights of 112 mm and 111 mm respectively. Both are dual-slot cards.

The release dates are far apart. The Quadro GP100 was released on 2016-09-30, while the PG506-232 was released on 2021-04-11. The predecessor and successor entries also differ: the Quadro GP100 follows Quadro Maxwell and precedes Quadro Volta, while the PG506-232 follows Tesla Turing and precedes Server Ada.

The Verdict

The data points to a clear split in use cases. For anyone running compute workloads that are captured by OpenCL benchmarks, the PG506-232 is the stronger card by a wide margin. Its 225,124 score versus 87,445 is not a close contest. The 157.4% lead means that the Ampere card can finish in roughly one third of the time of the Pascal card for the same OpenCL workload. It also places in the 99th percentile of all GPUs, surrounded by the A100 and L20 class of accelerators. For server deployments, headless compute nodes, or any environment where display output is irrelevant, the PG506-232 is the obvious choice.

The Quadro GP100 has its own niche. It is the only one of the two with display outputs, offering 1x DVI and 4x DisplayPort 1.4a. Its FP16 throughput of 20.69 TFLOPS is higher than the PG506-232's 10.32 TFLOPS, which could matter for workloads that rely on FP16 arithmetic without tensor cores. It also has a wider 4096-bit memory bus, though its effective bandwidth is lower. For a workstation that needs to drive displays and occasionally accelerate compute, the GP100 remains functional. Its 93rd percentile ranking and close rivalry with the RTX A4500 class suggest it is still competitive with mid-range cards of a later generation.

The PG506-232 wins on memory capacity (24 GB versus 16 GB), bandwidth (933.1 GB/s versus 732.2 GB/s), power efficiency (165 W versus 235 W), and interface generation (PCIe 4.0 versus PCIe 3.0). It also adds tensor cores, which are absent from the GP100. The Quadro GP100 wins on FP16 throughput, display support, and a wider memory bus, though the latter does not translate into higher bandwidth.

The verdict, strictly from the data: choose the PG506-232 for compute density, memory capacity, bandwidth, and modern architecture. Choose the Quadro GP100 only if display outputs or higher raw FP16 throughput without tensor hardware are required. In the single recorded head-to-head benchmark, the PG506-232 wins outright. The database records zero wins for the Quadro GP100. For most buyers comparing these two today, the PG506-232 is the more capable accelerator, and the 157.4% OpenCL delta is the headline number that explains why.

DETAILED SPECIFICATIONS

SPECIFICATION
PG506-232
Quadro GP100
Core Specs
Shading Units
3,584
3,584 0.0%
Shaders
3,584
3,584 0.0%
TMUs
224
224 0.0%
ROPs
96
96 0.0%
SM Count
56
56 0.0%
Clocks
Base Clock
930 MHz
1304 MHz
Boost Clock
1440 MHz
1443 MHz
Memory Clock
1215 MHz 2.4 Gbps effective
715 MHz 1430 Mbps effective
Memory
Memory Size
24 GB
16 GB
VRAM (MB)
24,576
16,384 -33.3%
Memory Type
HBM2
HBM2
Memory Bus
3072 bit
4096 bit
Bandwidth
933.1 GB/s
732.2 GB/s
Cache
L1 Cache
192 KB (per SM)
24 KB (per SM)
L2 Cache
24 MB
4 MB
Performance
Pixel Rate
138.2 GPixel/s
138.5 GPixel/s
Texture Rate
322.6 GTexel/s
323.2 GTexel/s
FP32 (TFLOPS)
10.32 TFLOPS
10.34 TFLOPS
FP64 (TFLOPS)
5.161 TFLOPS (1:2)
5.172 TFLOPS (1:2)
FP16 (TFLOPS)
10.32 TFLOPS (1:1)
20.69 TFLOPS (2:1)
AI/RT
Tensor Cores
224
Power
TDP
165 W
235 W
TDP (W)
165
235 +42.4%
Suggested PSU
450 W
550 W
Power Connectors
8-pin EPS
1x 8-pin
Architecture
Architecture
Ampere
Pascal
GPU Name
GA100
GP100
Generation
Server Ampere (Axx)
Quadro Pascal (Px000)
Process Size
7 nm
16 nm
Transistors
54,200 million
15,300 million
Die Size
826 mm²
610 mm²
Foundry
TSMC
TSMC
Density
65.6M / mm²
25.1M / mm²
API Support
DirectX
12 (12_1)
OpenGL
4.6
Vulkan
1.3
OpenCL
3.0
3.0
CUDA
8.0
6.0
Shader Model
6.0
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
112 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
1x DVI4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Turing
Quadro Maxwell
Successor
Server Ada
Quadro Volta
View PG506-232 Details View Quadro GP100 Details