NVIDIA P106-090 vs NVIDIA Quadro K4200 Comparison

NVIDIA
GEFORCE

NVIDIA P106-090

CORE STATE GP106
VRAM 3 GB
CLOCK SPEED 1531 MHz
TDP 75 W
BUS WIDTH 192 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017
VS
NVIDIA
GEFORCE

Quadro K4200

CORE STATE GK104
VRAM 4 GB
CLOCK SPEED 784 MHz
TDP 108 W
BUS WIDTH 256 bit
ARCHITECTURE Kepler
nm
PROCESS 28 nm
LAUNCH DATE 2014

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
509
N/A
geekbench_opencl
21,304
12,313
geekbench_vulkan
18,596
12,482

Analysis: NVIDIA P106-090 vs NVIDIA Quadro K4200

The NVIDIA Quadro K4200 and NVIDIA P106-090 are two very different products from the same company, aimed at entirely different tasks. The K4200 is a professional workstation card from the Kepler era, built for stability and display output, while the P106-090 is a mining-focused card from the Pascal era, stripped of video outputs entirely. The benchmark data shows a clear performance hierarchy, but the right choice depends entirely on whether you need to see a picture on a screen or just process data.

Where Each One Wins

The P106-090 wins every single head-to-head benchmark in this comparison. In Geekbench OpenCL, it scores 21,304 against the K4200's 12,313, a 42.2% advantage. In Geekbench Vulkan, it scores 18,596 versus 12,482, a 32.9% lead. The P106-090 also has a higher average benchmark score of 13,470 compared to the K4200's 12,398, placing it in the 54th percentile of all GPUs versus the K4200's 52nd percentile. There are no benchmark categories where the Quadro comes out ahead.

However, the K4200 wins in practical usability for any desktop application. It features display outputs (1x DVI and 2x DisplayPort 1.2), while the P106-090 has no outputs at all. The Quadro is a single-slot card and is designed for a PCIe 2.0 x16 interface, whereas the P106-090 is a dual-slot card with a severely limited PCIe 1.0 x1 interface. For any task that requires a monitor, the K4200 is the only functional choice, regardless of raw compute scores.

Architecture Differences

The two cards are built on different architectures from different manufacturing processes. The Quadro K4200 uses the GK104 chip on the Kepler architecture, fabricated on a 28 nm process at TSMC. It packs 3,540 million transistors on a 294 mm² die, yielding a transistor density of 12.0 million per square millimeter. The P106-090 uses the GP106 chip on the Pascal architecture, also from TSMC, but on a much smaller 16 nm process. It contains 4,400 million transistors on a 200 mm² die, achieving a higher density of 22.0 million per square millimeter.

The core configurations differ significantly. The K4200 has 1,344 shading units, 112 texture mapping units (TMUs), and 32 raster operation units (ROPs). The P106-090 has fewer shading units at 768, fewer TMUs at 48, but more ROPs at 48. Clock speeds are substantially higher on the P106-090, with a base clock of 1,354 MHz and boost clock of 1,531 MHz, compared to the K4200's modest 771 MHz base and 784 MHz boost. This leads to distinct performance characteristics: the K4200 achieves a pixel rate of 21.95 GPixel/s and texture rate of 87.81 GTexel/s, while the P106-090 reaches 73.49 GPixel/s and 73.49 GTexel/s respectively. The FP32 compute throughput favors the P106-090 at 2.352 TFLOPS versus 2.107 TFLOPS.

The P106-090 also supports newer API features. It supports DirectX 12 (12_1) and Vulkan 1.4, while the K4200 is limited to DirectX 12 (11_0) and Vulkan 1.2.175. Both support OpenGL 4.6. The P106-090 has a FP16 rate of 36.74 GFLOPS (1:64), a feature the K4200 does not list.

Head-to-Head Benchmarks

The data from the head-to-head comparison is unambiguous. In Geekbench OpenCL, the P106-090's score of 21,304 crushes the K4200's 12,313. This represents a delta of -42.2% for the Quadro, meaning it performs at less than 60% of the P106-090's level in this test. This is a massive gap, likely driven by the P106-090's much higher clock speeds and newer architecture.

The Geekbench Vulkan test shows a narrower but still decisive gap. The P106-090 scores 18,596 against the K4200's 12,482, a 32.9% deficit for the Quadro. Vulkan is a lower-level API that can benefit from newer hardware features, and the Pascal architecture's improvements over Kepler are evident here. The K4200's lower FP32 throughput and older API support (Vulkan 1.2.175 vs 1.4) contribute to this result.

Looking at relative positioning against rivals, the P106-090's average score of 13,470 sits just 0.3% behind the NVIDIA GeForce GTX 570 (13,515) and 0.5% ahead of the AMD Radeon Pro 555 (13,407). The K4200's average of 12,398 is 1.8% behind the NVIDIA Tesla K20Xm (12,625) and 3.3% ahead of the NVIDIA GeForce GTX 960A (11,998). These deltas show that while the P106-090 is in a higher performance bracket overall, both cards are competitive within their respective eras.

Specification Differences

The following specifications differ between the two cards:

  • Architecture: Kepler (K4200) vs Pascal (P106-090)
  • Process Node: 28 nm vs 16 nm
  • Transistors: 3,540 million vs 4,400 million
  • Die Size: 294 mm² vs 200 mm²
  • Transistor Density: 12.0M / mm² vs 22.0M / mm²
  • Base Clock: 771 MHz vs 1,354 MHz
  • Boost Clock: 784 MHz vs 1,531 MHz
  • Memory Clock: 1,350 MHz (5.4 Gbps effective) vs 2,002 MHz (8 Gbps effective)
  • Memory Size: 4 GB vs 3 GB
  • Memory Bus Width: 256 bit vs 192 bit
  • Memory Bandwidth: 172.8 GB/s vs 192.2 GB/s
  • Shading Units: 1,344 vs 768
  • TMUs: 112 vs 48
  • ROPs: 32 vs 48
  • Pixel Rate: 21.95 GPixel/s vs 73.49 GPixel/s
  • Texture Rate: 87.81 GTexel/s vs 73.49 GTexel/s
  • FP32: 2.107 TFLOPS vs 2.352 TFLOPS
  • FP16: Not listed vs 36.74 GFLOPS (1:64)
  • TDP: 108 W vs 75 W
  • Slot Width: Single-slot vs Dual-slot
  • Suggested PSU: 300 W vs 250 W
  • Bus Interface: PCIe 2.0 x16 vs PCIe 1.0 x1
  • Display Outputs: 1x DVI, 2x DisplayPort 1.2 vs No outputs
  • DirectX Support: 12 (11_0) vs 12 (12_1)
  • Vulkan Support: 1.2.175 vs 1.4
  • Release Date: 2014-07-21 vs 2017-07-30
  • Generation: Quadro Kepler (Kx200) vs Mining GPUs

FAQ

Q: Which card has higher raw compute performance?

A: The P106-090 has higher FP32 throughput at 2.352 TFLOPS compared to the K4200's 2.107 TFLOPS. It also wins both head-to-head benchmarks, with a 42.2% lead in Geekbench OpenCL and a 32.9% lead in Geekbench Vulkan.

Q: Can the P106-090 be used for normal desktop tasks with a monitor?

A: No. The P106-090 has no display outputs, making it impossible to connect a monitor directly. The K4200, with its 1x DVI and 2x DisplayPort 1.2 outputs, is the only one of the two that can drive a display.

Q: Why does the P106-090 have a higher average benchmark score despite fewer shading units?

A: The P106-090 has 768 shading units compared to 1,344 on the K4200, but it compensates with much higher clock speeds (1,531 MHz boost vs 784 MHz) and a newer Pascal architecture. Its average score of 13,470 is 8.6% higher than the K4200's 12,398.

Q: Which card is more power efficient?

A: The P106-090 has a lower TDP of 75 W versus 108 W for the K4200, and it also recommends a smaller PSU (250 W vs 300 W). This is notable given the P106-090's higher performance.

Q: Are there any API differences that matter?

A: Yes. The P106-090 supports DirectX 12 (12_1) and Vulkan 1.4, while the K4200 is limited to DirectX 12 (11_0) and Vulkan 1.2.175. Both support OpenGL 4.6.

Q: Which card has more memory?

A: The K4200 has 4 GB of GDDR5 memory on a 256-bit bus, while the P106-090 has 3 GB on a 192-bit bus. However, the P106-090 has higher memory bandwidth at 192.2 GB/s versus 172.8 GB/s due to its faster 8 Gbps effective memory clock.

The Verdict

The choice between these two cards is dictated entirely by the use case. The P106-090 is the superior compute performer by every measurable metric in this data. It wins both benchmarks decisively, has a higher average score (13,470 vs 12,398), higher pixel rate (73.49 vs 21.95 GPixel/s), higher FP32 throughput (2.352 vs 2.107 TFLOPS), and lower power consumption (75 W vs 108 W). It is also a newer design with a more advanced 16 nm process and better API support. If the workload is headless — such as compute tasks, mining, or server-side processing — the P106-090 is the clear choice based on performance data alone.

The K4200's only advantages are its display outputs, larger memory capacity (4 GB vs 3 GB), wider memory bus (256 bit vs 192 bit), and its single-slot form factor with a standard PCIe 2.0 x16 interface. The P106-090's PCIe 1.0 x1 interface is a severe bottleneck for any data transfer, though it does not affect raw compute scores. For any professional workstation task that requires a monitor, the K4200 is the only viable option. Its Kepler architecture offers proven driver stability in workstation environments, and the 4 GB memory capacity is better suited for large datasets than the P106-090's 3 GB.

In summary, the data says this: if you need to see your work, buy the K4200. If you never need a display and only care about raw throughput, the P106-090 is the faster card by a wide margin. The Quadro's professional pedigree and display support justify its existence, but the P106-090 wins the performance battle outright.

DETAILED SPECIFICATIONS

SPECIFICATION
P106-090
Quadro K4200
Core Specs
Shading Units
768
1,344 +75.0%
Shaders
768
1,344 +75.0%
TMUs
48
112 +133.3%
ROPs
48
32 -33.3%
SM Count
6
Clocks
Base Clock
1354 MHz
771 MHz
Boost Clock
1531 MHz
784 MHz
Memory Clock
2002 MHz 8 Gbps effective
1350 MHz 5.4 Gbps effective
Memory
Memory Size
3 GB
4 GB
VRAM (MB)
3,072
4,096 +33.3%
Memory Type
GDDR5
GDDR5
Memory Bus
192 bit
256 bit
Bandwidth
192.2 GB/s
172.8 GB/s
Cache
L1 Cache
48 KB (per SM)
16 KB (per SMX)
L2 Cache
1536 KB
512 KB
Performance
Pixel Rate
73.49 GPixel/s
21.95 GPixel/s
Texture Rate
73.49 GTexel/s
87.81 GTexel/s
FP32 (TFLOPS)
2.352 TFLOPS
2.107 TFLOPS
FP64 (TFLOPS)
73.49 GFLOPS (1:32)
87.81 GFLOPS (1:24)
FP16 (TFLOPS)
36.74 GFLOPS (1:64)
Power
TDP
75 W
108 W
TDP (W)
75
108 +44.0%
Suggested PSU
250 W
300 W
Power Connectors
1x 6-pin
1x 6-pin
Architecture
Architecture
Pascal
Kepler
GPU Name
GP106
GK104
Generation
Mining GPUs
Quadro Kepler (Kx200)
Process Size
16 nm
28 nm
Transistors
4,400 million
3,540 million
Die Size
200 mm²
294 mm²
Foundry
TSMC
TSMC
Density
22.0M / mm²
12.0M / mm²
API Support
DirectX
12 (12_1)
12 (11_0)
OpenGL
4.6
4.6
Vulkan
1.4
1.2.175
OpenCL
3.0
3.0
CUDA
6.1
3.0
Shader Model
6.8
6.5 (5.1)
Physical
Slot Width
Dual-slot
Single-slot
Length
250 mm 9.8 inches
241 mm 9.5 inches
Height
111 mm 4.4 inches
Outputs
No outputs
1x DVI2x DisplayPort 1.2
Bus Interface
PCIe 1.0 x1
PCIe 2.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Quadro Fermi
Successor
Quadro Maxwell
View P106-090 Details View Quadro K4200 Details