NVIDIA P106-100 vs NVIDIA Tesla K40m Comparison

NVIDIA
GEFORCE

NVIDIA P106-100

CORE STATE GP106
VRAM 6 GB
CLOCK SPEED 1709 MHz
TDP 120 W
BUS WIDTH 192 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017
VS
NVIDIA
GEFORCE

Tesla K40m

CORE STATE GK110B
VRAM 12 GB
CLOCK SPEED 876 MHz
TDP 245 W
BUS WIDTH 384 bit
ARCHITECTURE Kepler
nm
PROCESS 28 nm
LAUNCH DATE 2013

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
899
N/A
geekbench_opencl
35,951
19,885
geekbench_vulkan
32,897
N/A

Analysis: NVIDIA P106-100 vs NVIDIA Tesla K40m

Head-to-Head Benchmarks

The recorded data includes only one direct comparison between the NVIDIA P106-100 and the NVIDIA Tesla K40m, and it is a decisive victory for the newer card. In the Geekbench OpenCL test, the P106-100 scores 35,951 points against the Tesla K40m's 19,885 points. That is a delta of 80.8%, meaning the P106-100 delivers roughly 81% higher performance in this compute-oriented workload. This is not a marginal win; it is a generational leap in raw throughput for general-purpose GPU compute.

Looking at the broader benchmark database, the P106-100 also has results in two additional tests that the Tesla K40m does not appear in. In 3DMark Steel Nomad (DX12), the P106-100 scores 899 points. In Geekbench Vulkan, it scores 32,897 points. The Tesla K40m has no recorded scores for either of these tests, so no direct comparison is possible there. However, the average benchmark score across all recorded tests tells a similar story: the P106-100 averages 23,249 points, while the Tesla K40m averages 19,885 points. That puts the P106-100 about 16.9% ahead in average aggregate performance.

The delta between the two cards is consistent with what the architecture shift would suggest. The P106-100 is built on a much newer design with higher clocks and a more efficient process. The Tesla K40m compensates somewhat with more shading units and more memory bandwidth, but in the one workload where both were measured, it simply cannot keep pace. The 80.8% lead in OpenCL is the headline number for this matchup.

FAQ

Q: Which card has the higher average benchmark score?

A: The NVIDIA P106-100 has an average benchmark score of 23,249 points, compared to the Tesla K40m's 19,885 points. The P106-100 sits at the 68th percentile of all GPUs in the database, while the Tesla K40m sits at the 65th percentile.

Q: How much faster is the P106-100 in OpenCL compute?

A: In the Geekbench OpenCL test, the P106-100 scores 35,951 against the Tesla K40m's 19,885, a lead of 80.8%.

Q: Does the Tesla K40m have any benchmark wins over the P106-100?

A: No. In the head-to-head benchmark data, the P106-100 wins the only shared test (Geekbench OpenCL). The wins tally is 1 for the P106-100 and 0 for the Tesla K40m.

Q: What is the memory configuration difference?

A: The Tesla K40m has 12 GB of GDDR5 memory on a 384-bit bus, yielding 288.4 GB/s of bandwidth. The P106-100 has 6 GB of GDDR5 on a 192-bit bus, yielding 192.2 GB/s of bandwidth.

Q: Which card has better API support?

A: The P106-100 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The Tesla K40m supports DirectX 12 (11_1), OpenGL 4.6, and Vulkan 1.2.175. The P106-100 has the higher DirectX feature level and a newer Vulkan version.

Q: What is the launch MSRP of the Tesla K40m?

A: The Tesla K40m had a launch MSRP of 7,699 USD. The P106-100 has no recorded launch MSRP in the database.

Architecture Differences

The two cards come from completely different NVIDIA architectures, and the data reflects that split. The P106-100 is built on the Pascal architecture, using the GP106 chip on a 16 nm process from TSMC. It packs 4,400 million transistors into a 200 mm² die, giving a transistor density of 22.0 million per mm². The Tesla K40m is a Kepler architecture part, using the GK110B chip on a 28 nm process, also from TSMC. It holds 7,080 million transistors on a much larger 561 mm² die, with a lower transistor density of 12.6 million per mm².

Clock speeds tell the rest of the story. The P106-100 runs at a base clock of 1506 MHz and boosts to 1709 MHz. The Tesla K40m runs at just 745 MHz base and 876 MHz boost. That is a massive clock advantage for the Pascal card, which is why it can outperform the Kepler card despite having fewer than half the shading units. The P106-100 has 1,280 shading units, 80 texture mapping units, and 48 ROPs. The Tesla K40m has 2,880 shading units, 240 TMUs, and 48 ROPs. So the Kepler card has more raw execution resources, but the Pascal card's much higher clocks and architectural efficiency win out in practice.

The compute feature sets also differ. The P106-100 achieves 4.375 TFLOPS of FP32 performance and has a token FP16 rate of 68.36 GFLOPS (1:64). The Tesla K40m achieves 5.046 TFLOPS of FP32 and has no recorded FP16 capability. Despite the Tesla K40m's higher theoretical FP32 peak, the P106-100 still wins the real-world OpenCL test by a wide margin.

Both cards are compute-focused in that they have no display outputs. The P106-100 is classified in the database as part of the "Mining GPUs" generation, while the Tesla K40m belongs to the "Tesla Kepler (Kxx)" generation. Neither card has ray tracing cores or tensor cores. The Tesla K40m lists a predecessor of Tesla Fermi and a successor of Tesla Maxwell, while the P106-100 has no predecessor or successor listed.

Specification Differences

The two cards differ across nearly every major specification category. The process node is 16 nm for the P106-100 versus 28 nm for the Tesla K40m. Transistor count is 4,400 million versus 7,080 million. Die size is 200 mm² versus 561 mm². Transistor density is 22.0 million per mm² versus 12.6 million per mm².

Clocks: the P106-100 runs at 1506 MHz base and 1709 MHz boost. The Tesla K40m runs at 745 MHz base and 876 MHz boost. Memory clocks are 2002 MHz (8 Gbps effective) for the P106-100 versus 1502 MHz (6 Gbps effective) for the Tesla K40m.

Memory capacity is 6 GB versus 12 GB. Bus width is 192 bit versus 384 bit. Memory bandwidth is 192.2 GB/s versus 288.4 GB/s. The Tesla K40m has the clear advantage in memory capacity and bandwidth, which matters for workloads that are memory-bound, but it does not translate into a win in the recorded OpenCL test.

Shading units are 1,280 versus 2,880. TMUs are 80 versus 240. ROPs are identical at 48 for both. Pixel rate is 82.03 GPixel/s for the P106-100 versus 52.56 GPixel/s for the Tesla K40m. Texture rate is 136.7 GTexel/s versus 210.2 GTexel/s. FP32 is 4.375 TFLOPS versus 5.046 TFLOPS.

Power and physical specs differ as well. The P106-100 has a TDP of 120 W and requires a 300 W suggested PSU with a single 6-pin power connector. The Tesla K40m has a TDP of 245 W and requires a 550 W suggested PSU, with no power connector details listed. Both are dual-slot cards. The P106-100 is 250 mm (9.8 inches) long, while the Tesla K40m is 267 mm (10.5 inches) long.

The bus interface differs: the P106-100 uses PCIe 1.0 x16, while the Tesla K40m uses PCIe 3.0 x16. The API support differs in DirectX version (12_1 versus 11_1) and Vulkan version (1.4 versus 1.2.175). The release dates are far apart: the P106-100 launched in June 2017, the Tesla K40m in November 2013. Both are marked end-of-life in the database.

Where Each One Wins

The P106-100 wins where it matters most in this comparison: raw compute performance in the recorded benchmark. Its 80.8% lead in Geekbench OpenCL makes it the clear choice for general compute workloads that rely on OpenCL. It also holds a higher average benchmark score (23,249 versus 19,885), a higher percentile ranking (68th versus 65th), and additional benchmark results in 3DMark Steel Nomad and Geekbench Vulkan that the Tesla K40m simply does not have recorded. The Pascal architecture brings much higher clocks and a modern feature set, including DirectX 12 (12_1) and Vulkan 1.4 support. For anyone running compute tasks that can use these APIs, the P106-100 is the stronger card.

The Tesla K40m's advantages are in memory and raw throughput specifications. It has 12 GB of VRAM versus 6 GB, a 384-bit bus versus 192-bit, and 288.4 GB/s of bandwidth versus 192.2 GB/s. It also has more shading units (2,880 versus 1,280), more TMUs (240 versus 80), and a higher theoretical FP32 peak (5.046 TFLOPS versus 4.375 TFLOPS). These specs suggest the Tesla K40m could be preferable for workloads that are heavily memory-bound or that can fully utilize its massive shading unit count, if such workloads do not map well to the P106-100's architecture. Additionally, the Tesla K40m uses PCIe 3.0 x16, which is a more modern bus interface than the P106-100's PCIe 1.0 x16, a potential advantage for data transfer in some systems.

For pixel rate, the P106-100 wins at 82.03 GPixel/s versus 52.56 GPixel/s. For texture rate, the Tesla K40m wins at 210.2 GTexel/s versus 136.7 GTexel/s. So the split is not uniform: the P106-100 dominates in pixel throughput and compute tests, while the Tesla K40m leads in texture throughput and memory bandwidth.

The Verdict

The data points to the NVIDIA P106-100 as the better card for most purposes. Its 80.8% lead in the only shared benchmark, its higher average score, and its higher percentile ranking all favor the Pascal part. The P106-100 also has a lower TDP at 120 W versus 245 W, a smaller physical footprint, and a much later release date. For compute workloads that use OpenCL, Vulkan, or DirectX 12, the P106-100 is the sensible pick.

The Tesla K40m is not without merit. It offers double the VRAM (12 GB versus 6 GB), significantly more memory bandwidth (288.4 GB/s versus 192.2 GB/s), and a higher theoretical FP32 peak. Those specs could matter for specific memory-heavy tasks or for users who need the larger frame buffer. It also has PCIe 3.0 x16 support, which the P106-100 lacks. But none of these advantages show up in the recorded benchmark results, where the Tesla K40m loses the only direct comparison by a large margin. Its launch MSRP of 7,699 USD also reflects its enterprise positioning, though price is not part of this analysis.

For a builder choosing between these two end-of-life cards, the P106-100 is the default recommendation. It wins the measured performance tests, runs cooler, and supports newer APIs. The Tesla K40m should only be considered if the workload explicitly requires more than 6 GB of VRAM or if the memory bandwidth advantage is critical, since the benchmark data does not show it converting those specs into a win. The verdict, strictly from the data, is clear: the P106-100 wins the head-to-head, and the Tesla K40m remains a niche option for memory capacity needs.

DETAILED SPECIFICATIONS

SPECIFICATION
P106-100
Tesla K40m
Core Specs
Shading Units
1,280
2,880 +125.0%
Shaders
1,280
2,880 +125.0%
TMUs
80
240 +200.0%
ROPs
48
48 0.0%
SM Count
10
Clocks
Base Clock
1506 MHz
745 MHz
Boost Clock
1709 MHz
876 MHz
Memory Clock
2002 MHz 8 Gbps effective
1502 MHz 6 Gbps effective
Memory
Memory Size
6 GB
12 GB
VRAM (MB)
6,144
12,288 +100.0%
Memory Type
GDDR5
GDDR5
Memory Bus
192 bit
384 bit
Bandwidth
192.2 GB/s
288.4 GB/s
Cache
L1 Cache
48 KB (per SM)
16 KB (per SMX)
L2 Cache
1536 KB
1536 KB
Performance
Pixel Rate
82.03 GPixel/s
52.56 GPixel/s
Texture Rate
136.7 GTexel/s
210.2 GTexel/s
FP32 (TFLOPS)
4.375 TFLOPS
5.046 TFLOPS
FP64 (TFLOPS)
136.7 GFLOPS (1:32)
1.682 TFLOPS (1:3)
FP16 (TFLOPS)
68.36 GFLOPS (1:64)
Power
TDP
120 W
245 W
TDP (W)
120
245 +104.2%
Suggested PSU
300 W
550 W
Power Connectors
1x 6-pin
Architecture
Architecture
Pascal
Kepler
GPU Name
GP106
GK110B
Generation
Mining GPUs
Tesla Kepler (Kxx)
Process Size
16 nm
28 nm
Transistors
4,400 million
7,080 million
Die Size
200 mm²
561 mm²
Foundry
TSMC
TSMC
Density
22.0M / mm²
12.6M / mm²
API Support
DirectX
12 (12_1)
12 (11_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.2.175
OpenCL
3.0
3.0
CUDA
6.1
3.5
Shader Model
6.8
6.5 (5.1)
Physical
Slot Width
Dual-slot
Dual-slot
Length
250 mm 9.8 inches
267 mm 10.5 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 1.0 x16
PCIe 3.0 x16
Other
Launch Price
7,699 USD
Production
End-of-life
End-of-life
Predecessor
Tesla Fermi
Successor
Tesla Maxwell
View P106-100 Details View Tesla K40m Details