NVIDIA P104-100 vs NVIDIA T1000 Comparison
NVIDIA P104-100
T1000
PERFORMANCE BENCHMARKS
Analysis: NVIDIA P104-100 vs NVIDIA T1000
Head-to-Head Benchmarks
The recorded data contains two direct benchmark comparisons between the NVIDIA T1000 and the NVIDIA P104-100. Both are compute-oriented tests, and the results consistently favor the P104-100. In Geekbench OpenCL, the T1000 scores 37,704 points against 52,368 for the P104-100. That translates to a 28% deficit for the T1000, a substantial margin in raw compute throughput. The Vulkan test tells a similar story: the T1000 records 34,874 points, while the P104-100 reaches 45,165, a 22.8% gap. In both cases, the P104-100 is the clear winner, and the T1000 does not win a single head-to-head benchmark in the database.
Looking at the broader context, the T1000's average benchmark score is 36,289, placing it in the 80th percentile among all GPUs. Its nearest rivals include the AMD Radeon RX 5300M at 36,529 (0.7% higher), the NVIDIA GeForce GTX TITAN X at 36,530 (0.7% higher), the AMD Radeon Pro Duo at 35,860 (1.2% lower), and the NVIDIA Quadro GV100 at 35,520 (2.2% lower). This places the T1000 in a tight cluster where performance differences are minimal, often within a single percentage point. The P104-100, by contrast, averages 32,982 points and sits in the 77th percentile. Its nearest rivals are the NVIDIA T600 Mobile at 32,849 (0.4% higher), the NVIDIA T550 Mobile at 33,161 (0.5% lower), the NVIDIA GeForce RTX 3050 Mobile at 33,170 (0.6% lower), and the AMD Radeon Pro 570 at 33,207 (0.7% lower).
The interesting nuance here is that the P104-100 wins the direct comparison convincingly, yet its overall average score is lower than the T1000's. This is because the T1000 has only two recorded benchmarks, both in Geekbench, while the P104-100 has three, including a 3DMark Steel Nomad DX12 score of 1,413. That additional test point, likely a heavier workload, drags down the P104-100's average. The head-to-head data, however, is unambiguous: in the two tests where both cards appear, the P104-100 leads by roughly a quarter. The T1000's advantage in average score is an artifact of test selection, not a reflection of direct performance parity.
FAQ
Q: Which GPU wins the head-to-head benchmarks?
A: The NVIDIA P104-100 wins both recorded comparisons. It leads by 28% in Geekbench OpenCL (52,368 vs. 37,704) and by 22.8% in Geekbench Vulkan (45,165 vs. 34,874).
Q: How do the average benchmark scores compare?
A: The T1000 has a higher average score at 36,289, versus 32,982 for the P104-100. However, the P104-100 includes a 3DMark Steel Nomad DX12 result of 1,413 in its average, while the T1000 only has Geekbench OpenCL and Vulkan scores.
Q: What percentile does each card occupy among all GPUs?
A: The T1000 sits in the 80th percentile, while the P104-100 sits in the 77th percentile. Both are above the median, but the T1000 ranks slightly higher overall.
Q: Which card has higher compute throughput in raw FP32 performance?
A: The P104-100 is significantly ahead, with 6.655 TFLOPS compared to 2.500 TFLOPS for the T1000. This is a 2.66x advantage in raw shader performance.
Q: Do both cards support the same DirectX and Vulkan versions?
A: Yes, both support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The API feature sets are identical, despite the architectural differences.
Q: Which card has more memory bandwidth?
A: The P104-100 has double the bus width (256-bit vs. 128-bit) and uses GDDR5X memory, yielding 320.3 GB/s versus 160.0 GB/s for the T1000's GDDR6 on a 128-bit bus.
Architecture Differences
The two cards are built on different architectures from different generations. The T1000 uses the TU117 chip based on the Turing architecture, fabricated on a 12 nm process at TSMC. It integrates 4,700 million transistors on a 200 mm² die, giving a transistor density of 23.5 million per mm². The P104-100 uses the GP104 chip based on the older Pascal architecture, also from TSMC but on a larger 16 nm node. It packs 7,200 million transistors onto a 314 mm² die, with a density of 22.9 million per mm². The P104-100 has roughly 53% more transistors and a 57% larger die, but the T1000 achieves a marginally higher transistor density thanks to the more modern process.
Compute resources differ sharply. The T1000 has 896 shading units, 56 texture mapping units, and 32 ROPs. The P104-100 more than doubles the shading units to 1,920, with 120 TMUs and 64 ROPs. This explains the large gap in raw throughput: the T1000 delivers 2.500 TFLOPS FP32, while the P104-100 reaches 6.655 TFLOPS. Pixel and texture rates follow the same pattern: the T1000 records 44.64 GPixel/s and 78.12 GTexel/s, while the P104-100 hits 110.9 GPixel/s and 208.0 GTexel/s. The P104-100 is roughly 2.5x faster in pixel throughput and 2.7x faster in texture throughput.
Memory subsystems also diverge. The T1000 uses 4 GB of GDDR6 on a 128-bit bus, running at 1250 MHz with 10 Gbps effective speed, for a bandwidth of 160.0 GB/s. The P104-100 also has 4 GB, but it uses GDDR5X on a 256-bit bus, running at 1251 MHz with the same 10 Gbps effective speed, doubling bandwidth to 320.3 GB/s. Both cards lack dedicated ray tracing and tensor cores, and neither has a published game clock.
FP16 performance presents an interesting inversion. The T1000 supports FP16 at 5.000 TFLOPS with a 2:1 ratio relative to FP32, meaning it can double its rate when using half precision. The P104-100, by contrast, has heavily crippled FP16 at 104.0 GFLOPS, a 1:64 ratio, which is actually slower than its FP32 rate. For workloads that leverage FP16, the T1000 holds a decisive advantage, despite being the slower card overall in FP32.
The Verdict
The data points to a clear split based on workload type. If the task is general compute, especially FP32-heavy workloads, the P104-100 is the stronger choice. Its 6.655 TFLOPS FP32, 320.3 GB/s memory bandwidth, and double the shading units give it a commanding lead in the head-to-head Geekbench results, where it wins by 28% and 22.8%. The P104-100 also has a 2.7x advantage in texture rate and 2.5x in pixel rate, making it better suited for rendering tasks that push those units.
The T1000, however, is not without its strengths. It achieves a higher average benchmark score (36,289 vs. 32,982) and a higher percentile ranking (80th vs. 77th), which suggests it is more consistent across the test suite recorded. Its FP16 performance is dramatically better, at 5.000 TFLOPS versus 104.0 GFLOPS, and it does so on a lower power draw of 50 W compared to the P104-100's unlisted TDP. The T1000 is also a single-slot card with no external power connectors and a suggested PSU of 250 W, while the P104-100 is dual-slot, requires a single 8-pin connector, and has a suggested PSU of 200 W.
For a workstation context, the T1000 offers display outputs (four mini-DisplayPort 1.4a) and a PCIe 3.0 x16 interface. The P104-100 has no display outputs at all and only a PCIe 1.0 x4 interface, which severely limits data transfer to the host system. This makes the P104-100 unsuitable for any interactive or display-centric use case. It is a compute-only card, likely intended for mining or server-side computation where video output is irrelevant. The T1000, by contrast, is a proper workstation GPU with full display support and a modern bus interface.
Specification Differences
| Field | NVIDIA T1000 | NVIDIA P104-100 |
|---|---|---|
| Architecture | Turing | Pascal |
| Process Node | 12 nm | 16 nm |
| Transistors | 4,700 million | 7,200 million |
| Die Size | 200 mm² | 314 mm² |
| Base Clock | 1065 MHz | 1607 MHz |
| Boost Clock | 1395 MHz | 1733 MHz |
| Memory Type | GDDR6 | GDDR5X |
| Memory Bus Width | 128 bit | 256 bit |
| Memory Bandwidth | 160.0 GB/s | 320.3 GB/s |
| Shading Units | 896 | 1920 |
| TMUs | 56 | 120 |
| ROPs | 32 | 64 |
| Pixel Rate | 44.64 GPixel/s | 110.9 GPixel/s |
| Texture Rate | 78.12 GTexel/s | 208.0 GTexel/s |
| FP32 | 2.500 TFLOPS | 6.655 TFLOPS |
| FP16 | 5.000 TFLOPS (2:1) | 104.0 GFLOPS (1:64) |
| TDP | 50 W | Not specified |
| Slot Width | Single-slot | Dual-slot |
| Power Connectors | None | 1x 8-pin |
| Suggested PSU | 250 W | 200 W |
| Bus Interface | PCIe 3.0 x16 | PCIe 1.0 x4 |
| Display Outputs | 4x mini-DisplayPort 1.4a | No outputs |
| Length | 156 mm (6.1 inches) | 267 mm (10.5 inches) |
| Release Date | 2021-05-05 | 2017-12-11 |
Where Each One Wins
The P104-100 dominates in raw compute performance. Every major throughput metric favors it: FP32 is 2.66x higher, memory bandwidth is exactly 2x higher, shading units are 2.14x more numerous, and texture rate is 2.66x higher. In the head-to-head benchmarks, it wins both tests by substantial margins. If the workload is purely computational, such as batch rendering, scientific simulation, or any FP32-heavy task that does not require display output, the P104-100 is the data-backed choice.
The T1000 wins in several other areas that matter for practical deployment. Its FP16 throughput is 48x higher (5.000 TFLOPS vs. 104.0 GFLOPS), which is critical for workloads that use half-precision arithmetic, such as certain AI inference or image processing tasks. It consumes only 50 W, a fraction of the P104-100's unspecified but likely much higher power draw given the dual-slot cooler and 8-pin connector. The T1000 is single-slot, has no power connector requirement, and offers four mini-DisplayPort outputs, making it a drop-in solution for professional workstations. Its PCIe 3.0 x16 interface also provides far more host bandwidth than the P104-100's PCIe 1.0 x4, which is a severe bottleneck for any data transfer between GPU and CPU.
The T1000 also has a higher overall average benchmark score and percentile ranking, indicating better consistency across the tests recorded in the database. Its release date is also more recent (2021 vs. 2017), though both are now end-of-life. The P104-100's lack of display outputs alone disqualifies it for any interactive use. The T1000 is the only one of the two that can function as a standard graphics card. The verdict from the data is straightforward: the P104-100 is a compute accelerator with a narrow focus, while the T1000 is a versatile workstation GPU that trades raw FP32 power for efficiency, display capability, and modern features.