NVIDIA P104-100 vs NVIDIA T1000 8 GB Comparison
NVIDIA P104-100
T1000 8 GB
PERFORMANCE BENCHMARKS
Analysis: NVIDIA P104-100 vs NVIDIA T1000 8 GB
NVIDIA’s T1000 8 GB and P104-100 target completely different worlds, yet their benchmark data places them surprisingly close in overall standing. The T1000 is a professional Turing-based card aimed at workstations, while the P104-100 is a Pascal-based mining-oriented board with no display outputs. The data shows a clear split: the P104-100 dominates in raw compute and memory bandwidth, but the T1000 counters with efficiency, modern features, and a far more usable physical design.
Head-to-Head Benchmarks
The only directly comparable benchmark in the data is Geekbench Vulkan, and the results are decisive. The NVIDIA P104-100 scores 45,165, while the NVIDIA T1000 8 GB manages 34,561. That is a 23.5% deficit for the T1000, making the P104-100 the clear winner in this specific test. The gap is substantial and reflects the fundamental architectural differences between the two cards.
Looking at the broader picture, the P104-100’s average benchmark score across all tests is 32,982, while the T1000’s average is 34,561. This is an interesting inversion: the T1000 actually has a higher average score despite losing the Vulkan head-to-head. The reason is the P104-100’s other benchmark results. In 3DMark Steel Nomad DX12, the P104-100 scores 1,413, and in Geekbench OpenCL it scores 52,368. These additional data points pull the P104-100’s average down, likely because the Steel Nomad test is particularly demanding on its older architecture.
The T1000’s single benchmark result places it in the 79th percentile of all GPUs, while the P104-100 sits at the 77th percentile. That 2-percentile difference is small, but it tells a story: the T1000 is marginally better positioned relative to the entire GPU landscape. The nearest rivals for the T1000 confirm this. The AMD Radeon HD 7970 scores 34,541 (0.1% behind), the NVIDIA A2 scores 34,690 (0.4% ahead), the NVIDIA TITAN V scores 34,355 (0.6% behind), and the NVIDIA RTX A1000 scores 34,207 (1.0% behind). The T1000 is essentially in a dead heat with all of these cards, trading blows within a 1% margin.
For the P104-100, the nearest rivals are all mobile or older workstation parts. The NVIDIA T600 Mobile scores 32,849 (0.4% behind), the NVIDIA T550 Mobile scores 33,161 (0.5% ahead), the NVIDIA GeForce RTX 3050 Mobile scores 33,170 (0.6% ahead), and the AMD Radeon Pro 570 scores 33,207 (0.7% ahead). The P104-100’s average is thus bracketed by laptop-class GPUs, which is a telling indication of its niche positioning.
Architecture Differences
The T1000 is built on the TU117 chip using Turing architecture, fabricated on a 12 nm process at TSMC. It packs 4,700 million transistors into a 200 mm² die, yielding a transistor density of 23.5 million per square millimeter. The P104-100, in contrast, uses the GP104 chip with Pascal architecture, produced on a 16 nm process, also at TSMC. It contains 7,200 million transistors on a larger 314 mm² die, with a slightly lower density of 22.9 million per square millimeter. The P104-100 is the physically larger and more complex chip, but the T1000’s newer process node allows it to achieve higher density.
The compute capabilities diverge sharply. The T1000 has 896 shading units, 56 texture mapping units, and 32 raster output pipelines. The P104-100 more than doubles the shading units to 1,920, with 120 TMUs and 64 ROPs. This explains the raw performance gap: the P104-100 delivers 6.655 TFLOPS of FP32 compute versus the T1000’s 2.500 TFLOPS. The P104-100 also leads in pixel rate at 110.9 GPixel/s versus 44.64 GPixel/s, and in texture rate at 208.0 GTexel/s versus 78.12 GTexel/s. Every throughput metric favors the P104-100 by a wide margin.
Memory is another major divergence. The T1000 uses 8 GB of GDDR6 on a 128-bit bus, providing 160.0 GB/s of bandwidth. The P104-100 uses 4 GB of GDDR5X on a 256-bit bus, delivering 320.3 GB/s. Even though the P104-100 has half the capacity, its bandwidth is exactly double. This makes it far better suited for memory-intensive compute workloads, while the T1000’s larger pool is better for holding large datasets.
FP16 performance is a curiosity. The T1000 hits 5.000 TFLOPS with a 2:1 ratio, meaning it can double its FP32 throughput. The P104-100 is severely hamstrung here, managing only 104.0 GFLOPS with a 1:64 ratio. For any workload relying on half-precision math, the T1000 is the only viable option between the two.
The T1000 draws 50 W and requires no power connectors, with a suggested PSU of 250 W. The P104-100’s TDP is not listed, but it demands a single 8-pin connector and a 200 W suggested PSU. The T1000 is single-slot and short at 156 mm, while the P104-100 is dual-slot and long at 267 mm. Neither card has RT or tensor cores.
FAQ
Q: Which card has better raw compute performance?
A: The P104-100 is vastly superior, with 6.655 TFLOPS FP32 versus the T1000’s 2.500 TFLOPS. It also leads in pixel rate (110.9 GPixel/s vs 44.64 GPixel/s) and texture rate (208.0 GTexel/s vs 78.12 GTexel/s).
Q: How do the memory subsystems compare?
A: The P104-100 offers double the bandwidth at 320.3 GB/s over a 256-bit bus with 4 GB GDDR5X, while the T1000 has 8 GB GDDR6 on a 128-bit bus at 160.0 GB/s. The P104-100 wins on speed, the T1000 on capacity.
Q: Which card is more power-efficient?
A: The T1000 has a listed TDP of 50 W and requires no power connector, while the P104-100 needs a single 8-pin connector. The P104-100’s TDP is not listed, but its power requirements are clearly higher.
Q: Can either card be used for display output?
A: The T1000 has 4x mini-DisplayPort 1.4a outputs, while the P104-100 has no outputs whatsoever. The P104-100 is strictly a compute or mining card.
Q: Which card is better for FP16 workloads?
A: The T1000 is the clear choice, delivering 5.000 TFLOPS FP16 with a 2:1 ratio. The P104-100 manages only 104.0 GFLOPS with a 1:64 ratio, making it effectively unusable for half-precision tasks.
Q: How do their overall benchmark percentiles compare?
A: The T1000 sits at the 79th percentile of all GPUs, while the P104-100 is at the 77th percentile. The T1000’s average benchmark score of 34,561 is also higher than the P104-100’s 32,982.
Specification Differences
The two cards differ in nearly every measurable specification. The T1000 uses the TU117 chip on a 12 nm process, while the P104-100 uses GP104 on 16 nm. Transistor counts are 4,700 million versus 7,200 million, and die sizes are 200 mm² versus 314 mm². The T1000 has 896 shading units, 56 TMUs, and 32 ROPs; the P104-100 has 1,920 shading units, 120 TMUs, and 64 ROPs.
Clock speeds favor the P104-100, with a base of 1607 MHz and boost of 1733 MHz, versus the T1000’s 1065 MHz base and 1395 MHz boost. Memory clocks are nearly identical at 1250 MHz and 1251 MHz, both effective at 10 Gbps, but the bus widths differ: 128-bit for the T1000, 256-bit for the P104-100. Memory size is 8 GB GDDR6 versus 4 GB GDDR5X, and bandwidth is 160.0 GB/s versus 320.3 GB/s.
Pixel rate is 44.64 GPixel/s for the T1000 and 110.9 GPixel/s for the P104-100. Texture rates are 78.12 GTexel/s and 208.0 GTexel/s, respectively. FP32 performance is 2.500 TFLOPS versus 6.655 TFLOPS, and FP16 is 5.000 TFLOPS versus 104.0 GFLOPS.
The T1000 is single-slot, 156 mm long, 69 mm high, with no power connectors and a 250 W suggested PSU. The P104-100 is dual-slot, 267 mm long, with one 8-pin connector and a 200 W suggested PSU. The T1000 has 4x mini-DisplayPort 1.4a outputs; the P104-100 has none. Both support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The T1000 uses PCIe 3.0 x16, while the P104-100 is limited to PCIe 1.0 x4.
The Verdict
The data makes the choice straightforward for most users. The P104-100 wins decisively in every raw performance metric: FP32 compute, pixel rate, texture rate, and memory bandwidth. Its Vulkan score is 23.5% higher, and its OpenCL score of 52,368 dwarfs anything the T1000 can offer. If the workload is pure compute and does not require display output, the P104-100 is the stronger card by a significant margin.
However, the T1000 is not without merit. It has twice the memory capacity (8 GB vs 4 GB), a far more efficient FP16 implementation, a lower power draw, and proper display outputs. It also has a slightly higher average benchmark score (34,561 vs 32,982) and a better percentile ranking (79th vs 77th). For a workstation that needs to drive monitors, handle half-precision math, or fit in a small chassis, the T1000 is the sensible pick.
The P104-100’s lack of display outputs is a hard limitation. It cannot serve as a primary GPU in a workstation. Its PCIe 1.0 x4 interface is also a bottleneck that could hurt performance in real-world scenarios, despite its superior raw specs. The T1000’s modern PCIe 3.0 x16 interface and professional feature set make it a more versatile piece of hardware.
Where Each One Wins
NVIDIA P104-100: The P104-100 is the choice for headless compute tasks. It wins the Vulkan benchmark by 23.5%, and its OpenCL score of 52,368 indicates strong general-purpose compute capability. Its 320.3 GB/s of memory bandwidth and 6.655 TFLOPS FP32 make it ideal for tasks like mining or number-crunching that do not require a display. The 1,920 shading units and 64 ROPs provide massive throughput for parallel workloads. Its average score of 32,982 places it in the 77th percentile, and it trades closely with mobile RTX 3050-class parts.
NVIDIA T1000 8 GB: The T1000 wins on versatility and efficiency. Its 8 GB memory capacity is double that of the P104-100, allowing it to hold larger models or datasets. Its FP16 performance of 5.000 TFLOPS is orders of magnitude better than the P104-100’s 104.0 GFLOPS, making it the only option for half-precision workloads. The 50 W TDP and lack of power connectors mean it can run in low-power systems. Its 4x mini-DisplayPort outputs make it a functional workstation card. It also has a higher average benchmark score (34,561) and percentile (79th) than the P104-100, despite losing the Vulkan test.