NVIDIA GeForce RTX 5090 SE vs NVIDIA H100 CNX Comparison
NVIDIA GeForce RTX 5090 SE
H100 CNX
Analysis: NVIDIA GeForce RTX 5090 SE vs NVIDIA H100 CNX
Where Each One Wins
The recorded data separates these two NVIDIA accelerators by design intent rather than by raw capability alone. The GeForce RTX 5090 SE is a graphics-first product built around Blackwell 2.0 architecture, while the H100 CNX is a server-oriented Hopper part with no display outputs. The benchmark database shows no direct head-to-head measurements for this pair, so the analysis rests on architectural specifications and the recorded performance fields.
The RTX 5090 SE wins in pixel processing and texture throughput. Its pixel rate of 380.3 GPixel/s is dramatically higher than the H100 CNX's 44.28 GPixel/s, an advantage of roughly 8.6 times. Texture rate follows a similar pattern: the 5090 SE delivers 1,045.9 GTexel/s versus 841.3 GTexel/s for the H100 CNX, a 24% lead. These numbers point to the 5090 SE as the choice for rasterization-heavy workloads, display output, and real-time graphics pipelines.
The H100 CNX wins in memory capacity and bandwidth. It carries 80 GB of HBM2e across a 5120-bit bus, delivering 2.04 TB/s of memory bandwidth. The 5090 SE has 24 GB of GDDR7 on a 384-bit bus, yielding 1.34 TB/s. The H100 CNX's bandwidth advantage is 52%, and its capacity advantage is 233%. For large datasets, model weights, or in-memory compute tasks, the H100 CNX holds the edge.
In raw floating-point throughput, the split is nuanced. The 5090 SE posts 66.94 TFLOPS for FP32 and the same 66.94 TFLOPS for FP16 with a 1:1 ratio. The H100 CNX posts 53.84 TFLOPS for FP32, which is 19.6% lower, but its FP16 figure of 215.4 TFLOPS comes from a 4:1 ratio, meaning it is optimized for reduced-precision tensor work. The H100 CNX's FP16 throughput is 3.2 times the 5090 SE's FP16 figure, but that advantage only applies where the 4:1 ratio is acceptable. For FP32 workloads, the 5090 SE is faster.
The shading unit counts are close: the 5090 SE has 14,080 shading units and 440 TMUs, while the H100 CNX has 14,592 shading units and 456 TMUs. The H100 CNX's small lead in those counts does not translate to pixel rate because its ROP count is only 24 versus 160 for the 5090 SE. The tensor core counts are 440 for the 5090 SE and 456 for the H100 CNX, nearly identical, but the architecture differences in how those cores execute FP16 change the practical outcome.
Clock speeds favor the 5090 SE. Its base clock is 1740 MHz and boost is 2377 MHz. The H100 CNX runs a 690 MHz base and 1845 MHz boost. The 5090 SE's boost clock is 28.8% higher. That clock advantage, combined with the higher pixel and texture rates, makes the 5090 SE the stronger part for latency-sensitive graphics work.
The H100 CNX is the only one of the two with no display outputs, which eliminates it from any client-side graphics role. The 5090 SE provides 1x HDMI 2.1b and 3x DisplayPort 2.1b outputs. The API support also differs: the 5090 SE supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the H100 CNX lists no DirectX, OpenGL, or Vulkan support in the database.
The Verdict
The data indicates two distinct purchase rationales. The GeForce RTX 5090 SE is the pick for anyone running graphics applications, real-time rendering, or FP32 compute on a client platform. Its pixel rate of 380.3 GPixel/s, texture rate of 1,045.9 GTexel/s, and FP32 throughput of 66.94 TFLOPS are all higher than the H100 CNX's corresponding figures. It also carries a launch MSRP of 1,499 USD, which is the only pricing information available for either part.
The H100 CNX is the pick for server-side workloads that fit in its 80 GB HBM2e pool and can exploit its 2.04 TB/s bandwidth. Its FP16 throughput of 215.4 TFLOPS is the highest single number in either specification sheet, and the 5120-bit memory bus is the widest in this comparison. The lack of display outputs and graphics APIs makes it unsuitable for client graphics, but that is not its purpose.
A user whose work is dominated by FP32 compute and graphics should choose the 5090 SE. A user whose work is dominated by large-memory FP16 tensor operations should choose the H100 CNX. The two parts do not compete for the same socket or the same workload profile; they are complementary accelerators with one shared manufacturer and one shared process node.
Head-to-Head Benchmarks
The database contains no recorded head-to-head benchmark scores for this pair, so the comparison must be built from the specification-derived performance fields. The largest single advantage for the 5090 SE is in pixel rate. At 380.3 GPixel/s versus 44.28 GPixel/s, the 5090 SE is approximately 8.6 times faster. This is a direct consequence of its 160 ROPs compared to 24 ROPs on the H100 CNX. No amount of clock speed or memory bandwidth can compensate for that ROP deficit in pixel-heavy workloads.
The texture rate comparison is closer but still favors the 5090 SE. It delivers 1,045.9 GTexel/s against 841.3 GTexel/s, a 24.3% advantage. The H100 CNX actually has more TMUs (456 versus 440), but the 5090 SE's higher clock speed and architecture efficiency overcome that small unit count difference.
In FP32 compute, the 5090 SE delivers 66.94 TFLOPS versus 53.84 TFLOPS for the H100 CNX, a 24.3% advantage. This is the same percentage gap as the texture rate, which suggests the clock speed difference is the dominant factor there. The base clock of 1740 MHz on the 5090 SE is 2.5 times the H100 CNX's 690 MHz base, though the boost clocks are closer at 2377 MHz versus 1845 MHz.
The H100 CNX's biggest win is in FP16 throughput. Its 215.4 TFLOPS is 3.2 times the 5090 SE's 66.94 TFLOPS. This comes from the 4:1 FP16 ratio, which means the H100 CNX trades precision for speed. The 5090 SE runs FP16 at a 1:1 ratio, so its FP16 number equals its FP32 number. For applications that can tolerate the reduced precision, the H100 CNX is the faster part.
Memory bandwidth is the second major win for the H100 CNX. Its 2.04 TB/s is 52.2% higher than the 5090 SE's 1.34 TB/s. The H100 CNX also has 80 GB of memory versus 24 GB, a 3.3 times capacity advantage. The memory type differs as well: HBM2e on the H100 CNX against GDDR7 on the 5090 SE. The H100 CNX's bus width of 5120 bits is 13.3 times the 5090 SE's 384-bit bus, which is how it achieves higher bandwidth despite slower effective memory clocks (3.2 Gbps versus 28 Gbps).
Clock speed comparisons favor the 5090 SE across the board. Its boost clock of 2377 MHz is 28.8% higher than the H100 CNX's 1845 MHz. The base clock gap is larger at 152%, but the H100 CNX's low base clock is typical of server parts that operate in power-constrained environments.
FAQ
Q: Which GPU has higher FP32 performance?
A: The RTX 5090 SE delivers 66.94 TFLOPS FP32, which is 24.3% higher than the H100 CNX's 53.84 TFLOPS.
Q: Which GPU has more memory bandwidth?
A: The H100 CNX provides 2.04 TB/s from its HBM2e memory on a 5120-bit bus. The RTX 5090 SE provides 1.34 TB/s from GDDR7 on a 384-bit bus. The H100 CNX's bandwidth is 52.2% higher.
Q: Does the H100 CNX support display outputs?
A: No. The database lists "No outputs" for the H100 CNX. The RTX 5090 SE provides 1x HDMI 2.1b and 3x DisplayPort 2.1b.
Q: What is the FP16 performance difference?
A: The H100 CNX posts 215.4 TFLOPS FP16 using a 4:1 ratio, which is 3.2 times the RTX 5090 SE's 66.94 TFLOPS at a 1:1 ratio. The 5090 SE maintains the same FP16 and FP32 throughput, while the H100 CNX trades precision for speed.
Q: Which GPU has more memory capacity?
A: The H100 CNX has 80 GB of HBM2e, which is 3.3 times the RTX 5090 SE's 24 GB of GDDR7.
Q: Which GPU has higher pixel fill rate?
A: The RTX 5090 SE achieves 380.3 GPixel/s, approximately 8.6 times the H100 CNX's 44.28 GPixel/s. This difference stems from the 5090 SE's 160 ROPs versus 24 ROPs on the H100 CNX.
Architecture Differences
The two GPUs come from different architecture generations within the same manufacturer. The RTX 5090 SE uses the GB202 chip built on Blackwell 2.0 architecture. The H100 CNX uses the GH100 chip built on Hopper architecture. Both are fabricated by TSMC on a 5 nm process node, but the transistor counts differ: the GB202 has 92,200 million transistors on a 750 mm² die, while the GH100 has 80,000 million transistors on a larger 814 mm² die. The transistor density is consequently higher on the 5090 SE at 122.9M per mm² versus 98.3M per mm² for the H100 CNX.
The core configurations differ in several ways. The 5090 SE has 14,080 shading units, 440 TMUs, 160 ROPs, 110 RT cores, and 440 tensor cores. The H100 CNX has 14,592 shading units, 456 TMUs, 24 ROPs, no RT cores listed, and 456 tensor cores. The H100 CNX has no ray tracing hardware at all, which reinforces its server compute positioning. The 5090 SE includes 110 RT cores, making it the only one of the pair with hardware ray tracing support.
The memory subsystems use completely different technologies. The 5090 SE uses GDDR7 with a 384-bit bus and 24 GB capacity. The H100 CNX uses HBM2e with a 5120-bit bus and 80 GB capacity. The effective memory clocks are 28 Gbps for the 5090 SE and 3.2 Gbps for the H100 CNX, but the H100 CNX's much wider bus produces higher total bandwidth. The pixel rate difference (380.3 vs 44.28 GPixel/s) is the most visible architectural outcome, driven by the ROP count disparity.
API support also separates the two. The 5090 SE supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The H100 CNX lists no graphics API support in the database. This is consistent with the H100 CNX having no display outputs and no RT cores; it is not designed for any graphics rendering role.
The production status for both is listed as Active. The release dates differ by nearly three years: the H100 CNX was released in March 2023, while the RTX 5090 SE was released in December 2025. The predecessor and successor relationships also differ: the 5090 SE follows the GeForce 40 series and precedes the GeForce 60 series, while the H100 CNX follows Server Ada and precedes Server Blackwell.
Specification Differences
The two parts differ on nearly every measurable specification. The RTX 5090 SE has a base clock of 1740 MHz and boost of 2377 MHz. The H100 CNX has a base clock of 690 MHz and boost of 1845 MHz. The 5090 SE's boost clock is 28.8% higher. The H100 CNX's base clock is less than half the 5090 SE's, which reflects its server power envelope.
Power consumption differs substantially. The 5090 SE has a TDP of 500 W and requires a 900 W suggested PSU with a single 16-pin connector. The H100 CNX has a TDP of 350 W and requires a 750 W suggested PSU with an 8-pin EPS connector. The 5090 SE draws 42.9% more power at its TDP rating.
Memory specifications are the most divergent. The 5090 SE has 24 GB of GDDR7 on a 384-bit bus with 1.34 TB/s bandwidth. The H100 CNX has 80 GB of HBM2e on a 5120-bit bus with 2.04 TB/s bandwidth. The memory clock is 1750 MHz (28 Gbps effective) for the 5090 SE and 1593 MHz (3.2 Gbps effective) for the H100 CNX.
Shading units are close: 14,080 for the 5090 SE versus 14,592 for the H100 CNX, a 3.6% difference. TMUs are similarly close at 440 versus 456. The ROP count is the major differentiator: 160 versus 24, a 6.7 times difference. Tensor cores are 440 versus 456, nearly identical. RT cores exist only on the 5090 SE at 110 units.
Both cards are dual-slot and share the same dimensions in length (267 mm, 10.5 inches) and height (111 mm, 4.4 inches). The 5090 SE has a listed width of 40 mm (1.6 inches), while the H100 CNX has no width listed. Both use a PCIe 5.0 x16 bus interface.
The 5090 SE has a launch MSRP of 1,499 USD. The H100 CNX has no launch MSRP listed in the database. The 5090 SE provides display outputs (1x HDMI 2.1b, 3x DisplayPort 2.1b), while the H100 CNX has none. The 5090 SE supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4; the H100 CNX supports none of these APIs.
The transistor counts differ by 15.3%: 92,200 million for the 5090 SE versus 80,000 million for the H100 CNX. Die sizes differ in the opposite direction: 750 mm² for the 5090 SE versus 814 mm² for the H100 CNX. The 5090 SE packs more transistors into a smaller die, resulting in a 25% higher transistor density. The process node is the same for both: 5 nm at TSMC.