Intel Arc Pro B370 vs NVIDIA H20 NVL16 Comparison
Intel Arc Pro B370
H20 NVL16
Analysis: Intel Arc Pro B370 vs NVIDIA H20 NVL16
Intel Arc Pro B370 and NVIDIA H20 NVL16 occupy opposite ends of the hardware spectrum. The Intel part is a 25 W integrated graphics engine built for portable devices, while the NVIDIA part is a 400 W server-class accelerator with 96 GB of HBM3 memory. The recorded data shows no shared benchmark results, so the comparison relies entirely on architectural and specification differences.
The Verdict
The data separates these two products cleanly by use case and physical design. The Intel Arc Pro B370 is an integrated graphics processor (IGP) that shares system memory, draws 25 W, and fits into a portable device slot. The NVIDIA H20 NVL16 is a dedicated SXM module for servers, requires an 800 W suggested power supply, and provides no display outputs. Any user needing a display-capable, low-power graphics solution for mobile hardware would select the Intel part. Any workload requiring massive memory capacity, high compute throughput, and server integration would select the NVIDIA part. There is no overlap in target hardware: the Intel chip is an IGP with a 25 W thermal envelope, and the NVIDIA chip is a 400 W accelerator with a PCIe 5.0 x16 interface. The Intel part holds a 50th percentile ranking against all GPUs in the database, and the NVIDIA part holds the same 50th percentile ranking, but the absence of shared benchmarks means the percentile values reflect their respective categories rather than direct competition.
Architecture Differences
The Intel Arc Pro B370 uses the Xe3-LPG architecture on a 3 nm process fabricated by Intel. It is built on the Panther Lake chip and belongs to the Arc Graphics-WM (Panther Lake) generation. The NVIDIA H20 NVL16 uses the Hopper architecture on a 5 nm process fabricated by TSMC, built on the GH100 chip as part of the Server Hopper (Hxx) generation. The process node difference is direct: Intel uses 3 nm, NVIDIA uses 5 nm.
Transistor counts diverge sharply. The NVIDIA chip contains 80,000 million transistors on an 814 mm² die, with a transistor density of 98.3M per mm². The Intel chip lists its transistor count and die size as unknown in the database. The NVIDIA part includes 312 tensor cores and 312 texture mapping units, while the Intel part lists no tensor cores and 40 TMUs. The Intel part has 10 ray tracing cores; the NVIDIA part does not list a ray tracing core count.
API support also separates the two. The Intel Arc Pro B370 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The NVIDIA H20 NVL16 lists N/A for DirectX, OpenGL, and Vulkan. Display outputs follow the same pattern: the Intel part provides "Portable Device Dependent" outputs, while the NVIDIA part provides no outputs.
Head-to-Head Benchmarks
The database contains no head-to-head benchmark entries for these two products, and neither part has individual benchmark scores recorded. The comparison therefore rests on computed specification rates and raw throughput figures.
In FP32 compute, the NVIDIA H20 NVL16 delivers 39.54 TFLOPS, which is 6.4 times the Intel Arc Pro B370's 6.144 TFLOPS. In FP16, the NVIDIA part reaches 79.07 TFLOPS (2:1), while the Intel part reaches 12.29 TFLOPS (2:1), a ratio of approximately 6.4 to 1 as well. Texture rate tells a similar story: the NVIDIA part processes 617.8 GTexel/s versus 96.00 GTexel/s for the Intel part, a 6.4 times advantage. Pixel rate is nearly identical, 47.52 GPixel/s for NVIDIA versus 48.00 GPixel/s for Intel, meaning the Intel part actually holds a slight edge in pixel throughput.
Clock behavior differs by design. The Intel part runs a 300 MHz base clock and a 2400 MHz boost clock. The NVIDIA part runs a 1830 MHz base clock and a 1980 MHz boost clock. The Intel part has a higher boost clock by 420 MHz, but the NVIDIA part's massive parallel resources more than compensate in aggregate throughput.
Shading units favor NVIDIA decisively: 9984 versus 1280, a 7.8 times difference. Raster operation units are close, 24 for NVIDIA versus 20 for Intel. The NVIDIA part also lists 312 TMUs versus 40 for Intel, a 7.8 times difference that tracks the shading unit ratio.
Specification Differences
The two products differ in nearly every measurable field.
- Process node: Intel 3 nm, NVIDIA 5 nm
- Foundry: Intel, TSMC
- Transistors: Intel unknown, NVIDIA 80,000 million
- Die size: Intel unknown, NVIDIA 814 mm²
- Base clock: Intel 300 MHz, NVIDIA 1830 MHz
- Boost clock: Intel 2400 MHz, NVIDIA 1980 MHz
- Memory size: Intel System Shared, NVIDIA 96 GB
- Memory type: Intel System Shared, NVIDIA HBM3
- Memory bus width: Intel System Shared, NVIDIA 6144 bit
- Memory bandwidth: Intel System Dependent, NVIDIA 4.03 TB/s
- Shading units: Intel 1280, NVIDIA 9984
- TMUs: Intel 40, NVIDIA 312
- ROPs: Intel 20, NVIDIA 24
- RT cores: Intel 10, NVIDIA not listed
- Tensor cores: Intel not listed, NVIDIA 312
- Pixel rate: Intel 48.00 GPixel/s, NVIDIA 47.52 GPixel/s
- Texture rate: Intel 96.00 GTexel/s, NVIDIA 617.8 GTexel/s
- FP32: Intel 6.144 TFLOPS, NVIDIA 39.54 TFLOPS
- FP16: Intel 12.29 TFLOPS (2:1), NVIDIA 79.07 TFLOPS (2:1)
- TDP: Intel 25 W, NVIDIA 400 W
- Slot width: Intel IGP, NVIDIA SXM Module
- Power connectors: Intel None, NVIDIA not listed
- Suggested PSU: Intel not listed, NVIDIA 800 W
- Bus interface: Intel IGP, NVIDIA PCIe 5.0 x16
- Display outputs: Intel Portable Device Dependent, NVIDIA No outputs
- DirectX: Intel 12 Ultimate (12_2), NVIDIA N/A
- OpenGL: Intel 4.6, NVIDIA N/A
- Vulkan: Intel 1.4, NVIDIA N/A
- Release date: Intel 2026-01-26, NVIDIA 2025-09-01
- Predecessor: Intel HD Graphics-WM, NVIDIA Server Ada
- Successor: Intel not listed, NVIDIA Server Blackwell
The memory clock fields also differ. The Intel part lists "System Shared" for memory clock, while the NVIDIA part lists 1313 MHz with 5.3 Gbps effective.
FAQ
Q: Which GPU has higher FP32 compute?
A: The NVIDIA H20 NVL16 delivers 39.54 TFLOPS, which is 6.4 times the Intel Arc Pro B370's 6.144 TFLOPS.
Q: Can the NVIDIA H20 NVL16 output video to displays?
A: No. The NVIDIA part lists "No outputs" for display outputs, and its DirectX, OpenGL, and Vulkan APIs are all listed as N/A.
Q: What is the thermal design power of each part?
A: The Intel Arc Pro B370 has a TDP of 25 W, and the NVIDIA H20 NVL16 has a TDP of 400 W. The NVIDIA part also lists an 800 W suggested power supply.
Q: How much memory does each GPU use?
A: The Intel Arc Pro B370 uses system shared memory with system-dependent bandwidth. The NVIDIA H20 NVL16 has 96 GB of HBM3 memory on a 6144 bit bus with 4.03 TB/s bandwidth.
Q: Which GPU has more shading units?
A: The NVIDIA H20 NVL16 has 9984 shading units, while the Intel Arc Pro B370 has 1280. The NVIDIA part also has 312 tensor cores, which the Intel part does not list.
Q: Which GPU has a higher boost clock?
A: The Intel Arc Pro B370 has a 2400 MHz boost clock, while the NVIDIA H20 NVL16 has a 1980 MHz boost clock.
Q: Do the two GPUs share any API compatibility?
A: No. The Intel Arc Pro B370 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The NVIDIA H20 NVL16 lists N/A for all three APIs.
Where Each One Wins
The Intel Arc Pro B370 wins in pixel rate, 48.00 GPixel/s compared to 47.52 GPixel/s for the NVIDIA H20 NVL16. It also wins on boost clock, 2400 MHz versus 1980 MHz, and on integration simplicity: it requires no power connectors, uses no separate memory, and fits as an IGP. It provides display outputs, supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, and consumes only 25 W. The Intel part is the only option of the two for portable devices, given its IGP slot width and portable-device-dependent display outputs. Its release date of 2026-01-26 makes it the newer product, and its predecessor is HD Graphics-WM.
The NVIDIA H20 NVL16 wins decisively in raw compute and memory capacity. It delivers 39.54 TFLOPS FP32 and 79.07 TFLOPS FP16, versus 6.144 TFLOPS and 12.29 TFLOPS for the Intel part. Texture rate favors NVIDIA at 617.8 GTexel/s versus 96.00 GTexel/s. Memory capacity is 96 GB of HBM3 versus system shared memory, with 4.03 TB/s bandwidth versus system-dependent bandwidth. The NVIDIA part carries 9984 shading units, 312 TMUs, 312 tensor cores, and 80,000 million transistors on an 814 mm² die. It uses a PCIe 5.0 x16 bus interface and an SXM module slot, matching its server role. Its predecessor is Server Ada, and its successor is Server Blackwell. The data shows no scenario where the two products compete for the same socket or workload, so the selection depends entirely on whether the target system is a portable device or a server.