Intel Iris Xe MAX Graphics vs NVIDIA Tesla M4 Comparison

Intel
GPU

Intel Iris Xe MAX Graphics

CORE STATE DG1
VRAM 4 GB
CLOCK SPEED 1650 MHz
TDP 25 W
BUS WIDTH 128 bit
ARCHITECTURE Generation 12.1
nm
PROCESS 10 nm
LAUNCH DATE 2020
VS
NVIDIA
GEFORCE

Tesla M4

CORE STATE GM206
VRAM 4 GB
CLOCK SPEED 1072 MHz
TDP 50 W
BUS WIDTH 128 bit
ARCHITECTURE Maxwell 2.0
nm
PROCESS 28 nm
LAUNCH DATE 2015

PERFORMANCE BENCHMARKS

geekbench_opencl
14,315
16,932

Analysis: Intel Iris Xe MAX Graphics vs NVIDIA Tesla M4

Head-to-Head Benchmarks

The recorded data shows a single head-to-head benchmark result, and it is decisive in one direction. In the Geekbench OpenCL test, the NVIDIA Tesla M4 scores 16932, while the Intel Iris Xe MAX Graphics scores 14315. The Tesla M4 wins this comparison with a delta of 18.3% over the Intel part. That is not a marginal edge; it is a substantial performance gap that places the two GPUs in different tiers for compute workloads.

The Tesla M4's score of 16932 places it at the 60th percentile among all GPUs in the database. Its nearest rivals are tightly clustered around that figure: the AMD Radeon HD 7970M averages 17019 (0.5% higher), the NVIDIA GeForce GTX 690 averages 17037 (0.6% higher), the NVIDIA T400 4 GB averages 16792 (0.8% lower), and the AMD Radeon RX 7600 XT averages 17083 (0.9% higher). This clustering indicates that the Tesla M4 sits in a well-populated performance band, with no single rival dominating it by more than a fraction of a percent. The Intel Iris Xe MAX Graphics, by contrast, scores 14315, which places it at the 56th percentile. Its own nearest rivals are similarly close: the AMD Radeon Vega 11 averages 14352 (0.3% higher), the NVIDIA GeForce GTX 1070 Ti averages 14277 (0.3% lower), the NVIDIA GeForce GTX TITAN averages 14373 (0.4% higher), and the AMD Radeon RX Vega 11 averages 14385 (0.5% higher).

The practical interpretation of these numbers is straightforward: the Tesla M4 outperforms the Iris Xe MAX by a meaningful margin in raw OpenCL compute, and it does so while occupying a similar performance neighborhood as much more recent discrete GPUs like the RX 7600 XT. The Iris Xe MAX, meanwhile, lands in a band populated mostly by integrated-class and older discrete parts. In the head-to-head comparison, the data shows no scenario where the Intel part pulls ahead; it trails in the only recorded test.

Architecture Differences

The two GPUs come from different architectural lineages, and those differences explain the performance gap. The NVIDIA Tesla M4 uses the GM206 chip, built on the Maxwell 2.0 architecture, which is a mature design from the Tesla Maxwell generation. It is fabricated on a 28 nm process at TSMC, with 2,940 million transistors packed into a 228 mm² die, yielding a transistor density of 12.9M per mm². The Intel Iris Xe MAX Graphics uses the DG1 chip, based on the Generation 12.1 architecture, which is part of the Xe Graphics family. Intel fabricates this part on a 10 nm process, and the die measures 95 mm². No transistor count is recorded for the Intel chip, so a direct density comparison is not possible from the database.

The compute resources differ substantially. The Tesla M4 has 1024 shading units, 64 texture mapping units, and 32 render output units. The Iris Xe MAX has fewer of each: 768 shading units, 48 TMUs, and 24 ROPs. Despite having fewer resources, the Intel part achieves a higher peak pixel rate (39.60 GPixel/s versus 34.30 GPixel/s) and a higher texture rate (79.20 GTexel/s versus 68.61 GTexel/s). This is a consequence of the much higher boost clock on the Intel part: 1650 MHz versus 1072 MHz for the Tesla M4. The base clocks are also widely different, with the Tesla M4 running at 872 MHz and the Intel part at 300 MHz, but the boost behavior favors Intel substantially.

In raw FP32 throughput, the Intel Iris Xe MAX actually leads: 2.534 TFLOPS versus 2.195 TFLOPS for the Tesla M4. The Intel part also records FP16 performance of 5.069 TFLOPS (2:1 ratio), while the Tesla M4 has no recorded FP16 figure. This is notable because the benchmark result favors the Tesla M4 despite the Intel part's higher theoretical throughput. The explanation likely lies in memory bandwidth and driver efficiency, both of which are discussed below.

Memory systems also diverge. The Tesla M4 uses 4 GB of GDDR5 on a 128-bit bus, with a memory clock of 1375 MHz (5.5 Gbps effective) and a total bandwidth of 88.00 GB/s. The Intel Iris Xe MAX uses 4 GB of LPDDR4X on the same 128-bit bus, with a memory clock of 2133 MHz (4.3 Gbps effective) and a bandwidth of 68.26 GB/s. The Tesla M4 has a 29% bandwidth advantage, which is significant for compute workloads that are memory-bound. The Intel part compensates with a faster boost clock and higher pixel/texture rates, but the bandwidth deficit likely limits its OpenCL performance.

The bus interface also differs. The Tesla M4 uses PCIe 3.0 x16, which is a full-width connection from that era. The Intel Iris Xe MAX uses PCIe 4.0 x8, which provides similar total bandwidth but with a narrower lane count. Both parts have no display outputs, meaning they are intended for compute or render-farm use rather than direct display driving. The power envelope is also different: the Tesla M4 is rated at 50 W TDP, while the Intel part is rated at 25 W. The suggested PSU is 250 W for the Tesla M4 and 200 W for the Intel part.

FAQ

Q: Which GPU has a higher OpenCL benchmark score?

A: The NVIDIA Tesla M4 scores 16932 in Geekbench OpenCL, while the Intel Iris Xe MAX Graphics scores 14315. The Tesla M4 is 18.3% faster in this test.

Q: Does the Intel Iris Xe MAX have higher raw FP32 compute?

A: Yes. The Intel part records 2.534 TFLOPS FP32, while the Tesla M4 records 2.195 TFLOPS. Despite this, the Tesla M4 wins the OpenCL benchmark, indicating that the FP32 ceiling is not the limiting factor.

Q: How do the memory bandwidth figures compare?

A: The Tesla M4 offers 88.00 GB/s from GDDR5, whereas the Intel Iris Xe MAX offers 68.26 GB/s from LPDDR4X. The Tesla M4 has a 29% higher memory bandwidth, which likely contributes to its better benchmark result.

Q: What are the process node differences?

A: The Tesla M4 is built on a 28 nm process at TSMC, with a die size of 228 mm² and 2,940 million transistors. The Intel Iris Xe MAX is built on a 10 nm process at Intel, with a die size of 95 mm² and no recorded transistor count.

Q: Which GPU has more shading units?

A: The Tesla M4 has 1024 shading units, while the Intel Iris Xe MAX has 768. The Tesla M4 also has more TMUs (64 versus 48) and more ROPs (32 versus 24).

Q: Are both GPUs end-of-life?

A: Yes. Both are marked as end-of-life in the database. The Tesla M4 was released on 2015-11-09, and the Intel Iris Xe MAX was released on 2020-10-30.

Specification Differences

The following specifications differ between the two parts:

  • Chip: GM206 (NVIDIA) versus DG1 (Intel)
  • Architecture: Maxwell 2.0 versus Generation 12.1
  • Generation: Tesla Maxwell (Mxx) versus Xe Graphics
  • Process Node: 28 nm versus 10 nm
  • Foundry: TSMC versus Intel
  • Transistors: 2,940 million versus not recorded
  • Die Size: 228 mm² versus 95 mm²
  • Transistor Density: 12.9M / mm² versus not recorded
  • Base Clock: 872 MHz versus 300 MHz
  • Boost Clock: 1072 MHz versus 1650 MHz
  • Memory Clock: 1375 MHz (5.5 Gbps effective) versus 2133 MHz (4.3 Gbps effective)
  • Memory Type: GDDR5 versus LPDDR4X
  • Memory Bandwidth: 88.00 GB/s versus 68.26 GB/s
  • Shading Units: 1024 versus 768
  • TMUs: 64 versus 48
  • ROPs: 32 versus 24
  • Pixel Rate: 34.30 GPixel/s versus 39.60 GPixel/s
  • Texture Rate: 68.61 GTexel/s versus 79.20 GTexel/s
  • FP32: 2.195 TFLOPS versus 2.534 TFLOPS
  • FP16: Not recorded versus 5.069 TFLOPS (2:1)
  • TDP: 50 W versus 25 W
  • Slot Width: Single-slot versus IGP
  • Suggested PSU: 250 W versus 200 W
  • Bus Interface: PCIe 3.0 x16 versus PCIe 4.0 x8
  • Release Date: 2015-11-09 versus 2020-10-30
  • Predecessor: Tesla Kepler versus Graphics
  • Successor: Tesla Pascal versus Alchemist
  • Percentile vs All GPUs: 60 versus 56
  • Average Benchmark Score: 16932 versus 14315

The Verdict

The benchmark data is unambiguous: the NVIDIA Tesla M4 is the faster GPU in OpenCL compute workloads. It wins the only recorded head-to-head test with an 18.3% margin, which is a substantial difference in performance. Its score of 16932 places it at the 60th percentile overall, meaning it outperforms the majority of GPUs in the database. The Intel Iris Xe MAX Graphics, with a score of 14315, sits at the 56th percentile, which is lower both in absolute terms and in relative standing.

The architecture comparison explains why. The Tesla M4 has more shading units, more TMUs, more ROPs, and significantly higher memory bandwidth. It sacrifices peak clock speed and raw FP32 throughput, but those metrics do not translate into better OpenCL performance in the recorded test. The Intel part is more efficient in terms of power, with a 25 W TDP versus 50 W, and it has a higher pixel rate and texture rate, but the benchmark result favors the NVIDIA part.

For users choosing between these two end-of-life parts, the decision should be driven by the workload. If the task is OpenCL compute, the Tesla M4 is the clear choice based on the data. It also has the advantage of a full PCIe 3.0 x16 interface and higher memory bandwidth, which are beneficial for data transfer and memory-heavy operations. The Intel Iris Xe MAX Graphics, despite its newer process node and higher theoretical FP32, does not deliver a better benchmark result. Its advantages are limited to power consumption and peak pixel/texture rates, which are less relevant for the compute scenarios implied by these GPUs having no display outputs.

The database shows one winner in the head-to-head, and that winner is the NVIDIA Tesla M4. For any application that relies on OpenCL performance, the recorded data supports choosing the Tesla M4 over the Intel Iris Xe MAX Graphics.

DETAILED SPECIFICATIONS

SPECIFICATION
Iris Xe MAX Graphics
Tesla M4
Core Specs
Shading Units
768
1,024 +33.3%
Shaders
768
1,024 +33.3%
TMUs
48
64 +33.3%
ROPs
24
32 +33.3%
Execution Units
96
Clocks
Base Clock
300 MHz
872 MHz
Boost Clock
1650 MHz
1072 MHz
Memory Clock
2133 MHz 4.3 Gbps effective
1375 MHz 5.5 Gbps effective
Memory
Memory Size
4 GB
4 GB
VRAM (MB)
4,096
4,096 0.0%
Memory Type
LPDDR4X
GDDR5
Memory Bus
128 bit
128 bit
Bandwidth
68.26 GB/s
88.00 GB/s
Cache
L1 Cache
48 KB (per SMM)
L2 Cache
1024 KB
1024 KB
L3 Cache
16 MB
Performance
Pixel Rate
39.60 GPixel/s
34.30 GPixel/s
Texture Rate
79.20 GTexel/s
68.61 GTexel/s
FP32 (TFLOPS)
2.534 TFLOPS
2.195 TFLOPS
FP64 (TFLOPS)
633.6 GFLOPS (1:4)
68.61 GFLOPS (1:32)
FP16 (TFLOPS)
5.069 TFLOPS (2:1)
Power
TDP
25 W
50 W
TDP (W)
25
50 +100.0%
Suggested PSU
200 W
250 W
Architecture
Architecture
Generation 12.1
Maxwell 2.0
GPU Name
DG1
GM206
Generation
Xe Graphics
Tesla Maxwell (Mxx)
Process Size
10 nm
28 nm
Transistors
2,940 million
Die Size
95 mm²
228 mm²
Foundry
Intel
TSMC
Density
12.9M / mm²
API Support
DirectX
12 (12_1)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
5.2
Shader Model
6.6
6.8
Physical
Slot Width
IGP
Single-slot
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x8
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Graphics
Tesla Kepler
Successor
Alchemist
Tesla Pascal
View Iris Xe MAX Graphics Details View Tesla M4 Details