GEFORCE

NVIDIA Tesla K40m

NVIDIA graphics card specifications and benchmark scores

12 GB
VRAM
876
MHz Boost
245W
TDP
384
Bus Width

At a Glance

NVIDIA
VRAM 12 GB
Boost Clock 876 MHz
Shaders 2,880
Bus Width 384-bit
TDP 245W
Memory Type GDDR5
Architecture Kepler
nm
Process 28 nm
Released Nov 2013

NVIDIA Tesla K40m Specifications

GPU Core

Shader units and compute resources

The NVIDIA Tesla K40m GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
2,880
Shaders
2,880
TMUs
240
ROPs
48

Tesla K40m Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Tesla K40m's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla K40m by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
745 MHz
Base Clock
745 MHz
Boost Clock
876 MHz
Boost Clock
876 MHz
Memory Clock
1502 MHz 6 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's Tesla K40m Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla K40m's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
12 GB
VRAM
12,288 MB
Memory Type
GDDR5
VRAM Type
GDDR5
Memory Bus
384 bit
Bus Width
384-bit
Bandwidth
288.4 GB/s

Tesla K40m by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Tesla K40m, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
16 KB (per SMX)
L2 Cache
1536 KB

Tesla K40m Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla K40m against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
5.046 TFLOPS
FP64 (Double)
1.682 TFLOPS (1:3)
Pixel Rate
52.56 GPixel/s
Texture Rate
210.2 GTexel/s

Kepler Architecture & Process

Manufacturing and design details

The NVIDIA Tesla K40m is built on NVIDIA's Kepler architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla K40m will perform in GPU benchmarks compared to previous generations.

Architecture
Kepler
GPU Name
GK110B
Process Node
28 nm
Foundry
TSMC
Transistors
7,080 million
Die Size
561 mm²
Density
12.6M / mm²

Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Tesla K40m determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla K40m to maintain boost clocks without throttling.

TDP
245 W
TDP
245W
Suggested PSU
550 W

Tesla K40m by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Tesla K40m are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Dual-slot
Length
267 mm 10.5 inches
Bus Interface
PCIe 3.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Tesla K40m. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 (11_1)
DirectX
12 (11_1)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.2.175
Vulkan
1.2.175
OpenCL
3.0
CUDA
3.5
Shader Model
6.5 (5.1)

Tesla K40m Product Information

Release and pricing details

The NVIDIA Tesla K40m is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla K40m by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Nov 2013
Launch Price
7,699 USD
Production
End-of-life
Predecessor
Tesla Fermi
Successor
Tesla Maxwell

About NVIDIA Tesla K40m

The NVIDIA Tesla K40m is a Kepler-architecture compute card from late 2013, built for professional workloads rather than consumer gaming. It pairs 12 GB of GDDR5 memory with 2,880 shading units, and its benchmark data places it in a surprisingly competitive position against both older and newer hardware in raw compute throughput.

Benchmark Performance

The Tesla K40m’s sole recorded benchmark score is 19,519 points in Geekbench OpenCL. This places it at the 62nd percentile among all GPUs, meaning it outperforms roughly 62% of the tested database. The score is a measure of raw compute capability, not gaming frame rates, and it reflects the card’s design for parallel processing tasks.

The data shows a tight cluster of rivals around this score. The K40m sits just 0.5% above the AMD Radeon Pro 560X, which scores 19,426. That is a virtual tie — a margin of only 93 points. Against the NVIDIA Quadro K5200, the K40m is 0.5% slower, with the K5200 scoring 19,623. Again, this is within noise for a single benchmark run.

More interesting is the comparison to consumer gaming cards. The GeForce GTX 780 scores 19,405, placing the K40m 0.6% above it. The GeForce GTX 1080 Ti, a much newer and more powerful gaming card, scores 19,402 — also 0.6% below the K40m. The practical takeaway: in pure OpenCL compute, the aging Tesla K40m holds its own against hardware that is several generations newer. The 1080 Ti’s gaming dominance does not translate to a lead in this specific compute test.

The delta percentages are all under 1%, which means the K40m is effectively tied with all four rivals in this workload. The 62nd percentile ranking reinforces that picture: this is a mid-tier performer in the current database, not a top-tier card, but also not obsolete. The K40m’s 5.046 TFLOPS of FP32 compute is the architectural driver behind this score, and it explains why the card remains relevant in compute benchmarks despite its age.

How It Compares

AMD Radeon Pro 560X: The K40m edges out this mobile workstation card by 0.5%. Both are professional-oriented GPUs, but the Radeon is a much lower-power part. The K40m’s 12 GB VRAM and 384-bit bus give it a memory bandwidth advantage (288.4 GB/s vs. the Radeon’s unspecified figure), which matters for large datasets. In compute, they are functionally equivalent.

NVIDIA Quadro K5200: The K5200 leads by 0.5%, a margin of 104 points. This is the closest comparison in the pack. Both cards target professional visualization and compute, and the K5200’s slight edge likely comes from its higher boost clock architecture. The K40m counters with twice the VRAM (12 GB vs. the K5200’s unspecified amount), making it the better choice for memory-bound workloads.

NVIDIA GeForce GTX 780: The K40m is 0.6% faster than this consumer Kepler card. They share the same GK110 architecture family, but the K40m has more VRAM (12 GB vs. 3 GB) and a wider memory bus. The GTX 780 was designed for gaming; the K40m for compute. Their near-identical OpenCL scores show that the underlying silicon is similar, but the K40m’s memory subsystem gives it a durability edge for professional tasks.

NVIDIA GeForce GTX 1080 Ti: This is the most surprising result. The 1080 Ti, a 2017 flagship, scores 0.6% below the 2013 K40m. The 1080 Ti has far higher gaming performance, but in this specific OpenCL test, the K40m’s compute-oriented design (higher FP32 throughput relative to its shading units) keeps it competitive. The K40m’s 288.4 GB/s bandwidth and 5.046 TFLOPS are sufficient to match the newer card’s raw compute output.

Who Should Consider It

The K40m is not a gaming card — it has no display outputs, so you cannot connect a monitor directly. Its benchmark score of 19,519 in OpenCL indicates it is suited for compute workloads like scientific simulation, data processing, or machine learning inference where raw FP32 throughput matters. The 62nd percentile ranking means it will handle moderate compute tasks without issue, but it will lag behind modern high-end accelerators.

For users working with large datasets, the 12 GB VRAM is the standout feature. At 4K resolution or with high-resolution textures, memory capacity often becomes the bottleneck before compute power. The K40m’s 12 GB and 384-bit bus providing 288.4 GB/s bandwidth can hold substantial working sets. This makes it viable for tasks like rendering, video encoding, or medical imaging that need large memory pools.

The card’s 245 W TDP and dual-slot design mean it requires a proper workstation chassis with adequate cooling. The suggested PSU of 550 W is modest by modern standards. If your workload fits within the K40m’s 5.046 TFLOPS of FP32 compute, the benchmark data shows it performs on par with much newer cards. However, the end-of-life production status and lack of modern features like ray tracing or tensor cores mean it is not a future-proof investment.

FAQ

Q: How does the Tesla K40m perform in OpenCL benchmarks?

A: It scores 19,519 in Geekbench OpenCL, placing it at the 62nd percentile of all GPUs in the database.

Q: Is the K40m faster than a GTX 1080 Ti in compute?

A: In this specific OpenCL test, yes — the K40m scores 0.6% higher than the GTX 1080 Ti (19,519 vs. 19,402).

Q: Can I use the K40m for gaming?

A: No. The card has no display outputs, so it cannot connect to a monitor. It is designed for compute workloads only.

Q: How much VRAM does the K40m have?

A: It has 12 GB of GDDR5 memory on a 384-bit bus, providing 288.4 GB/s of bandwidth.

Q: What is the difference between the K40m and Quadro K5200?

A: The K5200 scores 0.5% higher in OpenCL (19,623 vs. 19,519). The K40m has 12 GB VRAM, while the K5200 has an unspecified amount, but the K40m’s memory configuration is its primary advantage.

Q: Is the K40m still worth using?

A: For compute tasks that fit within its 5.046 TFLOPS of FP32 performance and 12 GB memory, the benchmark data shows it performs on par with newer cards like the GTX 1080 Ti. However, it is end-of-life and lacks modern features like ray tracing or tensor cores.

Ray Tracing and Feature Set

The K40m has no ray tracing cores and no tensor cores. It is a pure compute card based on the Kepler architecture (GK110B chip), which predates hardware-accelerated ray tracing by several generations. The card’s feature set is focused on raw FP32 compute: 2,880 shading units, 240 texture mapping units, and 48 raster operation units deliver 5.046 TFLOPS of FP32 performance and 210.2 GTexel/s of texture fill rate.

API support is limited by its age. The K40m supports DirectX 12 (11_1), OpenGL 4.6, and Vulkan 1.2.175. The DirectX 12 support is partial (11_1 feature level), which means it cannot take full advantage of modern DirectX 12 features. OpenGL 4.6 and Vulkan 1.2.175 are more current, but the absence of dedicated ray tracing hardware means any ray-traced workload would run on general-purpose shaders, which is inefficient.

The card’s 28 nm process node and 7,080 million transistors on a 561 mm² die reflect its 2013 design era. The transistor density of 12.6 million per mm² is low by modern standards, but the architecture was optimized for compute throughput rather than power efficiency. With a 245 W TDP, it draws significant power for its performance level, but the 550 W suggested PSU requirement is manageable in a workstation.

Memory Subsystem

The K40m’s memory subsystem is its most distinctive feature. It packs 12 GB of GDDR5 memory on a 384-bit bus, yielding 288.4 GB/s of bandwidth. This configuration was exceptional for 2013 — most consumer cards had 3-4 GB — and it remains useful today for datasets that exceed the VRAM of newer cards.

The 12 GB capacity is the key differentiator. For compute workloads like deep learning inference or large-scale rendering, memory capacity often limits what models or scenes can be processed. The 384-bit bus provides high bandwidth for feeding the 2,880 shading units, and the 288.4 GB/s figure is sufficient to keep the GPU busy in memory-bound tasks. Memory operates at 1502 MHz (6 Gbps effective), which is standard for GDDR5 of that era.

The 48 ROPs deliver a pixel rate of 52.56 GPixel/s, which is modest by modern standards but irrelevant for a compute card with no display outputs. The texture rate of 210.2 GTexel/s aligns with the card’s compute focus. For high-resolution work, the 12 GB buffer is the selling point — it can hold textures, geometry, or data arrays that would cause out-of-memory errors on 8 GB cards. The bandwidth, while not class-leading today, is adequate for the K40m’s compute capability, as evidenced by its competitive OpenCL score against newer hardware.

Detailed benchmark scores and charts for the NVIDIA Tesla K40m are below.

Benchmark Scores

geekbench_openclSource

Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA Tesla K40m handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms.

geekbench_opencl #308 of 650
19,885
5%
Max: 388,405
Compare with other GPUs

Top 5 Performers

#1 NVIDIA RTX 6000D
388,405
#2 NVIDIA B300 SXM6 AC
369,831
#3 NVIDIA B200
345,482
#4 NVIDIA H200 NVL
334,891

Popular NVIDIA Tesla K40m Comparisons

See how the Tesla K40m stacks up against similar graphics cards from the same generation and competing brands.

Compare with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs