GPU Systems · Internship

GPU Systems Intern at Evore Labs.

Help extract maximum useful computation from every accelerator we run.

Stipend
£1,300 / month
Duration
2 months
Work
Remote
Location
Remote-first · UK
The Role

What you will own.

You will help extract maximum useful computation from every accelerator we run. That means profiling training and inference, understanding where time and memory actually go, and closing the gap between theoretical and achieved performance. This is a low-level, measurement-driven role for someone who enjoys reading a profiler trace as much as writing code.

Responsibilities

  • Profile training and inference workloads to find real bottlenecks.
  • Optimise kernels, memory movement and data pipelines.
  • Study GPU scheduling, occupancy and multi-GPU communication.
  • Benchmark rigorously and report reproducible numbers.
  • Collaborate with ML and compiler teams on performance.

Qualifications

  • Exposure to CUDA, GPU programming or accelerator internals.
  • Understanding of parallelism, memory hierarchies and roofline thinking.
  • Strong C++ and/or Python.
  • Measurement-driven; skeptical of unprofiled claims.
  • Patience for low-level debugging.

Nice to have

  • Triton, CUTLASS, or compiler / MLIR experience.
  • Distributed or mixed-precision training.
  • Familiarity with NCCL or interconnect tuning.
Learning

What this track teaches.

You will learn

  • Where performance is really won and lost on GPUs.
  • Profiling and kernel-level optimisation.
  • Multi-GPU communication and scaling.
  • Turning hardware into usable research throughput.
Tools

The stack around the work.

CUDA
C++
Triton
Nsight
PyTorch
NCCL

Apply

Send the work that proves the fit.

Tell us what you have built, tested or discovered. A researcher or engineer reads every application.