Responsibilities
- Profile training and inference workloads to find real bottlenecks.
- Optimise kernels, memory movement and data pipelines.
- Study GPU scheduling, occupancy and multi-GPU communication.
- Benchmark rigorously and report reproducible numbers.
- Collaborate with ML and compiler teams on performance.