Reinforcement Learning · Internship

Reinforcement Learning Intern at Evore Labs.

Explore sequential decision-making under uncertainty — sizing, timing and execution as learned policy.

Stipend
£1,300 / month
Duration
2 months
Work
Remote
Location
Remote-first · UK
The Role

What you will own.

You will explore decision-making under uncertainty, where the question is not 'what will happen?' but 'what should we do?'. You will help design environments, reward structures and policies for sequential decisions like sizing, timing and execution — then study how those policies behave when the world stops being stationary. Expect a mix of theory, careful simulation, and a healthy paranoia about the gap between backtest and reality.

Responsibilities

  • Build simulation environments and reward functions for market decisions.
  • Implement and evaluate policy-optimisation and offline-RL methods.
  • Study robustness to non-stationarity, regime shift and distribution drift.
  • Design safe bridges between simulation and live behaviour.
  • Analyse policies for stability, exploitability and failure modes.

Qualifications

  • Working knowledge of RL fundamentals — MDPs, policy gradients, value methods.
  • Ability to implement algorithms from papers in PyTorch or JAX.
  • Comfort with probability, optimisation and simulation.
  • Healthy skepticism about sim-to-real transfer.
  • Strong scientific communication.

Nice to have

  • Offline / batch RL or contextual-bandit experience.
  • Control theory or operations research background.
  • Publications or strong course projects at the frontier.
Learning

What this track teaches.

You will learn

  • How sequential decisions differ from prediction.
  • Designing rewards that do not get gamed.
  • Evaluating policies you cannot fully trust.
  • Safety patterns for autonomy that touches capital.
Tools

The stack around the work.

Python
PyTorch
JAX
Gymnasium
Ray RLlib
NumPy

Apply

Send the work that proves the fit.

Tell us what you have built, tested or discovered. A researcher or engineer reads every application.