Skip to main content

Reinforcement Learning Engineer

AI & Machine Learning Full Time Mid Level London, England, United Kingdom
£126,000–£210,000/year (£10,500–£17,500/month) — Negotiable

Job Description

Work on reinforcement-learning and preference-optimisation workflows for model behaviour improvement. The role focuses on reward modelling, PPO or DPO pipelines, diagnostics, and reproducible training runs.

Requirements

- Practical experience with reinforcement learning, preference optimisation, or RLHF-style workflows
- Strong Python and deep-learning framework skills
- Familiarity with reward modelling, PPO, DPO, training diagnostics, or evaluation design
- Ability to reason about data quality, policy drift, and training stability

Responsibilities

- Build and maintain RL or preference-optimisation training workflows
- Monitor training diagnostics and investigate instability or regressions
- Work with research and infrastructure teams on repeatable experiments
- Document model-behaviour changes and release implications

Benefits

- Work on applied model-improvement workflows with infrastructure support
- Collaborate with research, safety, and ML platform teams
- Help set standards for preference-data quality and training reliability

Job Overview

Employment Type Full Time
Experience Level Mid Level
Location London, England, United Kingdom
Vacancies 2

Apply Now