Reinforcement Learning Engineer
£126,000–£210,000/year (£10,500–£17,500/month) — Negotiable
Job Description
Work on reinforcement-learning and preference-optimisation workflows for model behaviour improvement. The role focuses on reward modelling, PPO or DPO pipelines, diagnostics, and reproducible training runs.
Requirements
- Practical experience with reinforcement learning, preference optimisation, or RLHF-style workflows
- Strong Python and deep-learning framework skills
- Familiarity with reward modelling, PPO, DPO, training diagnostics, or evaluation design
- Ability to reason about data quality, policy drift, and training stability
- Strong Python and deep-learning framework skills
- Familiarity with reward modelling, PPO, DPO, training diagnostics, or evaluation design
- Ability to reason about data quality, policy drift, and training stability
Responsibilities
- Build and maintain RL or preference-optimisation training workflows
- Monitor training diagnostics and investigate instability or regressions
- Work with research and infrastructure teams on repeatable experiments
- Document model-behaviour changes and release implications
- Monitor training diagnostics and investigate instability or regressions
- Work with research and infrastructure teams on repeatable experiments
- Document model-behaviour changes and release implications
Benefits
- Work on applied model-improvement workflows with infrastructure support
- Collaborate with research, safety, and ML platform teams
- Help set standards for preference-data quality and training reliability
- Collaborate with research, safety, and ML platform teams
- Help set standards for preference-data quality and training reliability