Volver

Research Scientist / Engineer – Reinforcement Learning Infrastructure

CompraTica Empleos

EMP:Technology
London, UK
Tiempo Completo
Remoto
0 vistas

Descripción

You'll build the systems that make reinforcement learning work at frontier scale — coupling policy optimization with large fleets of inference workers, agentic environments, and the reward and verification systems that turn model behavior into learning signal.

RL is how Luma's models go from capable to useful.

RL at scale is a full-loop systems problem: training, rollout generation, environment execution, and reward computation running concurrently across thousands of GPUs, all needing to stay fast, stable, and correct together.

It fits someone who has lived this — post-trained LLMs with RL, built environments and verifiers, and debugged asynchronous rollout pipelines at scale.

If you haven't operated RL at real scale, this will be deep water.

What You'll OwnDesign, build, and scale distributed RL post-training systems, orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs.

Build high-throughput rollout generation, integrating inference engines (vLLM, SGLang), weight synchronization, and asynchronous/off-policy schemes.

Design RL environments for agentic, multi-step tasks — sandboxed code execution, tool use, computer use, multimodal interaction — reproducible and scalable to millions of episodes.

Build reward infrastructure: verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking.

  • Develop the evaluation, monitoring, and debugging tooling that keeps large RL runs stable.

Advance training efficiency and stability, and turn new post-training ideas into production runs with researchers.

First 90 DaysOne way the first 90 could unfold.

Days 1–30 — Immerse & Diagnose: Learn the current RL stack and where throughput, stability, or correctness break.

Days 30–60 — Ship & Validate: Improve a piece of the loop (rollout throughput, reward infra, or an environment) and prove it on a real run.

Days 60–90 — Scale & Systemize: Harden the full loop across thousands of GPUs and asynchronous archit.

¿Te interesa? Aplicá ahora