We partner with enterprises tackling the hardest problems across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector, co-creating customized AI systems that they can run on their terms.
We are a dynamic, collaborative team passionate about AI and its potential to transform society.
Our diverse workforce thrives in competitive environments and is committed to driving innovation.
Our teams are distributed between Europe, North America, Asia and the Middle East.
We are creative, low-ego and team-spirited.
The RoleEvaluation is how we decide which models, checkpoints and recipes ship.
As a Research Engineer on the Eval Platform team, you will build the infrastructure every science team relies on to measure model quality, and make it reliable, reproducible and fast.
You don't need to have designed benchmarks before.
You do need to care about what a score means, and about when a difference between two runs is real.
What you will doBuild systems that keep eval results reproducible and comparable over time, as models, benchmarks and code evolve.
Run evaluations at scale across our GPU clusters, from model serving to scoring.
Make eval results easy to access, explore and trust, through APIs and dashboards that researchers use every day.
Catch broken or noisy evals before they mislead research decisions.
Work closely with researchers to turn new evaluation needs into robust, shared tooling.
What we're looking forMaster's or PhD in Computer Science, or equivalent experience.
4+ years building production-grade software, ideally large-scale ML codebases or distributed systems.
Excellent Python and strong software-design instincts: testing, code review, CI/CD.
Experience running workloads on GPU clusters (Slurm, Kubernetes, Ray or similar).
Familiarity with LLM in.