For OSWorld that's 369 desktop tasks, each run 3 times, with the steps, tokens and time of every attempt.
This role builds and runs the evaluation framework that produces runs like these.
H builds computer-use agents and the models behind them.
What this team ownsThe evaluation framework: orchestration, runtimes and observability.
Researchers and forward deployed engineers bring the benchmarks, across web apps, desktop applications and the command line.
Your job is to make the framework that runs them reliable, fast and cheap, and to make adding a new one quick.
Research uses the results to choose checkpoints and decide whether a model ships.
Product and the forward deployed engineers use them to measure agents on customer workflows.
It carries roughly 50 benchmarks now.
That number should be between 100 and 200 soon, and the framework has to keep up.
What you'd be doingIntegration support for researchers and forward deployed engineers bringing in a benchmark, with a shorter path each time.
Setting the standard for how a benchmark enters the framework, and building the checks that enforce it.
Scheduling and observability, so cluster capacity isn't left idle while evaluation jobs queue.
Reproducible results across trials, so a release decision rests on numbers that hold.
Whatever stack a benchmark calls for.
One week that's cluster tuning; the next it's a browser extension or desktop environments.
Time with customers, from single developers to large companies, to find out what they want measured, then automating it so the results flow back into our harnesses and models.
The first few monthsBy 3 months you'll have helped researchers or forward deployed engineers integrate 5 benchmarks, and started fixing what slows the.