Duration: 5 months | Hours: ~20/week.
Most AI companies aim to replace human work.
Retorio does the opposite, empowering enterprise sales and service teams through science-driven AI coaching.
The platform runs realistic client conversation simulations and behavioral analysis to build trusted advisors across global teams at companies like Vodafone, Merck, and Daimler Truck.
Tasks You build agents that run in production.
Not demos, not notebooks.
Our real-time conversation engine, content generation, and result generation are all agentic systems serving Fortune 500 customers every day, and you own pieces of that: design, deploy, scale, and prove they work.
Our product is grounded in behavioral science research, and we hold our engineering to the same standard.
Every change to an agent starts as a hypothesis and ends with a measurement.
That means the work is not only prompts and graphs.
You will write backend services, run your own deployments, and build the tooling that tells you whether your change actually improved anything.
Your Tasks Design and build agentic systems: multi-step reasoning, tool calling, MCP, memory, retrieval, structured outputs Run the research loop: read what is current, form a hypothesis, build the experiment, measure, then decide.
Kill your own ideas when the numbers say so Evaluate rigorously: eval datasets, offline and online scoring, LLM-as-judge, A/B tests, tracing and dashboards.
We ship on evidence, not vibes Build the backend around the agents: Python services and APIs, streaming interfaces, database and schema work Own the DevOps: containerize, deploy on GCP (Cloud Run, CI/CD), instrument logs, metrics, traces, and alerts, then debug your own production incidents Scale what you ship: latency, cost per conversation, model routing and fallbacks across providers, graceful failure when a model or a vendor misbehaves Take features from.