Data Scientist, AI/ML.
That’s why high-availability teams use Gremlin to find and fix reliability risks before they become incidents.
Gremlin Reliability Platform helps software teams proactively monitor and test their systems for common reliability risks, build and enforce reliability standards, and automate their reliability practices organization-wide.
As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the world’s largest organizations where high availability is non-negotiable.
You will be able to leverage your applied machine learning experience to inform product direction as well as solve complex technical problems that directly impact our customers (which range from the Fortune 500 to smaller organizations).
You will work closely with a small, talented engineering team focused on quality, delivery, and predictability with an emphasis on providing our customers a great user experience.
In this role, you’ll get to: Analyze Gremlin’s proprietary dataset of millions of chaos engineering experiments to identify failure patterns, root causes, and resilience signals across complex distributed systems Pretraining and fine-tuning machine learning models that automatically detect, classify, and explain failures observed during chaos experiments Build intelligent systems that deliver automated remediation recommendations, and eventually orchestration, by learning from historical experiment outcomes and system behavior Develop scalable data pipelines and feature stores to process,.