<div class="content-intro"><p><strong>.
Nebius is leading a new era in cloud infrastructure for the global AI economy
We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
</p> <p>Built by engineers, for engineers.
From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
</p> <p>Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel.
Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.We’re looking for a Senior HPC Cluster Engineer to join our team and play a key role in the development of our cutting-edge hyperscaler platform
The GPU & InfiniBand team is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack.
You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments.
The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system.
</p> <p> </p> <p><strong>In this position, you will be responsible for:</strong></p> <ul> <li><strong>Tuning the performance</strong> of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments.
</li> <li><strong>Analyzing and troubleshooting</strong> the root cause of issues related to GPUs and In.