Volver

ML Systems Engineer — Inference Acceleration

CompraTica Empleos

EMP:Technology
Paris Offices
Tiempo Completo
Remoto
0 vistas

Descripción

Meet Arago and the AragoniansArago’s mission is to re-engineer the foundations of computing from first principles.

The explosive growth of AI is pushing the industry to rethink how processors are built.

Arago is meeting that challenge with a proprietary technology that fuses optical and CMOS technologies to deliver an order-of-magnitude increase in performance.

Arago is the fastest, and currently the only, company to have built such a processor.

It's backed by leading deep-tech investors and some of the most respected figures in semiconductors and computing, including the CEO of Arm, the founder of macOS who worked directly with Steve Jobs at Apple, an Nvidia Fellow, the Head of Optics at Google, and many other industry leaders.

Our work is guided by three clear values: do great things, move with high velocity, and operate as one unit.

We work in a demanding environment where constant learning, ownership, and execution are expected, and where exceptional people have the opportunity to do their life’s work.

 What you’ll doOptimize the execution and serving of modern AI models on Arago's custom accelerator.

Work across kernels, model execution, multi-device distribution, runtime, and inference serving, while helping shape the software stack around the capabilities of Arago's hardware.

Habilidades

and ExperienceStrong experience in high-performance ML inference, GPU/accelerator programming, or ML systems engineering.Deep understanding of computer architecture, accelerator/GPU execution models, memory hierarchies, parallelism, and performance bottlenecks.Experience developing and optimizing custom kernels using CUDA, Triton, ROCm/HIP, or equivalent low-level programming environments.Experience with operator fusion, tiling, scheduling, data movement optimization, graph execution, and profiling of compute- and memory-bound workloads.Strong understanding of distributed model execution, including tensor, pipeline, sequence, and/or expert parallelism and communication/comp...

¿Te interesa? Aplicá ahora