Volver

AI Research Engineer (Kernel & Inference Optimization)

CompraTica Empleos

EMP:Technology
Switzerland
Tiempo Completo
Remoto
0 vistas

Descripción

This position is listed on behalf of a partner company, who manages all applications and next steps.

Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in Switzerland.

You will work at the intersection of AI research, systems engineering, and high-performance model inference.

Your focus will be on developing and optimizing model-serving architectures for advanced AI systems across a range of hardware environments.

You will tackle challenges involving latency, throughput, memory efficiency, and scalability, including deployment on resource-constrained mobile and edge devices.

The role combines hands-on research with low-level engineering, giving you the opportunity to develop novel inference strategies and GPU kernels.

You will work with complex architectures spanning text, image, audio, diffusion models, and vision transformers.

Your work will involve rigorous benchmarking, production testing, and iterative optimization to translate research into measurable performance improvements.

You will collaborate with cross-functional teams in a highly technical, remote environment focused on pushing the boundaries of efficient AI systems.

Accountabilities Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization.

  • Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms.

Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability.

Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates.

Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions.

Identify computational and memory bottlenecks across inferen.

¿Te interesa? Aplicá ahora