AGCO is designed to optimize LLMs across different GPUs and optimization targets, including very fast inference.
The team has 10 people, including 9 engineers and researchers and 4 PhDs.
Test it at playground.
Read the technical details on the Kog Labs blog.
WHAT YOU WILL WORK ONYou will work directly on AGCO.
The goal is to build a system that can explore ways to optimize LLM execution, generate changes, compile them, check correctness, run them on real hardware, measure the results, and use this feedback to guide the next optimization.
You will contribute to areas such as:Compiler and IR design for representing and transforming LLM computations.
Optimization passes, lowering, and code generation.
Search methods for exploring different implementations and execution strategies.
Verification and correctness checks for generated changes.
GPU execution, profiling, and performance optimization.
LLM inference across operators, memory, parallelism, and communication.
Optimization loops that connect generated changes to measurements on real GPUs.
One direction we are exploring combines an IR, a verifier, a compiler, and a search optimizer.
We plan to start with focused problems, build working prototypes, and extend the system from what we learn.
Your main area will.