Join the team
Training and Inference Infrastructure
Run the GPU clusters, distributed training stack, and inference services that everything else is built on.
Responsibilities
- Own the large-scale GPU-cluster infrastructure that supports LLM and world-model training, ensuring high availability and operational stability.
- Optimise distributed training frameworks across data, model, and pipeline parallelism, and identify and resolve communication, GPU-memory, and I/O bottlenecks.
- Build high-throughput, low-latency inference services covering key techniques such as quantisation, KV-cache optimisation, and continuous batching.
- Collaborate with algorithm teams to support rapid experimentation in training and inference, and evaluate and introduce new hardware and software stacks.
Qualifications
- At least five years of experience in infrastructure or systems engineering. Hands-on deployment experience with large-scale distributed training or inference is preferred.
- Deep understanding of GPU architecture, with familiarity with CUDA, NCCL, and performance-profiling toolchains.
- Source-level optimisation experience with at least one mainstream training framework, such as Megatron-LM, DeepSpeed, or PyTorch FSDP.
- Familiarity with inference-optimisation methods such as PagedAttention, speculative decoding, and continuous batching.
- Experience with Kubernetes, containerised deployment, and high-speed InfiniBand or RoCE networking.
- Strong Python and C++ programming skills, with the ability to independently read and modify low-level framework code.
Preferred Qualifications
- Experience building training or inference systems for world models, including video generation or physical simulation.
- Significant contributions to open-source inference frameworks such as vLLM, TensorRT-LLM, or SGLang.
- Publications at leading systems conferences such as MLSys, OSDI, SOSP, or SC.
Apply by email
Location: Hangzhou, Beijing, Switzerland. Send your CV to info@awomo.ch. Suggested subject line: “Name + Position” — the apply button fills in the role for you, so just replace “Name”.