Machine Learning Engineer
- $275,000 - $300,000
- San Francisco Bay Area
- Permanent
Machine Learning Engineer
Want to work on the systems that make LLMs fast, scalable and efficient in production?
A Stanford-spun AI company is hiring an AI Inference Engineer to help build and optimise the infrastructure powering its production AI platform.
This is a deeply technical role focused on squeezing more performance from GPUs, distributed systems and LLM inference at scale.
What you’ll be working on
- Architect high-performance infrastructure for LLM inference
- Optimise latency, throughput and cost across production workloads
- Scale inference across multi-GPU and multi-node environments
- Improve GPU scheduling, allocation and utilisation
- Profile bottlenecks across compute, memory and networking
- Optimise batching, quantisation, KV caching and parallelism
- Scale GPU workloads across Kubernetes infrastructure
What we’re looking for
You’ll come from an ML systems, distributed computing, HPC or performance engineering background and be comfortable working close to the hardware.
Strong Python and/or C++, PyTorch and GPU computing experience are key.
Experience with CUDA, NCCL, Triton, vLLM, TensorRT-LLM or SGLang would be particularly relevant.
You should be comfortable reasoning about how GPU utilisation, memory bandwidth, communication overhead and model architecture interact to impact inference performance.