AI Inference Engineer

LB160
  • $250,000 - $300,000
  • San Francisco Bay Area
  • Permanent

AI Inference Engineer


A Stanford-spun AI company in San Francisco is hiring an ML Inference Engineer to help redesign the infrastructure behind its production AI platform.


What you'll be solving

  • Architect high-performance systems for serving LLMs at production scale
  • Optimise inference for latency, throughput and cost efficiency
  • Build distributed execution across multi-GPU and multi-node environments
  • Improve how GPU resources are scheduled, allocated and utilised
  • Profile the inference stack to identify bottlenecks across compute, memory and networking
  • Scale GPU workloads across Kubernetes-based infrastructure
  • Optimise techniques including batching, quantisation, KV caching and parallelism
  • Make systems-level trade-offs where small improvements can have a major impact at scale


What you'll bring

  • You’ll have a strong background in ML systems, distributed computing, HPC or performance engineering, with experience working close to the infrastructure that runs modern ML models.
  • Strong Python and/or C++ skills are important, alongside a solid understanding of PyTorch and GPU computing.
  • Experience with technologies such as CUDA, NCCL or Triton would be highly relevant, as would hands-on work with inference frameworks including vLLM, TensorRT-LLM or SGLang.
  • You should be comfortable reasoning about performance at a systems level — understanding how GPU utilisation, memory bandwidth, communication overhead and model architecture interact to determine serving performance.


This is an opportunity to own meaningful parts of an inference stack being built from the ground up, rather than maintaining an established platform.

Anna Heneghan Senior ML Research & Engineering Recruiter

Apply for this role