Member of Technical Staff - ML Systems & Inference

LB80
  • $250,000 - $350,000
  • San Francisco Bay Area
  • Permanent

 Member of Technical Staff - ML Systems & Inference πŸš€


πŸ“ Bay Area, CA | Onsite πŸ“


Join a well-funded AI infrastructure startup building the orchestration layer for next-generation AI workloads


This role sits at the intersection of ML systems, inference, distributed systems, and performance engineering. You'll build production inference systems, optimize scheduling and memory management, improve KV cache efficiency, and work closely with compiler, kernel, and distributed systems engineers to push AI infrastructure forward


We're looking for engineers with:

  • Strong software engineering fundamentals
  • Experience with ML inference or model serving
  • Knowledge of distributed systems and performance optimization
  • Python and/or C++
  • Experience with vLLM, TensorRT-LLM, CUDA, or similar is a plus


This is an opportunity to help define how AI work

Zee Uddin Researcher

Apply for this role