Inference Serving - vLLM
LB137
Posted: 16/09/2026
- -
- Santa Clara, CA | Onsite
- Permanent
I'm hiring for a Machine Learning Engineer – LLM Inference Serving role on behalf of a stealth AI infrastructure startup 🚀
Must-have: deep hands-on experience with vLLM or SGLang 🔬
We’re looking for engineers who have:
- Made public GitHub contributions to vLLM/SGLang
- Refactored inference framework internals to improve performance
- Worked on KV Cache lifecycle management
- Used Ray, Dynamo, or similar inference orchestration platforms
📍 Santa Clara, CA | Onsite
Focus: LLM inference, distributed serving, throughput/latency optimisation, KV cache, scheduling
Please reach out with links to relevant GitHub contributions or inference framework work 📩
Zee Uddin
Researcher