AI Infrastructure Engineer
- -
- San Francisco, CA
- Permanent
Founding AI Infrastructure Engineer
🚨 Founding Infrastructure Engineer – GPU Neocloud & AI Infrastructure
📍 San Francisco, CA – On-site
I’m working with an early-stage infrastructure company building the first flex-power neocloud.
The company is building the hardware and software stack to keep GPU workloads reliable while managing compute around power availability. This could unlock approximately 70GW of existing capacity and bring AI compute online in months rather than waiting five or more years for new grid infrastructure.
This is a true founding engineering opportunity, working directly with the founders to build and operate production infrastructure for active customer workloads.
You’ll work on:
🖥️ Building and operating production GPU clusters across Kubernetes and Slurm
⚙️ Developing GPU orchestration, scheduling and lifecycle automation
💾 Designing distributed storage, NVMe and high-bandwidth networking systems
📊 Owning observability, reliability, workload performance and incident response
🤝 Working directly with customers to translate workload needs into infrastructure
Requirements:
✅ Hands-on experience building and operating Kubernetes or Slurm clusters from scratch
✅ Strong knowledge of GPU infrastructure, distributed storage and high-performance networking
✅ Experience owning production systems and resolving complex infrastructure issues
⭐ Experience with inference serving, distributed training or power-aware computing is highly valuable
The company was founded by experienced second-time (successful!) entrepreneurs from AI infrastructure, software, power, real estate and data centres.
Despite being at the founding-team stage, the business already has approximately $5M in annual contract value under contract, a further $40M in pipeline and ongoing partnership discussions with NVIDIA and AMD.
💰 Package: Competitive salary, cash bonus, benefits and meaningful founding-engineer equity.