Member of Technical Staff- Distributed Systems

LB73
  • $150,000 - $350,000
  • San Francisco, CA-
  • Permanent

Member of Technical Staff - Distributed Systems


📍 San Francisco, CA- 5 days per week onsite

💰 $150,000–$350,000 + equity


The Opportunity

Join a rapidly growing AI infrastructure company building a multi-silicon cloud platform for fast, efficient inference.

The future of AI inference will not run exclusively on GPUs. It will span CPUs, GPUs, and an expanding range of specialized accelerators. Our platform partitions, schedules, and routes workloads across this heterogeneous infrastructure while providing customers with a production-grade API that abstracts away the underlying complexity.

This is proven technology operating at meaningful scale. The company recently raised an $80 million Series A, has generated eight figures in revenue since emerging from stealth, and already supports a frontier AI lab and a hyperscaler.


The Role

You will design, build, and operate the distributed systems responsible for scheduling and coordinating AI workloads across thousands of nodes.

Your work will sit at the heart of the platform, with significant ownership over its architecture, reliability, and performance. You will help shape both the technical foundation and the engineering culture of an approximately 25-person company.


What You’ll Do

  • Build orchestration, scheduling, and control-plane systems for large-scale AI inference workloads
  • Design reliable distributed systems that operate across thousands of heterogeneous compute nodes
  • Improve fault tolerance, availability, observability, and operational resilience
  • Develop high-performance APIs for customers running mission-critical workloads
  • Diagnose performance bottlenecks and quantify the impact of improvements
  • Make pragmatic architectural decisions in a fast-moving, early-stage environment
  • Collaborate directly with founders and engineers from organizations including NVIDIA, Google AI, Intel, and Pixie Labs


What We’re Looking For

  • Strong software engineering and computer science fundamentals
  • Experience building or operating distributed systems in production
  • A deep understanding of concurrency, failure recovery, and architectural trade-offs
  • Evidence of personally building, optimizing, or scaling meaningful systems
  • The ability to clearly quantify the performance, reliability, or business impact of your work
  • Comfort operating with ambiguity and taking ownership in an early-stage company
  • Proficiency in Go, C++, Python, or comparable systems-oriented languages


Experience in the following areas is particularly relevant:

  • Kubernetes internals
  • Distributed schedulers and orchestration systems
  • RPC frameworks or messaging systems
  • Queues and asynchronous processing
  • High-throughput or latency-sensitive APIs
  • Cloud infrastructure or AI/ML systems


What Strong Candidates Demonstrate

We value direct technical contribution over credentials alone. The strongest candidates can explain what they personally built or improved, why they made specific architectural decisions, and how those decisions affected scale, performance, cost, or reliability.

Big-company pedigree is welcome, but it is not sufficient without clear evidence of hands-on ownership and impact.


Why Join?

🚀 This role combines genuine production scale with the ownership available only at an early-stage company. You will work on difficult distributed-systems problems, collaborate closely with an experienced founding team, and have meaningful influence over the platform’s architecture and the company’s engineering culture.

This is a five-day-per-week onsite role in San Francisco. Candidates should be excited by highly collaborative, in-person work and the opportunity to help shape an ambitious infrastructure company from an early stage.

Samuel Killick Researcher

Apply for this role