Software Engineer
- Up to $250,000
- San Francisco, CA
- Permanent
Forward Deployed Engineer - AI Inference
📍 San Francisco | Onsite
I’m working with a well-funded AI infrastructure startup building a high-performance inference cloud for open models.
The company optimizes the full path from model and serving engine down to kernels, hardware and production infrastructure, enabling customers to run inference faster and more cost-effectively across different accelerators.
This is not a traditional Solutions Engineer, Sales Engineer or project-management role.
You’ll own customer workloads from the initial technical handoff through evaluation, deployment, production and expansion:
→ Profile, benchmark and trace customer inference workloads
→ Identify bottlenecks across models, serving engines, kernels and infrastructure
→ Prove performance using customers’ own models and production traffic
→ Build and deploy the engineering required to bring workloads into production
→ Own latency, throughput, reliability and error-rate outcomes
→ Manage the technical relationship across three to six customer accounts
→ Turn recurring customer problems into improvements for the core platform
The strongest candidates will be exceptional systems engineers who can also work directly with demanding technical customers.
You should understand the fundamentals of LLM inference, including prefill, decode, latency, throughput and the trade-offs between them. Deep inference experience is valuable, but exceptional technical ability, strong engineering judgement and the capacity to learn quickly matter more than a narrow background.
Relevant experience could include:
• LLM inference and model serving
• Distributed systems and infrastructure
• Performance engineering and optimization
• GPU programming, CUDA, HIP or Triton
• Profiling, tracing and benchmarking
• Quantization and speculative decoding
• Heterogeneous accelerators and cluster operations
• Production reliability and observability
You must be comfortable opening a profiler, reading a trace, designing a meaningful benchmark and identifying the actual bottleneck.
You’ll also need to be credible in a room full of engineers, able to defend your methodology, explain technical constraints clearly and adjust your position when the evidence changes. The company currently serves trillions of tokens per month for mission-critical AI workloads and is hiring because customer demand is growing faster than the team’s capacity.
The environment is highly autonomous and genuinely intense. Engineers are expected to move across the stack, take responsibility for production outcomes and solve problems without waiting for a tightly defined specification.
This is an opportunity to join a small, talent-dense team where your work will directly influence customer performance, product direction and company growth.
Base salaries up to $250k + equity
Apply to Ethan to find out more
elewis@acceler8talent.com