Inference Engineer
- $180,000 - $300,000
- San Francisco Bay Area
- Permanent
Forward Deployed Engineer (Inference)
Location: San Francisco, CA
Employment Type: Full-time
Work Model: On-site
About the Company
A well-funded AI infrastructure company is building systems that automatically optimize inference workloads across different types of silicon.
The goal is to make accelerator capacity more fungible by optimizing each workload for the hardware best suited to serve it efficiently.
The company supports high-volume, mission-critical inference workloads for fast-growing AI companies and operates at significant production scale.
About the Role
Inference performance depends on far more than selecting a model and accelerator. Every production workload behaves differently, and understanding what can actually be delivered requires careful profiling, benchmarking, tracing, and systems engineering.
The work also continues well beyond a successful evaluation. Once a workload reaches production, it has to remain fast, reliable, and cost-efficient under real customer traffic.
As a Forward Deployed Engineer, you will own customer engagements from the moment the commercial team makes an introduction. You will determine what the platform can deliver, build the technical solution, deploy it into production, and keep it running.
This is not a project-management or traditional solutions-engineering role. You are the engineer responsible for both the system and the customer relationship.
What You’ll Do
- Own customer accounts end to end, from initial technical handoff through evaluation, deployment, production, and expansion.
- Profile, benchmark, and trace customer workloads to determine achievable performance and infrastructure requirements.
- Design meaningful benchmarks using real customer models and traffic.
- Lead technical evaluations and clearly explain benchmarking methodology, tradeoffs, and results to engineering teams.
- Implement the engineering required to bring inference workloads into production.
- Work across inference engines, model serving, infrastructure, observability, and performance tooling.
- Own production performance and reliability across a small portfolio of customer accounts.
- Participate in on-call responsibility when latency, reliability, throughput, or error rates regress.
- Debug production performance issues directly and drive them to resolution.
- Maintain trusted technical relationships with customer engineering teams.
- Turn repeated customer problems into product improvements and provide clear feedback to the core engineering organization.
- Make sound technical tradeoffs in ambiguous, high-urgency situations.
What We’re Looking For
- Strong systems-engineering fundamentals.
- Ability to operate directly with highly technical customers.
- Experience owning technical problems from investigation through production deployment.
- Understanding of LLM inference fundamentals, including:
- Prefill
- Decode
- Latency
- Throughput
- Performance tradeoffs
- Experience profiling, benchmarking, and tracing complex systems.
- Ability to read traces, investigate bottlenecks, and determine the actual source of a performance issue.
- Comfortable working across unfamiliar systems without a detailed specification.
- Strong technical communication skills with engineering audiences.
- Ability to explain benchmark methodology, technical constraints, and tradeoffs clearly.
- Strong production ownership and debugging instincts.
- Comfortable operating under ambiguity and urgency.
- Ability to turn incomplete customer requirements into a scoped technical plan and working implementation.
Strong Candidates May Also Have
- Experience with LLM inference engines.
- Model-serving infrastructure experience.
- GPU or accelerator performance optimization.
- Distributed inference or distributed-systems experience.
- Experience with inference benchmarking and performance analysis.
- Experience with CUDA, Triton, TensorRT-LLM, vLLM, SGLang, or similar systems.
- Kubernetes or production infrastructure experience.
- GPU profiling and tracing tools.
- Customer-facing engineering, solutions architecture, or forward-deployed engineering experience.
- Experience operating latency-sensitive production systems.
How Success Is Evaluated
The team values engineers who demonstrate:
- Resourcefulness
- Exceptional technical ability
- High standards
- Strong team orientation
- High emotional intelligence
- Rapid learning
- First-principles thinking
How the Team Works
This is a highly technical, hands-on environment with significant individual ownership.
Engineers are expected to operate as problem solvers across the stack rather than narrowly defined code contributors. You will have substantial autonomy over how problems are solved while working in a fast-moving production environment.
This role is not centered around meetings or handing technical implementation to another team. Most of your time will be spent benchmarking, profiling, debugging, implementing, and operating real inference workloads.
The defining feature of the role is breadth of ownership: you own the technical relationship, the production workload, and the engineering required to keep the customer successful.
Keywords
Forward Deployed Engineer, FDE, Inference Engineer, ML Systems, LLM Inference, AI Infrastructure, Model Serving, Inference Optimization, Performance Engineering, Systems Engineering, GPU Performance, Profiling, Benchmarking, Tracing, Performance Analysis, Latency, Throughput, Prefill, Decode, Time to First Token, TTFT, Inter-Token Latency, ITL, Tokens Per Second, vLLM, SGLang, TensorRT-LLM, CUDA, Triton, GPU, AI Accelerator, Distributed Inference, Distributed Systems, Kubernetes, Production Infrastructure, Observability, On-Call, Customer Engineering, Solutions Engineering, Technical Deployment, Performance Debugging.