Compiler Engineer
- $180,000 - $400,000
- San Francisco Bay Area
- Permanent
Member of Technical Staff - Compiler Engineer
Location: San Francisco, CA
Employment Type: Full-time
Work Model: On-site
About the Company
A fast-growing AI infrastructure company is building a next-generation compute platform for fast, efficient AI inference across heterogeneous hardware.
The platform combines large-scale compute infrastructure with an execution layer that partitions AI workloads and maps each stage to the hardware best suited to run it.
The team works with leading AI organizations on problems spanning frontier models, production infrastructure, and emerging accelerator architectures.
About the Role
As a Member of Technical Staff focused on compilers, you will help build the AI compiler stack that determines how machine learning workloads are represented, optimized, and executed across hardware with fundamentally different architectures, memory systems, and performance characteristics.
The compiler sits at the center of the inference stack, influencing scheduling, communication, memory movement, kernel execution, and end-to-end serving performance.
Unlike compilers optimized around a single hardware target, this stack must efficiently support CPUs, GPUs, and emerging accelerators without requiring a complete redesign for every new architecture.
The stack is built on open-source MLIR and LLVM infrastructure, with custom dialects, optimization passes, and lowering paths developed on top.
What You’ll Work On
- Scale compiler architecture across a growing set of heterogeneous compute targets.
- Design abstractions, MLIR dialects, compiler passes, and lowering paths that make it easier to support new CPUs, GPUs, and accelerators.
- Generate high-performance code competitive with native hardware toolchains.
- Build target-agnostic and target-specific compiler optimizations.
- Implement operator fusion, tiling, layout transformations, memory planning, and execution-planning optimizations.
- Optimize how workloads are partitioned across devices and how intermediate state moves between them.
- Improve end-to-end latency, throughput, utilization, and serving cost.
- Evaluate and integrate automatically generated kernels into compiler code-generation strategies.
- Develop systems that can choose between generated kernels, library implementations, and hand-written kernels based on performance.
- Enable new model architectures and inference techniques across multiple hardware targets.
- Work closely with kernel, runtime, ML systems, distributed-systems, and infrastructure engineers.
What Success Looks Like
In your first 12–18 months, you will:
- Ship compiler optimizations that measurably improve production latency, throughput, utilization, or cost.
- Help bring a new accelerator architecture into production.
- Reduce the engineering effort required to support additional hardware targets.
- Design execution strategies for partitioning workloads across heterogeneous hardware.
- Improve how the compiler selects between generated, library, and hand-written kernels.
- Enable new model architectures or inference techniques across multiple targets.
What We’re Looking For
- Experience building ML compilers or runtimes with a strong focus on performance optimization.
- Experience with MLIR, LLVM, or comparable compiler infrastructure.
- Experience designing intermediate representations.
- Experience writing compiler passes.
- Experience implementing lowering and code-generation pipelines.
- Strong C++ and/or Python skills.
- Strong understanding of SSA, memory systems, scheduling, and hardware efficiency.
- Experience with operator fusion, tiling, layout transformations, or memory planning.
- Experience using profilers to investigate performance across hardware/software boundaries.
- Familiarity with ML compiler or accelerator-programming frameworks such as IREE, XLA, TVM, Triton, or similar systems.
Nice to Have
- Experience optimizing ML inference or model-serving workloads.
- Kernel dispatch and launch API experience.
- Runtime-interface development.
- Memory allocator experience.
- Experience working across GPU and non-GPU accelerator architectures.
- Automated kernel generation.
- Autotuning or search-based optimization.
- Contributions to open-source compiler infrastructure.
- Experience with heterogeneous execution or distributed inference.
- Familiarity with speculative decoding, prefill/decode disaggregation, or similar inference optimizations.
Keywords:
MLIR, LLVM, ML Compiler, Compiler Engineer, Compiler Infrastructure, Intermediate Representation, IR, SSA, Compiler Passes, Lowering, Code Generation, Codegen, Operator Fusion, Tiling, Layout Transformation, Memory Planning, Scheduling, Runtime, Kernel Dispatch, CUDA, Triton, IREE, XLA, TVM, GPU Compiler, Accelerator Compiler, AI Compiler, ML Runtime, Heterogeneous Compute, AI Inference, Inference Optimization, Model Serving, Kernel Generation, Autotuning, Search-Based Optimization, C++, Python, Hardware Efficiency, GPU, AI Accelerator.