Kernel Engineer

LB173
  • $150,000 - $390,000
  • San Francisco Bay Area
  • Permanent

Member of Technical Staff - Kernels & GPU Performance

Location: San Francisco, CA

Employment Type: Full-time

Work Model: On-site



About the Company


A fast-growing AI infrastructure company is building a next-generation compute platform designed for high-performance, efficient AI inference across heterogeneous hardware.


The platform combines large-scale compute infrastructure with an execution layer that partitions AI workloads and maps each stage to the hardware best suited to run it.


The team works with leading AI organizations on technical challenges spanning frontier models, production infrastructure, and emerging accelerator architectures.



About the Role


As a Member of Technical Staff, you will build and optimize the low-level execution primitives that translate accelerator capability into production inference performance.


Rather than optimizing for a single hardware architecture, you will work across accelerators with different execution models, memory hierarchies, capabilities, and software stacks.


Your work will directly impact latency, throughput, hardware utilization, and efficiency across both established and emerging accelerator architectures.


You will work close to the hardware across kernel implementation, memory access, execution behavior, profiling, and performance validation, while partnering with compiler, ML systems, runtime, and distributed-systems engineers.



What You’ll Do


  • Build and optimize kernels for production AI workloads.
  • Improve latency, throughput, and hardware utilization.
  • Develop execution strategies across multiple accelerator architectures.
  • Optimize memory efficiency, scheduling behavior, and low-level execution characteristics.
  • Analyze accelerator execution models and memory hierarchies to identify bottlenecks.
  • Profile and validate performance across different hardware platforms.
  • Partner with compiler, runtime, ML systems, and distributed-systems engineers on end-to-end performance optimization.
  • Develop optimization approaches that account for architectural differences between accelerators.
  • Help establish performance-engineering standards and best practices across the execution platform.
  • Influence how heterogeneous compute hardware is deployed and utilized in production AI infrastructure.



What We’re Looking For


  • Strong software-engineering fundamentals.
  • Experience developing performance-critical systems close to hardware.
  • Strong understanding of low-level execution behavior.
  • Ability to reason about memory hierarchies, compute utilization, scheduling, and performance tradeoffs.
  • Experience profiling, debugging, and optimizing systems for latency and throughput.
  • Bachelor’s degree in a relevant technical discipline or equivalent practical experience.



Nice to Have


  • CUDA
  • Triton
  • CUTLASS
  • GPU kernel development
  • Accelerator programming models
  • GPU execution models including warps, wavefronts, blocks, and grids
  • Memory-access optimization and coalescing
  • Shared-memory optimization
  • Cache optimization
  • Occupancy tuning
  • Latency hiding
  • Instruction-level parallelism
  • GPU profiling and performance-analysis tools
  • Multi-GPU execution
  • Distributed execution
  • AI inference optimization
  • Heterogeneous accelerator experience



Keywords:

GPU Kernels, CUDA, Triton, CUTLASS, Kernel Optimization, GPU Performance, Performance Engineering, Accelerator Programming, AI Inference, Inference Optimization, Low-Level Systems, GPU Architecture, Memory Hierarchy, Memory Coalescing, Shared Memory, Cache Optimization, Occupancy, Latency Hiding, Instruction-Level Parallelism, ILP, Warp, Wavefront, Thread Block, Grid, Profiling, Nsight, Performance Analysis, Throughput Optimization, Latency Optimization, Hardware Utilization, Multi-GPU, Distributed Systems, Heterogeneous Compute, Accelerator Runtime, ML Systems, Compiler Runtime.

Kelly Dougherty Researcher

Apply for this role