Research Engineer - Mechanistic Interpretability

LB54
  • $150,000 - $350,000
  • San Francisco, CA | Onsite
  • Permanent


๐Ÿšจ Research Engineer โ€“ Interpretability Systems


๐Ÿ“ San Francisco, CA | Onsite


๐Ÿง  Early-stage AI research lab | Revenue-generating


An AI research lab working at the frontier of interpretability, alignment, and reinforcement learning is hiring Research Engineers focused on understanding whatโ€™s happening inside large language models


This role is for engineers who want to build the experimental systems that make interpretability research possible - not production ML, MLOps, or large-scale training infra


Youโ€™ll work on:

๐Ÿ” Activation tracing & mechanistic analysis

๐Ÿงช Custom RL-style environments for alignment research

๐Ÿง  Probing internal representations

๐ŸŽฏ Detecting latent concepts like deception, goals, uncertainty, or hidden objectives

๐Ÿ› ๏ธ Activation-level steering beyond prompting and fine-tuning

๐Ÿ“Š New benchmarks for model consistency and robustness


The work is fast, experimental, and greenfield: build custom tooling, test research ideas, get results, move on.


Ideal background:


โœ… Strong software engineering fundamentals

โœ… Experience with experimental ML / research systems

โœ… Comfort working close to model internals

โœ… Interest in interpretability, alignment, RL, or mechanistic understanding

โœ… PhD helpful, not required


This is not a role for scaling pipelines or maintaining production systems


Itโ€™s for people who enjoy ambiguous problems, fast research cycles, and building new tools from first principles


Interested? Apply & Drop me a message!




See curated AI tools for your jo

Zee Uddin Researcher

Apply for this role