Software Engineer

LB136
  • $250k base + equity + bonus
  • San Francisco, CA
  • Permanent

Software Engineer, Reinforcement Learning Environments

On-site | San Francisco, CA

$250k base + equity + bonus 


I’m working with a highly profitable AI research infrastructure company building reinforcement-learning environments for frontier model labs.

The team creates the tasks, simulations, reward signals, evaluation rubrics, and expert trajectories used to improve how advanced models reason and act. These are not generic annotation datasets. They are structured RL environments designed to expose failure modes, measure capability, and produce useful learning signals.


This role will directly influence model post-training by turning research objectives into environments and experiments that can run at scale.


You’ll work on problems such as:

  • Designing RL environments for coding, finance, and enterprise workflows
  • Creating tasks that expose meaningful model and agent failure modes
  • Building reward functions and evaluation rubrics for RLHF and RLVR
  • Analyzing trajectories to understand why agents succeed or fail
  • Improving the quality and reliability of training signals
  • Measuring how environment and dataset changes affect model capability
  • Developing real-world and synthetic-data pipelines



Looking for engineers who have:

  • Strong hands-on reinforcement-learning experience
  • Experience applying RL to LLM post-training, agents, or sequential decision-making
  • Experience building environments, reward functions, evaluation tasks, or scoring systems
  • Familiarity with RLHF, RLVR, policy optimization, or preference-based learning
  • The ability to design controlled experiments and interpret noisy results


This is an opportunity to build the reinforcement-learning environments and reward systems that directly shape how frontier models learn, reason, and improve.

The company operates in person from San Francisco and values speed, ownership, experimental judgment, and measurable results.

Anna Button Researcher

Apply for this role