Artificial Intelligence Engineer

LB144
  • $250,000 - $400,000
  • Bay Area, CA | Onsite
  • Permanent

Senior Pretraining Engineer – Video & Multimodal AI


Location: Bay Area, CA | Onsite

Focus: Large-scale pre-training, video, multimodal models, diffusion, flow matching


I'm working with an early-stage robotics company building a deeply integrated AI and robotics stack from first principles.


Their long-term goal is ambitious: create highly autonomous factories capable of manufacturing physical goods with dramatically less human manual labour. That means solving problems across robot learning, perception, simulation, rendering, GPU performance and large-scale multimodal model training.


They’re now hiring a Senior Pretraining Engineer to take ownership of large-scale video and multimodal pre-training.


This is not a role for someone who has only fine-tuned existing models or operated clean, established training pipelines.


You’ll be expected to understand what happens when big models, big datasets and big compute collide, including the failure modes that only become visible once training runs become genuinely expensive.


You’ll work across model architecture, training infrastructure, data and experimentation to build large-scale vision and multimodal systems that can ultimately contribute to robotic intelligence in complex physical environments.


  • Own large-scale video and multimodal pre-training runs
  • Train models across substantial distributed GPU infrastructure
  • Develop and improve generative approaches including flow matching and diffusion
  • Diagnose instability, convergence issues and large-scale training failures
  • Improve training efficiency, reliability and experiment velocity
  • Build safeguards based on lessons from failed or expensive training runs
  • Work closely with robotics, simulation, perception and systems engineers
  • Push research ideas into working systems rather than isolated experiments


Strong candidates will have experience with:

  • Large-scale video, vision or multimodal pre-training
  • Distributed training across substantial GPU compute
  • Training large models on large datasets
  • Debugging difficult training failures at scale
  • Understanding the interaction between models, data and compute
  • Flow matching, diffusion models or adjacent generative-model approaches
  • Strong software engineering and training-systems fundamentals
  • Designing experiments where compute cost makes mistakes consequential


Just as importantly, you should be able to talk openly about training runs that didn’t work: what failed, how you diagnosed it, what it cost, and what you changed afterward.


The wider engineering environment spans:

  • Robot learning and reinforcement learning
  • Video and multimodal foundation models
  • Perception for difficult real-world environments
  • Custom physics simulation
  • Rendering and light transport
  • GPU kernel optimization
  • Hardware-software co-design


That creates an unusually broad technical surface area. Your models won’t exist purely to improve benchmark scores. The longer-term objective is intelligence that can operate through physical systems and contribute to genuinely autonomous manufacturing.


The company is onsite in the Bay Area, with its R&D operation to be based in San Jose.


If you’ve personally taken large multimodal or video models through expensive pre-training runs, including the painful ones, this is one of the more unusual opportunities to apply that experience to physical-world AI.


Apply today!


Ethan Lewis

elewis@acceler8talent.com

Ethan Lewis Researcher

Apply for this role