Research Engineer - Vision Language Models / Multimodal AI / Computer Vision

LB152
  • -
  • Mountain View, CA
  • Permanent

🤖 Research Engineer (VLM / Multimodal AI) 🧠


I'm working with a stealth AI robotics startup building AI-powered observability for critical infrastructure 


Think: drones and quadrupeds autonomously monitoring solar farms, data centres, refineries, and other massive-scale environments in real time 🌎


They're hiring a Research Engineer to build the multimodal agentic layer that connects language, vision, and robotics


You'll be working on:

🧠 VLMs / Multimodal Models

🤖 Agentic AI Systems

👁️ Computer Vision

📡 Real-world robotics deployments

📊 Large-scale infrastructure analytics


Imagine enabling customers to ask:

➡️ "Inspect this fault"

➡️ "Analyze this live stream"

➡️ "Recommend the best next action"

And having AI coordinate robots, analyze visual data, and surface the right decisions in seconds


Why join?

🚀 Ex-Google & Meta robotics founders

💸 Seed funded with 2+ years runway

🏆 Founders with multiple successful exits ($115M & $350M)

🤝 Backed by top operators from YC, Dropbox, Nest & South Park Commons

🌍 Live deployments already operating in the field

📍 Mountain View, CA (Onsite)


Ideal background:

✅ Multimodal AI

✅ VLMs / VLAs

✅ Computer Vision

✅ Applied Research / Research Engineering


If you're excited by the intersection of multimodal AI, agentic systems, and real-world robotics, let's chat! 🙂

Zee Uddin Researcher

Apply for this role