Research Engineer - Vision Language Models / Multimodal AI / Computer Vision
- -
- Mountain View, CA
- Permanent
🤖 Research Engineer (VLM / Multimodal AI) 🧠
I'm working with a stealth AI robotics startup building AI-powered observability for critical infrastructure ⚡
Think: drones and quadrupeds autonomously monitoring solar farms, data centres, refineries, and other massive-scale environments in real time 🌎
They're hiring a Research Engineer to build the multimodal agentic layer that connects language, vision, and robotics
You'll be working on:
🧠 VLMs / Multimodal Models
🤖 Agentic AI Systems
👁️ Computer Vision
📡 Real-world robotics deployments
📊 Large-scale infrastructure analytics
Imagine enabling customers to ask:
➡️ "Inspect this fault"
➡️ "Analyze this live stream"
➡️ "Recommend the best next action"
And having AI coordinate robots, analyze visual data, and surface the right decisions in seconds
Why join?
🚀 Ex-Google & Meta robotics founders
💸 Seed funded with 2+ years runway
🏆 Founders with multiple successful exits ($115M & $350M)
🤝 Backed by top operators from YC, Dropbox, Nest & South Park Commons
🌍 Live deployments already operating in the field
📍 Mountain View, CA (Onsite)
Ideal background:
✅ Multimodal AI
✅ VLMs / VLAs
✅ Computer Vision
✅ Applied Research / Research Engineering
If you're excited by the intersection of multimodal AI, agentic systems, and real-world robotics, let's chat! 🙂