System Software Engineer

LB156
  • $200,000 - $350,000
  • Mountain View, CA
  • Permanent

System Software Engineer — Node & Cluster Management


Bay Area, CA



We’re working with an early-stage AI hardware company building a new generation of high-performance accelerator systems for large-scale AI workloads.

They’re looking for a System Software Engineer to help build the management and observability layer that makes new AI hardware usable at scale — from individual nodes through racks and clusters.


This is a hands-on systems role for someone who can work across infrastructure APIs, Linux services, hardware telemetry, BMC interfaces, and lower-level device software. You’ll be building the tooling operators rely on while also being comfortable dropping into drivers, firmware interfaces, and raw device access when something breaks.



What You’ll Work On


  • Design and build the node-level management plane for AI accelerator systems
  • Expose system health, inventory, telemetry, diagnostics, and control through HTTP/REST APIs
  • Build cluster management, failover, recovery, and availability mechanisms
  • Develop CLI tools used for diagnostics, firmware updates, device recovery, and system management
  • Create unified management across host-side software and BMC/out-of-band interfaces
  • Build fleet-wide health aggregation, inventory, alerting, and external management integrations
  • Develop management and telemetry daemons running on Linux hosts
  • Work with lower-level driver interfaces and direct hardware-access utilities when debugging or prototyping
  • Build tooling for provisioning, lab automation, test orchestration, and regression monitoring
  • Define interfaces between Linux services, BMC firmware, device software, and management layers
  • Debug issues spanning APIs, userspace daemons, kernel drivers, firmware, and hardware
  • Help define how a brand-new AI platform is managed from a single server through full-scale clusters



What We’re Looking For


  • 8+ years of experience in systems software, infrastructure software, platform software, or hardware management
  • Strong Linux systems programming experience
  • Experience building low-level userspace software and Linux daemons
  • Strong C skills plus experience with Go, Rust, C++, and/or Python
  • Experience building REST APIs and CLI tooling for hardware or infrastructure systems
  • Solid understanding of Linux systems and the hardware underneath them, including:
  • device drivers
  • PCIe devices
  • hardware telemetry
  • firmware interfaces
  • BMC-managed subsystems
  • Comfortable reading and debugging kernel and systems-level code
  • Experience troubleshooting issues across application, daemon, kernel, firmware, and hardware boundaries
  • Able to work closely with firmware and hardware engineers to create common management interfaces
  • Comfortable operating in an early-stage environment where specifications and platform capabilities are still evolving



Particularly Relevant Experience - Experience in one or more of the following would be especially valuable:


  • Redfish, OpenBMC, IPMI, or gNMI
  • GPU, accelerator, HPC, or datacenter fleet management
  • Node and cluster management
  • Hardware health monitoring and telemetry
  • Firmware update and recovery workflows
  • Hardware bring-up and lab infrastructure
  • Provisioning and test automation
  • Secure boot or device attestation
  • BMC and out-of-band management
  • Large-scale Linux infrastructure



Keywords: Systems Software, Linux, Node Management, Cluster Management, Fleet Management, Infrastructure Software, Platform Software, AI Infrastructure, GPU Infrastructure, Accelerator Infrastructure, Datacenter Systems, Redfish, OpenBMC, BMC, IPMI, gNMI, REST API, Linux Daemons, Telemetry, Observability, Hardware Management, Firmware Management, PCIe, Hardware Bring-Up, Lab Automation, Provisioning, Failover, Device Management, C, C++, Rust, Go, Python

Kelly Dougherty Researcher

Apply for this role