Hadrian Automation

Site Reliability Engineer, Robotics

$164K–$270KFull-time · Los Angeles, CA
✓ Verified live on the employer's own system · added 75 days ago
Save search

What this role involves

SlisSlosPrometheusDatadogKubernetesObservabilityTuning

Skills & tools

Industrial AutomationDevopsProgrammingOperationsEmbeddedJavascriptPythonGolang
Apply on company site ↗ See your fit → free

Full job description

- Own the reliability of our robotics systems, from PLCs through ROS2/middleware to Kubernetes.

- Build interfaces to our observability system to ingest telemetry from our controls and robotics systems. Leverage solutions such as Prometheus, Telegraf, OpenTelemetry, and Datadog.

- Write code frameworks and tools to support our controls and robotics systems, including diagnostic tools, shared libraries for telemetry data, and automated remediation.

- Partner with controls, robotics, and platform engineering teams to bake reliability in early. Review designs, develop SLOs and SLIs, introduce reliability release gates, and push for telemetry contracts to develop production-grade services.

- Ownership. Someone who has owned the reliability of a production system where downtime had physical or operational consequences (manufacturing line, autonomous vehicle, lab automation, network operations)

- Systems Thinker. Focused on understanding the relationship among various systems to design sustainable solutions, not one-time fixes.

- Problem Solver. Solving complex puzzles excites and motivates you to find an efficient solution.

- T-Shaped Skill Set. Comfortable with bare metal Kubernetes, networking, GitOps workflows, and Infrastructure as Code (IaC). Also skilled in programming in TypeScript, Python, Golang, or C++.

- Strong Communication. You can run a war room, write a post-mortem, and explain a reliability tradeoff to a stakeholder.

- You've built automated or self-healing remediation at scale. We want systems that remove humans from the loop.

- Background in edge/on-prem infrastructure. You’ve run Kubernetes at the edge (k3s, k0s, k0smotron), managing on-prem clusters, time-series at the edge, or air-gapped deployments. A deep understanding of Linux operating system fundamentals such as cgroups, sockets, and system tuning, is a big plus.

- Deep understanding of shipping and storing telemetry data at scale. Experience with Kafka/MQTT/RabbitMQ is a plus.

- Direct robotics experience. ROS/ROS2, OPC UA, EtherCAT, motion controllers, or fleet management for autonomous systems

- An individual who is self-directed and can deliver with high velocity.

For this role, the target salary range is $164,000 - $270,000 (actual range may vary based on experience).

More jobs at Hadrian Automation

Similar jobs near Los Angeles, CA

Tell me when more Site Reliability Engineer, Client Platform jobs post near Los Angeles, CA We re-check every listing against the employer’s own board — no résumé needed.

Search Site Reliability Engineer, Robotics jobs near Los Angeles, CA → Browse all live jobs

This posting was published by Hadrian Automation on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.