Experience: 4+ years
The SRE team owns reliability and infrastructure for Anduril's cloud deployments. We operate Kubernetes clusters, Terraform infrastructure, and observability platforms across 10+ production environments supporting active defense contracts. When platform services break under real operational load, we're the team that fixes them - often at the code level, not just the config level.
We are looking for a Senior Production Engineer to join our team in Costa Mesa, CA (or DC) . In this role, you will be responsible for diagnosing and fixing stability vulnerabilities in core platform services that cause cascading failures in multi-tenant cloud deployments. You will write production Go to implement resilience patterns - leader election, circuit breakers, failure domain isolation - directly in service code.
This will require deep experience with distributed systems, debugging complex failure modes across service boundaries, and writing production-quality Go. If you are someone who thrives on fixing hard reliability problems in live systems rather than building greenfield, this role is for you.
- Diagnose and fix stability vulnerabilities in core platform services that cause cascading failures under multi-replica, multi-tenant operation
- Implement resilience patterns (leader election, circuit breakers, failure domain isolation) directly in service code
- Design multi-replica support for services that currently assume single-instance operation
- Collaborate with service owners on contract testing and upgrade validation
- Trace cascading failures across service boundaries and drive them to root-cause fixes
- Contribute to observability platform improvements to support service stability
- Light infrastructure work: Terraform/Kubernetes changes to support service fixes (~20% of time)
- Production-quality Go - you'll be modifying core platform services, not writing scripts
- Practical experience with distributed systems: leader election, consensus, replication, failure modes
- Kubernetes - enough to understand how services run (not necessarily cluster administration)
- Debugging complex systems - tracing cascading failures across service boundaries
- 4+ years in SRE, platform engineering, or backend development roles
- Must be a U.S. Person due to required access to U.S. export controlled information or facilities
- Eligible to obtain and maintain an active U.S. Secret security clearance
- Experience fixing reliability problems in production services (not just building greenfield)
Search Senior Production Engineer jobs near Costa Mesa, CA +1 more → Browse all live jobs
This posting was published by Anduril Industries on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.