We are looking for an SRE Lead to serve as the senior technical leader and player-coach for ProdOps. This is a hybrid role: you will set the technical direction of the team and lead from the front during incidents, while also growing and managing a small group of exceptional engineers as the function scales.
As the calm center during a crisis, you will maintain a high-level mental model of the entire production ecosystem, freeing engineers to focus strictly on debugging and mitigation. You will own the reliability feedback loop end to end, which includes running blameless post-incident reviews, coordinating major incidents across Cloud systems, vehicle development pipelines, and Product Security, driving the systemic action-item backlog with TPMs and development teams, and verifying that fixes hold in production.
This role suits a senior systems engineer who still writes code, thinks in terms of systems and failure modes rather than single root causes, and wants to build a lean, automation-first reliability practice rather than staff a support queue.
Drive incident progress and coordinate cross-functional response as the central nervous system during a crisis. Own executive, customer, and parent-company communications, providing production expertise and communication leadership so engineers can concentrate on the technical problem. Maintain an accurate, high-level model of the full production ecosystem spanning Cloud, vehicle development pipelines, and Product Security.
Run post-incident reviews using Learning From Incidents (LFI) principles and HOWIE-style reporting. Move the organization away from the search for a single root cause and toward understanding how tooling, context, and multiple latent conditions combined to produce failure. Surface weaknesses in observability, process, testing, and tooling that impaired our ability to detect, mitigate, and recover.
Manage the backlog of action items generated by reviews. While ProdOps does not write the fixes itself, you will prioritize, track, and drive these items to closure in partnership with TPMs and development teams, keeping leadership focused on customer impact.
Once the fixes ship, measure and verify that they actually prevent recurrence. Drive accountability for outcomes and feed the results back into the Novel Incident Rate.
Design and build the systems, tooling, and AI-agent workflows that automate incident triage and administrative toil. Improve observability, including instrumentation, alerting, dashboards, and SLOs, so incidents are detected faster and understood more deeply. Write and review code where it multiplies the team's impact.
Set technical standards and operating rhythm for ProdOps. Mentor and develop engineers, and as the team scales, take on hiring and people management while preserving the minimal-headcount, maximum-automation philosophy.
We build the exceptional — and we believe the people doing that work should be rewarded accordingly. In addition to a competitive base salary, full-time positions may be is eligible to participate in our annual company performance bonus program.
Payments are discretionary and not guaranteed; actual amounts depend on company results and the terms of the plan in effect, and require active employment at the time of payout. This role is also eligible for equity in the form of Restricted Stock Units (RSUs), subject to board approval and the terms of our equity incentive plans, including applicable vesting requirements.
In addition to our compensation programs, we invest in our people with a comprehensive benefits package designed to support the health, wellbeing, and financial future for full-time employees — including health coverage, retirement savings, time off, and family planning programs. Offerings vary by country. Learn more about our global benefit programs.
External candidates can apply for this role through the Rivian and Volkswagen Group Technologies careers site (https://rivianvw.tech/#careers). If you are a current employee, please apply through our internal job board.
Rivian and Volkswagen Group Technologies is a joint venture between two industry leaders with a clear vision for automotive’s next chapter. From operating systems to zonal controllers to cloud and connectivity solutions, we’re addressing the challenges of electric vehicles through technology that will set the standards for software-defined vehicles around the world.
The road to the future is uncharted. By combining our expertise across connectivity, AI, security and more, we’ll map a new way forward. Working together, we’ll create a future that’s more connected, more intelligent, more sustainable for everyone.
Search Staff Site Reliability Engineer-Production Operations jobs near Palo Alto, California (Remote) → Browse all live jobs
This posting was published by Rivian and Volkswagen Group Technologies on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.