Drata

Senior Site Reliability Engineer

Full-time · Hybrid - San Francisco (Remote)
✓ Verified live on the employer's own system · added 105 days ago
Save search
Senior · 6+ yrs exp

Requirements

Experience: 6+ years

Skills & tools

DevopsProgrammingSystems EngineeringCloud PlatformsOperationsCustomer ServiceManagementGit

Benefits — mentioned in this posting

Equity / stock
Apply on company site ↗ See your fit → free

Full job description

Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE team where you grow your career, shape standards, and collaborate with peers - while also serving as the dedicated reliability partner for one of Drata's product engineering teams across the full lifecycle of their work.

This is a highly technical role at the intersection of software engineering and systems engineering. The best SREs at Drata are engineers first: they solve problems by building solutions, not by executing manual processes. Automation is a core value, and nowhere is that more visible than in how we approach reliability.

Our infrastructure runs on AWS across multiple accounts, defined entirely in Terraform. You'll work across a modern cloud-native stack to help Drata scale reliably for a rapidly growing customer base.

You are the reliability expert for your aligned product team. You engage early - during architecture reviews and design discussions - to surface risks before they become incidents.

- Lead Production Readiness Reviews (PRRs) before new services launch, with the authority to flag gaps and gate launches when critical reliability standards aren't met

- Partner with product engineering leads and staff engineers to define SLOs and SLIs for critical services, turning reliability from a vague goal into a measurable commitment

- Participate in team planning and architecture reviews to provide proactive reliability guidance

- Build reusable artifacts - SLO templates, observability checklists, alerting standards, reference dashboards - that raise the reliability floor across the team, not just the services you touch directly

You handle operational needs from your product team, but your job isn't to be a help desk. Your goal is to make each request the last of its kind. When an engineer needs something, your priority is: automate it so anyone can do it → document it so the team can self-serve → execute it manually only as a last resort.

- Build and maintain Datadog monitors, dashboards, and alert routing - enforcing infrastructure-as-code standards via Terraform so those resources are owned, versioned, and auditable

- Handle infrastructure requests: ECS task management, secret rotations, Terraform changes, capacity adjustments

- Identify repeated manual work and convert it into self-service tooling or runbooks

- Audit existing services for reliability anti-patterns and surface top risks before they cause incidents

Beyond your product team, you contribute to cross-cutting infrastructure, tooling, and standards that benefit every team at Drata. Recent examples include automated Datadog governance workflows, dynamic AWS account provisioning, and disaster recovery exercises.

- Design and build shared platform infrastructure - reusable Terraform modules, standardized observability stacks, service templates - so reliability improvements compound across the organization

- Participate in the on-call rotation and lead incident response when needed; conduct thorough post-incident reviews to drive lasting fixes

- Contribute to evolving SRE standards, tooling, and practices across the organization

- 6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building and maintaining scalable, resilient services

- Robust knowledge of cloud computing technologies: Terraform, Docker, Git, and Linux

- Hands-on experience with Datadog for monitoring, alerting, dashboards, SLO tracking, and distributed tracing

- Experience building software systems as a software engineer

- Experience developing tooling and automation in Python and/or Bash

- Experience with CI/CD pipeline automation, specifically GitHub Actions

- Experience with disaster recovery practices and incident management

- Strong understanding of observability concepts - monitoring, logging, distributed tracing, and metrics - and how to apply them to production systems

- Experience with container orchestration and deployment technologies including AWS ECS Fargate and/or Kubernetes

- Experience working with relational databases (MySQL proficiency is a plus)

- Ability to take ownership of problems and act on them independently in a constantly evolving environment

- Experience with AIOps - using AI/ML-based tooling for anomaly detection, predictive alerting, or automated incident triage

- Familiarity with the reliability characteristics of AI/ML-backed services (e.g., LLM inference latency, non-determinism, prompt pipeline observability)

- Familiarity with compliance frameworks like SOC 2, ISO 27001, or NIST

- Hands-on experience using AI-assisted development tools (e.g., GitHub Copilot, Cursor, or similar) to accelerate automation, scripting, or infrastructure work

- Demonstrated use of AI/AIOps capabilities for reliability tasks - anomaly detection, incident triage, runbook generation, or alert noise reduction

- Familiarity with the operational characteristics of AI/ML-backed services and what it means to make them observable and reliable in production

- Demonstrated passion for AI through personal projects, contributions, or continuous learning in the context of infrastructure or reliability engineering

This role will receive a competitive base salary, benefits, and stock, typically in the form of Restricted Stock Units (RSUs). The applicable salary range for this role is: $166,900 - $225,900.

More jobs at Drata

Similar jobs near Hybrid - San Francisco (Remote)

Tell me when more Senior Site Reliability Engineer, Fleet Management jobs post near Hybrid - San Francisco (Remote) We re-check every listing against the employer’s own board — no résumé needed.

Search Senior Site Reliability Engineer jobs near Hybrid - San Francisco (Remote) → Browse all live jobs

This posting was published by Drata on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.