Firstup

Director of Cloud Operations

$200K–$228KFull-time · Remote - US
✓ Verified live on the employer's own system · added 115 days ago
Save search
Senior

Skills & tools

OperationsCloud PlatformsTeam LeadershipSecurityProcess Improvement
Apply on company site ↗ See your fit → free

Full job description

Job Summary: We are seeking a Director of Cloud Operations (CloudOps) to lead and evolve our cloud infrastructure and operational practices across a globally distributed SaaS platform. This is a hands-on leadership role responsible for ensuring the reliability, scalability, and efficiency of our systems running across multiple AWS regions in the United States and Europe.

As part of the senior leadership team, you will partner closely with Engineering, Security, and Product to strengthen operational excellence, enhance system observability, and drive continuous improvement in how we build and run services. You will lead a distributed team of engineers across the US and UK, fostering a high-performing, collaborative, and growth-oriented environment.

This role is ideal for a leader who combines deep technical expertise with a pragmatic approach to improving systems, processes, and team capabilities.

Job Summary: We are seeking a Director of Cloud Operations (CloudOps) to lead and evolve our cloud infrastructure and operational practices across a globally distributed SaaS platform. This is a hands-on leadership role responsible for ensuring the reliability, scalability, and efficiency of our systems running across multiple AWS regions in the United States and Europe.

As part of the senior leadership team, you will partner closely with Engineering, Security, and Product to strengthen operational excellence, enhance system observability, and drive continuous improvement in how we build and run services. You will lead a distributed team of engineers across the US and UK, fostering a high-performing, collaborative, and growth-oriented environment.

This role is ideal for a leader who combines deep technical expertise with a pragmatic approach to improving systems, processes, and team capabilities.

Firstup expects the base salary for this role to be between $200,000-$228,000. The starting rate of pay may vary based on factors including, but not limited to, position offered, location, education, training, and/or experience.

Own the availability, performance, and resilience of our multi-region AWS platform.

Drive improvements in system reliability through well-defined SLIs/SLOs , error budgets, and proactive engineering practices.

Lead efforts to reduce MTTR and improve incident response effectiveness across the organization.

Guide architecture decisions for microservices, Kubernetes (EKS), and serverless workloads to ensure scalability and fault tolerance.

Advance our observability strategy using Datadog , ensuring actionable insights across infrastructure and applications.

Establish and refine incident management practices, including on-call processes, escalation paths, and post-incident reviews.

Act as an incident commander for critical events and contribute to the on-call rotation.

Elevate operational standards through automation, standardization, and adoption of modern best practices.

Drive cost optimization initiatives across AWS environments without compromising performance or reliability.

Leverage AI and automation to improve operational efficiency, accelerate root cause analysis, and enhance system insights.

Continuously improve CI/CD pipelines (CircleCI) and infrastructure-as-code practices (Terraform).

Lead, mentor, and support a distributed team of CloudOps engineers across the US and UK.

Foster a culture of accountability, learning, and continuous improvement.

Provide technical guidance while enabling the team to grow in ownership and capability.

Ensure stability and support for existing customers while maintaining clear operational boundaries with the cloud platform.

  • Experience

10+ years in cloud infrastructure, SRE, or DevOps roles, with 3+ years experience leading CloudOps/SRE teams .

Proven track record of leading operational or platform transformations in a SaaS environment.

Experience operating multi-region, customer-facing systems at scale .

Solid understanding of microservices and distributed systems design.

Familiarity with serverless architectures and modern cloud-native patterns.

Deep experience with incident management , on-call operations, and reliability engineering practices.

Strong understanding of SLO/SLI frameworks , monitoring strategies, and performance optimization.

Demonstrated ability to balance hands-on technical work with team leadership .

Collaborative, pragmatic leader who can influence across teams and functions.

Focus on continuous improvement, with a bias toward measurable outcomes.

More jobs at Firstup

Similar jobs near Remote - US

Tell me when more Director, Audience & Ad Solutions GTM jobs post near Remote We re-check every listing against the employer’s own board — no résumé needed.

Search Director of Cloud Operations jobs near Remote - US → Browse all live jobs

This posting was published by Firstup on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.