Salesforce

Director, Site Reliability Engineering

Full-time · New York - New York
✓ Verified live on the employer's own system · added 4 days ago
Save search
Senior · 10+ yrs exp

Requirements

Education: Bachelor's degree or related field

Experience: 10+ years

Skills & tools

OperationsDevopsSecurityTeam LeadershipFinancial AnalysisProgrammingRecordkeepingDistributed Systems
Apply on company site ↗ See your fit → free

Full job description

Job Title: Director, Site Reliability Engineering Location: New York, NY; San Francisco, CA; Dallas, TX

We are looking for a Director of Site Reliability Engineering to spearhead the evolution of our reliability, observability, and operational engineering capabilities.

In this role, you will transform our SRE function-moving our engineering organization from reactive incident response to a proactive, automated, and data-driven reliability culture. Partnering closely across Application Engineering, Platform, Architecture, Security, Infrastructure, and Product, you will ensure our services are resilient, observable, scalable, and production-ready long before they launch.

As an impactful people leader with sharp technical judgment, you will directly manage and empower a core team of ~6 engineers while driving cross-functional alignment across a complex organization. You won't just run existing playbooks; you will define the strategy, tooling, automation, and culture needed to mentor your team and scale system reliability enterprise-wide.

- Define and execute the long-term strategy and roadmap for Site Reliability Engineering.

- Establish a clear operating model for SRE, including team scope, engagement models, ownership boundaries, and success measures.

- Build and develop a high-performing team of site reliability and operations engineers.

- Modernize the SRE function through automation, AI-assisted operations, self-service capabilities, and engineering-first practices.

- Translate business priorities and customer impact into clear reliability investments and engineering outcomes.

- Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs.

- Establish service-level indicators, service-level objectives, error budgets, and reliability standards for critical services.

- Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems.

- Define what it means for a service to be operationally and observably ready for production.

- Develop readiness reviews and certification practices for high-impact services and launches.

- Drive improvements in system availability, performance, resiliency, and recovery.

- Ensure reliability requirements are incorporated throughout the software development lifecycle rather than addressed only after deployment.

- Define an enterprise observability strategy spanning metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry.

- Establish common instrumentation, telemetry, dashboards, alerting, and service-health standards.

- Reduce fragmented or duplicative observability implementations by promoting shared patterns and reusable capabilities.

- Improve end-to-end visibility across distributed systems, customer journeys, services, and infrastructure.

- Partner with engineering teams to ensure telemetry is actionable, contextual, and tied to customer and business outcomes.

- Establish governance and measurement to assess adoption and effectiveness of observability standards.

- Improve incident detection, response, mitigation, communication, and learning.

- Lead the transition from manual and reactive operations toward automated detection, diagnosis, remediation, and incident creation.

- Reduce mean time to detect, acknowledge, mitigate, and recover.

- Improve on-call practices, escalation paths, runbooks, and operational ownership.

- Establish blameless post-incident review practices that produce measurable engineering improvements.

- Identify recurring sources of operational toil and create plans to eliminate or automate them.

- Partner with engineering leaders to ensure actions from incidents are prioritized and completed.

- Develop a roadmap for intelligent operations, including anomaly detection, event correlation, automated triage, assisted root-cause analysis, and remediation.

- Evaluate opportunities to use agents and AI-assisted workflows across observability, incident response, capacity planning, and operational support.

- Build automation that reduces cognitive load and improves the speed and consistency of operational decisions.

- Ensure automation is safe, measurable, auditable, and designed with appropriate human oversight.

- Promote platform and self-service approaches that allow product teams to adopt reliability practices with minimal friction.

- Partner with engineering, DevOps, and business stakeholders.

- Influence teams that do not directly report into SRE and build shared accountability for production outcomes.

- Create clear service ownership models and operational expectations across teams.

- Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination.

- Communicate reliability posture, risks, trends, and investments to executive and technical audiences.

- Improved availability and reliability of critical services.

- Reduced time to detect, diagnose, mitigate, and recover from incidents.

- Increased percentage of services meeting observability and production-readiness standards.

- Reduced alert noise, operational toil, and manual incident-management activity.

- Increased adoption of service-level objectives and measurable reliability practices.

- Improved quality and completion rate of post-incident corrective actions.

- Increased automation across detection, triage, remediation, and operational workflows.

- Stronger ownership of production reliability across engineering teams.

- Comfortable challenging legacy operating models and assumptions.

- Able to move between technical detail and executive-level strategy.

- Builds trust through clarity, accountability, and strong partnership.

- Develops leaders and creates an inclusive, high-performance engineering culture.

- Treats incidents as opportunities to improve systems rather than assign blame.

- Brings urgency to operational risks while maintaining focus on sustainable solutions.

- Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field; Master's degree or MBA preferred.

- 10+ years of progressive engineering experience, including 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teams

- Proven experience building or transforming a reliability or operational engineering organization.

- Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery.

- Experience establishing observability, incident-management, service-level objective, and production-readiness practices.

- Demonstrated ability to improve reliability through engineering and automation rather than process alone.

- Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems.

- Strong understanding of modern telemetry, including metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring.

- Demonstrated success driving alignment and building consensus across cross-functional engineering teams and executive stakeholders

- Ability to balance immediate operational needs with long-term engineering transformation.

- Strong written, verbal, and executive communication skills.

- Strong experience operating large-scale systems in AWS or another major cloud environment.

- Proven track record with observability platforms such as New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, or OpenTelemetry.

- Demonstrated experience implementing OpenTelemetry or common instrumentation standards.

- Verified proficiency building internal developer platforms, paved roads, or self-service reliability capabilities.

- Experience applying AI, machine learning, or agent-based automation to operational workflows.

- Seasoned capability with chaos engineering, resilience testing, disaster recovery, capacity planning, and performance engineering.

- Software engineering experience and the ability to engage deeply in architecture and design discussions.

- Solid background in supporting high-profile launches, events, or systems with significant customer and business impact.

The typical base salary range for this position is $197,300 - $313,700 annually. In select cities within the San Francisco and New York City metropolitan area, the base salary range for this role is $237,700 - $344,700 annually.

More jobs at Salesforce

Similar jobs near New York - New York

Search Director, Site Reliability Engineering jobs near New York - New York → Browse all live jobs

This posting was published by Salesforce on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.