Experience: 10+ years
The Head of Enterprise Monitoring and Observability is responsible for defining and executing the enterprise observability strategy, operating model, and platform roadmap that enable proactive, predictive, and resilient technology operations.
This leader owns the enterprise capabilities for monitoring, telemetry, event intelligence, operational analytics, and observability engineering, ensuring technology services are observable, measurable, actionable, and continuously improving.
The role serves as the operational intelligence leader across Enterprise Technology, providing the platforms, standards, insights, and governance required to improve service reliability, accelerate issue detection and resolution, and support modern operational practices.
This leader partners closely with infrastructure, cloud, middleware, application, security, resiliency, and service management teams to build a unified enterprise-wide observability capability.
- Define and execute the enterprise observability vision, strategy, operating model, and multi-year roadmap. - Establish standards and governance for monitoring, logging, metrics, tracing, telemetry, alerting, and service observability. - Lead the evolution of observability capabilities from reactive monitoring toward proactive, predictive, and automated operations. - Drive enterprise adoption of modern observability practices, frameworks, and engineering standards.
- Own the strategy, architecture, lifecycle, and operation of enterprise monitoring and observability platforms. - Lead platform modernization, tool rationalization, capability expansion, and service offering improvements. - Manage vendor relationships, contracts, platform investments, and cost optimization efforts. - Ensure scalability, availability, resilience, and ongoing enhancement of observability tooling and services.
- Establish enterprise monitoring and instrumentation standards across infrastructure, cloud, middleware, applications, databases, and digital services. - Improve monitoring coverage, telemetry quality, alert effectiveness, signal quality, and operational visibility. - Drive adoption of distributed tracing, service dependency mapping, and end-to-end transaction monitoring. - Ensure service health and operational data are consistently available across the technology landscape.
- Lead implementation of operational analytics, event intelligence, anomaly detection, predictive insights, and AIOps capabilities. - Develop dashboards, scorecards, and reporting that provide actionable operational intelligence to technology leadership. - Enable intelligent event correlation, noise reduction, automated diagnostics, and automated operational workflows. - Partner with engineering and operations teams to use observability data for continuous improvement and outage prevention.
- Partner with technology teams to improve detection, response, recovery, and root cause analysis capabilities. - Establish observability practices that improve service reliability, operational resilience, and customer experience. - Drive post-incident learning and identification of automation opportunities. - Provide enterprise visibility into technology health, performance trends, capacity risks, and reliability risks.
- Build, lead, and develop a high-performing team of observability engineers, platform engineers, and operational intelligence specialists. - Foster a culture of innovation, accountability, operational excellence, continuous learning, and cross-functional partnership. - Lead organizational adoption of observability best practices through enablement, coaching, and stakeholder engagement. - Establish strong partnerships across technology, cybersecurity, architecture, engineering, operations, resiliency, and service management functions.
- 10+ years of experience in infrastructure operations, platform engineering, observability, reliability engineering, or related technology disciplines. - 5+ years of experience leading enterprise-scale engineering or operational teams. - Deep expertise in observability platforms, monitoring technologies, telemetry pipelines, logging, metrics, distributed tracing, and operational analytics. - Demonstrated experience implementing enterprise monitoring strategies and improving operational reliability at scale. - Strong understanding of cloud platforms, hybrid infrastructure, automation, DevOps practices, and modern operational architectures. - Experience leading organizational transformation and driving adoption of new operational capabilities. - Excellent leadership, communication, stakeholder management, and strategic planning skills.
- Increased enterprise monitoring and observability coverage. - Reduced mean time to detect (MTTD) and mean time to resolve (MTTR). - Improved signal quality and reduced alert fatigue. - Increased adoption of observability standards and platform capabilities. - Improved service reliability, operational resilience, and operational efficiency. - Expansion of predictive analytics, automation, and AIOps capabilities across Enterprise Technology.
Guardian Life is not currently or in the foreseeable future sponsoring employment-based visas (e.g., such as an H-1B). In order to be a successful applicant, you must be legally authorized to work in the United States, without the need for employer sponsorship/support now or at any time in the future.
Search Head of Enterprise Monitoring and Observability jobs near New York → Browse all live jobs
This posting was published by Guardianlife on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.