Education: Bachelor's degree
Experience: 4+ years
Snowflake's Incident Management team is expanding. We're looking for a Senior Incident Manager to help us deliver outstanding service to our customers in their most critical moments - when a major incident is impacting their business and they need a calm, decisive partner driving toward resolution.
As a Senior Incident Manager, you thrive in a high-performing, fast-paced environment and bring a high degree of tact, patience, and composure under pressure. You are results-oriented, using data and metrics to make sound operational and tactical decisions while a major incident is unfolding. You lead with a positive outlook and a high degree of integrity, accountability, attention to detail, and strong planning and execution skills.
A thorough understanding of how technical issues translate into customer and business impact is essential.
This is a fixed overnight shift role: 3:00 PM - 11:00 PM Pacific Time. We're looking for someone who is committed to working these hours long-term. In addition to the shift, you'll participate in a weekend on-call rotation (currently about one weekend every few months, subject to change).
There is no additional weekday off-hours on-call expectation.
- Lead the response to critical, customer-impacting major incidents from engagement through resolution, acting as both Primary and Secondary Incident Manager as the situation requires.
- Own command and control of the incident bridge - asking probing questions to establish a clear incident statement and business impact, dealing in facts over speculation, and keeping the call structured and moving.
- Drive the team toward resolution: coordinate the recovery plan, set and hold timelines for task execution, ensure validations take place, and confirm and capture service restoration.
- Ensure the right people are engaged and present - quickly identifying when additional resources are needed and holding required participants on the call until they're released.
- Maintain a clear, accurate, well-structured incident summary and action items in the incident tracker throughout the event, and provide comprehensive handovers to incoming regional Incident Managers as part of our follow-the-sun model.
- Manage customer-facing and internal communications throughout the incident - clearly explaining the nature of the disruption, its impact on customer workloads, and the actions underway to resolve it.
- Maintain disciplined, regular communication cadences, building credibility through timely, accurate updates and responsiveness for the duration of the incident.
- Translate complex technical information into clear business impact, risk, and status that both customers and executive stakeholders can readily understand, while adhering to established voice and style standards.
- Communicate incident status effectively - verbally and in writing - to executives, sales teams, and other stakeholders, including concise summaries of large volumes of information.
- Own incident close-out and governance follow-ups: capture follow-up actions, confirm the engineering postmortem owner, and set clear expectations and timelines for the RCA before releasing the team.
- Build strong partnerships across the company - with Engineering, Product Management, Sales, and other teams - to deliver the best possible customer experience.
- Meet deliverable timelines tied to scheduled activities and events, such as customer, team, and executive updates.
- B.S. or M.S. degree in CS, MIS, or an equivalent discipline.
- Technical competency in cloud environments, data warehouse architectures, and software development methods.
- 4+ years of incident management, technical operations, SRE, or related experience, with a proven track record of delivering business value and improvement.
- 4+ years of experience working with Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform (GCP), or a private cloud environment.
- Strong call-leadership skills - a calm, confident verbal presence with the ability to lead a room of engineers and stakeholders under pressure without relinquishing control of the call.
- Experience writing customer-facing root cause analysis or postmortem reports, and producing clear customer-facing incident communications (e.g., status page updates) with accurate detail and timing.
- The ability to shift levels of communication between technical and non-technical audiences with ease.
- Technical understanding of software/platform/infrastructure (SaaS/PaaS/IaaS) architectures, their use, and management.
- Excellent verbal, written, communication, and receptive listening skills.
- High levels of emotional intelligence (EQ), empathy, and proactivity, with the ability to advocate for both customers and internal teams while striving for mutually beneficial solutions.
- Successful experience working, collaborating, and building relationships with leadership, colleagues, and clients.
- The ability to adapt, stay flexible, and learn quickly in a dynamic environment.
- Excellent teaming skills, comfortable working with virtual and global cross-functional teams.
- Excellent abilities in business productivity applications (Google Workspace preferred) for documents, spreadsheets, and presentations.
Search Senior Incident Manager jobs near Remote, WA → Browse all live jobs
This posting was published by Snowflake on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.