Incident and Escalations Manager 3 US-West
San Francisco, CA - USA
Job Summary
The Incident and Escalation Management team (IEM) is part of Datadogs Global Support Engineering (GSE) organization. IEM exists to continuously improve Datadogs overall customer experience during incidents and other critical moments. Were looking for experts with a background in incident management and escalation handling to provide fast incident response clear ownership and calm confident communication for our global customer base.
In this individual contributor role you will play a key part in customer communication technical engagement and incident management for our global customers. You will also implement processes and automations use the same incident and case management tools we build for our customers and help shape how they evolve.
At Datadog we place value in our office culture the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.
What Youll Do:
- Drive incidents and escalations to resolution from start to finish including triage coordination with our support engineering and product teams. Own the response and keep customers and internal teams informed throughout.
- Independently coordinate cross-functional stakeholders through complex incidents and escalations making the call on prioritization response strategy and when to pull in additional teams or leadership.
- Own communications during active incidents including sensitive large-scale and publicly visible events adapting the message for customers internal teams and leadership.
- Lead projects that deliver operational improvements to how IEM detects manages and communicates during incidents and escalations as Datadog and its products evolve.
- Create maintain and improve incident and escalation documentation runbooks and training materials and lead retrospectives that turn individual incidents into durable process change.
- Mentor peers on incident and escalation management best practices and contribute to interviewing and hiring.
- Design and implement real-time and proactive monitoring of customer infrastructure to surface customer-impacting risks before they escalate.
- Identify recurring cloud infrastructure and product issues across incidents and escalations and partner with Engineering to feed findings back and reduce recurrence.
Who You Are:
- 5 years of related professional experience including experience in incident management and customer escalations with demonstrable ownership of complex incidents and critical customer situations; experience in a SaaS cloud or observability environment is a plus
- 2 years of experience in a hands-on technical role at a software or cloud company such as Support Engineering Site Reliability Engineering Solutions Architecture or a similar technical function
- Strong familiarity with cloud computing and modern software architectures; scripting ability in Python JavaScript or shell is a plus
- Ability to correlate behaviors across known system interdependencies assess customer and technical impact and make sound decisions while remaining calm in high-pressure and ambiguous situations
- Experience independently coordinating cross-functional stakeholders and driving complex incidents and escalations toward resolution
- Experience driving projects from conception through delivery with strong problem-solving skills and the ability to operate effectively in a fast-paced environment
- Strong written and verbal English communication skills with the ability to communicate clearly with both technical and customer-facing audiences
- Bachelors degree in Computer Science Information Science/Technology Engineering or equivalent practical experience
Datadog values people from all walks of life. We understand not everyone will meet all the below qualifications on day one. Thats okay. If youre passionate about technology and want to grow your skills we encourage you to apply.
Benefits and Growth:
- Generous and competitive US benefits
- New hire stock equity (RSUs) and employee stock purchase plan
- Continuous career development and pathing opportunities
- Product training to develop an in-depth understanding of our product and space
- Best in breed onboarding
- Internal mentor and buddy program cross-departmentally
- Friendly and inclusive workplace culture
Benefits and Growth listed above may vary based on the country of your employment and the nature of your employment with Datadog.
#LI-Hybrid
Required Experience:
Manager