Director, Site Reliability Engineering & Service Enablement
Santa Clara County, CA - USA
Job Summary
Team:
Our Site Reliability Engineering (SRE) team consists of highly skilled engineers responsible for maintaining and enhancing the reliability scalability and performance of the ServiceNow infrastructure. Our SREs are empowered to resolve technical issues across the entire technology stack from hardware to applications. Additionally they work to improve the platforms operability aiming to reduce the number of incidents and minimize Mean Time to Recovery (MTTR). To achieve this the team combines software development networking database and systems engineering skills to tackle complex problems striving to maintain our platform operating for our customers.
Role:
We are looking for a Director of Site Reliability Engineering to lead the next phase of our reliability transformation as ServiceNow modernizes toward a cloud-agnostic cloud-ready production platform.
This leader will own key elements of the SRE operating model across Reliability Engineering Service Enablement Service Registry SLI/SLO standards reliability governance automation AI-enabled operations and production readiness. The role will lead a global engineering organization and partner across Product Engineering Infrastructure Architecture Security Release Engineering and Customer Support to establish consistent reliability practices across ServiceNow products and services.
The Director will play a critical role in evolving the organization from reactive operations toward an engineering-led SRE model focused on prevention automation resilience and continuous improvement.
What you get to do in this role:
- Define and execute the SRE strategy and operating model across reliability engineering service enablement observability automation incident learning and production readiness.
- Lead and develop a global organization of engineering managers technical leaders and SREs.
- Establish enterprise reliability standards for service ownership tiering golden signals SLIs/SLOs error budgets alerting on-call practices and service health reviews.
- Lead the Service Enablement strategy by establishing minimum reliability requirements and maturity standards for critical services.
- Own the Service Registry strategy improving service ownership dependency visibility maturity tracking and impact-aware operational decision-making.
- Drive adoption of SLIs SLOs error budgets and burn-rate alerting across critical services ensuring teams consistently use reliability signals to manage customer impact.
- Build a culture of engineering away toil by turning recurring operational work and incident patterns into automation self-service and systemic fixes.
- Establish the AI-enabled SRE roadmap including change-risk assessment operational insights remediation recommendations and policy-driven automation.
- Drive reliability and production-readiness strategy across AWS Azure and GCP by establishing cloud-agnostic patterns while addressing hyperscaler-specific operational requirements.
- Partner with product and platform engineers to design launch and operate reliable services throughout the production lifecycle.
- Establish launch and production-readiness practices that validate availability latency performance capacity dependencies rollback and recovery before customer impact.
- Drive sustainable operations by scaling self-service capabilities automation platforms and systemic reliability improvements across engineering teams.
- Lead incident response blameless postmortems and corrective actions that convert production failures into lasting reliability improvements.
- Measure reliability through SLIs SLOs error budgets golden signals change failure rate MTTR capacity health and toil reduction.
- Influence architecture and platform direction to simplify operating models and improve reliability across ServiceNows global infrastructure.
- Partner with executive and engineering leaders to prioritize reliability investments and drive adoption beyond the direct SRE organization.
Qualifications :
To be successful in this role you have:
- Experience in leveraging or critically thinking about how to integrate AI into work processes decision-making or problem-solving. This may include using AI-powered tools automating workflows analyzing AI-driven insights or exploring AIs potential impact on the function or industry.
- 12 years of significant leadership experience in Site Reliability Engineering Production Engineering Platform Engineering Cloud Infrastructure or large-scale distributed systems with a Bachelors degree; or 8 years and a Masters degree; or a PhD with 5 years experience; or equivalent experience.
- Proven success leading managers and senior technical leaders across geographically distributed engineering organizations.
- Demonstrated success leading SRE infrastructure or reliability transformation at scale.
- Strong understanding of SLIs/SLOs error budgets observability incident management reliability governance and on-call practices.
- Experience with service catalogs service registries service ownership models Backstage CMDB dependency mapping or service topology.
- Strong background in cloud infrastructure and modernization across AWS Azure and/or GCP.
- Understanding of Kubernetes distributed systems networking databases infrastructure automation and cloud-native architecture.
- Experience driving automation through orchestration Infrastructure as Code self-service platforms and auto-remediation.
- Familiarity with AI-assisted operations autonomous remediation or agentic technologies is highly desirable.
- Experience establishing production-readiness practices for releases resilience disaster recovery infrastructure changes and cloud migrations.
- Ability to use incident reliability and operational data to prioritize engineering work and drive systemic improvements.
- Strong cross-functional influence and executive communication skills.
- Ability to operate effectively through ambiguity organizational transformation and large-scale technical change.
For positions in this location we offer a base pay of $221200 - $387100 plus equity (when applicable) variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline and individual total compensation will vary based on factors such as qualifications skill level competencies and work location. We also offer health plans including flexible spending accounts a 401(k) Plan with company match ESPP matching donations a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.
Additional Information :
Work Personas
We approach our distributed world of work with flexibility and trust. Work personas (flexible remote or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race color religion sex sexual orientation national origin age disability gender identity veteran status or any other category protected by addition all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process or are unable to use this online application and need an alternative method to apply please contact for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations including the U.S. Export Administration Regulations (EAR) ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. 2026 Fortune Media IP Limited. All rights reserved. Used under license.
Remote Work :
No
Employment Type :
Full-time
About Company
Learn here. Grow here. Make a difference here. At ServiceNow, our cloud?based platform and solutions deliver digital workflows that create great experiences and unlock productivity for employees and enterprises. Were growing fast, innovating even faster, and making an impact on our c ... View more