Sr Mgr, Reliability Engineering Mgmt
Job Summary
Team :
Our Site Reliability Engineering (SRE) team consists of highly skilled engineers responsible for maintaining and enhancing the reliability scalability and performance of the ServiceNow infrastructure. Our SREs are empowered to resolve technical issues across the entire technology stack from hardware to applications. Additionally they work to improve the platforms operability aiming to reduce the number of incidents and minimize Mean Time to Recovery (MTTR). To achieve this the team combines software development networking database and systems engineering skills to tackle complex problems striving to maintain our platform operating for our customers.
Role
As a Sr Manager Site Reliability Engineering at ServiceNow youll lead a team of SRE leaders managers and engineers focused on ensuring the reliability and availability of critical enterprise platforms and applications while helping drive our cloud modernization journey through automation operational excellence resilience and continuous improvement.
Do you:
have a technical background in roles like systems engineering or devops or site reliability engineering
know operating systems in various levels of troubleshooting and diagnostics
have experience with cloud technologies and hyperscalers such as AWS GCP or Azure and a passion for driving cloud modernization and cloud-native transformation
have low tolerance to repetitive tasks and automate your way through work
have experience in leading a team of engineers and exposure to people management
If you Answered yes to these questions we want to hear from you. Hit the Apply button and lets have a chat about the role and your skills and experiences.
As a Sr Manager of the SRE team your responsibilities will be:
Lead and develop a global team of SRE leaders managers and engineers with accountability for talent development performance prioritization succession planning and execution.
Define and drive the SRE strategy and operating model across reliability observability automation incident response production readiness and continuous improvement.
Own and evolve observability capabilities across metrics logs traces alerting SLI/SLOs error budgets and service health practices to improve detection diagnosis and overall production reliability.
Lead the strategy and execution for production-like staging environments ensuring critical services and releases can be validated in environments that closely represent production.
Partner with Engineering and Release teams to strengthen release pipelines and confidence gates for Now releases including smoke integration resiliency performance rollback and production-readiness validation.
Drive a culture of eliminating repetitive operational work through automation orchestration self-healing and AI-assisted operations shifting reliability practices earlier in the software-development lifecycle.
Lead cloud modernization initiatives modernizing legacy infrastructure tooling and operational practices toward cloud-native architectures and scalable engineering patterns.
Drive adoption and operational excellence across AWS GCP and Azure hyperscalers including architecture scaling resiliency availability cost efficiency and operational readiness.
Advance containerization and Kubernetes-based operating models including workload resiliency scalability orchestration deployment patterns observability and lifecycle management.
Partner with Product Engineering Platform Engineering Release Engineering Security and Infrastructure teams to improve the reliability and availability of critical enterprise platforms and services.
Establish reliability standards and measurable outcomes around SLIs/SLOs error budgets availability MTTR change failure rate automation and operational readiness.
Provide leadership during major incidents while ensuring incident and problem-management learnings are translated into durable engineering improvements automation and preventative controls.
Establish follow-the-sun and on-call SRE practices that keep engineers close to production signals and use operational learnings to drive shift-left improvements.
Evaluate existing platforms processes and technologies and drive simplification modernization standardization and operational efficiency.
Identify and adopt emerging technologies that support the organizations reliability and cloud-modernization goals while building the skills and capabilities required to operate them at scale.
Build a global engineering culture that values technical excellence accountability collaboration continuous learning and diverse perspectives.
Qualifications :
To be successful in this role you have:
Significant experience leading Site Reliability Engineering Production Engineering DevOps Platform Engineering or Cloud Infrastructure organizations in large-scale production environments.
5 years of people-management experience including experience leading managers senior technical leaders and geographically distributed engineering teams.
Strong experience with cloud modernization and migration including modernizing legacy platforms and tooling into cloud-native architectures.
Deep working knowledge of one or more major hyperscalers - AWS GCP or Azure with experience designing and operating resilient scalable highly available production systems.
Strong understanding of Kubernetes containers orchestration service networking autoscaling and modern cloud-native architecture patterns.
Experience designing and operating observability platforms at scale including metrics logging tracing alerting dashboards golden signals SLI/SLOs and error budgets.
Experience building or operating production-like staging/test environments and establishing validation strategies for reliability resiliency performance integration and release readiness.
Experience with CI/CD and release pipelines including automated confidence gates progressive/phased deployments zero-downtime deployment strategies rollback mechanisms and production-readiness controls.
Experience applying automation orchestration and infrastructure-as-code to reduce operational toil and improve repeatability and reliability.
Experience leveraging or critically evaluating AI and AI-assisted operations to automate workflows accelerate diagnosis and remediation and improve engineering productivity.
Strong technical foundation across Linux distributed systems databases networking systems troubleshooting scripting and software engineering fundamentals.
Experience operating high-scale software platform and infrastructure-as-a-service environments with demanding availability and reliability requirements.
Strong understanding of incident management problem management operational readiness and continuous improvement practices.
Demonstrated ability to influence organizational boundaries and drive alignment among Engineering Product Architecture Security Infrastructure and Operations teams.
Excellent written and verbal communication skills with the ability to translate complex technical topics into clear business outcomes and leadership decisions.
Ability to lead effectively through ambiguity and change while maintaining a strong focus on execution customer impact and engineering excellence.
We also have pluses. These are not a must but please highlight them on your resume if you have:
RHCE CCNA ITIL or other industry certifications
Additional Information :
Work Personas
We approach our distributed world of work with flexibility and trust. Work personas (flexible remote or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race color religion sex sexual orientation national origin age disability gender identity veteran status or any other category protected by addition all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process or are unable to use this online application and need an alternative method to apply please contact for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations including the U.S. Export Administration Regulations (EAR) ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. 2026 Fortune Media IP Limited. All rights reserved. Used under license.
Remote Work :
No
Employment Type :
Full-time
About Company
Learn here. Grow here. Make a difference here. At ServiceNow, our cloud?based platform and solutions deliver digital workflows that create great experiences and unlock productivity for employees and enterprises. Were growing fast, innovating even faster, and making an impact on our c ... View more