Software Engineering Manager Site Reliability Center
Pittsburgh, PA - USA
Job Summary
As a Software Engineering Manager for PNCs Site Reliability Engineering Center you will work within PNCs Information Technology Group and be located at one of our IT Hubs: Cleveland Ohio; Birmingham Alabama; Pittsburgh Pennsylvania; Dallas Texas; Denver Colorado or Phoenix Arizona and manage the daylight shift.
The Site Reliability Center (SRC) is focused on establishing a culture of operational excellence by ensuring infrastructure platforms and applications adhere to SRC onboarding standards that improve reliability enable proactive issue resolution and reduce customer impact. This role supports the vision of building a collaborative technology organization across application infrastructure and security teams to deliver a stable reliable and secure environment. Key responsibilities include driving customer-centric service improvements implementing proactive and preventative reliability practices fostering cross-functional collaboration enhancing monitoring and observability capabilities promoting a blameless culture of continuous learning and reducing operational toil through automation. The ideal candidate will help improve service performance strengthen operational resiliency and advance automation and observability initiatives that enhance the overall customer experience.
As a Software Engineering Manager Site Reliability Engineering (SRE) you will lead a team responsible for ensuring the reliability scalability and operational excellence of mission-critical platforms that power PNCs digital experiences. This role blends technical leadership hands-on problem solving and people management driving both production stability and continuous improvement across complex distributed systems. You will.
Manage SRE and related Teams; lead coach and develop a team of SRE engineers; set clear goals drive accountability and foster a culture of ownership and excellence; partner with cross-functional stakeholders to align technology and business objectives; support talent development performance management and succession planning; encourage innovation continuous learning and DevOps/SRE best practices.
Provide after-hours operational leadership and on-call support. Participate in an on-call leadership rotation supporting critical production services major incidents and high-severity customer-impacting events. Availability outside standard business hours including evenings weekends and holidays may be required to support incident response change events escalations and business continuity needs.
Lead incident management & remediation; manage and actively participate in end-to-end incident response for major (P1/P2) incidents; guide real-time triage diagnostics and troubleshooting across application infrastructure and network layers; ensure rapid execution of remediation actions and service restoration; provide clear timely communication to stakeholders during incidents; oversee post-incident analysis reporting and documentation to drive improvements.
Provide technical leadership in production support; serve as an escalation point for complex production issues; guide troubleshooting across: applications infrastructure (Linux/Windows) databases (Oracle SQL) middleware and integrations; ensure efficient log metric and system analysis; oversee batch/ETL monitoring and recovery processes; foster strong collaboration across engineering infrastructure and vendor teams.
Drive problem management & root cause resolution; lead root cause analysis (RCA) efforts for major and recurring incidents; ensure ownership and resolution of problem records; drive permanent fixes and systemic improvements to eliminate repeat issues identify trends and patterns to reduce risk and improve stability; partner with engineering teams to resolve code defects and system gaps and promote knowledge sharing via runbooks knowledge articles and error catalogs.
Oversee change management & release execution; ensure safe and compliant execution of production changes and releases; validate change readiness testing rollback strategies and risk assessments; represent the team in CAB reviews providing technical risk evaluation; oversee post-implementation reviews (CPIR) and ensure follow-through and drive
improvements in change success rate and reduction in production defects.
Advance monitoring alerting & observability; lead efforts to build and optimize monitoring dashboards and alerting frameworks champion use of tools such as Dynatrace BigPanda Logscale and enterprise platforms improve signal-to-noise ratio through alert tuning; enable proactive issue detection before customer impact; strengthen event management and observability practices.
Provide technical leadership in production support; serve as an escalation point for complex production issues; guide troubleshooting across: applications infrastructure (Linux/Windows) databases (Oracle SQL) middleware and integrations; ensure efficient log metric and system analysis; oversee batch/ETL monitoring and recovery processes; foster strong collaboration across engineering infrastructure and vendor teams.
Champion resiliency stability & availability; lead efforts to ensure high availability of critical systems; oversee disaster recovery failover and continuity testing; identify and eliminate single points of failure and drive improvements in MTTR uptime and service reliability.
Enable scalability & performance optimization; guide capacity planning and performance tuning strategies; ensure systems scale effectively under peak demand; partner with development teams for performance-driven design improvements; optimize system configurations to improve efficiency and throughput.
Lead a Global 24x7 Operation; manage distributed teams supporting critical systems around the clock. Participate in leadership escalation and on-call rotations to support major incidents critical production events and operational continuity.
Drive Automation & Operational Efficiency; identify and prioritize opportunities to reduce manual effort through automation; implement automation across: Incident remediation monitoring and alerting deployment and validation promote standardized runbooks and automation frameworks and improve operational metrics and reduce toil.
Ensure Governance Risk & Compliance; maintain adherence to enterprise policies and regulatory standards; support audits vulnerability remediation and risk controls; ensure accurate documentation and operational procedures and champion security access management and data governance practices
Qualifications:
5 years of related experience and 3 years of management experience.
Strong experience in Site Reliability Engineering Production Support or DevOps.
Proven ability to lead teams in high-availability enterprise environments
Deep understanding of incident problem and change management frameworks
Hands-on knowledge of monitoring tools cloud/infrastructure platforms and automation
Experience improving system reliability observability and operational maturity
Strong communication skills with the ability to lead during high-pressure situations.
Experience with OCP under infrastructure (Linux/Windows OCP)
MongoDB Cassandra under databases (Oracle SQL MongoDB Cassandra) and working knowledge of Elasticsearch Redis MQ and Kafka is a plus.PNC is an in-office company that fosters a supportive culture where employees can thrive and achieve balance. We encourage candidates to connect with their recruiter and hiring manager to understand workplace expectations and ensure the role aligns with their goals.PNC will not provide sponsorship for employment visas or participate in STEM OPT for this position.
- Manages development projects development teams and application support functions.
- Oversees multiple application programming and analysis projects which include development installation and maintenance of application programs.
- Monitors and maintains adherence and compliance to quality standards on an ongoing basis.
- Maximizes staff contribution through professional growth and development to increase teamwork and more effectively meet business needs.
- Analyzes applications to ensure that all systems that are developed meet business needs and specifications.
PNC Employees take pride in our reputation and to continue building upon that we expect our employees to be:
- Customer Focused - Knowledgeable of the values and practices that align customer needs and satisfaction as primary considerations in all business decisions and able to leverage that information in creating customized customer solutions.
- Managing Risk - Assessing and effectively managing all of the risks associated with their business objectives and activities to ensure they adhere to and support PNCs Enterprise Risk Management Framework.
PNC also has fundamental expectations of our people managers. As a manager of talent in PNC you will be expected to:
- Include Intentionally - Cultivates diverse teams and inclusive workplaces to expand thinking.
- Live the Values - Role models our values with transparency and courage.
- Enable Change - Takes action to drive change and innovation that will transform our business.
- Achieve Results - Takes personal ownership to deliver results. Empowers and trusts others in decision making.
- Develop the Best - Raises the bar with every talent decision and guides the achievement of all employees and customers.
Successful candidates must demonstrate appropriate knowledge skills and abilities for a role. Listed below are skills competencies work experience education and required certifications/licensures needed to be successful in this position.
To learn more about these and other programs including benefits for full time and part-time employees visit.
If an accommodation is required to participate in the application process please contact us via email at . Please include accommodation request in the subject line title and be sure to include your name the job ID and your preferred method of contact in the body of the email. Emails not related to accommodation requests will not receive responses. Applicants may also call and say Workday for accommodation assistance. All information provided will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations.
At PNC we foster an inclusive and accessible workplace. We provide reasonable accommodations to employment applicants and qualified individuals with a disability who need an accommodation to perform the essential functions of their positions.
PNC provides equal employment opportunity to qualified persons regardless of race color sex religion national origin age sexual orientation gender identity disability veteran status or other categories protected by law.
This position is subject to the requirements of Section 19 of the Federal Deposit Insurance Act (FDIA) and for any registered role the Secure and Fair Enforcement for Mortgage Licensing Act of 2008 (SAFE Act) and/or the Financial Industry Regulatory Authority (FINRA) which prohibit the hiring of individuals with certain criminal history.
Required Experience:
Manager