Incident Analyst Change Manager In Office
Philadelphia, PA - USA
Job Summary
Analyst III Reliability Operations
Role Summary
The Analyst III is a senior operational leader who acts as Incident Commander during major outages leads problem management efforts and reviews/approves complex changes for operational readiness. This role partners closely with SRE and engineering teams on reliability strategy.
Key Responsibilities
Incident Management
Lead high severity P1/P0 incidents as Incident Commander.
Coordinate cross functional engineering teams in real time.
Drive rapid troubleshooting impact assessment and resolution decisions.
Ensure high-quality incident documentation executive-ready summaries and follow through
Apply Technical knowledge of Application architecture flows in driving the incident towards mitigation
Problem Management
Lead problem investigations for major or recurring incidents.
Perform deep root cause analysis with engineering teams.
Validate corrective actions and track long-term problem remediation.
Present problem findings and preventive strategies to leadership.
Change Management
Review and approve high-risk or complex changes for operational readiness.
Participate in Change Advisory Board (CAB) when needed.
Validate rollback strategies and operational safety measures.
Lead change execution for major maintenance or reliability events.
Reliability Leadership
Drive reliability initiatives to reduce MTTA MTTR and incident volume.
Mentor junior analysts on incident handling and operational maturity.
Partner with SRE teams to expand observability automation and resilience.
Partner with SRE/Engineering teams on service reliability initiatives.
Lead maintenance events failover tests and resilience validation exercises.
Review and enhance runbooks automation workflows and monitoring strategies
Contribute to Automation & AI Ideas for improving efficiency & reduce MTTM /MTTR
Reliability Engineering Contributions
Perform deep post incident analysis to identify systemic issues.
Contribute to automation solutions and self healing systems.
Own service reliability dashboards and operational KPIs.
Leadership & Mentoring
Provide technical leadership to Analyst I/II team members.
Serve as a point of escalation for complex incidents.
Lead operational readiness reviews and training sessions.
Required Qualifications
35 years in SRE Operations Incident Management DevOps or related fields.
Expert knowledge of monitoring tools logging systems and incident response.
Strong troubleshooting skills: networking Linux/Windows servers cloud services.
Strong communication and leadership skills in high-pressure situations.
Preferred Qualifications
Hands-on experience with cloud platforms (AWS Azure GCP).
Automation experience (Python Go Bash or PowerShell).
Familiarity with microservices containers and distributed architectures.
Required Experience:
Manager
About Company
As a global leader, Wipro blends consulting and AI expertise across design, engineering and operations to accelerate business transformation and deliver future-ready technology.