Incident Lead
Job Summary
About the Job: We are seeking an experienced Incident Lead with Scrum Master capabilities to lead the response coordination and governance of production incidents across cross-functional technology teams. The successful candidate will own critical incident execution SWAT queue health stakeholder communication and service restoration while applying Agile practices to improve team flow accountability and continuous improvement.
Office Location: Toronto
Employment Type: Permanent
Role Type: New position - current requirement
Work Arrangement: Hybrid (2 days in office per week)
Position Responsibilities:
Incident Leadership & Response Management
Lead the end-to-end management of critical production incidents from initial triage through service restoration stakeholder communication root-cause review and closure.
Establish incident command confirm severity and business impact assign clear ownership and coordinate application engineering infrastructure security product and vendor teams.
Drive timely resolution of critical tickets within agreed SLAs and escalate risks blockers and resource constraints appropriately.
Maintain accurate incident timelines decisions actions dependencies and recovery updates throughout the incident lifecycle.
Remove production support bottlenecks and enable rapid decision-making during high-priority incidents.
Own daily ticket triage and the SWAT queue ensuring incidents and support tickets are correctly categorized prioritized assigned and progressed.
Monitor ticket ageing stalled work recurring issues capacity constraints and ownership gaps to maintain a manageable backlog.
Balance urgent restoration work with defects service requests technical debt and preventive improvement initiatives.
Improve ticket throughput and backlog hygiene while maintaining quality compliance and operational controls.
Facilitate daily SWAT stand-ups sprint planning backlog refinement retrospectives service reviews and operational governance meetings.
Coach support and engineering teams on Scrum and Agile practices suited to production support and interrupt-driven work.
Partner with product owners and service owners to maintain a prioritized transparent backlog with clear acceptance criteria and ownership.
Identify and remove team impediments manage dependencies support capacity planning and improve delivery flow across teams.
Use retrospectives and operational data to implement measurable improvements in incident response and support delivery.
Track and report SLA compliance mean time to acknowledge mean time to resolution (MTTR) ticket ageing throughput backlog health critical incident volume and recurrence trends.
Prepare dashboards and scorecards that provide leadership with clear visibility into service performance operational risks bottlenecks and improvement actions.
Facilitate incident and operational governance reviews ensuring decisions escalations risks and action items are documented and closed on time.
Promote cross-team accountability through clear owners target dates escalation paths and transparent follow-through.
Lead post-incident reviews and root-cause analysis for major and recurring incidents without creating a blame-focused environment.
Ensure corrective and preventive actions are prioritized tracked and implemented to reduce recurring incidents.
Identify trends and systemic weaknesses then partner with technology teams to improve resilience monitoring automation and support readiness.
Drive continuous improvement in incident processes escalation models runbooks communications and service management practices.
Required Qualifications:
8 years of experience leading production support and incident management teams including coordinating the triage prioritization and resolution of software incidents in an enterprise technology environment.
Demonstrated Scrum Master experience including facilitation of Agile ceremonies backlog governance impediment removal coaching and continuous improvement.
Proven ability to coordinate high-severity incidents across application engineering infrastructure security product business and vendor teams.
Hands-on experience with ticket triage incident queues escalation management root-cause analysis and corrective-action tracking.
Working knowledge of SLA MTTR ticket ageing throughput backlog health and other production support metrics.
Strong stakeholder communication facilitation decision-making conflict-resolution and executive reporting skills.
Ability to remain composed establish accountability and drive outcomes in high-pressure and time-sensitive situations.
Experience managing cross-functional and geographically distributed teams.
Preferred Qualifications:
Experience supporting enterprise applications microservices integrations and cloud environments such as AWS Microsoft Azure or Google Cloud Platform.
Familiarity with ITIL practices DevOps CI/CD pipelines observability monitoring and modern production support workflows.
Experience building operational dashboards and scorecards using data from service management and delivery platforms.
Proficiency with tools such as ServiceNow Confluence or similar incident and collaboration platforms.
Preferred certifications include ITIL Certified Scrum Master (CSM) Professional Scrum Master (PSM) SAFe Scrum Master PMP or PRINCE2.
Salary Range: $90000 to $93000 CAD/ year
The final compensation offered will depend on local market conditions and geographic location as well as job-related factors such as the candidates knowledge skills qualifications relevant experience and education/training. Compensation may also include additional components such as benefits and/or other incentives where accordance with new employment standards requirements we retain copies of this job posting and applicant information for three (3) years after the posting is removed. We do not use AI technology; all applications are also reviewed by our recruitment team.
Infoya is an equal opportunity employer committed to diversity and inclusion. We welcome applications from all qualified individuals regardless of race color religion sex sexual orientation gender identity national origin age disability protected veteran status aboriginal status or any other legally protected factors.