Enter a job title or keyword

Site Reliability Engineering Lead, Product Enablement


Job Location:

Toronto - Canada

Monthly Salary: Not provided by the employer
Posted: 7 October 2026 (2 days ago)
Application Deadline: 4 January 2027
Vacancies: 1 Vacancy

Job Summary

The opportunity


The Site Reliability Engineering Lead (SREL) is accountable for the reliability health and operational sustainability of a broader portfolio of applications data pipelines platforms and services. The role proactively identifies systemic and recurring risks and leads or directs technical initiatives including the design and implementation of monitoring automation operational support tooling and selected code to reduce incidents manual intervention cost and operational burden. The SREL serves as a senior technical advisor and escalation point translating complex issues options risks and trade-offs into clear business language for stakeholders.



Who youll work with

  • Reports to: Director Product / Data Enablement
  • Works closely with: Site Reliability Engineering Specialists Product Engineering Leads (PEL) Data Solutions Leads Data Platform Leads Enterprise Architecture Security Technology Services business stakeholders and third-party vendors.
  • Collaborates regularly with Product Engineering Data Solutions and Data Platform teams to support smooth transitions from build to run and to identify and resolve recurring reliability and supportability issues. Depending on the nature and complexity of the change the Lead may implement or direct selected code configuration automation or process changes through development testing release and post-implementation validation or work with the appropriate team to secure prioritization ownership and completion.
  • Partners with Technology Services teams including Infrastructure Cloud and Platform teams to troubleshoot complex and cross-service issues and to design and implement or direct the implementation of observability automation resilience and operational support tooling that improves service quality and reduces operational burden.
  • Frequent communication with business stakeholders is required to provide transparency into system health incident status systemic risks and improvement priorities adapting the level of technical detail to the audience including senior business stakeholders when required.
What youll do

Service Reliability Strategy and Ownership

  • Own and continuously improve the reliability supportability and operational efficiency of a broad portfolio of applications data pipelines platforms and services including availability latency performance resilience and cross-service dependencies.
  • Define and govern Service Level Availability (SLAs) and related reliability metrics for the portfolio in accordance with enterprise standards and business expectations and drive changes where performance or operational risk warrants.
  • Establish and maintain a prioritized reliability roadmap incorporating application data pipeline and service health plans balancing business criticality operational risk support effort capacity cost and technology priorities
  • Use deep technical and trend analysis across incidents alerts service performance capacity storage cost and support effort to identify systemic failure modes and lead cross-service improvement initiatives through implementation and outcome validation.

Operational Excellence

  • Lead response to high-impact complex or cross-service incidents coordinating technical teams and ensuring timely audience-appropriate stakeholder communication.
  • Drive Root Cause Analysis and problem-management activities beyond minimum process requirements identify systemic technical architectural monitoring process and organizational contributing factors and challenge recurring issues that have become normalized as operational support work. Ensure corrective or preventive actions are completed and validated through permanent remediation automated mitigation or appropriately authorized risk acceptance.
  • Own the end-to-end resolution of selected recurring production defects and supportability issues. Where appropriate implement or direct code configuration scripting automation and process changes through technical review testing release and post-implementation validation and drive changes requiring Product or Data ownership through prioritization and portfolio-level capacity performance resilience and recovery analysis and improve Mean Time to Detect (MTTD) and Mean Time to Restore through monitoring diagnostics runbooks automation and operational learning.
  • Contribute to Disaster Recovery (DR) and resilience plans and participate in recovery testing.

Observability & Automation

  • Design and implement or direct the implementation of portfolio-wide monitoring observability event-management diagnostic automation self-healing and operational support tools that improve support effectiveness and service quality.
  • Assess and rationalize existing alerts automated emails service checks dashboards automated restarts runbooks and integrations to reduce duplication and noise simplify support processes and clarify ownership.
  • Define target-state standards and reusable patterns for telemetry dashboards alerting event correlation diagnostics and automation and oversee adoption technical quality documentation and sustainable support arrangements.
  • Measure the effectiveness of monitoring and automation improvements through reductions in non-actionable alerts recurring incidents manual intervention support effort Mean Time to Detect Mean Time to Restore and avoidable technology cost.

Production Readiness & Governance

  • Establish minimum portfolio standards for production readiness supportability monitoring resilience runbooks operational documentation service ownership and recovery capabilities.
  • Lead or provide technical assurance for complex or high-risk production-readiness reviews engaging Product Engineering Data Solutions Data Platform Architecture Security and Technology Services teams as appropriate during design reviews to identify and address reliability resilience and supportability risks.
  • Ensure unresolved production-readiness gaps have documented remediation plans compensating controls or appropriately authorized risk acceptance before release.
  • Ensure changes implemented or led by the SRE function follow established source-control peer-review testing security change-management release rollback documentation and production-validation requirements.
  • Monitor and govern vendor service performance against defined SLAs driving escalation and service performance plans through the appropriate vendor and portfolio governance channels.

Stakeholder Management

  • Present portfolio health reliability trends systemic risks operational burden and improvement progress in team and governance forums.
  • Act as the senior technical escalation point for major or cross-service production incidents.
  • Translate complex technical risks business impacts solution options costs dependencies and trade-offs into clear recommendations for business and technology stakeholders including senior stakeholders when required.
  • Build trusted relationships and use operational evidence to influence Product Data Technology Services Architecture and vendor teams to prioritize corrective and preventive reliability work.
  • This role does not carry formal people management accountability but provides functional and technical leadership across the production support function including setting reliability priorities and standards coordinating work and reviewing technical approaches and outcomes.
  • The SREL provides work direction coaching and mentoring to Site Reliability Engineering Specialists and operational support resources strengthening diagnostic automation and independent problem-solving capability and assuring the technical quality and completion of reliability work performed under its direction. Formal performance management and employment decisions remain with the applicable people leader.

The role has defined autonomy to determine the appropriate remediation approach for reliability and supportability issues.


This may include operational process changes monitoring automation configuration changes code changes platform improvements or escalation into a Product or Data team backlog including:


  • Establishing monitoring alerting automation diagnostic and support-tooling standards and determining technical approaches for approved reliability initiatives.
  • Prioritizing portfolio reliability and operational burden reduction initiatives based on business criticality service risk recurring incidents support effort capacity cost and available resources.
  • Leading incident and problem management strategies including determining when recurring issues require permanent remediation automated mitigation compensating controls or formal risk acceptance.
  • Determining whether selected changes can be implemented directly or through SRE and operational support resources or require Product Data Technology Services Architecture Security or vendor ownership and driving the agreed change through development testing release and validation.
  • Escalating material financial regulatory architectural funding or enterprise-risk decisions through the appropriate leadership and governance channels.

What youll need

  • Bachelors degree in Computer Science Engineering or related field (or equivalent combination of education and experience).
  • 8 years of experience in Site Reliability Engineering production engineering platform engineering application support or technology service delivery roles.
  • Advanced understanding of application data and platform architectures cloud platforms and enterprise systems.
  • Strong technical leadership collaboration communication facilitation and influencing skills with ademonstrated ability to deliver outcomes across multiple teams without direct reporting authority and tailor communications for technical business and senior stakeholder audiences.
  • Ability to navigate ambiguity manage competing priorities exercise sound technical judgement and lead complex high-impact outcomes under pressure.
  • Advanced knowledge of Site Reliability Engineering principles and practices including Service Level Availability (SLAs) observability incident and problem management capacity and performance engineering resilience disaster recovery operational automation and reduction of repetitive operational work.
  • Advanced knowledge of enterprise monitoring observability event-management diagnostic automation and operational support-tooling capabilities and integration patterns including tools such as Dynatrace.
  • Strong understanding of source control CI/CD pipelines automated testing deployment automation change management release management rollback and production-validation practices.
  • Deep understanding of application data database platform cloud infrastructure integration and distributed-system architectures and their common failure modes.
  • Knowledge of capacity management data and technology lifecycle considerations cloud storage and cost optimization resilience and recovery concepts.
  • Hands-on experience operating and improving highly available production systems across multiple applications data pipelines platforms or services.
  • Experience leading high-impact and cross-service incident response deep technical investigation root cause analysis problem elimination and corrective-action programs.
  • Experience designing and implementing or directing the implementation of monitoring observability automation diagnostic and operational support capabilities across multiple services.
  • Experience implementing or leading code configuration scripting or automation changes through source control technical review testing release and post-release validation.
  • Experience leading ambiguous cross-service technical initiatives from problem definition and requirements clarification through solution analysis implementation adoption and benefits validation including influencing Product Data Technology Services Architecture Security and vendor teams without formal authority.
  • Demonstrated experience applying and guiding the effective use of approved AI-assisted engineering tools including IDE-integrated assistants and Model Context Protocol (MCP)-enabled integrations to perform complex cross-service investigations; correlate application database code telemetry and infrastructure context; develop and validate remediation strategies; and improve incident response technical quality and engineering productivity.
  • Expertise in delivery methodologies (Agile Waterfall DevOps) and IT service management frameworks (ITIL COBIT).

#LI-OTPP #LI-AP1

What were offering

  • Numerous opportunities for professional growth and development

  • Comprehensive employer paid benefits coverage

  • Retirement income through a defined benefit pension plan

  • The opportunity to invest back into the fund through our Deferred Incentive Program

  • A flexible work environment combining in office collaboration and remote working

  • Competitive time off

  • Our Flexible Travel Program gives you the option to work abroad in another region/country for up to a month each year

  • Employee discount programs including Edvantage and Perkopolis

At Ontario Teachers diversity is one of our core strengths. We take pride in ensuring that the people we hire and the culture we create reflect and embrace diversity of thought background and experience. Through our Diversity Equity and Inclusion strategy and our Employee Resource Groups (ERGs) we celebrate diversity and foster inclusion through events for colleagues to connect for professional development networking & mentoring. We are building an inclusive and equitable workplace where our talent is respected accepted and empowered to be themselves. To learn more about our commitment to Diversity Equity and Inclusion check out Life at Teachers.

How to apply

Are you ready to pursue new challenges and take your career to the next level Apply today! You may be invited to complete a pre-recorded digital interview as part of your application.

Accommodations are available upon request () for candidates with a disability taking part in the recruitment process and once hired.

Candidates must be legally entitled to work in the country where this role is located.

Ontario Teachers may use AI-based tools to assist in screening and assessing applicants for this position. These tools may help us identify candidates whose skills and experience align with Ontario Teachers objectives by analyzing information provided in resumes and applications. Our use of AI does not replace human decision-making.

To learn more about how Teachers uses AI with your personal information please visit our Privacy Centre.

Functional Areas:

Information Technology

Vacancy:

Current

Requisition ID:

7284

#LI-AP1

Required Experience:

Senior IC