Enter a job title or keyword

Senior Manager, I&IT Site Reliability Engineering & Operations

Metrolinx


Job Location:

Toronto - Canada

Monthly Salary: Not provided by the employer
Posted: 6 October 2026 (14 hours ago)
Application Deadline: 3 January 2027
Vacancies: 1 Vacancy

Job Summary

Description
Metrolinx is connecting communities across the Greater Golden Horseshoe. Metrolinx operates GO Transit and UP Express as well as the PRESTO fare payment system. We are also building new and improved rapid transit including GO Expansion Light Rail Transit routes and major expansions to Torontos subway system to get people where they need to go better faster and easier. Metrolinx is an agency of the Government of Ontario.

At Metrolinx equity diversity and inclusion are essential to living our values of serving with passion thinking forward and playing as a team.

I&IT ONLY: Metrolinxs Innovation and Information Technology group supports female team members via Go Tech Women an affinity group for women in Information Technology led by our Chief Information Officer.

If you enjoy technology and innovation value diversity appreciate work/balance and are looking for an opportunity to make a better world via public service Metrolinx would like to hear from you!

The SRE and Operations embed resilience recoverability and operational excellence across Metrolinx I&IT services. Through continuous delivery infrastructure as code proactive observability and tested disaster recovery capabilities the function reduces operational risk and enables rapid recovery from major outages cyber incidents and disruptions.

What will I be doing

  • Drive SRE & Operations Leadership: Lead coach and direct engineering teams responsible for day-to-day infrastructure operations platform reliability data center management and disaster recovery.
  • Hands-On Leadership Role:Operate as a hands-on technical leader who remains actively engaged in architecture engineering decisions operational problem-solving incident response and delivery execution while leading and developing the team.
  • Establish Operating Models & Standards: Build and govern the organizations SRE operating model priorities and cloud reliability standards covering availability scalability recoverability and operational support across public and hybrid cloud environments.
  • Govern Cloud & Infrastructure Reliability: Define production readiness criteria and operational acceptance guardrails for new or materially changed services ensuring architecture monitoring logging alerting security controls and recovery automation are fully validated prior to release.
  • Manage Service Health: Define Service-Level Indicators and Service-Level Objectives for critical services to balance deployment velocity with platform stability.
  • Direct Incident Response & Post-Incident Learning: Lead major incident response efforts to restore disrupted services rapidly and facilitate post-incident reviews to ensure root causes are remediated and tracked through engineering backlogs.
  • Day-to-Day Operations: Oversee day-to-day data centre operational management including HPE Synergy Hyper-V VMware SAN/NAS storage configuration F5 NetScaler file servers database enablement (Oracle Data Guard SQL Always On Oracle ExaCC) data backup networking integration (Cisco ACI and Firepower Palo Alto Firewalls) and physical infrastructure tuning.
  • Proactively Manage Infrastructure Risk: Direct reliability risk assessments capacity forecasting and controlled resilience testing including but not limited to multi-region availability-zone network and identity failure scenarios to eliminate single points of failure.
  • Steward Budgets: Support the governance and allocation of an annual capital budget managing operational and capital budgets to meet product roadmap objectives within financial constraints.
  • Govern Vendor Relationships: Direct vendor deliverables evaluate technology proposals participate in contract negotiations and establish Service Level Agreements (SLAs) and Operational Level Agreements (OLAs) with strategic partners.
  • Build a Culture of Innovation & Excellence: Build high-performance team capabilities in automated operations infrastructure as code and SRE disciplines.
  • Flexible Hours: Preparedness to occasionally work outside regular business hours including evenings and weekends during emergency major incidents or urgent technology deployments.

What Skills and Qualifications Do I Need

Education Completion of a degree in Computer Science Information Systems Business Administration Engineering or a related discipline.

Certifications

  • Microsoft Certified Azure Solutions Expert Azure DevOps Engineer or Equivalent is an asset.
  • Cisco Certified Network Professional (CCNP) is an asset.
  • Certified Information System Security Professional is an asset.
  • Certified Kubernetes Administrator is an asset.
  • Information Technology Infrastructure Library (ITIL) Foundations is an asset.
  • Agile certification (Agile Certified Professional Certified Scrum Product Owner) is an asset.
  • Project Management Professional (PMP) or PRINCE2 is an asset.

Professional Experience & Technical Mastery

  • Progressive Leadership:Demonstrated years of progressively responsible experience in infrastructure operations platform engineering cloud engineering or SRE within a large-scale service delivery environment including leadership of technical teams and complex operational programs.
  • Large-Scale Azure Cloud Migration:Hands-on experience leading large-scale cloud migrations from on-premises environments to Microsoft Azure including migration planning architecture execution risk management cutover and operational transition.
  • SRE & Cloud Practice: Deep knowledge of Site Reliability Engineering frameworks dynamic resource management service quotas autoscaling infrastructure as code (IaC) policy as code and cloud cost optimization.
  • Relationship Management: Exceptional communication facilitation and negotiation skills to brief executive leadership influence internal stakeholders manage vendor contracts and introduce modern methodologies (DevOps Lean Agile).

Core Technologies (Including but not limited to)

  • Cloud Platforms:Microsoft Azure Copilot Exchange Online Teams SharePoint etc.
  • DC Infrastructure:VMware Hyper-V HPE compute and storage and Brocade SAN.
  • Network & Application Delivery:Palo Alto Firewalls F5 Citrix NetScaler Fortinet SD-WAN and Infoblox.
  • Identity & Security:Microsoft Active Directory ZscalerMicrosoft Entra ID Okta Microsoft Defender and CyberArk Privilege Cloud.
  • Cloud Operations Automation & Observability:Grafana ServiceNow Terraform Ansible and GitHub.
  • Operating Systems & Management Platforms:Red Hat Enterprise Linux Windows Server SCVMM SCCM and Red Hat Satellite.
  • DR Solutions: Hands-on experience implementing enterprise recovery technologies including Azure Cloud DR Cisco ACI multisite Zerto Hyper-V Oracle Data Guard SQL Always On F5 application delivery gateway and related replication solutions.
  • Process Acumen: Understanding of project management practices ITIL frameworks risk and business impact analysis (BIA) etc.

Dont Meet Every Requirement
If youre excited about working with Metrolinx but your past experience doesnt quite align with every qualification of this posting we encourage you to apply. You just might be the right candidate for this or other roles. We are always looking for great talent to join our team.

We invite all interested individuals to apply and encourage applications from members of equity-deserving communities including those who identify as Indigenous Black racialized women people with disabilities and people with diverse gender identities expressions and sexual orientations.

Accommodation:
Metrolinx is an equal opportunity employer and committed to a diverse and inclusive workforce. We are also committed to offering reasonable accommodation to job applicants in accordance with applicable legislation. If you require assistance or an accommodation at any point during the hiring process please contact us at: or email

Application Process:
All applicants must be legally entitled to work in Canada. Metrolinx will be using email to communicate with you for all job competitions. It is your responsibility to include an updated email address that is checked daily and accepts emails from unknown users. As we send time-sensitive correspondence we recommend that you check your email regularly. If no response is received we will assume you are no longer interested in pursuing the opportunity. Please be advised that a Criminal Record Check may be required of the successful candidate.

For Internal applicants with the recent implementation of the Internal Mobility Policy the internal recruitment process has changed for non-union roles. Candidates must be in their current role for 12 months prior to applying for another role and each applicant must be in good standing (not participating in a Performance Improvement Plan). Please review all provisions of the policybefore submitting your application.

Should it be determined that any background information provided is misleading inaccurate or incorrect Metrolinx reserves the right to discontinue with the consideration of your application.

We thank all applicants for their interest however only those selected for further consideration will be contacted.

WE ARE AN EQUITABLE AND INCLUSIVE EMPLOYER.




Required Experience:

Senior Manager