Site Reliability Engineer Kubernetes


Job Location:

Toronto - Canada

Monthly Salary: Not Disclosed
Experience Required: 5years
Posted on: 12 days ago
Vacancies: 1 Vacancy

Job Summary

Enterprise Kubernetes SRE (Python GitOps API Container Cloud MongoDB Postgres)

Toronto ON - Hybrid (4 Days WFO)
12 months


We are seeking an experienced Site Reliability Engineer to join our Enterprise
Kubernetes Platform team at a leading financial services organization. Youll
be responsible for ensuring the reliability performance and scalability of
our enterprise-grade Kubernetes platform that powers mission-critical
applications across the organization.

This role places a strong emphasis on automation with the expectation that
the successful candidate will continuously identify and eliminate manual toil
through intelligent tooling self-healing systems and AI-assisted operational
workflows. You will work alongside platform engineers DevOps teams and
application developers to build and maintain a world-class container
orchestration platform.


WHAT YOULL DO


PLATFORM RELIABILITY & OPERATIONS

Ensure 99.9% availability SLA for platform services across 60 Kubernetes
clusters spanning production DR UAT QA and development environments
Manage and operate enterprise Kubernetes distributions across on-premises
and cloud-hosted environments
Implement and maintain disaster recovery patterns across multi-AZ
architectures and geographically distributed data centres
Design and execute capacity planning resource optimization and cluster
scaling strategies
Support full cluster lifecycle operations including provisioning upgrades
patching and decommissioning
Automate cluster health checks and validation workflows for continuous
reliability assurance
Manage multi-tenant cluster environments with strict isolation and RBAC
enforcement

INCIDENT MANAGEMENT & ON-CALL

Participate in on-call rotation for platform infrastructure support with
sub-15-minute MTTR targets
Lead incident response troubleshooting and root cause analysis for
platform issues
Conduct blameless post-incident reviews and implement preventive measures
Develop and maintain runbooks troubleshooting guides and operational
playbooks
Coordinate with application teams during incidents affecting workloads
Build automated incident detection and response systems to reduce manual
intervention
Integrate AI-assisted triage tools for faster incident classification and
resolution


Enterprise Kubernetes SRE (Python GitOps API Container Cloud MongoDB Postgres)Toronto ON - Hybrid (4 Days WFO) 12 months We are seeking an experienced Site Reliability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization. Youll be responsible for en...