Site Reliability Engineer Kubernetes
Posted on:
12 days ago
Vacancies:
1 Vacancy
Job Summary
Enterprise Kubernetes SRE (Python GitOps API Container Cloud MongoDB Postgres)
Toronto ON - Hybrid (4 Days WFO)
12 months
We are seeking an experienced Site Reliability Engineer to join our Enterprise
Kubernetes Platform team at a leading financial services organization. Youll
be responsible for ensuring the reliability performance and scalability of
our enterprise-grade Kubernetes platform that powers mission-critical
applications across the organization.
This role places a strong emphasis on automation with the expectation that
the successful candidate will continuously identify and eliminate manual toil
through intelligent tooling self-healing systems and AI-assisted operational
workflows. You will work alongside platform engineers DevOps teams and
application developers to build and maintain a world-class container
orchestration platform.
WHAT YOULL DO
PLATFORM RELIABILITY & OPERATIONS
Ensure 99.9% availability SLA for platform services across 60 Kubernetes
clusters spanning production DR UAT QA and development environments
Manage and operate enterprise Kubernetes distributions across on-premises
and cloud-hosted environments
Implement and maintain disaster recovery patterns across multi-AZ
architectures and geographically distributed data centres
Design and execute capacity planning resource optimization and cluster
scaling strategies
Support full cluster lifecycle operations including provisioning upgrades
patching and decommissioning
Automate cluster health checks and validation workflows for continuous
reliability assurance
Manage multi-tenant cluster environments with strict isolation and RBAC
enforcement
INCIDENT MANAGEMENT & ON-CALL
Participate in on-call rotation for platform infrastructure support with
sub-15-minute MTTR targets
Lead incident response troubleshooting and root cause analysis for
platform issues
Conduct blameless post-incident reviews and implement preventive measures
Develop and maintain runbooks troubleshooting guides and operational
playbooks
Coordinate with application teams during incidents affecting workloads
Build automated incident detection and response systems to reduce manual
intervention
Integrate AI-assisted triage tools for faster incident classification and
resolution
12 months
We are seeking an experienced Site Reliability Engineer to join our Enterprise
Kubernetes Platform team at a leading financial services organization. Youll
be responsible for ensuring the reliability performance and scalability of
our enterprise-grade Kubernetes platform that powers mission-critical
applications across the organization.
This role places a strong emphasis on automation with the expectation that
the successful candidate will continuously identify and eliminate manual toil
through intelligent tooling self-healing systems and AI-assisted operational
workflows. You will work alongside platform engineers DevOps teams and
application developers to build and maintain a world-class container
orchestration platform.
WHAT YOULL DO
PLATFORM RELIABILITY & OPERATIONS
Ensure 99.9% availability SLA for platform services across 60 Kubernetes
clusters spanning production DR UAT QA and development environments
Manage and operate enterprise Kubernetes distributions across on-premises
and cloud-hosted environments
Implement and maintain disaster recovery patterns across multi-AZ
architectures and geographically distributed data centres
Design and execute capacity planning resource optimization and cluster
scaling strategies
Support full cluster lifecycle operations including provisioning upgrades
patching and decommissioning
Automate cluster health checks and validation workflows for continuous
reliability assurance
Manage multi-tenant cluster environments with strict isolation and RBAC
enforcement
INCIDENT MANAGEMENT & ON-CALL
Participate in on-call rotation for platform infrastructure support with
sub-15-minute MTTR targets
Lead incident response troubleshooting and root cause analysis for
platform issues
Conduct blameless post-incident reviews and implement preventive measures
Develop and maintain runbooks troubleshooting guides and operational
playbooks
Coordinate with application teams during incidents affecting workloads
Build automated incident detection and response systems to reduce manual
intervention
Integrate AI-assisted triage tools for faster incident classification and
resolution