Staff Site Reliability Engineer
Job Summary
This is a hands-on senior engineering role focused on improving production resilience strengthening security driving operational excellence and enhancing the developer experience across the organization.
In this role you will design build and evolve the foundational systems tooling and operational practices that enable engineering teams to ship secure reliable and scalable software with confidence. You will help establish reliability standards define service level objectives (SLOs) improve observability automate operational processes and drive incident management and post-incident learning practices that strengthen platform stability over time.
Partnering closely with Engineering Security Platform and Product teams you will architect scalable distributed systems optimize Kubernetes and AWS-based infrastructure and build automated delivery pipelines that support rapid and safe software releases. You will play a key role in reducing operational toil improving system performance increasing platform reliability and ensuring that our infrastructure can support continued business growth.
This is a full-time permanent position
This is an existing vacancy
Location: This is a remote location open to candidates legally authorized to work in Canada.
- Drive reliability engineering initiatives and operational excellence for mission-critical services running on AWS and Kubernetes.
- Design implement and continuously improve deployment release and rollback strategies across complex distributed systems.
- Establish secure-by-default CI/CD pipelines with robust automation governance and policy-driven controls.
- Enhance platform observability through metrics logs tracing and actionable alerting to improve system visibility and operational efficiency.
- Define implement and mature Service Level Indicators (SLIs) Service Level Objectives (SLOs) and reliability standards across the organization.
- Lead response efforts for high-severity incidents ensuring timely resolution effective communication and meaningful post-incident reviews that drive continuous improvement.
- Partner closely with engineering teams to strengthen platform standards improve service resilience optimize runtime performance and embed reliability best practices.
- Mentor and guide engineers on cloud-native technologies site reliability engineering principles and operational excellence practices fostering a culture of continuous learning and accountability.
- 8 years of experience in Site Reliability Engineering (SRE) Platform Engineering DevOps or related cloud-native engineering roles.
- Deep expertise in AWS services including EKS IAM VPC Lambda CloudFront S3 and cloud networking/security best practices.
- Advanced experience operating and scaling production Kubernetes environments.
- Strong hands-on experience with Istio service mesh including traffic management security observability and resiliency.
- Proven expertise with Infrastructure as Code (IaC) preferably using AWS CDK.
- Experience building and managing CI/CD pipelines using GitHub Actions or similar platforms.
- Strong troubleshooting performance optimization and incident management experience in distributed systems.
- Excellent communication collaboration and technical leadership skills.
- Experience designing and operating monitoring logging tracing and alerting solutions for cloud-native platforms.
- Strong knowledge of AWS CloudWatch OpenTelemetry AWS X-Ray and Kubernetes observability tooling.
- Experience defining and operationalizing SLIs SLOs alerting strategies runbooks and reliability metrics.
- Proven ability to leverage observability data to improve service reliability reduce incident impact and optimize operational performance.
- Strong proficiency in TypeScript and for platform engineering automation and operational tooling.
- Experience building and maintaining scalable backend services APIs and event-driven systems.
- Deep understanding of Kubernetes architecture controllers Gateway API ingress management and service networking.
- Experience implementing zero-trust architectures mTLS and service-to-service security controls.
- Commitment to high-quality engineering practices including automated testing code reviews and observability-driven development.
- Strong understanding of resilience engineering including autoscaling disruption management failure testing and safe deployment strategies.
- Experience with progressive delivery practices such as canary blue/green and feature-flag-based deployments.
- Experience working in regulated compliance-driven or security-sensitive SaaS environments.
- Familiarity with FinOps principles and cost optimization strategies for cloud platforms.
- Experience building internal developer platforms and self-service engineering tooling.
- Cloud-native certifications such as CKA CKAD CKS KCSA or KCNA.
- Kubestronaut certification or equivalent advanced Kubernetes expertise is highly regarded.
Salary Range:
The annual base salary for this position is between $140000 CAD and $155000 CAD per year.
This role is also eligible for discretionary bonus and/or commission as well as other benefits. Actual pay within the listed range will be determined based on factors such as transferable skills relevant experience market conditions and primary work location. The posted range is subject to change and may be updated periodically.
Required Experience:
Staff IC
About Company
Accounting, audit, analytics and compliance software built by seasoned accountants. Manage your audit and financial reporting more efficiently with less risk.