Enter a job title or keyword

Site Reliability Engineer (SRE) – Azure AKS, Observability & Terraform


Job Location:

Vancouver - Canada

Monthly Salary: K 10 - 10
Experience Required: 5years
Posted: 30 May 2026 (30+ days ago)
Application Deadline: 27 August 2026
Vacancies: 1 Vacancy
The job posting is outdated and position may be filled

Job Summary

Site Reliability Engineer (SRE) Azure AKS Observability & Terraform

Remote Role

Key Responsibilities

  • Observability SRE DevOps roles with expertise in infrastructure and application reliability
  • Dynatrace ELK Splunk PagerDuty
  • SLI/SLO frameworks
  • Azure Kubernetes Service (AKS) Terraform Azure managed services

What will you do

  • Design and implement observability-as-code solutions using Terraform for monitoring pipelines dashboards and alerting across distributed systems
  • Drive observability improvements using Dynatrace ELK Splunk PagerDuty for real-time performance insights and system visibility
  • Instrument applications for end-to-end observability including distributed tracing metrics collection and log aggregation across microservices and event-driven architectures
  • Troubleshoot complex production incidents across service layers databases caches and APIs using SLI/SLO frameworks
  • Investigate and resolve Azure Kubernetes Service (AKS) infrastructure issues ensuring reliability and scalability of containerized workloads using Terraform and Azure services (SQL MI Redis Functions Event Grid)
  • Translate business requirements into observable resilient systems aligned to SLIs/SLOs
  • Automate operational tasks using Infrastructure-as-Code and CI/CD to reduce toil and improve resilience
  • Lead incident response and remediation for critical systems including blameless postmortems and chaos engineering practices
  • Collaborate with development platform and business teams to improve availability scalability and operational excellence

What do you need to succeed

Must-have

  • 8 years experience in SRE DevOps or Observability roles focused on infrastructure and application reliability
  • Strong expertise in Dynatrace ELK Splunk PagerDuty and observability principles (instrumentation correlation IDs SLIs/SLOs)
  • Advanced proficiency in Azure Kubernetes Service (AKS) Terraform and Azure managed services (SQL MI Redis Functions Event Grid)
  • Hands-on experience with observability instrumentation (distributed tracing metrics logs) across microservices and event-driven systems
  • Strong troubleshooting skills across distributed systems (services databases caches APIs) in production environments
  • Incident management expertise using PagerDuty and ServiceNow including high-severity incident resolution and RCA
  • Knowledge of incident problem and change management SRE principles blameless postmortems and chaos engineering
  • Strong communication and leadership skills for cross-functional coordination and incident handling



Required Skills:

Experience (Years): 8-10