Site Reliability Engineer (SRE) – Azure AKS, Observability & Terraform
Posted:
30 May 2026 (30+ days ago)
Application Deadline:
27 August 2026
Vacancies:
1 Vacancy
Job Summary
Site Reliability Engineer (SRE) Azure AKS Observability & Terraform
Remote Role
Key Responsibilities
- Observability SRE DevOps roles with expertise in infrastructure and application reliability
- Dynatrace ELK Splunk PagerDuty
- SLI/SLO frameworks
- Azure Kubernetes Service (AKS) Terraform Azure managed services
What will you do
- Design and implement observability-as-code solutions using Terraform for monitoring pipelines dashboards and alerting across distributed systems
- Drive observability improvements using Dynatrace ELK Splunk PagerDuty for real-time performance insights and system visibility
- Instrument applications for end-to-end observability including distributed tracing metrics collection and log aggregation across microservices and event-driven architectures
- Troubleshoot complex production incidents across service layers databases caches and APIs using SLI/SLO frameworks
- Investigate and resolve Azure Kubernetes Service (AKS) infrastructure issues ensuring reliability and scalability of containerized workloads using Terraform and Azure services (SQL MI Redis Functions Event Grid)
- Translate business requirements into observable resilient systems aligned to SLIs/SLOs
- Automate operational tasks using Infrastructure-as-Code and CI/CD to reduce toil and improve resilience
- Lead incident response and remediation for critical systems including blameless postmortems and chaos engineering practices
- Collaborate with development platform and business teams to improve availability scalability and operational excellence
What do you need to succeed
Must-have
- 8 years experience in SRE DevOps or Observability roles focused on infrastructure and application reliability
- Strong expertise in Dynatrace ELK Splunk PagerDuty and observability principles (instrumentation correlation IDs SLIs/SLOs)
- Advanced proficiency in Azure Kubernetes Service (AKS) Terraform and Azure managed services (SQL MI Redis Functions Event Grid)
- Hands-on experience with observability instrumentation (distributed tracing metrics logs) across microservices and event-driven systems
- Strong troubleshooting skills across distributed systems (services databases caches APIs) in production environments
- Incident management expertise using PagerDuty and ServiceNow including high-severity incident resolution and RCA
- Knowledge of incident problem and change management SRE principles blameless postmortems and chaos engineering
- Strong communication and leadership skills for cross-functional coordination and incident handling
Required Skills:
Experience (Years): 8-10