Enter a job title or keyword

Director of Site Reliability Engineering

EPAM Systems


Job Location:

London - UK

Monthly Salary: Not provided by the employer
Posted: 9 September 2026 (13 hours ago)
Application Deadline: 7 December 2026
Vacancies: 1 Vacancy

Job Summary

Were looking for a Director of Site Reliability Engineering to join our team in London United Kingdom in a hybrid working mode.

This role is responsible for driving reliability engineering and operational excellence across global technology platforms while leading the adoption of AI-enabled solutions for automation and efficiency. The position combines strategic leadership with hands-on governance to ensure highly available resilient systems that align with business and regulatory requirements. As a technology thought leader you will influence engineering standards enhance operational frameworks and foster a culture of continuous improvement across mission-critical environments.

Responsibilities
  • Lead and scale a global SRE organization focusing on engineering excellence and team empowerment
  • Collaborate with product platform operations and security teams to embed reliability within SDLC practices
  • Define and monitor KPIs for system reliability performance and operational efficiency
  • Advance automation Infrastructure as Code approaches and promote self-healing systems using AI/ML techniques
  • Develop robust incident management frameworks and lead major incident response activities for critical systems
  • Implement blameless postmortems and deliver systemic improvements across production environments
  • Establish observability strategies with standardized tooling for metrics logs and tracing to support distributed systems
  • Adopt and enforce SRE practices including SLIs SLOs SLAs and error budgets across services
  • Drive resilience strategies with highly available architectures and disaster recovery readiness
  • Champion an automation-first culture leveraging CI/CD pipelines and operational tooling to reduce manual processes
Requirements
  • Strong background in Site Reliability Engineering DevOps or platform operations in complex distributed environments
  • Expertise in observability platforms troubleshooting distributed systems and telemetry-driven insights
  • Hands-on experience with automation Infrastructure as Code (Terraform or CloudFormation) and CI/CD practices
  • Deep understanding of incident management processes ITSM standards and ITIL principles
  • Knowledge of resilience design patterns high availability and fault-tolerant architectures
  • Familiarity with AI/ML-driven approaches for operational efficiency and system reliability
  • Ability to lead transformation influence across teams and foster continuous improvement in culture
Nice to have
  • Experience in financial services or other highly regulated mission-critical environments
  • Certifications in cloud technologies such as AWS
  • Exposure to AIOps platforms or advanced observability tooling

Required Experience:

Director