Director of Site Reliability Engineering
Job Summary
Were looking for a Director of Site Reliability Engineering to join our team in London United Kingdom in a hybrid working mode.
This role is responsible for driving reliability engineering and operational excellence across global technology platforms while leading the adoption of AI-enabled solutions for automation and efficiency. The position combines strategic leadership with hands-on governance to ensure highly available resilient systems that align with business and regulatory requirements. As a technology thought leader you will influence engineering standards enhance operational frameworks and foster a culture of continuous improvement across mission-critical environments.
- Lead and scale a global SRE organization focusing on engineering excellence and team empowerment
- Collaborate with product platform operations and security teams to embed reliability within SDLC practices
- Define and monitor KPIs for system reliability performance and operational efficiency
- Advance automation Infrastructure as Code approaches and promote self-healing systems using AI/ML techniques
- Develop robust incident management frameworks and lead major incident response activities for critical systems
- Implement blameless postmortems and deliver systemic improvements across production environments
- Establish observability strategies with standardized tooling for metrics logs and tracing to support distributed systems
- Adopt and enforce SRE practices including SLIs SLOs SLAs and error budgets across services
- Drive resilience strategies with highly available architectures and disaster recovery readiness
- Champion an automation-first culture leveraging CI/CD pipelines and operational tooling to reduce manual processes
- Strong background in Site Reliability Engineering DevOps or platform operations in complex distributed environments
- Expertise in observability platforms troubleshooting distributed systems and telemetry-driven insights
- Hands-on experience with automation Infrastructure as Code (Terraform or CloudFormation) and CI/CD practices
- Deep understanding of incident management processes ITSM standards and ITIL principles
- Knowledge of resilience design patterns high availability and fault-tolerant architectures
- Familiarity with AI/ML-driven approaches for operational efficiency and system reliability
- Ability to lead transformation influence across teams and foster continuous improvement in culture
- Experience in financial services or other highly regulated mission-critical environments
- Certifications in cloud technologies such as AWS
- Exposure to AIOps platforms or advanced observability tooling
Required Experience:
Director