Senior Infrastructure Engineer (SRE)
Johannesburg - South Africa
Job Summary
- Design implement and maintain highly available and resilient infrastructure for data analytics and AI platforms.
- Define and drive site reliability engineering (SRE) practices including service-level objectives (SLOs) service-level indicators (SLIs) and service-level agreements (SLAs).
- Implement observability frameworks including monitoring logging tracing and alerting across all platforms and services.
- Ensure proactive detection diagnosis and resolution of infrastructure and application reliability issues.
- Lead disaster recovery (DR) planning testing and readiness activities across cloud and on-prem environments.
- Automate infrastructure health checks failover processes and operational runbooks.
- Collaborate with Cloud Engineers Platform Engineers and Solution Architects to ensure resilient architecture design.
- Support capacity planning performance tuning and scalability improvements across systems.
- Manage incident response root cause analysis (RCA) and post-incident reviews to drive continuous improvement.
- Ensure infrastructure complies with security governance and enterprise architecture standards.
- Drive reliability engineering best practices across DevOps Data Engineering and Platform teams.
- Support CI/CD and deployment reliability in collaboration with Platform Engineering teams.
- Bachelors degree in Computer Science Information Technology Engineering or a related discipline.
- 610 years experience in infrastructure engineering DevOps or Site Reliability Engineering roles.
- Strong experience with cloud platforms (Azure AWS or Google Cloud Platform).
- Experience implementing observability stacks (e.g. Prometheus Grafana ELK/EFK Azure Monitor or equivalent).
- Strong understanding of distributed systems high availability and fault tolerance.
- Experience with automation and scripting (Python Bash PowerShell or similar).
- Experience with container orchestration technologies such as Kubernetes is advantageous.
- Strong experience in incident management root cause analysis and production support.
- Knowledge of disaster recovery strategies backup systems and business continuity planning.
- SRE or cloud certifications (e.g. Google SRE Azure Administrator/Architect AWS SysOps) are advantageous.
- Site Reliability Engineering (SRE) principles
- Infrastructure reliability and resilience
- Observability and monitoring
- Incident management and root cause analysis
- Disaster recovery and business continuity
- Cloud infrastructure engineering
- Automation and scripting
- Performance and capacity management
- CI/CD reliability support
- Cross-functional collaboration
Required Experience:
Senior IC
About Company
With 15 years of experience in the Telecommunications and Financial Services industry in Africa and the Middle East, we are experts at supporting our clients through their digital transformation journeys. We believe in blending traditional methods with cutting-edge strategies to drive ... View more