Site Reliability Engineer (Global Support)
Durban - South Africa
Job Summary
Job Family: Operations Function: Information Systems / Computer Operations Location: Durban Salary: Rneg up to R45k p/m (depending on experience)
About the Role
Were looking for an Engineer: Global Operations to join our 24/7 production operations team supporting global iGaming this role youll ensure system stability reliability and availability through incident management root cause analysis and technical problem-solving — balancing reactive incident response with proactive improvement.
Youll play a key part in leading post-incident reviews supporting operational readiness assessments validating AI-assisted operational decisions and mentoring junior colleagues to build team capability.
What Youll Do
- Participate in 24/7 shift rotations managing incident response and triage — owning alert acknowledgement prioritisation and investigation decisions resolving issues within established procedures and escalating to specialists only when needed.
- Conduct root cause analysis (RCA) investigations determine technical solutions collaborate with engineering teams on complex problems and take accountability for resolution outcomes.
- Coordinate and facilitate severity-tiered post-incident reviews extracting lessons learned and driving organisational learning.
- Own and maintain incident knowledge and documentation — knowledge base structure standards and accessibility for the team.
- Mentor and coach junior team members transferring technical knowledge and developing team proficiency in investigation and response procedures.
- Identify operational bottlenecks and inefficiencies propose and implement workflow refinements reduce alert noise and optimise incident response procedures.
- Support AI-assisted operational decision validation — establishing validation processes escalation criteria curating incident data for agent training and optimising human-in-the-loop model performance.
- Assess operational readiness for new features and system changes including SLI/SLO implications and operational risk identification.
- Evaluate and recommend operational tools and platforms assessing supportability reliability and integration options.
- Drive automation and toil-reduction initiatives — identifying repetitive manual tasks and implementing efficiency improvements.
What Were Looking For
Education
- Advanced Diploma or Bachelors Degree in Information Technology Computer Science Computer Engineering Information Systems Cybersecurity or an equivalent technical field.
Experience
- 24 years of progressive experience in IT operations technical support cybersecurity or systems administration with demonstrated capability in incident investigation troubleshooting and process improvement.
Skills
- Incident Management Troubleshooting Root Cause Analysis System Administration Technical Documentation Problem-Solving (DevelopingIntermediate)
- Communication Skills (IntermediateAdvanced)
- Process Improvement Knowledge Management Escalation Management (DevelopingIntermediate)
- Automation Agile Methodology Quality Assurance Change Management Influencing Skills ITIL 4 (AwarenessDeveloping)
Knowledge
- Broad knowledge across incident management post-incident learning process improvement emerging operational technologies and operational readiness assessment
- Advanced understanding of incident response practices — alert triage investigation methodology root cause analysis escalation procedures and documentation standards
- Demonstrated technical knowledge of production systems application architecture and infrastructure components
- Developing proficiency with AI-assisted operational tools and automation opportunities
- Working knowledge of operational readiness assessment SLOs and quality standards
Why Join Us
- Be part of a global operations team supporting platforms that operate across multiple countries and continents.
- Work at the intersection of traditional operations and AI-assisted automation helping shape how human-in-the-loop processes evolve.
- Grow your technical and leadership skills through mentoring cross-functional collaboration with engineering teams and exposure to a broad range of operational challenges.
- Contribute directly to platform reliability and customer trust in a regulated high-availability environment.
This is a full-time role requiring participation in a 24/7 shift rotation.