Enter a job title or keyword

Senior Reliability Engineer (SRE) – DevOpsCloud Infrastructure GautengHybrid ISB551356

ISanqa Resourcing


Job Location:

Midrand - South Africa

Monthly Salary: Not provided by the employer
Posted: 8 June 2026 (30+ days ago)
Application Deadline: 5 September 2026
Vacancies: 1 Vacancy
The job posting is outdated and position may be filled

Job Summary

Were seeking a strategic Senior Reliability Engineer with 5 years experience in DevOs cloud infrastructure or reliability engineering roles and proven expertise in DevOps cloud infrastructure operations and SRE practices.

The ideal candidate has deep expertise in DevOps and cloud infrastructure operations with strong automation skills for build test and deployment pipelines (CI/CD) and Infrastructure as Code (Terraform ARM templates or similar).

Demonstrable experience with containerization and orchestration (Docker Kubernetes Helm) monitoring observability and alerting tooling (Prometheus Grafana or equivalent) and tier-3 incident management is essential.

Strong systems engineering mindset with ability to influence without direct command drive root cause analysis and support global GROUP Group Customer Companion product is critical.

Own and improve operational reliability and availability of critical applications across their lifecycle building and optimizing CI/CD pipelines to enable fast reliable and repeatable releases.

Become the Senior Reliability Engineer guiding mission-critical Customer Companion operations where your technical vision will develop monitoring alerting and observability solutions lead incident response and deliver transformational value across AWS/Azure platforms.

POSITION: Contract: 01 August 2026 to 31 December 2028

EXPERIENCE: Minimum 5 years experience in DevOps cloud infrastructure or reliability engineering roles

COMMENCEMENT: 01 August 2026

LOCATION: Hybrid: Midrand/Menlyn/Rosslyn/Home Office Rotation

TEAM: DevOps - Operations Engineering

Qualifications and Experience

  • IT degree or equivalent qualification and solid background in systems engineering DevOps or cloud operations
  • Minimum 5 years experience in DevOps cloud infrastructure or reliability engineering roles
  • Relevant cloud certification(s) and proven experience implementing automation and monitoring at enterprise scale (AWS/Azure preferred)

Essential Skills Requirements

Technical:

  • Deep expertise in DevOps and cloud infrastructure operations
  • Strong automation skills for build test and deployment pipelines (CI/CD)
  • Experience with Infrastructure as Code (Terraform ARM templates or similar)
  • Proficiency in at least one scripting/programming language (Python Bash)
  • Strong knowledge of containerization and orchestration (Docker Kubernetes Helm)
  • Experience with monitoring observability and alerting tooling (Prometheus Grafana or equivalent)
  • Solid understanding of systems engineering concepts reliability and availability best practices
  • Familiarity with security compliance and certificate management in enterprise environments
  • Any additional responsibilities assigned in the Agile Working Model AWM Charter

Agile and DevOps:

  • Execution according to the Agile Methodology and attending of all team meetings including Standups Sprint Review Sprint Retrospectives Sprint Planning meetings
  • Daily use of the Agile Tool Chain as per the updates required by the respective feature team or teams
  • JIRA/Confluence knowledge

Stakeholder Management:

  • Excellent communication and stakeholder management skills to coordinate across development support and business teams.
  • Influence (Effect change without direct exercise of command where persuasion is required and reaching buy-in might be difficult).
  • Collaborate closely with development teams to improve observability testability and deployability of applications.

Soft Skills:

  • Strong incident management and problem-solving skills for tier-3 escalations
  • Self-motivated and keen attention to detail or time management
  • Completely independent worker that will only escalate tasks that are complex and outside of their span of control
  • May need to lead individual members of a team of Entry and Advanced level experts

Advantageous Skills Requirements

  • AWS or Azure certifications (Solution Architect / DevOps / Developer)
  • Experience with Grafana/Prometheus ecosystems and advanced observability patterns
  • Knowledge of Kustomize Helm and Kubernetes deployment strategies
  • Background in enterprise middleware and integration (SAP exposure advantageous)
  • Experience implementing certificate management and secure service communication
  • Familiarity with Agile development and DevOps cultural practices
  • Exposure to cloud-native toolchains and platform engineering approaches
  • Experience creating runbooks operational documentation and user/operational manuals
  • Knowledge of monitoring development process metrics and reliability KPIs
  • Prior work in regulated corporate environments with strong process and compliance requirements

Role Requirements

Operations and Support:

  • Own and improve operational reliability and availability of critical applications across their lifecycle
  • Develop monitoring alerting and observability solutions to detect and prevent incidents
  • Lead incident response for escalated production issues and drive root cause analysis and remediation

Delivery and Configuration:

  • Design implement and maintain automation for infrastructure provisioning and deployments (IaC)
  • Build and optimize CI/CD pipelines to enable fast reliable and repeatable releases
  • Drive performance tuning capacity planning and availability engineering activities
  • Plan and execute upgrades migrations and infrastructure improvements with minimal downtime
  • Specify new products processes standards based on organization strategy or set short- to mid-term operational plans
  • The requests are often unclear and will need substantial investigation to ensure that the specifics are well understood by all parties

Continuous Improvement:

  • Improve a product or a system that already exists by making conceptual changes and enhancements
  • Can manage the solution for complex problems that may require simple solutions but affect multiple systems
  • Implement and enforce operational best practices runbooks and playbooks for the team

Documentation and Knowledge:

  • Ensure security compliance and certificate management practices are applied to platforms and services
  • Mentor and coach junior DevOps and operations team members to uplift team capability
  • Contribute to business case development risk identification and operational documentation
  • May need to lead individual members of a team of Entry and Advanced level experts

NB: South African citizens or residents are preferred. Applicants with valid work permits will also be considered. By applying you consent to be added to the database and to receive updates until you unsubscribe. If you do not receive a response within 2 weeks please consider your application unsuccessful.

#JobID1356 #ReliabilityEngineer #Senior #SRE #DevOps #Kubernetes #Terraform #Prometheus #Grafana #AWS #Azure #ITHub #NowHiring #fuelledbypassionintegrityexcellence

iSanqa is your trusted Level 2 BEE recruitment partner dedicated to continuous improvement in delivering exceptional service. Specializing in seamless placements for permanent staff temporary resources and efficient contract management and billing facilitation iSanqa Resourcing is powered by a team of professionals with an outstanding track record. With over 100 years of combined experience we are committed to evolving our practices to ensure ongoing excellence.