Are you ready to unleash your potential
At Deloitte our purpose is to make an impact that matters for our clients our people and the communities we serve.
We believe we have a responsibility to be a force for good andWorldImpactis our portfolio of initiatives focused on making a tangible impact on societys biggest challenges and creating a better future. We strive to advise clients on how to deliver purpose-led growth and embed more equitable inclusive as well as sustainable business practices.
Hence we seek talented individuals driven to excel and innovate working together to achieve our shared goals.
We are committed to creating positive work experiences that foster a culture of respect and inclusion where diverse perspectives are celebrated and everyone is recognised for their contributions.
Ready to unleash your potential with us Join the winning team now!
Work youll do:
As a Site Reliability Engineer (SRE) you will play a key role in maintaining the reliability and performance of critical services. Your expertise will help bridge the gap between development and operations ensuring robust scalable and responsive infrastructure. This role emphasizes strong system architecture and design principles focusing on key SRE practices such as Service Level Objectives (SLOs) Service Level Indicators (SLIs) and the reduction of operational toil. You will collaborate closely with diverse teams to drive reliability improvements and foster a culture of continuous learning and accountability.
Key Responsibilities:
- Monitor and maintain system performance to ensure the stability and reliability of applications and infrastructure.
- Design and implement resilient system architectures that support high availability and scalability.
- Develop automation tools and scripts to enhance operational efficiency and reduce manual effort.
- Define track and analyze SLOs and SLIs to ensure reliability and performance meet business needs.
- Conduct thorough post-mortem analyses following incidents driving continuous improvement through root cause identification and solution implementation.
- Collaborate with development and operations teams to establish best practices in system reliability and incident management.
- Troubleshoot and resolve issues related to database performance network connectivity and deployment failures including diagnosing problems at the underlying platform level (e.g. Kubernetes virtual machines).
- Ensure that issues are resolved within the stipulated Service Level Agreements (SLAs) maintaining high standards of service delivery.
- Identify and troubleshoot performance bottlenecks in applications and infrastructure providing actionable recommendations for enhancements.
- Maintain detailed documentation of processes and incident responses to support knowledge sharing and compliance.
- Improve monitoring solutions to proactively identify and mitigate issues before they impact services.
- Assist in the deployment and configuration of new applications and services ensuring adherence to best practices.
- Participate in on-call rotations and respond to critical incidents as they arise.
- Analyze system logs and metrics to identify trends and potential areas for improvement.
Your role as a leader:
At Deloitte we believe in the importance of empowering our people to be leaders at all levels. We connect our purpose and shared values to identify issues as well as to make an impact that matters to our clients people and the communities. Additionally Senior Consultants across our Firm are expected to:
- Monitor and maintain system performance to ensure the stability and reliability of applications and infrastructure.
- Design and implement resilient system architectures that support high availability and scalability.
- Develop automation tools and scripts to enhance operational efficiency and reduce manual effort.
- Define track and analyze SLOs and SLIs to ensure reliability and performance meet business needs.
- Conduct thorough post-mortem analyses following incidents driving continuous improvement through root cause identification and solution implementation.
- Collaborate with development and operations teams to establish best practices in system reliability and incident management.
- Troubleshoot and resolve issues related to database performance network connectivity and deployment failures including diagnosing problems at the underlying platform level (e.g. Kubernetes
- virtual machines).
- Ensure that issues are resolved within the stipulated Service Level Agreements (SLAs) maintaining high standards of service delivery.
- Identify and troubleshoot performance bottlenecks in applications and infrastructure providing actionable recommendations for enhancements.
- Maintain detailed documentation of processes and incident responses to support knowledge sharing and compliance.
- Improve monitoring solutions to proactively identify and mitigate issues before they impact services.
- Assist in the deployment and configuration of new applications and services ensuring adherence to best practices.
- Participate in on-call rotations and respond to critical incidents as they arise.
- Analyze system logs and metrics to identify trends and potential areas for improvement.
Requirements:
- Strong experience with Linux systems and distributed computing fundamentals.
- Proven experience in troubleshooting application issues with a focus on performance and connectivity.
- Familiarity with networking concepts and effective troubleshooting techniques.
- Experience in Bash/Shell scripting or automation for system administration tasks.
- Experience in programming languages such as Python Golang Java or similar focusing on operational efficiency would be an added advantage.
- Demonstrated experience in system architecture and design prioritizing reliability and scalability.
- Understanding of SRE principles including SLOs SLIs toil reduction and incident post-mortems would be an added advantage.
- Hands-on experience with cloud environments (e.g. AWS Azure Google Cloud) and their operational management.
- Excellent problem-solving abilities and a proactive approach to operational challenges.
- Ability to work independently while effectively collaborating within a team environment.
- Open to a rotational shift schedule across different time slots with reasonable schedules shared in advance.
- Able to communicate effectively in Mandarin would be an added advantage.
Preferred Skills:
- Observability & Monitoring: Prometheus Grafana Alertmanager Loki Jaeger/Tempo OpenTelemetry
- Containerization & Orchestration: Kubernetes Helm service mesh (Istio/Linkerd)
- Big Data & Streaming: Apache Flink Kafka Spark; experience with large-scale data pipelines
- Infrastructure as Code & Automation: Terraform Ansible CI/CD pipelines
- Cloud Platforms: AWS Azure GCP; multi-region and hybrid cloud deployments
- Programming & Scripting: Python Go Bash or Java for automation and tooling
- Resiliency & Reliability Engineering: Incident response RCA chaos engineering disaster recovery (DR/BCP)
Due to volume of applications we regret that only shortlisted candidates will be notified.
Please note that Deloitte will never reach out to you directly via messaging platforms to offer you employment opportunities or request for money or your personal information. Kindly apply for roles that you are interested in via this official Deloitte website.
#LIMH