Enter a job title or keyword

Site Reliability Engineer


Job Location:

Lahore - Pakistan

Monthly Salary: Not provided by the employer
Posted: 2 September 2026 (10 days ago)
Application Deadline: 30 November 2026
Vacancies: 1 Vacancy

Job Summary

Requirements:

  • 5 years of experience in systems infrastructure or SRE engineering operating production systems at scale.
  • Bachelors degree in Computer Science.
  • Deep Linux troubleshooting skills across the OS networking storage and performance with hands-on experience working as root on production systems.
  • Experience with Ubuntu is highly relevant as it is used almost exclusively.
  • Hands-on experience operating GPU servers in production including troubleshooting driver device and hardware-level issues rather than only the workloads running on top of them.
  • Practical network troubleshooting experience including diagnosing physical-layer faults.
  • Strong automation mindset with programming skills in Python or a comparable language.
  • Experience with configuration management node provisioning and infrastructure-as-code (IaC) using Ansible Terraform or similar tools.
  • Experience building observability and alerting solutions using Grafana and Prometheus.
  • Experience operating GPU clusters or AI infrastructure at production scale.
  • Production experience with Kubernetes or Slurm; experience with both is a bonus.
  • Background in HPC or research computing.
  • Familiarity with the NVIDIA GPU stack InfiniBand/RDMA and NCCL.
  • Experience with CLI-based AI coding agents such as Claude Code rather than browser-based assistants alone.
  • Contributions to open-source projects within the cloud-native HPC or AI infrastructure ecosystem.

Responsibilities:

  • Own the reliability availability and performance of production Linux GPU clusters covering the operating system drivers GPUs high-speed networking and storage.
  • Lead deep end-to-end troubleshooting of complex distributed systems GPU nodes networking and storage issues.
  • Troubleshoot and resolve GPU rail and NCCL performance issues across multi-GPU and multi-node collective communication paths.
  • Diagnose network faults end to end including configuration routing and physical-layer issues such as cabling transceivers and link errors.
  • Configure and maintain workload managers that schedule customer jobs including Kubernetes Slurm or both along with the identity storage and networking services they depend on.
  • Build automation and tooling to eliminate operational toil using a modern language such as Python or Go to design solutions review implementations and redirect approaches when needed.
  • Use AI-assisted engineering tools such as Claude to accelerate automation runbook development and incident analysis.
  • Automate provisioning image deployment configuration and remediation using Ansible and infrastructure-as-code.
  • Design and operate observability using Grafana Prometheus and Loki while tuning alerts for meaningful signals and building self-healing capabilities that reduce the need for human intervention.
  • Lead incident response on-call activities blameless postmortems and reliability improvements that maintain customer SLAs.
  • Partner with Platform and Systems Engineering teams on capacity planning rollouts and continuous improvement.

Working Hours:

8 PM - 4 AM