Site Reliability Engineer
Posted:
2 September 2026 (10 days ago)
Application Deadline:
30 November 2026
Vacancies:
1 Vacancy
Job Summary
Requirements:
- 5 years of experience in systems infrastructure or SRE engineering operating production systems at scale.
- Bachelors degree in Computer Science.
- Deep Linux troubleshooting skills across the OS networking storage and performance with hands-on experience working as root on production systems.
- Experience with Ubuntu is highly relevant as it is used almost exclusively.
- Hands-on experience operating GPU servers in production including troubleshooting driver device and hardware-level issues rather than only the workloads running on top of them.
- Practical network troubleshooting experience including diagnosing physical-layer faults.
- Strong automation mindset with programming skills in Python or a comparable language.
- Experience with configuration management node provisioning and infrastructure-as-code (IaC) using Ansible Terraform or similar tools.
- Experience building observability and alerting solutions using Grafana and Prometheus.
- Experience operating GPU clusters or AI infrastructure at production scale.
- Production experience with Kubernetes or Slurm; experience with both is a bonus.
- Background in HPC or research computing.
- Familiarity with the NVIDIA GPU stack InfiniBand/RDMA and NCCL.
- Experience with CLI-based AI coding agents such as Claude Code rather than browser-based assistants alone.
- Contributions to open-source projects within the cloud-native HPC or AI infrastructure ecosystem.
Responsibilities:
- Own the reliability availability and performance of production Linux GPU clusters covering the operating system drivers GPUs high-speed networking and storage.
- Lead deep end-to-end troubleshooting of complex distributed systems GPU nodes networking and storage issues.
- Troubleshoot and resolve GPU rail and NCCL performance issues across multi-GPU and multi-node collective communication paths.
- Diagnose network faults end to end including configuration routing and physical-layer issues such as cabling transceivers and link errors.
- Configure and maintain workload managers that schedule customer jobs including Kubernetes Slurm or both along with the identity storage and networking services they depend on.
- Build automation and tooling to eliminate operational toil using a modern language such as Python or Go to design solutions review implementations and redirect approaches when needed.
- Use AI-assisted engineering tools such as Claude to accelerate automation runbook development and incident analysis.
- Automate provisioning image deployment configuration and remediation using Ansible and infrastructure-as-code.
- Design and operate observability using Grafana Prometheus and Loki while tuning alerts for meaningful signals and building self-healing capabilities that reduce the need for human intervention.
- Lead incident response on-call activities blameless postmortems and reliability improvements that maintain customer SLAs.
- Partner with Platform and Systems Engineering teams on capacity planning rollouts and continuous improvement.
Working Hours:
8 PM - 4 AM