Systems Engineer
Posted:
2 September 2026 (11 days ago)
Application Deadline:
30 November 2026
Vacancies:
1 Vacancy
Job Summary
Requirements:
- 5 years of experience in Linux systems administration or infrastructure engineering at scale.
- Strong knowledge of Linux internals with hands-on troubleshooting skills across the OS kernel storage and system performance.
- Solid understanding of TCP/IP networking fundamentals with hands-on network troubleshooting experience.
- Hands-on experience operating NFS and other shared storage systems in production environments.
- Proven experience with configuration management using Ansible.
- Experience with monitoring tools such as Grafana and ticketing tools such as ServiceNow and JIRA.
- Proficiency in Bash and Python scripting for automation.
- Bachelors degree in Computer Science Engineering or an equivalent field/experience.
- Experience operating GPU servers HPC or AI infrastructure.
- Familiarity with Kubernetes Slurm or other cluster schedulers.
- Exposure to storage technologies such as Lustre GPFS/Spectrum Scale or Ceph.
- Knowledge of InfiniBand or RDMA and high-performance networking.
- Familiarity with or working knowledge of AI coding tools such as Claude OpenAI or other similar tools to accelerate automation and troubleshooting.
- Relevant certifications such as RHCSA/RHCE CCNA or cloud provider certifications.
Responsibilities:
- Administer patch and harden a large fleet of Linux servers including RHEL Rocky and Ubuntu across bare-metal and cloud environments.
- Troubleshoot complex OS kernel performance and hardware issues to identify and resolve root causes.
- Design configure and manage TCP/IP networking including routing VLANs DNS DHCP firewalls and network bonding.
- Deploy and manage NFS storage and other shared file systems for high-throughput workloads.
- Automate provisioning configuration and remediation using Ansible and infrastructure-as-code practices.
- Build and maintain monitoring alerting and dashboards using Grafana and related observability tools.
- Manage operational workflows through ticketing systems such as ServiceNow (SNOW) and JIRA.
- Own capacity planning OS lifecycle management standardized system builds and image management.
- Collaborate with Platform Engineering and SRE teams on reliability security and the rollout of new services.
- Document standards runbooks and procedures while continuously reducing operational toil.
Working Hours:
8 PM - 4 AM