HPC Engineer (AL-FNC1)
Job Summary
We are looking for HPC Engineers to design deploy and operate high-performance computing infrastructure supporting research simulation and data-intensive workloads.
You will manage compute clusters parallel storage high-speed interconnects and job scheduling environments.
The role requires deep knowledge of HPC hardware and software alongside strong operational skills for cluster administration performance optimisation and user support.
- Design deploy and manage HPC compute clusters and GPU infrastructure
- Administer job scheduling platforms such as Slurm PBS Pro or LSF
- Maintain and optimise parallel file systems including Lustre BeeGFS or IBM Spectrum Scale (GPFS)
- Manage high-speed interconnect technologies such as InfiniBand and Omni-Path
- Deploy and support scientific computing software MPI frameworks compilers and research applications
- Monitor cluster health utilisation job performance and storage capacity
- Troubleshoot system application and infrastructure-related issues
- Implement automation and configuration management using tools such as Ansible
- Collaborate with researchers faculty members IT teams and technology vendors to support research workloads
- Maintain technical documentation capacity planning records and operational procedures
- Support security compliance and governance requirements within the HPC environment
Requirements
- Degree in Computer Science Engineering Information Technology or a related discipline
- 2 to 5 years of experience in Linux systems administration infrastructure operations or HPC environments
- Experience with cluster management job schedulers or large-scale Linux platforms
- Knowledge of storage networking and performance tuning concepts
- Familiarity with automation and scripting tools
- 5 to 10 years of experience supporting HPC scientific computing or large-scale distributed infrastructure
- Strong expertise in HPC cluster architecture and operations
- Experience supporting GPU environments and accelerator technologies
- Hands-on experience with parallel file systems and high-speed interconnect fabrics
- Proven track record in performance optimisation capacity planning and enterprise-scale operations
Experience in several of the following areas:
- Linux Administration (Red Hat / Rocky Linux / CentOS)
- Slurm PBS Pro LSF
- NVIDIA GPU Platforms
- MPI (OpenMPI Intel MPI)
- InfiniBand / Omni-Path
- GPFS / Spectrum Scale Lustre BeeGFS
- Ansible Infrastructure Automation
- Singularity / Apptainer Containers
- Spack Environment Modules Lmod
- Performance Monitoring and Capacity Management
Required Experience:
IC
About Company
The best of Xcellink today is the result of having evolved through more than 2 decades of Enterprise ICT Operations management experience and capabilities development as a trusted vendor partner to high-growth global companies, estalished local enterprises and government-linked corpor ... View more