High-Performance Computing (HPC) Engineer
El Segundo, CA - USA
Job Summary
The Aerospace Corporation is the trusted partner to the nations space programs solving the hardest problems and providing unmatched technical expertise. As the operator of a federally funded research and development center (FFRDC) we are broadly engaged across all aspects of space delivering innovative solutions that span satellite launch ground and cyber systems for defense civil and commercial customers. When you join our team youll be part of a special collection of problem solvers thought leaders and innovators. Join us and take your place in space.
The Aerospace Corporation is seeking a talented High-Performance Computing (HPC) Engineer (Site Reliability Engineer Staff III/IV) to join our Computational Services this role you will develop implement and optimize HPC clusters that support both on-premises and cloud environments. You will work alongside rocket scientists and engineers tackling complex space enterprise challenges while having direct impact on critical national security missions. We value a collaborative proactive mindset and a shared commitment to engineering excellence.
The selected candidate will be required to work full-time on-site at our facility in El Segundo CA or Chantilly VA.
What Youll Be Doing
- Collaborate with scientists and engineers on diverse projects supporting mission-critical technical analysis for national space assets
- Lead cross-functional teams and mentorjunior engineers.
- Design and implement HPC solutions that optimize resource utilization across diverse workloads in both classified and unclassified settings.
- Manageon premise10000-core classified cluster and a5000-core unclassified cluster to ensure peak performance.
- Deliver high-quality HPC infrastructure design and system configuration.
- Develop and deploy automation solutions using tools such asClush.
- Manage infrastructure using Infrastructure-as-Code and GitOps practices
- Implement supportandoptimizeGPU computing.
- Monitor analyze and tune HPC system performance utilization and resource allocation to maintain operational efficiency.
- Develop cost-efficientHPC service offerings that align with mission and business objectives.
- Harden Linux systems to meet stringent security requirements
What You Need to be Successful
Minimum Requirements for theSite Reliability Engineer Staff III:
- Bachelors degree in Computer Science Engineering or equivalent experience.
- Minimum of7years experience in Linux system administration within an enterprise HPC environment.
- Experience supporting technical software (compilers mod&sim tools languages COTS GOTs) including the development of environment modules.
- In-depth knowledge of Linux networking and HPC systems.
- Experience withInfrastructure-as-Code andGitOps
- Proven experience in managing the Slurm scheduler and setting up HPC systems for both interactive and batch workloads.
- Experience provisioning and supporting AI & NVIDIA GPU technologies(e.g. CUDA)
- Proficiency in scripting and competence with automation tools such asClush.
- Experience hardening Linux systems to meet security requirements
- Experience with hardware and infrastructure automation in environments using server vendors such as HPE or Cisco.
- Strong communication skills with an ability to work both independently and as part of a geographically distributed team.
- CompTIA Security CE certification or equivalent that meets DoD 8570.01-m requirements for IAT Level II personnel
- Ability to obtain and maintain a TS/SCI clearance (U.S. citizenship required).
- Demonstrated ability to lead cross-functional teams and mentor junior engineers.
In addition to the above the minimum requirements for theSite Reliability Engineer Staff IVinclude:
- 9 years of experience in an enterprise100 serverHPCclusteroperations and administration
- Experiencewith performance analysis and optimizationwithcustom developedtechnical softwarein collaboration with scientists and engineers.
- Expertise in optimizing and customizing Slurm partitions qualities of serviceand priorityto balance utilization and reduce job wait times.
- Experience performing in-place upgrades of Slurm.
- Implementing visualization of live system telemetry
- Experiencedevelopingand architectingclusterconfiguration management
- AdvancedInfrastructure-as-CodeGitOps(e.g. multi-branchpipelines)
- Skill in provisioning and supporting AI & NVIDIA GPU technologies(e.g. CUDANsight) with expertise in GPU integration resource allocation and scheduling using Slurm.
How You Can Stand Out
It would be impressive if you have one or more of these:
- An active TS/SCI clearance with CI Polygraph.
- Experience integrating Slurm with SELinux
- Experienceimplementing DISA STIG compliance
- ExperienceintegratingKubernetes and Slurm e.g. Slinky
- Experience supportingand managingdiverse HPC workloads includingcomputational fluid dynamicsMonteCarlo structural analysis
- Experienceintegratinguser web portalstolaunch & manageworkloads ( OnDemandanddeveloping plugins for session persistence and VSCode)
- Experienceimplementingutilization dashboards ()with Slurm
- Experience managing parallel file systems such as Lustre.
- Experience developing solutions that optimize data storage
- Experienceimplementingor supporting Slurm REST API
- Experience withautomated provisioning Kickstart PXE
- Knowledge of NVLINKand DCGMfor optimizing GPU workflows.
- Familiarity with Prometheus and Grafana for monitoring and performance visualization.
- Background in containerization within an HPC context usedfor data processing and technical analysis
- Experiencepackaging custom software(e.g. RPMs)
- Proficiency with automation tools such as Ansible for HPC
- Hands-on background with cloud HPC services
- Experience with AWS Parallel Computing Service (AWS ParallelCluster).
We offer a competitive compensation package where youll be rewarded based on your performance and recognized for the value you bring to our business. The grade-based pay range for this job is listed below. Individual salaries within that range are determined through a wide variety of factors including but not limited to education experience knowledge and skills.
(Min - Max)
$135200.00 - $202800.00Pay Basis: AnnualLeadership Competencies
Our leadership philosophy is simple: every employee regardless of level and role can demonstrate leadership. At Aerospace our commitment is our people. To cultivate our talent and ensure that we have a strong pipeline of future leaders we want individuals who:
- Operate Strategically
- Lead Change
- Engage with Impact
- Foster Innovation
- Deliver Results
Ways We Reward Our Employees
During your interview process our team will provide details of our industry-leading benefits.
Benefits vary and are applicable based on Job Type. A few highlights include:
Comprehensive health care and wellness plans
Paid holidays sick time and vacation
Standard and alternate work schedules including telework options
401(k) Plan Employees receive a total company-paid benefit of 8% 10% or 12% of eligible compensation based on years of service and matching contributions; employees are immediately eligible and vested in the plan upon hire
Flexible spending accounts
Variable pay program for exceptional contributions
Relocation assistance
Professional growth and development programs to help advance your career
Education assistance programs
An inclusive work environment built on teamwork flexibility and respect
We are all unique from various backgrounds and all walks of life yet one thing bonds all of us to each otherthe belief that we can make a difference. This core belief empowers us to do our best work at The Aerospace Corporation.
Equal Opportunity Commitment
The Aerospace Corporation is an equalopportunity employer. All qualified applicants will receive consideration for employment and will not be discriminated against on the basis of race age sex (including pregnancy childbirth and related medical conditions) sexual orientation gender gender identity or expression colorreligiongeneticinformation marital status ancestry national origin protected veteran status physical disability medical condition mental disability or disability status and any other characteristic protected by state or federal law. If youre an individual with a disability or a disabled veteran who needs assistance using our online job search and application tools or need reasonable accommodation to complete the job application process please contact us by phone at 310.336.5432 or by emailat .You can also reviewKnow Your Rights: Workplace Discrimination is Illegal.
Required Experience:
IC
About Company
Aerospace operates the only federally funded research and development center (FFRDC) committed exclusively to the space enterprise. Our technical experts span every discipline of space-related science and engineering.