HPC Orchestration Architect
Dallas, TX - USA
Job Summary
Location: Dallas TX
Work Arrangement: Hybrid
Employment Type: Direct Hire
Relocation: Available for qualified non-local candidates
Compensation: Competitive Base Salary Performance Bonus
Benefits: 100% Company-Paid Family Benefits
Our client is seeking an experienced HPC Orchestration Architect to define and evolve the orchestration architecture supporting a large-scale high-performance computing and AI infrastructure platform.
This role will own architecture across compute orchestration virtualization containers Kubernetes and resource management helping ensure highly scalable and highly available infrastructure can support demanding research AI/ML simulation and scientific computing workloads.
The HPC Orchestration Architect will work closely with compute storage networking and platform engineering teams to develop cohesive reference architectures evaluate emerging technologies and translate complex workload requirements into scalable production solutions.
This is a highly technical architecture role for someone who understands how compute GPUs storage networking virtualization and Kubernetes must work together as an integrated HPC platform.
- Own the architecture and long-term strategy for compute orchestration across large-scale HPC and accelerated computing environments.
- Design highly scalable resilient and distributed systems supporting demanding research AI/ML and scientific workloads.
- Define Kubernetes and container orchestration architectures supporting large-scale compute platforms.
- Evaluate existing and emerging HPC virtualization container and orchestration technologies against business performance and operational requirements.
- Develop reference architectures and technical standards for implementation by infrastructure and engineering teams.
- Partner with compute storage and networking architects to ensure orchestration decisions align with the overall HPC platform architecture.
- Evaluate resource placement scheduling workload isolation utilization and performance tradeoffs across CPUs GPUs and other accelerators.
- Guide engineering teams through architecture implementation and ensure solutions meet scalability reliability and engineering quality standards.
- Identify opportunities to improve infrastructure automation developer velocity operational efficiency and platform reliability.
- Establish data-driven methods for measuring and improving platform performance capacity utilization and operational constraints.
- Evaluate new technologies through technical research proof-of-concept testing benchmarking and production validation.
- Support architecture decisions involving Kubernetes clusters virtualization platforms containers GPU infrastructure and distributed systems.
- Work directly with technical stakeholders and customers to understand workload requirements and translate them into scalable infrastructure designs.
- Stay current on emerging HPC AI infrastructure cloud-native Kubernetes and accelerated computing technologies.
- Strong experience designing or architecting large-scale HPC AI infrastructure cloud or distributed computing environments.
- Deep understanding of compute architecture including CPUs GPUs accelerators and their role within large-scale distributed systems.
- Strong hands-on experience with Kubernetes including building deploying scaling and operating production Kubernetes clusters.
- Deep knowledge of container technologies container orchestration virtualization and workload management.
- Understanding of Kubernetes resource management scheduling placement isolation and performance optimization.
- Experience designing highly available and scalable distributed systems.
- Strong Linux expertise including system-level optimization performance profiling troubleshooting and kernel-level tuning.
- Experience with HPC clusters parallel computing environments or large-scale accelerated compute platforms.
- Understanding of how compute networking and storage architecture interact to influence workload performance.
- Familiarity with distributed and parallel storage platforms such as VAST Data Lustre or similar technologies.
- Experience evaluating infrastructure technologies through benchmarking proof-of-concept environments and performance analysis.
- Ability to translate technical and business requirements into scalable infrastructure architectures.
- Strong communication skills with the ability to work across architecture engineering operations vendors and technical customers.
Experience with one or more of the following is highly valued:
- Large-scale GPU or AI/ML clusters
- NVIDIA GPU infrastructure and accelerated computing
- Kubernetes scheduling and advanced resource management
- HPC workload orchestration and scheduling platforms
- Bare-metal and virtualized compute environments
- Infrastructure automation and platform engineering
- Distributed and parallel storage
- Scientific computing simulation or research workloads
- Multi-cluster or geographically distributed Kubernetes environments
- Cloud-native infrastructure and hybrid cloud architectures
Bachelors or Masters degree in Computer Science Computer Engineering Electrical Engineering Physics or a related technical discipline preferred. Equivalent advanced professional experience in HPC distributed systems platform engineering or large-scale infrastructure will also be considered.
- Competitive base salary
- Performance bonus
- 100% company-paid family medical benefits
- Relocation assistance available for qualified non-local candidates
- Opportunity to architect infrastructure supporting large-scale HPC AI and accelerated computing environments