HPC Performance & Validation Engineer
Dallas, TX - USA
Job Summary
Location: Dallas TX
Our client is seeking an experienced HPC Performance & Validation Engineer to help ensure large-scale GPU infrastructure is ready for production and consistently delivers the performance required for demanding AI machine learning and research workloads.
This role combines GPU performance engineering infrastructure benchmarking automated validation and observability across compute storage and networking. The engineer will establish testing standards build repeatable validation frameworks and investigate performance bottlenecks across a distributed HPC environment.
Working within the Architecture organization this individual will partner with engineering infrastructure and research teams to turn benchmark results into practical improvements and guide future infrastructure decisions.
- Design and implement validation frameworks to verify GPU node readiness utilization and production performance.
- Establish repeatable methodologies for evaluating AI/ML workload performance across large-scale GPU clusters.
- Develop and execute industry-standard and workload-specific benchmarks across compute storage and networking.
- Investigate performance bottlenecks identify root causes and coordinate improvements with the appropriate engineering teams.
- Establish baseline performance metrics and continuous validation practices to measure reliability and efficiency as infrastructure evolves.
- Build scalable validation tools and micro-benchmarking frameworks using Python Go and Kubernetes.
- Integrate automated testing and benchmarking into CI/CD pipelines.
- Implement monitoring and dashboards to track cluster health utilization and performance using Prometheus Grafana OpenTelemetry and ELK.
- Define an observability strategy that supports performance analysis troubleshooting and ongoing infrastructure validation.
- Translate benchmark findings into recommendations for infrastructure design tuning and capacity decisions.
- Document testing methodologies hardware evaluations and performance findings in technical reports.
- Partner with engineering infrastructure and research teams to align validation efforts with workload requirements and business priorities.
- Evaluate emerging hardware tools and architectures to inform long-term infrastructure planning.
- Experience profiling and tuning large-scale GPU clusters and accelerator-based infrastructure.
- Strong knowledge of NVIDIA ClusterKit Nsight NVIDIA validation tools MLPerf and DCGM including GPU and DPU performance assessment.
- Experience benchmarking and optimizing network and storage performance across InfiniBand and RoCE environments using ClusterKit iPerf or comparable tools.
- Hands-on experience with Linux system benchmarking including the Phoronix Test Suite or equivalent.
- Strong proficiency developing automation and micro-benchmarking tools using Python Go and Kubernetes in an Ubuntu Linux environment.
- Experience supporting HPC workloads across geographically distributed environments and using performance data to guide architectural decisions.
- Strong knowledge of OpenTelemetry Prometheus ELK and Grafana with experience defining observability practices for HPC infrastructure.
- Ability to evaluate emerging technologies and incorporate relevant findings into infrastructure strategy.
- Demonstrated ability to lead complex technical initiatives communicate findings and influence decisions across engineering and research teams.
Bachelors degree in Computer Science Engineering or a related field or equivalent professional experience.