Senior Software Engineer – Fleet Automation GPU Infrastructure
Dallas, TX - USA
Job Summary
Location: Dallas TX
Work Model: Hybrid 3 days onsite
Relocation: Available
Compensation: $170000$220000 base bonus
Benefits: 100% company-paid benefits
Employment Type: Direct Hire
Our client is seeking a Senior Software Engineer to join its Fleet Automation team supporting large-scale HPC and GPU infrastructure.
This team builds the software automation and internal platforms used to provision configure monitor and manage hundreds of high-performance GPU and CPU compute nodes. The environment sits at the intersection of software engineering infrastructure and distributed systems with a strong focus on eliminating manual operational work through scalable automation.
This is a hands-on engineering role for someone who enjoys building backend services while also understanding how those services interact with Linux systems physical hardware networking storage and GPU infrastructure.
- Design and build fleet automation platforms for provisioning configuration validation and lifecycle management of GPU and CPU compute nodes.
- Develop internal services and APIs that automate hardware deployment imaging remediation and decommissioning.
- Build reliable backend services using Go C# and/or TypeScript.
- Design data models and persistent state for automation workflows using relational and NoSQL databases.
- Develop and maintain CI/CD pipelines for infrastructure and configuration changes.
- Automate hardware validation and testing across large-scale compute environments.
- Build observability monitoring dashboards and alerting using tools such as Prometheus Grafana Alertmanager and ELK.
- Partner closely with Infrastructure Network Operations and Research teams to identify operational pain points and automate repeatable processes.
- Participate in on-call rotations and support incident response root-cause analysis and post-incident reliability improvements.
- Identify systemic infrastructure issues and develop engineering solutions that improve fleet reliability efficiency and scalability.
- 5 years of software engineering experience building production backend services infrastructure platforms or automation tooling.
- Strong development experience with at least one of the following:
- Go
- C#
- TypeScript
- Experience designing APIs backend services and distributed or stateful systems.
- Strong experience with relational and/or NoSQL databases.
- Solid Linux systems knowledge including:
- Networking
- Storage
- Process management
- System troubleshooting
- Ubuntu and/or RHEL environments
- Experience building and maintaining CI/CD pipelines.
- Hands-on experience with production monitoring and observability platforms such as:
- Prometheus
- Grafana
- Alertmanager
- ELK
- Strong troubleshooting and problem-solving skills across both software and infrastructure environments.
- Experience supporting GPU HPC AI/ML or large-scale compute infrastructure.
- Familiarity with NVIDIA technologies such as:
- DCGM
- nvidia-smi
- NVIDIA Container Toolkit
- Experience with bare-metal provisioning and hardware lifecycle automation.
- Exposure to event-driven architectures and messaging platforms such as Kafka.
- Experience working with infrastructure network SRE or platform engineering teams.
- Bachelors degree in Computer Science Software Engineering or equivalent practical experience.
- Work directly on large-scale GPU and HPC infrastructure supporting advanced AI workloads.
- Build automation platforms that have direct impact on infrastructure reliability and scalability.
- Highly technical environment combining software engineering systems engineering and infrastructure automation.
- Competitive $170K$220K base salary bonus.
- 100% company-paid benefits.
- Relocation assistance available for candidates moving to Dallas.
- Hybrid schedule with three days per week onsite.