Enter a job title or keyword

Hardware Diagnostics Engineer Infrastructure

TensorWave


Job Location:

Las Vegas, NV - USA

Monthly Salary: Not provided by the employer
Posted: 6 October 2026 (Yesterday)
Application Deadline: 3 January 2027
Vacancies: 1 Vacancy

Department:

Engineering

Job Summary

About TensorWave

Our mission is simple: deliver seamless secure reliable and resilient AI compute at scale. Weve built a versatile cloud platform that eliminates infrastructure barriers empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas not infrastructure.

About the Role

We are looking for a Hardware Diagnostics Engineer to run burn-in triage what fails work servers out-of-band and own RMAs end to end. If you like hardware that misbehaves in ways that take real work to explain this is a good seat.

Before any GPU server carries a customer workload it has to prove it works under load at temperature for hours. When it doesnt somebody has to figure out why get replacement hardware in and send the failed part back to the vendor.

What Youll Do

  • Run server and GPU burn-in and stress testing interpret the results and decide whether hardware is production-ready

  • Triage failures across GPUs memory drives NICs PSUs and cabling: reproduce the failure isolate the faulty component and document what proved it

  • Work servers out-of-band through IPMI and Redfish for power control boot configuration BIOS settings and sensor and event log collection

  • Apply firmware updates across the fleet following the teams qualified baselines and rollout process

  • Drive RMAs with vendors from ticket through replacement installation and return of the failed part

  • Keep asset serial and replacement history accurate in NetBox so we know whats actually in every rack

  • Track failure patterns across the fleet and raise them when the same part or firmware version keeps turning up

  • Improve the runbooks you work from and script the steps you find yourself repeating

  • Partner with datacenter operations on hands-on work during turn-ups and expansions

  • Take part in an on-call and escalation rotation for hardware issues

Who You Are

Required Qualifications

  • 36 years in datacenter operations systems administration hardware support or infrastructure engineering

  • Hands-on experience with enterprise server hardware: component replacement POST and boot failures and reading hardware behavior at the rack

  • Practical experience with BMCs and out-of-band management: IPMI Redfish iDRAC iLO or equivalent

  • Strong Linux troubleshooting: boot process driver and device issues and diagnostic tools such as dmesg lspci ipmitool and SMART

  • Comfort reading sensor data event logs and thermal and power telemetry well enough to tell a real failure from noise

  • Working scripting ability in Bash or Python enough to automate a repetitive task and read someone elses tooling

  • Experience running hardware RMAs with vendors or a clear track record of driving issues to closure with outside parties

  • A methodical troubleshooting habit: you isolate variables you dont change three things at once and you can say what evidence led to your conclusion

  • Clear written communication for tickets runbooks and vendor cases

Preferred Qualifications

  • GPU server experience especially AMD GPUs and ROCm

  • Burn-in stress testing or node validation tooling in a GPU or HPC environment

  • Familiarity with firmware update processes and why fleet-wide changes get staged

  • NetBox or other DCIM and IPAM tooling

  • Ansible or Python against REST APIs

  • Prior work in a high-volume hardware environment: hyperscaler colo integrator or manufacturing test

First Six Months

  • By 90 days youll run burn-in cycles and triage failures independently from our runbooks and youll have driven at least one RMA to closure. By six months youre the person who spots the pattern before anyone else this batch this firmware this part and youve automated at least one step you used to do by hand.

What We Offer

  • Stock Options

  • 100% paid Medical Dental and Vision insurance for Employees

  • Company Health Savings Account Contributions

  • 100% paid Short Term and Long Term Disability Insurance for Employees

  • Life and Voluntary Supplemental Insurance Options

  • Other Insurance Options such as Pet & Legal Insurance

  • Various Supplementary Health Benefits such as discounted Virtual Healthcare Appointments and Serious Illness Support

  • Flexible Spending Account

  • 401(k)

  • Employee Assistance Program

  • Flexible PTO

  • Paid Holidays

  • Parental Leave

  • Other In-Office Perks

Equal Employment Opportunity

TensorWave is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of any protected status under applicable law.

Reasonable Accommodations

TensorWave provides reasonable accommodations in accordance with applicable laws. If you require accommodation during the hiring process please contact

Employment Eligibility

All offers of employment are contingent upon verification of identity and authorization to work in United States as required by law.

Background Checks

Where permitted by law employment may be contingent upon the successful completion of a job-related background check.

Data Privacy Notice

By submitting an application you acknowledge that TensorWave may collect use and retain your personal information for recruiting and employment-related purposes in accordance with applicable data privacy laws.


Required Experience:

IC