Shape the future of AI infrastructure! Join RRL and develop system-level test solutions for our innovative ML acceleration hardware deployed across our global server fleet.
Were seeking highly skilled and motivated Product Test Engineers to join our RRL test & repair this role youll be at the forefront of validating and ensuring that test infrastructure is deployed and operating at required scale. Youll be responsible for designing and implementing comprehensive system-level test strategies that cover the full spectrum of our ML acceleration products from individual components to fully integrated systems. This position requires a unique blend of hardware knowledge software expertise and systems thinking as youll be working at the intersection of custom silicon complex firmware and high-performance ML workloads.
Youll collaborate closely with cross-functional teams including hardware designers software engineers and operations specialists to develop robust test solutions that can scale to meet the demands of RRLs global infrastructure. Your work will be crucial in identifying and resolving integration issues optimizing system performance and ultimately ensuring that our ML acceleration products meet the highest standards of reliability and efficiency in real-world data center environments. If youre passionate about pushing the boundaries of ML hardware testing and have a knack for solving complex system-level challenges we want you on our team.
Key job responsibilities
* Design and implement system-level test strategies for ML acceleration products
* Develop comprehensive functional and performance tests for complete ML systems
* Create and maintain scalable test infrastructure for high-volume product validation
* Implement product bring-up and first-boot test procedures
* Drive improvements in test coverage product quality and manufacturing efficiency
* Collaborate with hardware and software teams to ensure end-to-end product validation
* Analyze system-level test data to identify and resolve integration issues
* Debug complex hardware/software interactions in a production environment
* Develop and maintain documentation for system test procedures and manufacturing processes
* Optimize test workflows to balance thoroughness with production efficiency
- 4 years of site reliability engineering (SRE) systems engineering systems administration DevOps security administration or network administration experience
- 5 years of Linux experience
- 5 years of systems engineering experience
- Bachelors degree in Systems Engineering Computer Science or related field or relevant work experience
- Experience in site reliability engineering (SRE) systems engineering systems administration DevOps security administration or network administration
- Experience working with Linux
- Experience in systems engineering
- Experience in any of the following: Python Java Perl PHP Ruby Bash Shell or equivalent
- Knowledge of TCP/IP and networking protocols such as HTTP and DNS
- Experience designing and developing scripts to automate operational burdens and reviewing scripting changes to ensure they meet the standards for maintainability scalability and security
- Experience working in 24/7 production environment
- Experience with service-oriented architecture and web services
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status disability or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process including support for the interview or onboarding process please visit
for more information. If the country/region youre applying in isnt listed please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience qualifications and location. Amazon also offers comprehensive benefits including health insurance (medical dental vision prescription Basic Life & AD&D insurance and option for Supplemental life plans EAP Mental Health Support Medical Advice Line Flexible Spending Accounts Adoption and Surrogacy Reimbursement coverage) 401(k) matching paid time off and parental leave. Learn more about our benefits at KY Florence - 104500.00 - 160000.00 USD annually