Enter a job title or keyword

Software Development Engineer, Infrastructure Reliability Engineering

Amazon


Job Location:

Arlington, TX - USA

Yearly Salary: USD 136500 - 184700
Posted: 29 September 2026 (16 hours ago)
Application Deadline: 27 December 2026
Vacancies: 1 Vacancy

Department:

Software Development

Job Summary

Join us in building Reflex the agentic incident-response platform for Amazons fulfillment network. Youll design and ship production software and AI agents on Amazon Bedrock AgentCore that triage high-severity incidents generate real-time call intelligence draft stakeholder communications and produce structured post-incident records shifting incident response from a manual pull-based model to an intelligent push-based one.

Amazons network of fulfillment centers is the infrastructure that Amazon Robotics runs on. When it degrades robots stop and packages stop moving. The team manages thousands of high-severity incidents every year with Incident Managers assembling context across many systems under time pressure before resolution work can even begin. The software you build puts that context in front of responders in seconds and removes the repetitive work that extends incidents today. Youll design the data evaluation and feedback mechanisms that make the system measurably better with every incident it touches.

This role sits on a software team within Operations Infrastructure Services (OIS) part of Amazon Robotics. Youll stay close to live operations through incident reviews workflow observation and call shadowing and turn what you learn into durable software. Reflex is the immediate focus; as it matures the same foundations are expected to extend toward shared incident context across organizations coordinated agent workflows and carefully guarded automation of recovery validation and repeatable response actions with every step gated by measurable confidence.


Key job responsibilities
- Design build test deploy and operate production services and AI agents on AWS (Amazon Bedrock AgentCore serverless compute event-driven pipelines) that automate incident triage call intelligence communications and post-incident documentation and reporting
- Own features end-to-end: from discovery with Incident Managers and resolver teams through design implementation evaluation deployment and production operation
- Build the foundations that gate agent autonomy: LLM output evaluation observability and alerting for agents in production and identity and access controls aligned with Amazon standards
- Design the feedback loops through which agents learn: capturing human reviews corrections approvals and incident outcomes as evaluation signal and turning resolved incidents into structured history that improves recommendations over time
- Integrate with the incident lifecycle (ticketing chat telemetry detection feeds and live call transcription) and model consistent incident state across those systems
- Raise the bar on software quality security testing and operational excellence for AI systems acting inside production incident workflows

A day in the life
You start by reviewing overnight agent evaluation results and fixing a class of inaccurate triage recommendations shipping an improvement responders see on the next incident. Later you pair with an Incident Manager to see how they corrected an agent-drafted call summary then turn that correction into an automated evaluation case and an agent fix. You might close the day reviewing a design for representing incident state across source systems or shadowing part of a live bridge call to spot where responders lose time. The operational insight you gather today becomes the software that shortens tomorrows incident.

Amazon offers a full range of benefits to support you and eligible family members including domestic partners and children. Benefits can vary by location the number of regularly scheduled hours you work length of employment and job status such as seasonal or temporary employment. The benefits that generally apply to regular full-time employees include:
1. Medical Dental and Vision Coverage
2. Maternity and Parental Leave Options
3. Paid Time Off (PTO)
4. 401(k) Plan

If you are not sure that every qualification on the list above describes you exactly wed still love to hear from you! At Amazon we value people with unique backgrounds experiences and skillsets. If youre passionate about this role and want to make an impact on a global scale please apply!

About the team
Were a software engineering team within Operations Infrastructure Services part of Amazon Robotics. We build incident-response software for major incident management and for the engineering and operations teams that keep Amazons fulfillment infrastructure healthy. Our users work in high-pressure environments where missing context unclear ownership and repetitive manual work directly extend incidents and we work backward from those problems.

- 3 years of non-internship professional software development experience
- 2 years of non-internship design or architecture (design patterns reliability and scaling) of new and existing systems experience
- 1 years of software development engineer or related occupational experience
- 1 years of designing and developing large-scale multi-tiered multi-threaded embedded or distributed software applications tools systems and services using: C# C Java or Perl experience
- 1 years of Object Oriented Design experience
- Bachelors degree or foreign equivalent in Computer Science Engineering Mathematics or a related field
- Experience programming with at least one software programming language

- 3 years of full software development life cycle including coding standards code reviews source control management build processes testing and operations experience
- Knowledge of Machine Learning and LLM fundamentals including transformer architecture training/inference lifecycles and optimization techniques

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status disability or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process including support for the interview or onboarding process please visit for more information. If the country/region youre applying in isnt listed please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience qualifications and location. Amazon also offers comprehensive benefits including health insurance (medical dental vision prescription Basic Life & AD&D insurance and option for Supplemental life plans EAP Mental Health Support Medical Advice Line Flexible Spending Accounts Adoption and Surrogacy Reimbursement coverage) 401(k) matching paid time off and parental leave. Learn more about our benefits at TN Nashville - 136500.00 - 184700.00 USD annually
USA VA Arlington - 143700.00 - 194400.00 USD annually


Required Experience:

IC


About Company

Company Logo

Free shipping on millions of items. Get the best of Shopping and Entertainment with Prime. Enjoy low prices and great deals on the largest selection of everyday essentials and other products, including fashion, home, beauty, electronics, Alexa Devices, sporting goods, toys, automotive ... View more

View Profile View Profile