Senior Machine Learning Engineer, Alexa-Conv A Modeling&Learning
Bellevue, WA - USA
Department:
Job Summary
We are looking for a Senior Machine Learning Engineer to build and own core systems in this agentic platform. You will take one of its foundational areas - agentic evaluation infrastructure reinforcement learning training systems self-learning pipelines or agentic inference serving - and own it end to end: the design the implementation the operational bar and the interfaces that scientists and partner teams build on. You will work directly with applied scientists work backwards from committed product launches and turn research prototypes into infrastructure that runs unattended at scale.
The work is concrete. Agents are evaluated in sandboxed recreatable environments at hundreds of concurrent trials and every source of infrastructure noise you remove is a model decision the organization can trust. They are trained on long-horizon multi-turn trajectories where the rollout and learner engines have to agree token for token. They are served under latency budgets measured in hundreds of milliseconds. And they improve week over week only if the pipeline that turns production experience into training data actually holds. You will own a piece of that loop make it reliable and make it fast.
This is a platform role with room to grow. The systems you own serve every Alexa agent program rather than a single product and the engineer who makes them dependable becomes the person the organization routes its hardest cross-system problems to.
Key job responsibilities
Design build and operate major components of the agentic AI platform: evaluation harnesses sandboxed environments and mocked resources RL and post-training pipelines self-learning data pipelines or inference serving for agentic traffic
Lead the design work in your area: write the design documents drive them through review and make the build-versus-adopt calls within your scope
Own reliability and performance: instrument your systems drive down the failure modes that make results untrustworthy (process management resource contention unreliable external calls) and report platform health in metrics rather than anecdotes
Partner with applied scientists to turn research code into production infrastructure and expose it through interfaces other teams can use without your involvement
Scale what you build: hundreds of concurrent evaluation trials long-context multi-turn training jobs and large GPU formations on shared company infrastructure
Raise the engineering bar through code and design reviews operational excellence practices and deep dives on cross-system problems
Mentor engineers earlier in their careers and help set technical direction for your team
A day in the life
You might spend the morning making the evaluation platform reproducible under high concurrency tracking down why scores drift when a hundred trials share a host midday pairing with a scientist to get a long-context training job to converge identically across the rollout and learner engines and the afternoon in a design review deciding how environment snapshots should be versioned and served to partner teams. You work daily with applied scientists and other engineers and your systems are the reason their results are trustworthy and shippable.
About the team
Our organization owns the applied science and platform engineering for Alexas agentic experiences. We operate at the intersection of large language models reinforcement learning with verifiable rewards agentic architectures and large-scale distributed systems serving customers across dozens of languages and device types. Our platform provides the shared evaluation training self-learning and serving foundation for Alexas flagship agent programs and the broader agent portfolio behind them.
- 5 years of non-internship professional software development experience
- 5 years of programming with at least one software programming language experience
- 5 years of leading design or architecture (design patterns reliability and scaling) of new and existing systems experience
- Experience as a mentor tech lead or leading an engineering team
- 5 years of full software development life cycle including coding standards code reviews source control management build processes testing and operations experience
- Bachelors degree in computer science or equivalent
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status disability or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process including support for the interview or onboarding process please visit for more information. If the country/region youre applying in isnt listed please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience qualifications and location. Amazon also offers comprehensive benefits including health insurance (medical dental vision prescription Basic Life & AD&D insurance and option for Supplemental life plans EAP Mental Health Support Medical Advice Line Flexible Spending Accounts Adoption and Surrogacy Reimbursement coverage) 401(k) matching paid time off and parental leave. Learn more about our benefits at WA Bellevue - 168100.00 - 227400.00 USD annually
Required Experience:
Senior IC
About Company
Free shipping on millions of items. Get the best of Shopping and Entertainment with Prime. Enjoy low prices and great deals on the largest selection of everyday essentials and other products, including fashion, home, beauty, electronics, Alexa Devices, sporting goods, toys, automotive ... View more