Enter a job title or keyword

Principal AIML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp

Amazon


Job Location:

Seattle, WA - USA

Yearly Salary: USD 182800 - 247300
Posted: 13 September 2026 (18 hours ago)
Application Deadline: 11 December 2026
Vacancies: 1 Vacancy

Job Summary

As part of the AWS Applied AI Solutions organization we have a vision to provide business applications leveraging Amazons unique experience and expertise that are used by millions of companies worldwide to manage day-to-day operations. We will accomplish this by accelerating our customers businesses through delivery of intuitive and differentiated technology solutions that solve enduring business challenges. We blend vision with curiosity and Amazons real-world experience to build opinionated turnkey solutions. Where customers prefer to buy over build we become their trusted partner with solutions that are no-brainers to buy and easy to use.

Are you ready to transform how businesses leverage artificial intelligence and machine learning at scale Join our team and become a strategic partner in delivering Amazon AI/ML solutions that empower global enterprises to innovate optimize and achieve unprecedented operational excellence.

Amazon Web Services (AWS) is seeking an experienced Principal AI/ML HPC Specialist to join our Technical Account Manager (TAM) team.

Youll be at the forefront of solving complex AI HPC implementation challenges guiding NAMER Resarch labs to enterprise customers through their most ambitious machine learning transformation journeys. By combining deep technical expertise with collaborative problem-solving youll help organizations unlock the full potential of artificial intelligence and machine learning technologies from distributed model training on GPU clusters to production-grade inference at scale.

AWS Support includes experts from across AWS who help our customers design build operate and secure their cloud environments. Customers innovate with AWS Professional Services upskill with AWS Training and Certification optimize with AWS Support and Managed Services and meet objectives with AWS Security Assurance Services. Our expertise and emerging technologies include AWS Partners AWS Sovereign Cloud AWS International Product and AI/ML-native solutions. Youll join a diverse team of technical experts in dozens of countries who help customers achieve more with the AWS cloud.

Key job responsibilities
Deliver Strategic Technical Engagements Lead comprehensive technical deep-dives and performance optimization for enterprise AI/ML workloads including distributed training cluster architecture using AWS Parallel Computing Service (PCS) and AWS ParallelCluster the latest GPU-accelerated computing (i.e. P6/P6e G7/G7e instances) AWS Trainium-based training (Trn3 UltraServers) and multi-node NCCL communication tuning over EFAs SRD protocol.

Architect and Validate Innovative Solutions Design and implement production-grade AI/ML training and inference solutions leveraging Slurm-based job scheduling distributed training frameworks (PyTorch FSDP DDP DeepSpeed Megatron-LM) SageMaker HyperPod for managed GPU clusters with automated health checks and node replacement high-performance parallel storage (Amazon FSx for Lustre) and container runtimes on Deep Learning AMIs (DLAMIs) against reference architectures and HPC lens to ensure performance reliability and cost governance at scale.. Architect solutions using P6e UltraServers for multi-trillion parameter frontier models and Trn3 with the AWS Neuron SDK for cost-optimized training and inference.

Enable Customer Success Support customers in implementing business-critical HPC capabilities including the development of large language model (LLM) (Llama GPT-class models) physics-informed neural networks (PINNs) and surrogate models MLOps pipelines simulation-ML hybrid architectures orchestrated by AWS Step Functions and AWS Batch distributed data processing cluster observability and governance controls for GPU/Trainium-intensive workloads.

Enable Business Critical Outcomes Partner with with service teams to enhance model training throughput optimize NCCL collective communications improve GPU/Trainium utilization across multi-node UltraClusters and drive operational efficiency through proactive monitoring automated failure recovery (HyperPod health checks) and capacity planning (EC2 Capacity Blocks for ML). Contribute to product roadmap PFR share refrerence architecture performance and benchmarks with broader TAM and Technical communities

Serve as Trusted Advisor and Advocate Develop and nurture technical partnerships with enterprise stakeholders serving as the trusted advisor for AI/ML infrastructure decisions spanning compute networking (Elastic Fabric Adapter with SRD) storage orchestration and the HPC-to-AI convergence journey.

A day in the life
Your day will be dynamic and impactful involving deep technical consultations on distributed training architectures strategic solution design for GPU and Trainium cluster deployments and collaborative problem-solving across multi-node ML environments. Youll engage with technical leaders architect innovative AI/ML implementations from Slurm-managed PCS clusters and SageMaker HyperPod to PyTorch FSDP/DeepSpeed training jobs and Neuron SDK compilation workflows and provide expert guidance that bridges machine learning infrastructure with business objectives.

You will partner with TAMs SAs and service teams to provide customers with AWS AI/ML best practice guidance diving deep into machine learning infrastructure services (PCS ParallelCluster HyperPod Batch) promoting customers AI/ML workloads to production developing regional AI/ML strategies advising on HPC-to-AI convergence patterns (simulation-surrogate loops physics-informed neural networks) and training field teams on distributed training patterns GPU/Trainium cluster operations and the use cases and benefits of artificial intelligence and machine learning at scale.

About the team
We are a collaborative group of technical innovators dedicated to pushing the boundaries of cloud computing and artificial intelligence. Our team thrives on solving complex challenges from optimizing NCCL all-reduce operations across hundreds of GPUs to architecting elastic training clusters that scale with customer demand. We believe in continuous learning mutual support and driving technological advancement.

Diverse Experiences

Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description we encourage candidates to apply. If your career is just starting hasnt followed a traditional path or includes alternative experiences dont let it stop you from applying.

Why AWS

Amazon Web Services (AWS) is the worlds most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating thats why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.

Work/Life Balance

We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home which is why flexible work hours and arrangements are part of our culture. When we feel supported in the workplace and at home theres nothing we cant achieve in the cloud.

Inclusive Team Culture

Here at AWS its in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences including our Conversations on Race and Ethnicity and AmazeCon conferences inspire us to never stop embracing our uniqueness.

Mentorship and Career Growth

Were continuously raising our performance bar as we strive to become Earths Best Employer. Thats why youll find endless knowledge-sharing mentorship and other career-advancing resources here to help you develop into a better-rounded professional.

- Bachelors degree
- 8 years of experience in AI/ML distributed computing or GPU-accelerated infrastructure (e.g. model training inference systems HPC for ML)
- 3 years of hands-on experience designing implementing or consulting on large-scale ML training or inference architectures in a customer-facing role
- 10 years of IT development or implementation/consulting in the software cloud computing or AI/ML industries
- Experience with at least one major deep learning framework (PyTorch TensorFlow JAX) in a production or research environment
- Demonstrated ability to serve as a trusted technical advisor to enterprise customers

- Deep experience with distributed training techniques including data parallelism model parallelism pipeline parallelism and Fully Sharded Data Parallel (PyTorch FSDP)
- Experience with distributed training frameworks such as PyTorch DDP DeepSpeed and Megatron-LM for multi-node model training
- Hands-on experience with GPU/accelerator cluster infrastructure: NVIDIA Blackwell (GB200 B200 B300) H100/H200 GPUs AWS Trainium (Trn3/Trn2) NVLink/NVSwitch InfiniBand or Elastic Fabric Adapter (EFA) and NCCL collective communications tuning
- Experience with AWS Neuron SDK (torch-neuronx neuronx-nemo-megatron) for compiling and optimizing models on Trainium and Inferentia (Inf2) instances
- Familiarity with SageMaker HyperPod for managed distributed training clusters including automated health checks node replacement and checkpoint-based recovery
- Experience with HPC job schedulers (Slurm PBS LSF) for orchestrating multi-node ML training workloads
- Experience with high-performance parallel file systems (Amazon FSx for Lustre GPFS/Spectrum Scale) for ML data pipelines
- Familiarity with AWS Parallel Computing Service (PCS) AWS ParallelCluster AWS Batch or equivalent managed HPC/ML cluster services
- Experience training or fine-tuning large language models (LLMs) such as Llama GPT or similar transformer architectures at multi-billion parameter scale
- Understanding of HPC-AI convergence patterns: simulation-surrogate loops physics-informed neural networks (PINNs) graph neural networks for molecular property prediction and data format interoperability (HDF5 VTK NetCDF to ML-ready tensors)
- Knowledge of ML Ops tooling container orchestration for training (Docker Enroot Pyxis) and Deep Learning AMIs (DLAMIs)
- Experience with cluster observability and monitoring for GPU/Trainium utilization training throughput and job performance (CloudWatch Prometheus Grafana)
- Experience with EC2 Capacity Blocks for ML Capacity Reservations or similar GPU capacity planning strategies
- Experience with pipeline orchestration using AWS Step Functions for simulation-ML workflows
- Experience with containers EKS and ECS
- Track record of driving operational excellence and proactive risk mitigation for mission-critical AI/ML workloads
- AWS certifications (Solutions Architect Professional Machine Learning Specialty) preferred

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status disability or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process including support for the interview or onboarding process please visit for more information. If the country/region youre applying in isnt listed please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience qualifications and location. Amazon also offers comprehensive benefits including health insurance (medical dental vision prescription Basic Life & AD&D insurance and option for Supplemental life plans EAP Mental Health Support Medical Advice Line Flexible Spending Accounts Adoption and Surrogacy Reimbursement coverage) 401(k) matching paid time off and parental leave. Learn more about our benefits at TX Austin - 182800.00 - 247300.00 USD annually
USA TX Dallas - 182800.00 - 247300.00 USD annually
USA VA Herndon - 182800.00 - 247300.00 USD annually
USA WA Seattle - 182800.00 - 247300.00 USD annually


Required Experience:

Staff IC


About Company

Company Logo

Free shipping on millions of items. Get the best of Shopping and Entertainment with Prime. Enjoy low prices and great deals on the largest selection of everyday essentials and other products, including fashion, home, beauty, electronics, Alexa Devices, sporting goods, toys, automotive ... View more

View Profile View Profile