Principal Software Engineer AI Infrastructure
Redmond, WA - USA
Department:
Job Summary
Responsibilities
- Define and drive the architecture and technical roadmap for platform reliability observability and operational health.
- Design unified health and telemetry capabilities that connect customer impact with application capacity dependency infrastructure and deployment signals.
- Develop platform safeguards for overload protection capacity management routing integrity configuration consistency and automated isolation and recovery.
- Establish engineering practices for production validation progressive delivery regression detection fault testing and automated rollback.
- Advance end-to-end request tracing and diagnostics across distributed services including routing retries failover and asynchronous operations.
- Lead cross-team architecture efforts mentor engineers and turn production learnings into reusable platform capabilities and measurable reliability improvements.
Qualifications
Required Qualifications:
- Bachelors Degree in Computer Science or related technical field AND 6 years technical engineering experience with coding in languages including but not limited to C C C# Java JavaScript or Python
- OR equivalent experience.
Preferred Qualifications:
- Masters Degree in Computer Science or related technical field AND 8 years technical engineering experience with coding in languages including but not limited to C C C# Java JavaScript or Python
- OR Bachelors Degree in Computer Science or related technical field
- AND 12 years technical engineering experience with coding in languages including but not limited to C C C# Java JavaScript or Python
- OR equivalent experience.
- Experience with AI/ML serving platforms high-performance computing accelerator-based infrastructure or other compute-intensive distributed systems.
- Experience with distributed tracing capacity management load balancing admission control retries backpressure and graceful degradation.
- Proficiency in Azure Monitoring systems in the Azure ecosystem would be a plus.
- Experience designing building and operating large-scale distributed systems cloud services or other complex production platforms.
- Experience with reliability engineering observability service health telemetry and production incident response.
- Demonstrated ability to lead complex technical initiatives and drive alignment across engineering teams and organizational boundaries.
- Strong written and verbal communication skills including the ability to explain technical strategy and architectural decisions to engineers and senior leaders.
Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142800 - $274800 per year. There is a different range applicable to specific work locations within the San Francisco Bay area and New York City metropolitan area and the base pay range for this role in those locations is USD $188000 - $304200 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
position will be open for a minimum of 5 days with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age ancestry citizenship color family or medical care leave gender identity or expression genetic information immigration status marital status medical condition national origin physical or mental disability political affiliation protected veteran or military status race ethnicity religion sex (including pregnancy) sexual orientation or any other characteristic protected by applicable local laws regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process read more about requesting accommodations.
Required Experience:
IC