Software Engineer, Platform Reliability Engineering, AiDP
Sunnyvale, CA - USA
Job Summary
Were seeking an experienced software engineer to join our Platform Reliability Engineering team and drive the design operation and optimization of large-scale distributed systems that power our GenAI ML and big data platforms. Youll leverage cutting-edge open source technologies in hybrid cloud environments to build resilient infrastructure that enables seamless inference data processing and machine learning this role youll own mission-critical platform components respond to production incidents and collaborate across teams to shape the future of our data and AI infrastructure.n
Design build and maintain scalable multi-tenant systems that support diverse workloads and technologies at enterprise scalenOwn the full lifecycle of infrastructure and platform projectsfrom architectural design and implementation through deployment monitoring and optimizationnOperate and optimize high-throughput mission-critical services to ensure reliability performance and cost-efficiencynParticipate in on-call rotations to respond to production incidents; diagnose root causes implement rapid fixes and drive post-incident improvementsnLead cross-functional collaboration with engineering teams to define requirements validate designs and deliver customer-impacting features and improvementsnProactively identify operational bottlenecks and systemic issues; implement preventive measures to reduce incident frequency and improve system resiliencenEstablish observability practices and continuously refine operational excellence standards across the platform
Bachelors degree in Computer Science Computer Engineering or equivalent professional experiencenProficiency in at least one systems programming language (Python Go Java or similar)nStrong expertise in distributed systems architecture with deep knowledge of reliability scalability and containerization principlesnHands-on experience with cloud platforms and data processing infrastructure (Kubernetes Spark Flink Ray Trino or equivalent technologies)n
7 years of experience in SRE DevOps or infrastructure engineering with demonstrated expertise managing distributed systems at in diagnosing and resolving complex production incidents and performance bottlenecks in large-scale distributed with open source codebases; ability to read understand and explain complex system implementationsnStrong understanding of system architecture and proven ability to collaborate effectively across engineering teamsnHands-on experience with big data technologies (Spark Flink Iceberg) and/or ML/AI platforms (Ray MLflow model serving infrastructure).nStrong foundational knowledge of Linux databases and security principlesnProactive mindset with demonstrated commitment to optimizing reliability and uptime for mission-critical servicesnExcellent written and verbal communication skills with ability to articulate technical concepts and strategies to both engineering teams and non-technical leadershipnDemonstrated track record of designing and operating systems at scalen
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more