Site Reliability Engineer, AiDP Production Engineering

Apple


Job Location:

Austin, TX - USA

Monthly Salary: Not Disclosed
Posted on: 2 hours ago
Vacancies: 1 Vacancy

Job Summary

The Production Engineering team within the AI and Data Platform (AiDP) organization manages a wide array of real-time near real-time and batch analytical solutions. These platforms are integral to core business functions across Apple. These include sales operations finance AppleCare marketing and services and are instrumental in driving critical data-driven decisions. To build these solutions we leverage a combination of proprietary and leading open-source technologies such as Kafka Spark Iceberg and Airflow. A key part of our mission is to enable AI-centric automations that enhance the overall efficiency and intelligence of the platform. We are looking for passionate engineers who thrive on solving complex infrastructure challenges at scale both on-premises and in the cloud. If you are dedicated to optimizing scalable maintainable and user-friendly systems you will find compelling opportunities to make a significant impact at AiDP.

The Service Reliability Engineer (SRE) role within AiDP Production Engineering is a dynamic position that blends strategic architectural design with hands-on technical execution. As an SRE you will be responsible for configuring tuning and ensuring the resilience of complex multi-tiered systems to achieve optimal application performance stability and availability. Our team manages critical data pipelines and applications across both bare-metal and cloud computing platforms delivering essential data processing for all of Apples key business functions. We operate at an immense scale handling exabytes of data petabytes of memory and tens of thousands of jobs to enable predictable and performance data analytics that power features and inform decisions across the company. If you are passionate about designing building and running data infrastructure that has a direct and significant impact on Apples global business operations this is the ideal opportunity for you.

Ability to understand the application requirements (Performance Security Scalability etc.) and assess the right services/topology on AWS Baremetal u0026 automation to enable self-healing tools to monitor high performance u0026 alert the low latency to troubleshoot application specific core network system u0026 performance in challenging and fast paced projects supporting Apples business by delivering innovative with engineering teams to prioritize and fix production knowledge transition from engineering teams for changes being rolled out in incidents based on the impact devise and implement mitigation steps to unblock the RCA log defects and partner with engineering team for java based applications u0026 Spark/Flink jobs on Baremetal AWS u0026 on-call rotation with other team members to support apps and services in scope.

4 years experience in cloud-native services including ETL frameworks like Apache Spark and Flink.n4 years experience in messaging systems (Kafka) and cloud infrastructure u0026 services AWS GCP Kubernetes.n4 years of experience in modern u0026 distributed databases such as Snowflake Cassandra SingleStore and SAP HANA.n4 years of programming experience in Python or in computer science or equivalent experience.

Solid understanding of system design data structures and incident management best be able to understand complex architectures and be comfortable working with multiple tools (e.g: Prometheus Grafana CloudWatch).nAbility to conduct performance analysis and troubleshoot large scale distributed be highly proactive with a keen focus on improving uptime/availability of our mission critical expertise in troubleshooting complex production problem solving critical thinking and communication ability to resolve incidents perform root cause analysis and drive system reliability using GenAI or automation tools for issue detection alerting or in data visualization tools such as Tableau Business Objects ThoughtSpot.

Required Experience:

IC

The Production Engineering team within the AI and Data Platform (AiDP) organization manages a wide array of real-time near real-time and batch analytical solutions. These platforms are integral to core business functions across Apple. These include sales operations finance AppleCare marketing and se...

About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile