Site Reliability Engineer
Madrid, IA - USA
Job Summary
Apple is looking for an SRE (or SWE) to join a small passionate team that builds monitors automates and maintains a sophisticated system for a critical and unique customer-facing Apple service. This is a rare opportunity to help build and run systems and software for a service that customers will rely on every day on a team with a no-ops culture. Were looking for people who like to solve operational problems using software rather than shell prompts as we scale Apples services. You should be a forever learner with a bias toward action and positive energy. Job duties include participating in on-call rotations occasionally. Help us build the Apple experience on a global scale!
Software development with networked services on Linux - Experience building and deploying Linux-based applications with network communication protocolsnCloud-native development practices - Proficiency with containerization and orchestration (Kubernetes Docker) and experience with at least one major cloud platform (AWS GCP or Azure)nMonitoring and observability - Hands-on experience with monitoring stacks (Prometheus Alertmanager Grafana) and understanding of distributed system observabilitynDistributed systems fundamentals - Understanding of distributed architectures microservices patterns and challenges inherent to cloud-native environments
Proven expertise designing AWS security controls: IAM roles and policies landing zones SCPs and permission boundaries enforcing least-privilege auditable systems design - Practical experience designing and debugging distributed systems at scale including consistency fault tolerance and performance optimizationnModern cloud development practices - Familiarity with infrastructure-as-code GitOps workflows service mesh technologies and cloud-native development patternsnInfrastructure automation - Experience with infrastructure-as-code tools (Terraform Ansible CloudFormation) for repeatable scalable deploymentsnProgramming proficiency - Strong capabilities in at least one of: Python Go or C for building cloud-native servicesnCI/CD and DevOps expertise - Advanced knowledge of CI/CD pipelines automated testing and deployment automation in cloud environmentsnRoot cause analysis and resilience - Persistent in identifying systemic issues understanding failure modes in distributed systems and driving solutions to completionnPerformance and reliability focus - Analytical mindset toward observing end-to-end service performance system health and user impact in production environmentsnDynamic environment adaptability - Comfort working in fast-growing evolving environments with changing priorities and emerging technologies
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more