Site Reliability Engineer (SRE) Elastic Disk
Seattle, WA - USA
Job Summary
We are looking for seasoned software and systems engineers to join the Elastic Disk SRE team at Apple. The role involves tremendous amount of individual responsibility and influence over the direction the platform shaping its use by many critical Apple Cloud services for years to come. You are solution-oriented and have a passion for software delivered as a service to improve reuse efficiency and simplicity. Your work will affect hundreds of millions of users and be essential to the success of some of the most visible current and future Apple role involves understanding the teams priorities; taking ownership of projects or deliverables; crafting solutions and building buy-in for those designs; and successful delivery of those designs in order to meet the project goal. The role involves giving technical feedback to colleagues to assist them in the delivery of their designs features and projects as well as driving technical standards across the two-site team in collaboration with other senior members of the team has an on-call rota including the week-ends and the successful candidate should expect to handle alerts and other critical issues in order to maintain a high level of availability and functionality for our provided Apple Cloud we run a mix of open source vendor licensed and internally developed tools to perform functions such as system configuration management provisioning software development u0026 deployment logging and monitoring. Youll learn these tools and have opportunities to improve them. We think critically and strive to balance the best solution with the need to get things done for each engineering challenge we face. Good ideas are heard and results are rewarded.
4 years experience in building operating and scaling distributed storage systems in a private public or hybrid cloud environment or working with other large-scale stateful systems such as distributed ability to design author understand and release code in languages like Go (preferred) Java Python or understanding of block object and file storage solutions in Linux (such as LVM XFS ext4 S3 Ceph Gluster NFS).nUnderstanding of Linux internals standard networking protocols and distributed with provisioning data migration backup u0026 recovery at-scale testing disaster recovery and capacity planning.
Acute drive to automate manual operations and to improve them through repeated of best practices for deployment of storage systems - implication of physical and virtual deployment models to change management. failure domains hardware lifecycle management -on experience managing large numbers of diverse systems with configuration management or software delivery platforms (such as Puppet Chef Ansible and Spinnaker).nExperience with deploying supporting and monitoring new and existing services platforms and application with microservices architecture and container orchestration with with relational u0026 non-relational databases (such as Cassandra Postgres u0026 RocksDB)
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more