Enter a job title or keyword

Software Development Engineer, Compute Platform

Apple


Job Location:

Seattle, WA - USA

Monthly Salary: Not provided by the employer
Posted: 24 August 2026 (18 hours ago)
Application Deadline: 21 November 2026
Vacancies: 1 Vacancy

Job Summary

People at Apple dont just build products they craft the kind of experience that has revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple and help us leave the world better than we found Apple Services Engineering (ASE) team builds and provides systems and infrastructure that power Apples services (such as iCloud iTunes Siri and Maps). We are the foundation on which Apples software developers build the products that our customers love. Our services have to scale globally stay highly available and just work. If you love designing engineering and running systems and infrastructure that will help millions of customers then this is the place for you!nnApple Services Engineering (ASE)s Compute team is seeking a software engineer comfortable across data and systems to help our batch-focused compute platform make better decisions. The platform runs millions of jobs a day across tens of thousands of hosts and retains a detailed record of how each one behaved. You will turn that record into predictions optimizations and services the platform can act on improving both the efficiency of the fleet and the reliability of the workloads that run on work is end to end: you will explore the data build and validate the model take it to production and demonstrate the gains on live clusters. The platform is technically deep and a prediction only pays off if you understand the systems that will act on it.n

In this role you will develop debug and maintain data-driven features of a large-scale batch focussed compute platform. You will:nn* Own features end to end that turn job behavior into decisions the platform acts on from the initial data analysis through the model or heuristic the APIs production serving and the measurement that proves it workedn* Implement your own changes across the stack from platform services and control plane paths down to node-level resource limits and configurationn* Design and run validation in production: shadow mode staged rollouts explicit success metrics and a rollback plann* Have full ownership for the decisions your models make. When a prediction misbehaves diagnose it bound the impact and help fix the system that trusted itn* Write and review code generate and review design documentationn* Engage directly with customers and partner teams on their compute needs: understand their workloads analyze cluster utilization and capacity and turn that analysis into demand forecasts node pool configuration and quota decisions that balance customer demand against fleet efficiencyn* Bring quantitative analysis to open questions across the team sizing planning and prioritization decisions where good data changes the answern* Participate in software qualifications and rollouts to production clustersn* Participate in an on-call rotation where engineers respond to platform issues for same-day resolutionn* Work with a wide range of software and hardware engineering teams across Apple to support their workflows or integrate their technology into our platformn* Hold yourself and others to a high quality standard expected of Apple productsnn

Strong programming skills in a general-purpose language (Go Python C or similar) and demonstrated ability to design debug and test complex software systemsnWorking knowledge of distributed systems and operating system fundamentalsnExperience instrumenting large-scale infrastructure and analyzing the data it produces reasoning quantitatively about how CPU memory I/O or network consumption behaves and how it scales and turning that analysis into a decision or a measurable improvement in productionnComfort reasoning quantitatively about production data distributions and tails not just averages and about the risk a wrong estimate creates for a running workloadnTrack record of owning a project from an ambiguous problem statement through productionnStrong communication and organizational skills including the ability to work directly with customers and to make quantitative results legible and actionable for people who are not data specialistsnBS in Computer Science / related fields and 3 years of experience or MS with 1 years of experience or PhD

Experience building or operating a large-scale compute platform with responsibility for its capacity efficiency or performancenExperience improving utilization on a production platform right-sizing requests from historical usage oversubscription bin packing co-locating batch alongside latency-sensitive work or reclaiming unallocated capacity. Comparable work on Kubernetes VPA on runtime or demand prediction feeding an HPC scheduler or on an in-house equivalent is equally relevantnExperience characterizing workloads at fleet scale for example grouping jobs into behavioral classes or building the telemetry to do sonExperience shipping a model or heuristic that made automated decisions in production and owning the outcomenExperience with forecasting regression or uncertainty estimation applied to operational time seriesnExperience modelling how a systems resource consumption scales for example projecting the network storage or I/O demand created by growing a compute footprint and using that to inform capacity or sizing decisionsnExperience with Kubernetes / Slurm or a comparable batch scheduling system

Required Experience:

IC


About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile