Engineering Manager, ML Infrastructure, London
London, KY - USA
Job Summary
Private Cloud Compute is the system that lets Apple Intelligence reach beyond the device without compromising a users privacy: generative AI inference running in Apples cloud with verifiable privacy guarantees no other large-scale AI platform offers. It is the server software behind Apple Intelligence and it is the critical function this organisation exists to stack is deep. On-device client frameworks hand requests to a cloud service that attests routes and orchestrates them; an inference engine serves them; and model runtimes execute across heterogeneous hardware platforms from Apple silicon to industry-standard accelerators each with different performance characteristics and constraints. Cutting across all of it are the problems that decide whether the platform is fast affordable and operable: context and cache management model asset management and lifecycle throughput and latency observability and the developer and test infrastructure that everything else is built would own set of components in this stack. The generative-AI landscape is a rapidly evolving and we are looking for managers with an agile mindset that are energized by change. You can hold a clear technical direction while the ground shifts and have a strong desire to help define our to day you will hire grow and lead a team of engineers; own delivery against a roadmap you help set; lead design reviews and make architectural calls yourself when your team needs a decision; run a healthy on-call and incident practice; and partner across time zones with ML research hardware and platform teams security and privacy SRE and the product teams that depend on you. You will work with teams in London Cupertino and Seattle whose work spans low-level operating systems and accelerator runtimes through data-centre services network protocols and public should be technically credible you do not need to be the strongest individual contributor on the team but you must be able to hold your own in a design review read the code when it matters and tell a good argument from a confident one. You should be able to absorb shifting priorities on behalf of your team rather than passing them along. And you should care about the privacy promise this platform makes to users; much of what makes the engineering here hard and interesting is that the usual shortcuts are not available to us.
Build coach and retain a high-performing inclusive team: hiring onboarding growth feedback and performance ownership of one or more areas of the inference stack and be accountable for their delivery quality and technical and defend a technical roadmap that stays credible as priorities and platforms change balancing near-term delivery against longer-term across engineering research hardware security and privacy SRE and product to deliver outcomes that span team a rigorous engineering bar: honest measurement reproducible results operational readiness incident review and your teams work to senior leadership and advocate for the resources and direction it technical leads within your team delegating real architectural ownership rather than retaining the team applies privacy-by-design and secure-by-design principles throughout particularly around what may and may not be observed or logged in a system handling user content.
Experience managing software engineers including hiring coaching feedback and performance strong software engineering background in systems backend distributed systems or platform work with the ability to engage deeply and specifically in design ownership of delivery on an infrastructure or platform team: roadmap sequencing cross-team dependencies and shipped agile mindset and a track record of operating effectively in ambiguity able to absorb rapidly shifting priorities without losing execution discipline or the teams written communication and effective working habits across geographies and time /US collaboration particularly genuine security and privacy mindset for systems handling sensitive user content.
Any strong combination of the following is interesting to us we do not expect all of them:nnExperience with LLM inference or model serving at scale: batching and scheduling KV-cache reuse paged attention prefix caching disaggregated serving speculative decoding quantisation or model with GPU or custom-accelerator performance work and with the internals of an ML runtime or leading teams that own a platform other engineers build on including API and compatibility stewardship across versions and hardware with production operations for latency-sensitive services: SLOs and error budgets observability capacity planning canary and rollback in developer experience and build or test infrastructure and a view on how to reduce cycle time without lowering knowledge of Swift; systems-language experience (C Rust Go) and Python tooling experience are all valuable with privacy-preserving security-sensitive or attested systems and with reasoning rigorously about what may be logged or growing a team from a small senior core and developing engineers into technical leadership.
Required Experience:
Manager
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more