Systems Architect Apple Pay
New York City, NY - USA
Job Summary
The Architect will own the technical direction for Apple Pays distributed systems driving reliability engineering practices observability strategy and scalability architecture across a high-throughput payment-critical infrastructure. Day to day this means reviewing designs for resilience gaps writing code and prototypes to validate architectural decisions and leading incident reviews and postmortems that drive systemic fixes. The role also involves close collaboration with client server devops and product engineering teams to ensure architecture decisions account for real-world operational constraints at Apples global scale.
Guide the evolution of Apple Pays distributed systems with a clear point of view on trade-offs (consistency vs. availability latency vs. durability cost vs. redundancy) appropriate to payment-critical infrastructurenDefine and drive adoption of reliability engineering practices across the org: SLIs/SLOs/SLAs error budgets capacity planning and failure-mode analysis tailored to systems where correctness and availability directly affect transaction successnEstablish the observability and metrics strategy needed to operate payment systems reliably at scale including latency traffic errors distributed tracing and alerting that reflects real transaction impactnPartner with teams across Apple Pay to review designs for scalability bottlenecks points of failure and resilience gapsnWrite code and prototypes where it matters diving into the codebase to validate designs unblock teams or resolve production issuesnLead technical reviews and postmortems for major incidents affecting Apple Pay systems driving systemic fixes rather than one-off patchesnMentor engineers on distributed systems fundamentals: consensus replication partitioning idempotency exactly/at-least-once semantics and trade-offs relevant to payment processingnCollaborate with client server devops and product engineering teams to ensure architecture decisions account for real-world operational constraints (deployment rollback multi-region failover capacity headroom) at Apples global scale
Proven track record designing and operating distributed systems at scale in high-throughput low-latency production environmentsnDeep expertise in distributed systems fundamentals: replication partitioning/sharding consensus protocols eventual vs. strong consistency idempotency and failure handling across network partitionsnHands-on experience defining and instrumenting metrics for availability and scalability (e.g. SLOs error budgets distributed tracing) and using them to drive concrete engineering decisionsnStrong background in the building blocks of scalable systems: load balancing caching layers message queues/event streaming database sharding/replication service mesh and rate limiting/backpressure mechanismsnDemonstrated ability to set technical direction across multiple teams without direct management authority and to communicate trade-offs clearly to both engineers and leadershipnTrack record of leading incident reviews or reliability programs and turning postmortem findings into durable systemic improvementsnSkilled at operating with ambiguity across large complex system landscapes and driving alignment across multiple teams and stakeholders
Experience with multi-region or globally distributed systems including failover and disaster recovery design at scalenPrior experience in payments fraud/risk systems or other regulated high-availability financial infrastructurenContributions to open-source distributed systems projects or published technical writing on the subjectnExperience establishing or maturing a devops/reliability practice within an organization
Required Experience:
Staff IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more