Director, Site Reliability Engineering
Job Summary
At Klaviyo we value the unique backgrounds experiences and perspectives each Klaviyo (we call ourselves Klaviyos) brings to our workplace each and every day. We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements. If youre a close but not exact match with the description we hope youll still consider applying. Want to learn more about life at Klaviyo Visit see how we empower creators to own their own destiny.
At Klaviyo we empower creators to own their destiny. Our SaaS platform provides businesses of all sizes with powerful easy-to-use marketing automation tools making first-party data accessible and actionable like never before. We are a team of ambitious bold peers who are insatiably curious and meticulous in our craft. As we scale globally we are seeking a foundational leader to build and lead our Site Reliability functions from our new strategic hub in Dublin.
About the Team
The SRE team is the bedrock of trust for the 200000 businesses that rely on Klaviyo to power their growth. This team is responsible for the secure reliability scalability and performance of our entire platform: from our core data processing pipelines to our cutting-edge AI services. As a foundational leader in our new Dublin office you will build this team from the ground up establishing a center of excellence for platform integrity and playing a pivotal role in Klaviyos European expansion. You will be the ultimate custodian of the trust our customers place in us every day.
How Youll Make a Difference
- Architect and Lead: Define the vision strategy and roadmap for a unified SRE organization. You will build and mentor a world-class multi-disciplinary team of engineers in Dublin fostering a culture of operational excellence proactive problem-solving and blameless learning.
- Champion Secure Reliability: Drive a secure and reliable by design philosophy across all of engineering. Partner with product and platform teams to embed reliability principles into the entire software development lifecycle from initial design to production deployment.
- Own Platform Integrity: Take ownership of the availability latency performance efficiency change management monitoring emergency response and capacity planning for Klaviyos global platform.
- Drive Automation: Build the paved road for Klaviyo engineers by developing automated self-service tooling and infrastructure that enables teams to move fast no shortcuts.
- Ensure Global Compliance: Serve as the on-the-ground technical leader for GDPR and other international data protection regulations. You will be responsible for implementing and verifying the technical controls that safeguard customer data and ensure compliance.
- Lead Through Incidents: Command reliability incidents guiding teams to rapid resolution while fostering a culture of blameless post-mortems that drive meaningful systemic improvements.
Who You Are
- You are a proven engineering leader with a track record of building and scaling high-performing geographically distributed teams in a fast-paced SaaS environment.
- You are intellectually curious and a continuous learner with a deep understanding of modern reliability principles.
- You are a Driver who is biased toward action. You dont wait for problems to find you; you proactively identify risks and rally teams to solve them.
- You are a cross-functional influencer and an exceptional communicator capable of articulating complex technical concepts to diverse audiences from engineers to executives to enterprise customers.
- You are a player-coach who can dive deep into technical details with your team while also developing and articulating a long-term strategic vision.
- You are passionate about building a strong inclusive team culture and have experience integrating new teams with existing ones across different time zones and cultures.
- You thrive in an environment of high ambiguity and rapid change embodying a 1% done mindset and seeing opportunity in every challenge.
What Youll Need
- 12 years of experience in software engineering with at least 5 years in a leadership role managing SRE DevOps or Security Engineering teams.
- Proven experience building and scaling engineering teams ideally with experience establishing a new team or office.
- Deep hands-on experience with cloud-native production systems at scale (AWS preferred).
- Strong technical background in distributed systems observability container orchestration (Kubernetes) infrastructure as code (Terraform) and CI/CD principles.
- Experience managing geographically distributed teams and fostering a strong unified culture across multiple time zones.
- Youve already experimented with AI in work or personal projects and youre excited to dive in and learn fast. Youre hungry to responsibly explore new AI tools and workflows finding ways to make your work smarter and more efficient.
Required Experience:
Director
About Company
Klaviyo unifies AI-powered email marketing and SMS to drive growth, retention, and measurable results. Build personalized, omnichannel experiences across WhatsApp, ecommerce, and more with K:AI Agents.