Director of Engineering, Infrastructure
Boston, MA - USA
Job Summary
At Klaviyo we value the unique backgrounds experiences and perspectives each Klaviyo (we call ourselves Klaviyos) brings to our workplace each and every day. We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements. If youre a close but not exact match with the description we hope youll still consider applying. Want to learn more about life at Klaviyo Visit see how we empower creators to own their own destiny.
Klaviyos mission is to empower businesses to independently drive their growth and the Engineering Departments contribution to this mission is crucial. As Director of Production Infrastructure youll helm the creation and management of a high-performance platform designed to support the rapid innovation demanded by our R&D teams. This role is all about defining the pillars of our infrastructure compute storage networking observability and setting a robust set of principles that guide their use.
In this position youll be entrusted with the responsibility of developing and maintaining platform primitives that empower our engineering teams to bring ideas to life seamlessly. Collaborating with industry leaders across engineering security and finance your decisions will shape the infrastructure blueprint that underpins our scalable secure and cost-effective operations. As a leader your mission is to foster a culture of ownership innovation and productivity while steering teams toward achieving critical reliability and performance metrics. Your role will span across defining clear service contracts instituting capacity plans and honing our developer enablement strategy to reduce friction and enhance developer velocity.
How Youll Make a Difference
- Lead the definition of platform primitives such as compute runtimes storage options and service networking ensuring they are scalable secure and aligned with Klaviyos standards for operational excellence.
- Create and disseminate golden paths and decision trees that simplify the technological choices for R&D teams enhancing consistency and self-sufficiency across engineering efforts.
- Drive initiatives that enhance the reliability of production systems focusing on incident prevention transparent response protocols and proactive capacity planning.
- Coordinate with product teams to identify and eliminate infrastructure bottlenecks aiding in improving the time-to-market for new services and increasing developer satisfaction.
- Establish frameworks for cost-effective infrastructure management balancing financial discipline with flexibility and efficiency to maximize value delivery.
- Mentor and develop high-performing teams fostering a culture of inclusivity and ownership while setting clear impactful goals that align with business priorities.
- Collaborate with cross-functional partners to manage platform investments clarify ownership and safely implement infrastructure changes that drive strategic outcomes.
- Track and report critical performance metrics such as system reliability developer productivity and infrastructure costs enabling data-driven decision-making and accountability.
- Optimize the use of AI to enhance infrastructure management and development processes pioneering innovative workflows that keep Klaviyo at the forefront of technological advancement.
- Champion operational readiness by establishing robust SLAs and SLIs ensuring all infrastructure components meet defined performance thresholds conforming to Klaviyos quality standards.
- Facilitate a culture of continuous learning and experimentation with AI tools deploying enhancements that intelligently streamline engineering workflows.
- Lead a disciplined approach to incident management and postmortems establishing a blameless culture of learning and innovation to minimize future disruptions.
Who You Are
- You have over 10 years of experience in infrastructure SRE platform engineering or security engineering with at least 5 years managing managers and senior ICs demonstrating strong leadership and team-building skills.
- Your leadership style fosters inclusivity and empowerment setting high standards that inspire your teams to achieve ambitious goals aligned with Klaviyos mission.
- You possess a deep understanding of SRE principles with proven expertise in designing effective SLOs and SLIs managing incidents and ensuring capacity planning and operational continuity.
- You have experience with LiveSite (the enterprise web platform) and understand the infrastructure requirements integration patterns and operational demands that come with supporting enterprise-grade web properties at scale.
- You own the live site you bring urgency clear judgment and a structured approach to production incidents and you build teams that treat uptime and reliability as first-order concerns not afterthoughts. You have a proven track record of standing up or strengthening LiveSite culture inside engineering orgs.
- You excel at simplifying complexity and creating logical clarity using your strong communication skills to drive organizational changes across product data and security domains.
- Your decision-making is driven by outcomes and youre comfortable prioritizing initiatives that deliver the most significant business impact even if it means narrowing the scope.
- Your hands-on experience with AI makes you AI-curious and youre eager to leverage it to drive smarter more efficient infrastructure operations.
- Your technical expertise is both broad and deep you command public cloud services container orchestration service meshing data storage and observability at a systems level and you remain hands-on when it counts able to roll up your sleeves and debug a complex distributed-systems problem alongside your team.
- You build and nurture a culture of ownership and continuous improvement encouraging your teams to innovate and excel while maintaining a strong focus on customer value.
- You possess a keen ability to balance strategic thinking with tactical execution ensuring infrastructure solutions not only meet current needs but are future-proof and scalable.
- You are an org builder you know how to scale a team thoughtfully develop high-potential engineers into managers and create the conditions for leaders to emerge and grow.
- You raise the bar relentlessly you set high standards for engineering excellence hold yourself and your teams accountable to those standards and invest continuously in growing the people around you.
- You are an org builder you know how to scale a team thoughtfully develop high-potential engineers into managers and create the conditions for leaders to emerge and grow.
- You raise the bar relentlessly you set high standards for engineering excellence hold yourself and your teams accountable to those standards and invest continuously in growing the people around you.
- Committed to fostering a learning environment you continually seek ways to improve team skills processes and technologies to align with Klaviyos growth and innovation objectives.
- You bring a security-first mindset to infrastructure decisions with experience partnering closely with security engineering teams to build platforms that are secure by design and operationally hardened.
- You bring a security-first mindset to infrastructure decisions with experience partnering closely with security engineering teams to build platforms that are secure by design and operationally hardened.
Nice to Have
- Previous experience in transforming internal platforms into productized services complete with SLAs and developer experience-focused roadmaps.
- Background in architecting data-centric and event-driven systems at a large scale particularly within a high-growth environment.
- Established partnerships with centralized data platforms defining clean ownership boundaries and integrating efficiently with existing workflows.
- Proven track record of optimizing both cost-to-serve and reliability metrics in a scaling SaaS company driving significant impact on the bottom line.
- Familiarity with the challenges and opportunities characteristic of a high-growth SaaS milieu with a focus on leveraging those to drive platform innovation and efficiency.
Why This Role Matters
The Director of Production Infrastructure at Klaviyo is a mission-critical role that directly influences the core foundations of our technology stack. Ensuring that our infrastructure is robust scalable and future-proof enables our engineering teams to accelerate product development effectively and safely cementing Klaviyos position as an industry leader in empowering business growth.
Your leadership will impact our ability to innovate rapidly and maintain the high standards of reliability our customers expect. By driving infrastructure improvements that reduce developer friction enhance system reliability and lower operational costs you ensure Klaviyo can offer exceptional value to our clients while scaling our operations sustainably.
For a top candidate this role represents a unique opportunity to lead a team in redefining infrastructure excellence against the backdrop of a dynamic business landscape. Youll play a pivotal role in shaping the technological trajectory of Klaviyo working alongside passionate professionals committed to creating powerful solutions that enable our customers to achieve significant growth.
Required Experience:
Director
About Company
Klaviyo unifies AI-powered email marketing and SMS to drive growth, retention, and measurable results. Build personalized, omnichannel experiences across WhatsApp, ecommerce, and more with K:AI Agents.