Member of Technical Staff Site Reliability
New York City, NY - USA
Department:
Job Summary
AI is transforming how every company operates but most enterprises are stuck. They want to move fast with AI Agents tools and workflows but they cant do it safely. Were fixing that.
Our team built AI Actions for OpenAI shipped Zapier Agents to millions of users and launched the first remote MCP server with Anthropic. We helped establish the protocol and now were building the platform enterprises need to actually put MCP to work.
Runlayer is one platform for MCPs Skills and Agents: purpose-built security fine-grained governance and complete observability so organizations can go all-in on AI across the entire company without the risk. We just raised a $30M Series A led by Felicis with participation from Khosla Ventures bringing our total raised to $42M. Already trusted by Gusto Instacart Opendoor dbt Labs Cursor and Decagon.
Were a team of 35 mostly engineers shipping fast. Agent operators on Runlayer grew 5x in four weeks while human operators stayed flat. The shift to agentic work is already showing up in our own usage.
As our Site Reliability Engineer youll own the reliability performance and scalability of Runlayers infrastructure as we grow to serve enterprise customers across cloud and on-prem environments.
Impact: Build the infrastructure foundation for the enterprise MCP platform directly enabling AI adoption at scale
Excellence: Work closely with founders and a small senior engineering team shipping fast in a high-growth environment
Ownership: Own reliability end-to-end from database performance to incident response to CI/CD pipelines
Own reliability and performance of our cloud infrastructure across AWS (ECS Aurora CloudWatch) and GCP
Manage and optimize Kubernetes clusters and container orchestration
Drive database reliability engineering including performance tuning and scaling
Build and maintain CI/CD pipelines for rapid safe deployments
Run incident response and on-call rotations
Partner with product engineers to design scalable resilient systems
Strong AWS experience particularly ECS Aurora and CloudWatch
GCP experience as we expand cross-cloud
Kubernetes and container orchestration expertise
DBRE experience with database performance tuning
CI/CD pipeline ownership and incident response experience
Background at a B2B SaaS company serving enterprise customers ideally in infrastructure
Experience deploying and supporting on-prem or hybrid environments
Python backend familiarity (our platform is Python-based)
Experience at an early-stage or high-growth company
We provide a competitive package designed to attract and retain top talent who can work effectively with enterprise customers.
Competitive salary and equity compensation that reflects your expertise and customer-facing responsibilities.
Paid time off paid vacation paid sick leave and paid parental leave.
Professional development budget for conferences courses and certifications in AI enterprise software and customer success.
Top-tier equipment your choice of laptop and accessories to create your ideal work environment.
Health benefits comprehensive health dental and vision coverage.
Customer interaction opportunities work directly with innovative companies and see the immediate impact of your work.
Not quite the right fit Reach out to with details about your experience and interests.
Required Experience:
Staff IC