Site Reliability Engineer (SRE)
Job Location:
San Francisco, CA - USA
Yearly Salary:
$ 163710 - 306000
Posted:
24 August 2026 (19 hours ago)
Application Deadline:
21 November 2026
Vacancies:
1 Vacancy
Department:
Job Summary
ABOUT RETOOL
Nearly every company in the world runs on custom software for critical operations like tracking performance metrics handling support workflows building admin dashboards and countless processes you might never have thought of. But most companies dont have the resources to properly invest in these tools leading to a lot of old clunky internal software or worse teams still stuck in manual and spreadsheet workflows.
AI has changed who gets to build software. The definition of developer now includes analysts operators and domain experts creating solutions directlyand the tools they reach for are multiplying by the week. Thats both an opportunity and a challenge: as more people build with more AI tools the risk of shipping ungoverned software into production grows just as fast.
At Retool were building the platform that makes all of it safe to ship. Build with any AI tool you want then deploy into one place that connects to your real business data enforces enterprise policies automatically and lets teams create once and reuse everywhere with shared trusted components. The cost of building software has collapsed. The cost of governing it hasntand thats the problem we solve.
Developers and domain experts have already automated over 100 million hours of work on our platform freeing them to focus on creative problem-solving and strategic work that drives real business value. The people closest to the problem can now build the software to solve it safely and within enterprise guardrails.
Lets build the future together.
WHY WERE LOOKING FOR YOU:
Good software has to run where customers need it. For many of Retools largest customers that means running Retool in their own infrastructure behind their own controls with the reliability and operational clarity they would expect from any critical system.
Retools Core Infrastructure team owns the systems that make this possible: Retool Cloud managed single tenant environments BYOC (bring-your-own-cloud) environments Kubernetes and Helm deployments Docker Compose and the migration paths between them. It is a broad surface area and it is one of the biggest levers we have for making Retool work for enterprise customers.
The work is not clean-room infrastructure. Customers run different clouds different versions different deployment models and different levels of operational maturity. A bad upgrade experience can leave a customer many versions behind. A manual Terraform run can become the bottleneck during a launch or incident.
We are hiring SREs who want to turn that mess into leverage. You will help us reduce customer toil automate upgrades and infrastructure changes build reliability tooling across Retool Cloud and customer-owned environments and make Retool easier to deploy and operate at enterprise scale. The strongest candidates are comfortable debugging Kubernetes Terraform AWS Postgres networking and deployment problems then stepping back and building the automation or product surface that prevents the same problem from happening again.
What youll do:
- Own reliability across Retool Cloud managed single tenant BYOC and self-hosted deployment paths including provisioning upgrades migrations configuration changes and production escalations.
- Build the automation that turns todays manual infrastructure work into repeatable systems: Terraform runs customer environment updates upgrade workflows secret rotations and migration steps.
- Improve observability for Retool Cloud self-hosted customers and internal operators. We care less about exposing every metric and more about turning health signals into clear status likely causes and recommended actions.
- Design safer deployment upgrade and rollback paths so Cloud and managed customers can stay current
- Help move customers from legacy or less-supported deployment models toward supported paths such as Retools official deployment paths (Blueprints Kubernetes and Helm) with migration flows that are repeatable enough for customers Support and TAMs to trust.
- Partner with product engineers on infrastructure requirements for new Retool products especially when they introduce new dependencies
- Lead through ambiguity make careful risk calls and communicate clearly while things are moving quickly.
- Write the docs runbooks design notes and migration guides that make complex systems understandable to other engineers and to customers.
What were looking for:
Infrastructure fundamentals
- Deep experience operating production infrastructure in AWS.
- Experience improving reliability for customer-facing SaaS systems.
- Strong Kubernetes fundamentals.
- Real Terraform or infrastructure-as-code experience.
- Good operational judgment around databases especially Postgres.
Reliability and automation
- Experience building or operating observability systems.
- Programming ability in a language such as Go Python TypeScript Java or Ruby.
- A bias toward automation. If you find yourself doing the same operational task twice you should start thinking about the interface workflow or tool that eliminates the third time.
Customer and team judgment
- Clear written communication.
- Comfort working directly with customer-facing teams and when useful customers themselves.
What makes SREs successful here:
You will do well here if you like infrastructure that sits close to real customer pain. Some days that means debugging a specific customer environment. Other days it means improving Retool Cloud reliability or designing the migration path so the next 25 customers do not need that same debugging session.
We value SREs who are ambitious curious energetic and careful with the details. Retool moves quickly priorities can change and the systems are not always as clean as we want them to be. The work needs SREs who can get their hands dirty tell the truth about tradeoffs and leave the system better than they found it.
Retool offers generous benefits to all employees and hybrid work location. For more information please visit the benefits and perks section of our careers page!
Retool is currently set up to employ all roles in the US and specific roles in the UK. To find roles that can be employed in the UK please refer to our careers page and review the indicated locations.
Required Experience:
IC
About Company
Build, deploy, and manage internal tools with Retool’s unified engine. Connect to any database, API, or LLM. Leverage AI throughout your business.