Director of DevOps & SRE
Job Summary
- Strategy and roadmap. The DevOps/SRE strategy for the InvestorFlow platform aligned to measurable outcomes across reliability delivery throughput and cost.
- The team. Leading hiring and developing technical leads senior and mid-level engineers across DevOps and SRE and setting priorities across infrastructure CI/CD security and production support.
- Cloud platform ownership. Reliability performance and cost-efficiency across DEV QA STG and PRD -compute data secrets management caching and networking/edge security. Azure is our primary platform with a supporting AWS and Snowflake footprint.
- CI/CD and release engineering. GitHub Actions end to end multi-region deployment automation advanced deployment strategies (blue/green canary rolling) and rollback plus security scanning compliance checks and vulnerability management embedded in the pipeline.
- Infrastructure as code. Terraform practice and standards that standardise provisioning and reduce manual ticket-driven infrastructure work.
- Automation-first operating standard. Setting and holding the expectation that repeatable work is codified rather than performed across provisioning releases access remediation and reporting and driving developer enablement through self-service tooling that increases engineering capacity while maintaining controls.
- Identity and access. Enterprise directory services role-based and privileged access controls Auth0 machine-to-machine credentials and least-privilege access across engineering and QA.
- Incident response and resilience. Escalation paths root-cause analysis and postmortems production readiness reviews automated failover and our disaster recovery and RTO/RPO commitments. You are the escalation point for high-severity incidents and major client environment changes ensuring appropriate change management and CAB governance with occasional off-hours support for the teams.
- Observability and service levels. SLO/SLA and monitoring strategy across Grafana and our wider observability tooling building alert coverage that is meaningful rather than noisy.
- AI in engineering operations. The operational rollout of AI-assisted engineering and support tooling: automated triage agent-based workflows and developer-facing AI capability.
- Architectural review and technical governance. Reviewing significant infrastructure and platform change with a security- and compliance-first lens ensuring designs meet our obligations by default rather than by exception and owning the shared-responsibility operating model between DevOps/SRE and product engineering.
- Vendor and cost management. Relationships across cloud security CI/CD and observability platforms including spend and renewal negotiation.
- 8 years in DevOps SRE or infrastructure engineering including 3 years leading and developing engineering teams.
- Deep hands-on Microsoft Azure experience at production scale: compute data secrets identity services networking and WAF and cost/capacity management. Azure depth is essential for this role.
- Strong CI/CD and release engineering background (GitHub Actions and/or Azure DevOps Pipelines) and infrastructure as code with Terraform including module development and state management.
- Production support and incident response experience for a multi-tenant SaaS platform.
- Containerisation and orchestration in production (Docker Kubernetes).
- Solid grounding in identity and access management (role-based and privileged access SSO/SAML OAuth/Auth0) and secrets management practices.
- A genuine automation-first instinct backed by strong scripting (Python Bash PowerShell or similar) and a track record of replacing manual process with reliable supportable tooling.
- Networking fundamentals (TCP/IP DNS load balancing) and the ability to partner effectively with network and security teams.
- Observability experience with Grafana Prometheus or equivalent and a track record of building signal rather than noise.
- Comfort operating in a security and compliance-conscious environment: financial services private markets or another regulated industry including audit penetration testing and data retention requirements.
- Excellent written and verbal communication able to translate infrastructure risk and trade-offs for engineering peers executives and clients.
- The judgement to know when to go deep technically and when to lead through others.
Nice to have: AWS exposure alongside Azure Snowflake or another cloud data platform CDN and edge platforms (Cloudflare or similar) Redis event-driven or message-oriented middleware (Kafka Azure Service Bus) and prior experience in private equity private credit real assets or capital-markets SaaS.
Required Experience:
Director
About Company
InvestorFlow is a leading provider of front-office software applications for private equity, real estate and hedge fund investment firms.