Service Operations Manager
Job Summary
dunnhumby is the global leader in Customer Data Science partnering with the worlds most ambitious retailers and brands to put the customer at the heart of every decision. We combine deep insight advanced technology and close collaboration to help our clients grow innovate and deliver measurable value for their customers.
dunnhumby employs nearly 2500 experts in offices throughout Europe Asia Africa and the Americas working for transformative iconic brands such as Tesco Coca-Cola Nestlé Unileverand Metro.
We are looking for a highly motivated Engineering Manager to lead and evolve our Service Operations function into a modern observability-led engineering-focused capability. This role is responsible for ensuring operational excellence across production platforms through proactive monitoring incident management service reliability engineering (SRE) practices automation and continuous service improvement. The Engineering Manager will lead a team of SevOps and Observability specialists partnering closely with Product Engineering Platform Infrastructure and Support teams to ensure services are resilient recoverable scalable and aligned with business objectives. The role will drive the transition from traditional operational processes to a telemetrydriven automation-first operating model.
What youll need:
- 15 years of experience in Engineering with 7 years in platform engineering/DevOps/SRE leadership roles.
- Proven success leading large-scale platform transformations in cloud-native environments (preferably GCP and Azure).
- Hands-on and strategic experience with Kubernetes CI/CD GitOps Terraform Crossplane Docker Infrastructure-as-Code and multi-tenant platform design.
- Deep expertise in platform observability developer self-service golden paths and IDPs such as Backstage.
- Advanced understanding of DevSecOps compliance automation and security bydesign principles.
- Proven experience operating large-scale production environments in cloud and hybrid infrastructure.
- Demonstrated success in defining and driving engineering OKRs metrics based decision making and cost accountability.
- Strong ability to balance technical depth with cross-functional influence managing senior stakeholders and C-level engagement.
- A builders mindset with a focus on automation resilience scalability and simplification.
- Excellent communication and stakeholder management skills; able to collaborate effectively with teams in India and internationally and to balance ambition feasibility and risk.
Key Responsibilities:
Service Reliability & Operational Excellence
- Own the day-to-day operational reliability and availability of business-critical and mature SRE practices including Service Level Objectives (SLOs) Service Level Indicators (SLIs) error budgets and reliability reporting.
- Drive proactive identification and mitigation of operational risks through telemetry observability and data-driven insights. Lead the continual improvement of Incident Problem Change and Major Incident Management processes. Ensure effective operational readiness reviews for new products features and platform changes.
Observability & Telemetry Strategy
- Lead the organizations observability strategy across infrastructure applications and cloud platforms. Drive adoption and optimization of monitoring logging tracing and alerting capabilities.
- Ensure operational dashboards provide actionable insights and meaningful service health visibility. Establish standards for alert quality ownership runbooks service mapping and dependency monitoring.
Incident & Crisis Management
- Own Major Incident Management standards and execution. Lead high-severity incident response and stakeholder communications during service disruptions.
- Drive root-cause analysis (RCA) culture and ensure corrective actions are tracked to completion. Reduce recurring incidents through automation resilience engineering and preventative improvements.
Automation & Platform Operations
- Promote automation-first operational practices across monitoring remediation reporting and support activities. Partner with Platform Engineering and Infrastructure teams to reduce manual operational effort.
- Drive implementation of self-healing predictive alerting and AI-assisted operational capabilities. Continuously improve operational workflows and service efficiency through tooling enhancements.
Engineering Partnership & Release Readiness
- Act as a strategic partner to Product Owners Engineering Managers and Platform teams. Ensure operational requirements are incorporated into engineering delivery processes.
- Ensure monitoring alerting support documentation and ownership are in place before production releases.
Governance Compliance & Risk Management
- Ensure operational processes comply with internal governance security audit and regulatory requirements. Maintain operational standards policies and controls.
- Support ISO27001 audit resilience and business continuity initiatives. Monitor service risks and develop mitigation plans.
Leadership & People Management
- Lead coach and develop a high-performing team of SevOps and Observability specialists. Build a culture of accountability collaboration continuous learning and operational excellence.
- Drive workforce planning succession planning and capability development. Foster adoption of modern operational practices SRE principles and engineering led service management.
Vendor & Stakeholder Management
- Manage relationships with service providers technology partners and operational tool vendors. Work closely with senior stakeholders to align operational priorities with business objectives.
- Provide regular reporting on service health reliability metrics operational risks and improvement initiatives.
What you can expect from us
We wont just meet your expectations. Well defy them. So youll enjoy the comprehensive rewards package youd expect from a leading technology company. But also a degree of personal flexibility you might not expect. Plus thoughtful perks like flexible working hours and your birthday off.
Youll also benefit from an investment in cutting-edge technology that reflects our global ambition. But with a nimble small-business feel that gives you the freedom to play experiment and learn.
And we dont just talk about diversity and inclusion. We live it every day with thriving networks including dh Gender Equality Network dh Proud dh Family dh One dh Enabled and dh Thrive as the living proof. We want everyone to have the opportunity to shine and perform at your best throughout our recruitment process. Please let us know how we can make this process work best for you.
Our approach to Flexible Working
At dunnhumby we value and respect difference and are committed to building an inclusive culture by creating an environment where you can balance a successful career with your commitments and interests outside of work.
We believe that you will do your best at work if you have a work / life balance. Some roles lend themselves to flexible options more than others so if this is important to you please raise this with your recruiter as we are open to discussing agile working opportunities during the hiring process.
For further information about how we collect and use your personal information please see our Privacy Notice which can be found (here)
Required Experience:
Manager
About Company
Global leader in Customer data science, retail media and analytics, experts in working with brands, grocery retail, retail pharmacy, and retailer financial services.