Team Lead Site Reliability Engineering (all genders)
Job Summary
FACT-Finder builds product discovery technology for eCommerce and is trusted by leading online shops across Europe with its two products Next Generation and Infinity. We are actively modernizing our hosting toward Kubernetes on Harvester as an on-prem hybrid with the option to scale fully into the cloud in the mid-term. As Team Lead Site Reliability Engineering (all genders) you own the reliability scalability and cost of our hosting environments drive this transformation end-to-end and lead the team that delivers it.
- You own the operational health of our hosting across on-premise (Frankfurt Stockholm) and cloud availability performance and incident management.
- You actively drive the modernization toward Kubernetes on Harvester: cluster topology storage (Longhorn) networking (VLAN load balancing ingress) backup and disaster recovery.
- You build a production-grade k8s platform: lifecycle upgrades RBAC secrets GitOps (Argo CD / Flux) observability and policy guardrails.
- You shape the NG Search Operator (custom Kubernetes operator) and solve auto-scaling (HPA VPA KEDA cluster autoscaler) for the current architecture.
- You concretely define our on-prem hybrid model: which workloads run where how we burst into the cloud how we keep latency and cost under control while keeping the architecture portable enough for a future cloud-only move.
- You own capacity planning and hosting cost and turn cost into a deliberate managed lever.
- You lead and develop our currently 4-person Hosting team own performance and technical direction and set the standards and ownership culture.
- You make AI a core part of our operations: diagnosis automation monitoring and insight.
- Strong background in infrastructure or platform engineering across on-premise and cloud.
- Hands-on depth with Kubernetes in production: cluster lifecycle upgrades networking storage RBAC observability GitOps delivery.
- Proven people leadership experience excellent communication and stakeholder management skills.
- Ideally practical experience with Harvester or comparable HCI/virtualization platforms (KubeVirt vSphere/ESXi OpenStack).
- Experience leading a real migration from bare metal / classic VMs to a k8s-based platform including stateful workloads storage migration cutover and rollback.
- Comfort designing or operating Kubernetes operators (custom controllers / CRDs) ideally for stateful systems like search databases or streaming.
- Solid grasp of auto-scaling primitives (HPA VPA cluster autoscaler KEDA) and how they interact with capacity planning on-prem and in the cloud.
- Experience with on-prem hybrid architectures and owning reliability capacity and cost for production systems.
- Hands-on fluency with AI tools in day-to-day operations.
- Fluent English; German is a plus.
- Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.
- Leadership with real scope: You lead an established team and shape our platform in a decisive phase of our transformation.
- Modern tech stack: Kubernetes Harvester GitOps auto-scaling and an exciting path toward the cloud with room to build things right.
- AI-first mindset: We use AI as a real part of our daily work not as a buzzword.
- Ownership & growth: Clear responsibility short decision paths and the opportunity to actively shape your role.
- Flexible work: Hybrid work model with a focus on outcomes.
- Strong team: Experienced engineers an open feedback culture and an environment where reliability is treated as a real engineering discipline.
- Attractive benefits: Competitive salary modern equipment learning budget and regular team events.
Berlin Munich Pforzheim or Stockholm (hybrid)
About Company
FactFinder empowers eCommerce businesses to achieve more – driving higher conversion rates and customer loyalty through AI-powered product discovery and search. Trusted by more than 2,000 B2B and B2C online shops worldwide, we help millions of shoppers quickly find the most relevant p ... View more