Enter a job title or keyword

Senior Site Reliability Engineer SRE – Kubernetes & Hybrid Cloud (mfd)

FactFinder


Job Location:

Berlin - Germany

Monthly Salary: Not provided by the employer
Posted: 4 September 2026 (14 hours ago)
Application Deadline: 2 December 2026
Vacancies: 1 Vacancy

Job Summary

Introduction
At a glance
  • Location & work model: Berlin hybrid
  • Tech stack: Kubernetes on our own servers Harvester (KubeVirt) Argo CD/Flux Prometheus/Grafana Longhorn/Ceph
  • Team: A growing SRE team you report to our CTPO for now and to the Team Lead SRE were hiring next; two system administrators in Pforzheim run the physical hardware
  • Process: Intro call take-home task (2h) 90-min tech interview with our developers leadership conversation meet the team
  • Languages: Fluent English required; German is a plus not a must

Why this role is special

Most SRE jobs today mean clicking around a managed cloud console. This one doesnt. We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester on-prem by default with elastic burst into the public cloud and the option to go cloud-only later. You wont inherit a finished SRE practice: youll help define it side by side with our Berlin development teams and you wont do it alone a Team Lead SRE hire is coming next.

SRE here is an enabling discipline: you build what our developers need to ship reliably while two system administrators in Pforzheim run the physical hardware. And the impact is direct our product discovery technology powers more than 2000 European online shops (Intersport SPAR Douglas and more) handling billions of shopper queries a year. When product discovery is slow or down our customers lose revenue in real time.

Your first 90 days

You get to know both products join the on-call rotation with a buddy and own your first reliability topic SLOs for one product alerting that actually helps at 3 a.m. or automating away a piece of toil. By day 90 youve shipped visible improvements and know where you want to take the platform next.

Your mission
  • Define and own SLOs SLIs and error budgets; drive data-informed reliability decisions
  • Lead incident response end-to-end: fast detection clear communication blameless postmortems and reduce whole classes of incidents structurally not case by case
  • Eliminate toil through automation and GitOps; evolve our observability (metrics logs traces alerting runbooks) across two different stacks
  • Help build our custom Kubernetes operator (CRDs) that makes stateful search clusters declarative self-healing and safely upgradable and roll out the auto-scaling (HPA/VPA KEDA cluster auto scaler) todays architecture makes hard
  • Plan capacity performance and cost across on-premises and cloud including the large-catalogue and peak-season loads our merchants care about and use AI tools wherever they measurably speed up diagnosis and operations
Your profile
Must-haves:
  • Kubernetes in production built not just used: youve set up and maintained clusters on your own servers (e.g. kubeadm RKE2 k3s) and know cluster lifecycle and upgrades managed-only experience isnt enough for this role
  • Lived SRE practice: SLOs error budgets incident management on-call
  • Hands-on experience with GitOpsor comparable infrastructure/deployment automation experience with Argo CD or Flux is a strong plus
  • Solid observability skills metrics logs traces alerting that people trust
  • A strong automation instinct youd rather fix a problems cause than repeat its workaround
  • A collaborative enabling mindset you see SRE as a service to our developers: you ask what they need discuss trade-offs openly and dont fall in love with your own solution
Nice-to-haves (genuinely optional well teach you the rest):
  • Harvester KubeVirt vSphere/ESXi OpenStack or similar virtualization/HCI platforms
  • Container storage (Longhorn Ceph) and datacenter networking (load balancing ingress VLAN)
  • Auto-scaling (HPA VPA KEDA cluster auto scaler) and capacity/cost planning
  • Experience building Kubernetes operators/CRDs
  • German language skillsCertifications (CKA CKS) are welcome but no substitute for hands-on experience in the tech interview well ask about what youve actually built and operated.

You dont tick every box or your title was never SRE Apply anyway. If youve owned production systems handled incidents and worked deeply with Kubernetes we want to hear from you production experience and engineering mindset matter more to us than titles or buzzwords.

THE JOY OF WORKING WITH US
  • Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.
  • Modern tech stack: Kubernetes Harvester GitOps auto-scaling and an exciting path toward the cloud with room to build things right.
  • AI-first mindset: We use AI as a real part of our daily work not as a buzzword.
  • Ownership & growth: Clear responsibility short decision paths and the opportunity to actively shape your role.
  • Flexible work: Hybrid work model three office days per week with a focus on outcomes.
  • Strong team: Experienced engineers an open feedback culture and an environment where reliability is treated as a real engineering discipline.
Job Location

Berlin Munich Pforzheim or Stockholm (all Hybrid)


About Company

Company Logo

FactFinder empowers eCommerce businesses to achieve more – driving higher conversion rates and customer loyalty through AI-powered product discovery and search. Trusted by more than 2,000 B2B and B2C online shops worldwide, we help millions of shoppers quickly find the most relevant p ... View more

View Profile View Profile