Platform Engineer
Job Summary
At Mako we believe in the power of collaboration to drive innovation in pursuit of our collective ambition; excellence in trading. Our diverse community is connected through a commitment to being the best we can be with the highest standards of integrity.
Were looking for a Platform Engineer to help design build and operate our on-premises Kubernetes platform running on a self-managed Linux VM fabric (KVM/libvirt-based hypervisor layer). Youll own the infrastructure that sits beneath our application workloads from the hypervisor and VM provisioning up through Kubernetes cluster lifecycle networking storage and the golden-path tooling that application teams use to ship software.
This is a hands-on deeply technical role for someone who enjoys operating infrastructure at the systems level not a cloud-managed-service consumer role. The VM fabric itself is still being designed and built out so youll have real influence over its architecture not just its day-to-day operation. Youll be responsible for keeping the fabric and the clusters running on top of it healthy secure and performant without the safety net of a hyperscalers managed control plane.
What youll be involved in:
- Design the Linux VM fabric underpinning the platform from the ground up: hypervisor architecture (KVM/libvirt) host networking topology storage backing for VM disks and how the fabric will scale as workload demand grows
- Operate and maintain the hypervisor layer across multiple physical hosts including host patching live migration/evacuation and failure response with minimal workload disruption
- Design and maintain a distributed shared storage system such as Ceph or an alternative
- Design and maintain VM templating and golden-image pipelines so Kubernetes nodes are provisioned consistently and can be rebuilt or rotated on demand
- Automate the VM lifecycle end-to-end provisioning scaling patching decommissioning via infrastructure-as-code
- Manage compute memory and storage capacity planning across the fabric including host-level oversubscription strategy and headroom for node failure or maintenance
- Own virtual networking within the fabric host networking VLANs/overlay networks and design how it hands off cleanly into the Kubernetes CNI layer above it
- Design build and maintain the full lifecycle of on-prem Kubernetes clusters: bootstrapping version upgrades node scaling and decommissioning using tooling such as kubeadm Cluster API or Kubespray
- Manage the control plane end-to-end including etcd operations (backup/restore performance tuning disaster recovery) since theres no managed control plane to fall back on
- Configure and tune cluster networking: CNI selection network policy enforcement and on-prem load balancing
- Stand up and manage ingress and internal DNS for workloads across environments
- Own persistent storage integration for stateful workloads via CSI drivers
- Define and enforce multi-tenancy patterns across both layers tenant isolation and resource allocation on the VM fabric (compute storage network) as well as namespace/resource quota strategy RBAC and policy enforcement (OPA/Gatekeeper or Kyverno) at the Kubernetes layer
- Build and maintain GitOps-based delivery for both cluster configuration and workloads (ArgoCD or Flux) treating cluster and infrastructure state as code
- Harden hosts and clusters against security baselines (CIS benchmarks for Linux and Kubernetes) manage secrets (Vault sealed-secrets) and keep the container runtime and node OS patched
- Build observability across the full stack from hypervisor/host health up through cluster metrics and logs (we currently use Prometheus Grafana OpenSearch and Checkmk; open to alternatives) with particular focus on the capacity and failure signals a managed cloud provider would normally surface for you
- Plan and execute Kubernetes version upgrades and node OS/kernel upgrades with minimal workload disruption
- Design and maintain disaster recovery and backup strategy spanning both layers VM snapshots/backups and etcd/cluster state so the platform can be rebuilt from bare infrastructure if required
- Troubleshoot incidents across the entire stack from a misbehaving pod down through kubelet container runtime and CNI into the underlying VM and hypervisor layer when needed
- Participate in an on-call rotation for platform-level incidents; drive root-cause analysis and post-incident reviews
- Partner with application teams to define and support a smooth developer experience (self-service namespaces CI/CD integration internal developer platform tooling)
- Coordinate with datacentre/facilities and network teams on physical host provisioning rack capacity and hardware refresh cycles
- Contribute to the platform roadmap: capacity growth tooling upgrades and reducing operational toil through automation
What we need from you:
Essential
- Solid production experience running Kubernetes in a self-managed on-premises context (not just EKS/GKE/AKS) you understand what breaks when theres no managed control plane and youve operated etcd and the control plane yourself
- Hands-on experience designing and operating a Linux KVM/libvirt-based VM fabric as the foundation for Kubernetes host architecture templating and provisioning automation ideally from relatively early stage rather than just inheriting a mature environment
- Strong Linux systems administration background (networking storage service management kernel tuning troubleshooting under pressure)
- Practical in-depth knowledge of Kubernetes networking (CNI internals service meshes a plus) and storage (CSI drivers distributed storage systems such as Ceph/Longhorn)
- Experience with infrastructure-as-code and configuration management (Terraform Ansible Packer)
- Experience with GitOps workflows and CI/CD pipelines
- Comfortable with observability stacks (Prometheus/Grafana ELK/Loki) and using them to diagnose infra issues without cloud-native tooling
- Security-conscious: familiar with hardening standards RBAC network segmentation and secrets management
- Strong troubleshooting instincts across the full stack hypervisor OS network container runtime Kubernetes control plane
- Good written and verbal communication; comfortable working with distributed/hybrid teams
Desirable
- Experience with bare-metal Kubernetes provisioning
- Background in a regulated or air-gapped/restricted-network environment
- Contributions to open-source infrastructure tooling
- Experience using AI tooling (e.g. AI coding assistants LLM-based automation) to accelerate development were keen to use AI to speed up our development cycle
- Experience running AI infrastructure.
Why This Role
Youll have real ownership over infrastructure end-to-end from the hypervisor to the pod with no black-box managed services standing between you and root cause. If you like understanding systems all the way down and want to shape how a platform team operates outside the public cloud this is that role.
Were a FOSS-first company for our on-prem infrastructure: we build our platform on open-source tooling where it fits rather than defaulting to proprietary or vendor-locked products. That means fewer licensing constraints on how you design solutions and direct access to the source when something needs to be understood or fixed at depth.
We are Mako
At Mako we are welcoming inclusive and collaborative. We work fast and smart in a supportive and dress-down environment that allows colleagues to be themselves and achieve great things. We uphold the principles of a flat structure that offers unrivalled engagement with senior leadership and career development opportunities. We have a comprehensive benefits package including:
- Flexible leave and hybrid working policies
- Private health and dental insurance
- Generous pension scheme
- Free access to the Mako gym
- Employee wellbeing guidance and support
- Opportunity to become involved in the rewarding work of the Mako Foundation
About Mako
Mako is a leading options market maker with a global trading footprint. It has been at the forefront of options market making since 1999 from the open outcry trading pits to screen trading and automated algorithmic execution strategies that are driving the future of the industry.
From offices in London Dublin Amsterdam Singapore Sydney Brisbane and Chengdu Mako offers the best-in-class liquidity solutions across Equities Fixed Income Commodities and FX derivatives markets and prides itself in its entrepreneurial collaborative and philanthropic culture.
If you require any reasonable adjustments or assistance during the recruitment process please email and we will arrange this.
For further information on the Mako Group please refer to our website: .
Mako does not accept unsolicited CVs or candidate details from recruiters or search firms and will not pay any fees to such firms without a signed agreement.
Required Experience:
IC
About Company
Mako provides liquidity to global derivatives markets, primarily through options market making. Our proprietary technology, developed and refined over the last two decades, helps us stay ahead of the curve.