The Engineering Experience & Platforms domain aims to make life easier for engineers across NN. To strengthen reliability and improve observability across key platforms NN is looking for an interim Site Reliability Engineer to design implement and embed a scalable observability foundation.
What you are going to do The assignment focuses on setting up an OpenTelemetry collector layer for sources including Azure Databricks AWS Kubernetes Azure Kubernetes AWS and Azure. You will define and implement SLIs SLOs and SLAs for the Portable Stack domain and AI domain and enable teams to use the new observability capabilities effectively.
What we offer you Our people are the driving force behind our organisation. We value the knowledge and expertise you bring. We believe that your temporary commitment can take our organisation to a higher level. We offer you:
Competitive hourly rate depending on your knowledge and experience
Project starts 1st of August 2026 and has a duration of 6 to 9 months
Hybrid way of working partly from home and partly from the office
International working environment with loads of knowledge sharing
Who you are You are an experienced Site Reliability Engineer with strong production experience in Kubernetes and containerized workloads. You have hands-on cloud engineering experience in Azure and/or AWS including infrastructure as code and GitOps-based deployments. You bring deep observability expertise across metrics logs and traces using tools such as Grafana Prometheus Loki Tempo or similar. You are comfortable defining and managing SLOs SLIs and error budgets and you have a structured approach to incident management root cause analysis and reliability improvements. Automation skills in Python Bash or Go are expected as well as solid knowledge of CI/CD and safe deployment practices. You communicate clearly and work effectively with development teams to embed reliability into the software delivery lifecycle.
Who you will work with You will work closely with the newly formed observability team the domain architect and the Principal Engineer of the Kubernetes domain. You will also align with other teams in the domain to deliver a practical scalable solution that raises SRE maturity across NN.
The Engineering Experience & Platforms domain aims to make life easier for engineers across NN. To strengthen reliability and improve observability across key platforms NN is looking for an interim Site Reliability Engineer to design implement and embed a scalable observability foundation.What you a...
The Engineering Experience & Platforms domain aims to make life easier for engineers across NN. To strengthen reliability and improve observability across key platforms NN is looking for an interim Site Reliability Engineer to design implement and embed a scalable observability foundation.
What you are going to do The assignment focuses on setting up an OpenTelemetry collector layer for sources including Azure Databricks AWS Kubernetes Azure Kubernetes AWS and Azure. You will define and implement SLIs SLOs and SLAs for the Portable Stack domain and AI domain and enable teams to use the new observability capabilities effectively.
What we offer you Our people are the driving force behind our organisation. We value the knowledge and expertise you bring. We believe that your temporary commitment can take our organisation to a higher level. We offer you:
Competitive hourly rate depending on your knowledge and experience
Project starts 1st of August 2026 and has a duration of 6 to 9 months
Hybrid way of working partly from home and partly from the office
International working environment with loads of knowledge sharing
Who you are You are an experienced Site Reliability Engineer with strong production experience in Kubernetes and containerized workloads. You have hands-on cloud engineering experience in Azure and/or AWS including infrastructure as code and GitOps-based deployments. You bring deep observability expertise across metrics logs and traces using tools such as Grafana Prometheus Loki Tempo or similar. You are comfortable defining and managing SLOs SLIs and error budgets and you have a structured approach to incident management root cause analysis and reliability improvements. Automation skills in Python Bash or Go are expected as well as solid knowledge of CI/CD and safe deployment practices. You communicate clearly and work effectively with development teams to embed reliability into the software delivery lifecycle.
Who you will work with You will work closely with the newly formed observability team the domain architect and the Principal Engineer of the Kubernetes domain. You will also align with other teams in the domain to deliver a practical scalable solution that raises SRE maturity across NN.