Your Responsibilities:
- Ensure the stability and reliability of cloud-native applications deployed on GCP containerized with Docker and orchestrated via Kubernetes.
- Define implement and monitor SLOs SLAs and SLIs to measure system performance and user experience.
- Automate infrastructure provisioning using Terraform and manage Kubernetes configurations with Kustomize and Helm.
- Develop and maintain monitoring and alerting systems using Datadog and GCP-native tools.
- Conduct incident analysis and postmortems to drive continuous improvement.
- Collaborate with development teams to integrate reliability practices into CI/CD pipelines using GitHub Actions.
- Manage and troubleshoot database systems particularly PostgreSQL and Cassandra.
- Apply networking knowledge and Linux system administration skills to troubleshoot and optimize system connectivity and performance.
Qualifications :
- Educational background in Computer Science Software Engineering or equivalent practical experience.
- 5 years of experience in Site Reliability Engineering.
- Proven experience designing and operating elastic resilient systems in cloud environments.
- Strong understanding of GCP Kubernetes and container orchestration.
- Proficiency in infrastructure as code and configuration management tools (Terraform Helm Kustomize).
- Experience with monitoring and observability tools (Datadog GCP Monitoring).
- Solid scripting skills in bash and familiarity with automation frameworks.
- Experience with CI/CD pipelines especially using GitHub Actions.
- Familiarity with networking fundamentals and troubleshooting.
- Strong coding skills and ability to develop reliability-focused tooling.
- Strong problem-solving skills and a process-oriented mindset.
- Ability to work independently and collaboratively in a fast-paced environment.
- Passion for clean code automation and continuous improvement.
- Experience working within Agile/Scrum development teams.
- Very good fluency in English (written and spoken).
- Nice-to-Have: Availability for oncall.
Additional Information :
This resonates with you Apply now!
What we offer at
- Flexible and remote work:create your own schedule!
Flexibility defines the way we work and interact with each other. you have thepossibilitytowork remotely andadapt your working hoursin a very flexible way.
Wewantyou to become the best version of yourselfwithindividual and company-wideprograms andtrainingsfor people otherondevelopmentleadershipappreciation ...its timeto upskill yourcareer.
Life is full of surprises full of challenges and we want to support you whenever YOU need - at an individual level and duringevery stage of your life.
Want to know more about all our benefits Discover more here.
Lets connect soon. Apply for the role now!
Position grade within our career framework: Site Reliability Engineer T4 (Md8).
Remote Work :
No
Employment Type :
Part-time
Your Responsibilities: Ensure the stability and reliability of cloud-native applications deployed on GCP containerized with Docker and orchestrated via Kubernetes.Define implement and monitor SLOs SLAs and SLIs to measure system performance and user experience.Automate infrastructure provisioning us...
Your Responsibilities:
- Ensure the stability and reliability of cloud-native applications deployed on GCP containerized with Docker and orchestrated via Kubernetes.
- Define implement and monitor SLOs SLAs and SLIs to measure system performance and user experience.
- Automate infrastructure provisioning using Terraform and manage Kubernetes configurations with Kustomize and Helm.
- Develop and maintain monitoring and alerting systems using Datadog and GCP-native tools.
- Conduct incident analysis and postmortems to drive continuous improvement.
- Collaborate with development teams to integrate reliability practices into CI/CD pipelines using GitHub Actions.
- Manage and troubleshoot database systems particularly PostgreSQL and Cassandra.
- Apply networking knowledge and Linux system administration skills to troubleshoot and optimize system connectivity and performance.
Qualifications :
- Educational background in Computer Science Software Engineering or equivalent practical experience.
- 5 years of experience in Site Reliability Engineering.
- Proven experience designing and operating elastic resilient systems in cloud environments.
- Strong understanding of GCP Kubernetes and container orchestration.
- Proficiency in infrastructure as code and configuration management tools (Terraform Helm Kustomize).
- Experience with monitoring and observability tools (Datadog GCP Monitoring).
- Solid scripting skills in bash and familiarity with automation frameworks.
- Experience with CI/CD pipelines especially using GitHub Actions.
- Familiarity with networking fundamentals and troubleshooting.
- Strong coding skills and ability to develop reliability-focused tooling.
- Strong problem-solving skills and a process-oriented mindset.
- Ability to work independently and collaboratively in a fast-paced environment.
- Passion for clean code automation and continuous improvement.
- Experience working within Agile/Scrum development teams.
- Very good fluency in English (written and spoken).
- Nice-to-Have: Availability for oncall.
Additional Information :
This resonates with you Apply now!
What we offer at
- Flexible and remote work:create your own schedule!
Flexibility defines the way we work and interact with each other. you have thepossibilitytowork remotely andadapt your working hoursin a very flexible way.
Wewantyou to become the best version of yourselfwithindividual and company-wideprograms andtrainingsfor people otherondevelopmentleadershipappreciation ...its timeto upskill yourcareer.
Life is full of surprises full of challenges and we want to support you whenever YOU need - at an individual level and duringevery stage of your life.
Want to know more about all our benefits Discover more here.
Lets connect soon. Apply for the role now!
Position grade within our career framework: Site Reliability Engineer T4 (Md8).
Remote Work :
No
Employment Type :
Part-time
View more
View less