Enter a job title or keyword

Cloud Engineer


Job Location:

Chennai - India

Monthly Salary: Not provided by the employer
Experience Required: 2years
Posted: 28 August 2026 (Yesterday)
Application Deadline: 25 November 2026
Vacancies: 1 Vacancy

Job Summary

Culture at CloudifyOps :

Working at CloudifyOps is a rewarding experience! Great people a work environment that thrives on creativity and the opportunity to take on roles beyond a defined job description are just some of the reasons you should work with us.

About the Role :

Were looking for someone who genuinely wants to understand why systems fail not just respond to alerts. This role sits at the crossroads of cloud infrastructure and production reliability. Youll own monitoring handle on-call and be the person who digs in when things go wrong. At the same time were building an AI-powered pipeline monitoring tool and need someone curious enough to contribute to shaping it not just watching over it.


What youll do:

  • Handle the on-call rotation and own incidents end-to-end triage mitigation escalation where needed and clean resolution. You dont pass the baton and disappear.

  • Write clear structured RCAs after every significant incident what happened when why and what changes going forward. These go to clients so they need to work for both an engineer and a non-technical reader.

  • Maintain and improve the monitoring stack across environments dashboards alerting rules log pipelines and distributed traces. Treat noisy alerts as a problem to fix not something to mute.

  • Provision and manage cloud infrastructure on AWS using Terraform. This is a hands-on role not just reviewing what others set up.

  • Work with Kubernetes across multiple environments debugging pod and node issues.

  • Monitor CI/CD pipeline health via Jenkins and support teams using Rancher for workload and cluster management.

  • Track application performance using APM tooling and JVM metrics: spot anomalies investigate degradation and flag systemic issues before they become incidents.

  • Contribute to the AI monitoring tool initiative: prototype test iterate. This is early-stage work and needs someone willing to figure things out not just execute a finished design.

Tech Stack:

  • Cloud & Infrastructure : AWS Kubernetes (K8s) Terraform Linux

  • Observability & Metrics : Prometheus Grafana APM (Datadog / New Relic / Kfuse) JVM Metrics & GC Analysis ELK / EFK Stack Distributed Tracing

  • CI/CD & Platform : Jenkins ArgoCD Rancher Git Docker

  • Good to Have(Not Mandatory) : Python / Bash scripting OpenTelemetry Zenduty / OpsGenie ML / AI basics

Expectations:


On-call here is happen outside business hours and when they do its your responsibility to pick them up and drive them forward. Thats not unusual for this type of role but we want to be direct about it upfront.

Client expectations are high. Youll produce RCAs incident timelines and status communications that clients read closely. Your writing needs to be clear structured and free of vagueness. We investigated and fixed the issue isnt good enough. What was the issue why did it happen what was the business impact and what prevents recurrence.

We expect precision regarding your own work. After a change an incident or a deployment you should be able to clearly explain what you did and why without being prompted. Ownership doesnt end when the alert clears.


Who were looking for:


  • The ideal candidate should have 2.5 years to 5 years of work experience.

  • Strong fundamentals. You understand how distributed systems actually behave under load not just that a dashboard went red. You can read logs metrics and traces together.

  • Ownership without prompting. If you find a gap in monitoring coverage you close it. If an RCA feels incomplete you go back and make it precise. You dont wait to be asked.

  • Writes clearly under pressure. During an incident your updates should help not add noise. After one your documentation should be good enough that anyone picking it up six months later understands what happened.

  • Curious about what comes next. The AI tooling initiative needs someone interested in figuring it out not just waiting for a ticket. Some comfort with experimentation and ambiguity goes a long way here.


Equal opportunity employer

CloudifyOps is proud to be an equal opportunity employer with a global culture that embraces diversity. We are committed to providing an environment free of unfair discrimination and harassment. We do not discriminate based on age race color sex religion national origin disability pregnancy marital status sexual orientation gender reassignment veteran status or other protected category.




Required Skills:

AWS Azure GCP Kubernetes Docker Jenkins Monitoring Tools Ticketing Tools Scripting.