Enter a job title or keyword

Site Reliability Engineer (SRE) Lead


Job Location:

Bengaluru - India

Monthly Salary: Not provided by the employer
Posted: 17 July 2026 (30+ days ago)
Application Deadline: 14 October 2026
Vacancies: 1 Vacancy

Job Summary

Greetings from Maneva!

Job Description

Job Title - Site Reliability Engineer (SRE) Lead

Experience - 7 - 10 Years

Location - PAN India

Notice - Immediate Joiner

Requirements:

A Senior SRE drives enterprise-wide reliability observability automation and operational excellence across large-scale hybrid environments. This level requires architectural judgement leadership in high-severity incidents and the ability to mature SRE practices. operational functions performed by Site Reliability Engineering teamsto ensure availability reliability performance and resilienceof applications and infrastructure.

Key Responsibilities

1. Reliability Engineering & Service Governance

- Define SLIs/SLOs/SLAs and govern error budgets.

- Lead architecture reviews for reliability and resilience.

- Drive SRE adoption and reduce toil.

- Maintain and improve service availability resilience and SLO/SLI governance.

- Apply SRE principles (risk management error budget tracking elimination of toil).

- Evaluate application readiness and architecture for reliability standards.

2. Observability & AIOps

- Architect observability platforms (AppDynamics Datadog Prometheus Grafana Splunk Dynatrace ELK).

- Implement logging metrics and tracing.

- Improve visibility reduce noise and optimize MTTD/MTTI/MTTR.

3. Incident Problem & Change Management

- Lead major incident response and RCA.

- Execute capacity security and change control processes.

- Improve operational processes (capacity security change).

4. Automation & Platform Engineering

- Build automation for deployments monitoring remediation.

- Design CI/CD pipelines and IaC.

- Drive AIOps adoption.

5. Leadership & Collaboration

- Mentor teams and lead reliability initiatives.

- Influence roadmaps and enforce reliability standards.

Required Technical Expertise

- Strong hands-on experience in Unix/Linux Shell scripting Python/Java/Go/NodeJS.

- Deep expertise in monitoring & observability tools

- Distributed systems knowledge.

- Incident leadership and RCA.

- Proven CI/CD pipeline and IaC experience.

- Dashboard/KPI/metric creation and tracking.

Experience Requirements

- 5 years in SRE/DevOps/Production Engineering.

- Experience with mission-critical large-scale systems.

- Ability to work with architects and leadership.

- Experience operating large-scale distributed systems using SRE practices.

If you are excited to grab this opportunity please apply directly or share your CV atand.