SRE – Site Reliability Engineer
Job Location:
Indianapolis, IN - USA
Monthly Salary:
Not provided by the employer
Posted:
11 August 2026 (9 hours ago)
Application Deadline:
8 November 2026
Vacancies:
1 Vacancy
Job Summary
Key Responsibilities:
- Design and maintain highly available and scalable applications and infrastructure.
- Monitor production systems and respond to incidents and outages.
- Develop automation to reduce manual operational work.
- Build and maintain CI/CD pipelines.
- Implement monitoring logging alerting and observability solutions.
- Perform root-cause analysis and implement permanent fixes.
- Manage cloud infrastructure across AWS Azure or GCP.
- Support Kubernetes and containerized environments.
- Define and track SLIs SLOs and SLAs.
- Participate in on-call rotations and incident management.
- Improve system performance scalability security and reliability.
- Work with development teams to improve application reliability.
Required Technical Skills:
| Category | Technologies |
|---|---|
| Cloud | AWS / Azure / GCP |
| Containers | Docker Kubernetes |
| CI/CD | Jenkins GitHub Actions GitLab CI Argo CD |
| Infrastructure as Code | Terraform CloudFormation |
| Scripting | Python Bash Shell |
| Monitoring | Prometheus Grafana Datadog New Relic |
| Logging | ELK/Elastic Stack Splunk |
| Version Control | Git GitHub/GitLab/Bitbucket |
| OS | Linux Unix |
| Networking | TCP/IP DNS HTTP/HTTPS Load Balancers |
| Databases | SQL NoSQL |
| Reliability | SLI SLO SLA Error Budgets |
| Incident Management | PagerDuty ServiceNow Jira |
Preferred Skills:
- Strong Linux administration experience.
- Kubernetes production experience.
- Experience with AWS/Azure/GCP.
- Knowledge of microservices and distributed systems.
- Experience with Terraform and Infrastructure as Code.
- Strong troubleshooting and debugging skills.
- Experience with production incident response.
- Understanding of security and networking concepts.
- Experience building observability platforms.