Site Reliability Engineer (SRE) – Observability & Monitoring
Woonsocket, RI - USA
Job Summary
Job Description
About the Role
We are seeking a Site Reliability Engineer (SRE) with strong expertise in observability monitoring and distributed tracing to join our SRE team. The ideal candidate will help us design build and scale an observability framework that provides end-to-end visibility into our systems and applications. A strong focus will be places on OpenTelemetry as we continue to standardize our telemetry pipeline across logs metrics and traces.
Responsibilities
- Design implement and maintain observability solutions using OpenTelemetry Prometheus Grafana AppDynamics and Splunk.
- Build and manage telemetry pipelines (metrics logs traces) ensuring reliable data collection transformation and export.
- Lead initiatives to improve incident detection response and post-incident analysis with a strong emphasis on RCA (Root Cause Analysis).
- Define and maintain SLIs SLOs and error budgets to measure and improve system reliability.
- Partner with development and operations teams to instrument applications and services for better monitoring and tracing coverage.
- Develop dashboards alerts and visualizations to provide actionable insights into system health and performance.
- Contribute to automation and self-healing practices that improve uptime and reduce operational toil.
- Stay current with trends in observability and advocate best practices across the engineering organization.
Requirements
- 7 years of SRE/ Devops/ Cloud/ Infrastructure engineering experience with a focus on monitoring and observability.
- Hands-on experience with OpenTelemetry SDKs collectors and exporters.
- Proficiency with observability stacks such as Prometheus Grafana Loki Tempo Elastic Stack or Splunk Observability (Splunk/AppDynamics).
- Strong knowledge on cloud platforms (GCP or Azure).
- Hand-on experience on container orchestration using Kubernetes (OCP GKE AKS)
- Familiarity with CI/CD pipelines like (Jenkins and Github actions) infrastructure as code (Terraform/Ansible/ARM/CloudFormation).
- Experience provisioning infrastructure and capacity planning.
- Hands-on skills in programming languages like Java and Python.