SRE Engineer


Job Location:

Washington, DC - USA

Monthly Salary: Not Disclosed
Posted on: 1 hour ago
Vacancies: 1 Vacancy

Job Summary

Partnering with a premier client in the Washington D.C. area to find a talented Site Reliability Engineer (SRE) to champion system availability performance and automation across their enterprise cloud this role you will bridge the gap between development and operations by implementing robust CI/CD pipelines and Infrastructure-as-Code (IaC) while heavily leveraging the Dynatrace observability platform to drive deep-dive distributed tracing build intelligent dashboards and tune anomaly detection. As a core member of the reliability team you will apply formal SRE principles-such as defining SLIs/SLOs and managing error budgets-to optimize capacity and resiliency ensure strict security compliance and participate in an on-call rotation using ITIL frameworks to minimize incident response times.

Key Responsibilities

Observability & Monitoring: Standardize and automate Dynatrace installations integrate telemetry collection into CI/CD pipelines enforce tagging/metadata standards configure distributed tracing with context propagation and optimize custom dashboards and anomaly alerts.

Deployment & Automation: Design implement and maintain CI/CD pipelines using GitHub Actions AWS CodePipeline or Jenkins; provision scalable cloud infrastructure using Terraform CloudFormation or AWS CDK.

Incident & Problem Management: Serve as a production on-call responder using ITIL frameworks and ServiceNow; lead troubleshooting efforts conduct deep root-cause analysis (RCA) and author comprehensive knowledge base articles.

Reliability Engineering: Champion SRE metrics including Service Level Indicators (SLIs) Service Level Objectives (SLOs) and error budgets; design and execute resiliency test plans and support performance testing.

Performance & Capacity Optimization: Drive operational cost optimization initiatives across cloud environments and configure robust auto-scaling policies and thresholds.

Security & Compliance: Manage service accounts access permissions and digital certificates; respond rapidly to security incidents and execute remediation protocols.

Qualifications & Requirements

Education & Experience: Bachelors degree in Computer Science Engineering or a related technical field paired with 2 to 4 years of hands-on experience in SRE DevOps or infrastructure-focused roles.

Cloud & Containerization: Practical hands-on experience managing multi-tenant environments within AWS and Azure alongside a solid understanding of container technologies like Docker Kubernetes and Amazon ECS.

Automation & Scripting: Mid-level proficiency in Python (or similar scripting languages) and practical experience with configuration management tools like Ansible to build automated self-service tools.

Systems & Networking Architecture: Strong foundational knowledge of Linux systems engineering core networking concepts and navigating relational cloud-native and NoSQL databases.

Professional Competencies: Excellent written and verbal communication skills for cross-functional collaboration a proven ability to work independently and the flexibility to participate in an on-call rotation outside standard business hours.

Partnering with a premier client in the Washington D.C. area to find a talented Site Reliability Engineer (SRE) to champion system availability performance and automation across their enterprise cloud this role you will bridge the gap between development and operations by implementing robust CI/CD ...