Enter a job title or keyword

SRE Leader

Staffworxs


Job Location:

Atlanta, GA - USA

Monthly Salary: Not provided by the employer
Posted: 8 September 2026 (2 days ago)
Application Deadline: 6 December 2026
Vacancies: 1 Vacancy

Job Summary

At Staffworxs we dont just connect talent we power transformation. Headquartered in Frisco TX with teams in Bengaluru and Hyderabad we combine global reach with deep expertise. Our Digital & Data Analytics practice drives growth and innovation for some of the worlds top brands who continue to retain us as their trusted partner. If youre ready to make an impact youre in the right place.


Job Details:

Job Title: SRE Leader
Work Location: Atlanta GA (Alternative week onsite)
Duration: 12 months Contract

Job Description: -

Role Summary:
As an SRE Architect you will be a pivotal technical leader responsible for designing building and evolving the foundational systems and practices that ensure the reliability scalability performance and efficiency of our critical services. Moving beyond day-to-day operations you will focus on the strategic architectural direction of SRE function defining standards blueprints and frameworks that enable development teams and fellow SRE operations team to build and operate highly resilient systems. Leverage deep expertise in software engineering distributed systems cloud infrastructure and SRE principles to influence technology choices establish best practices and foster a proactive culture of reliability across the organization and much beyond observability pillar.

Key Responsibilities:
Reliability Strategy & Design: Architect and design highly available scalable secure and cost-effective infrastructure and application patterns on AWS
Define and evangelize SRE best practices standards and blueprints for service design deployment monitoring and operational readiness across the engineering organization
Review current observability implementation to identify gaps and define steps to reach next level maturity of observability setup to provide deep insights into system health and behaviour
With overall maturity lead the definition and implementation strategy for Service Level Indicators (SLIs) Service Level Objectives (SLOs) and Error Budgets for critical services

Platform Architecture & Automation: Design solutions to systematically reduce operational toil through automation and improved system design
Evaluate current SRE tools and automation frameworks (e.g. CI/CD pipelines Infrastructure as Code modules automated incident remediation chaos engineering platforms) and suggest enhancement that will help overall enhancement of capability
Evaluate prototype and recommend new technologies tools and methodologies to enhance system reliability developer productivity and operational efficiency


Technical Leadership & Consultation: Act as a senior technical advisor and subject matter expert on reliability scalability and performance for development and platform teams
Provide architectural guidance during the design phase of new services and features to ensure reliability principles are embedded early (shift-left)
Mentor and coach other SREs and engineers fostering technical excellence and adherence to SRE principles
Lead architectural reviews and production readiness assessments for critical systems

  • Resilience:
    • Lead blameless postmortems for significant incidents ensuring root causes are identified and systemic architectural improvements are prioritized and implemented
    • Architect and advocate for resilience patterns (e.g. circuit breaking rate limiting graceful degradation chaos engineering) within applications and infrastructure.
  • AI-Driven Operations & Intelligence
  • Define and drive the organizations AIOps strategy leveraging AI/ML and Generative AI capabilities to improve observability incident management root cause analysis capacity forecasting and operational efficiency.
  • Design and implement intelligent operational platforms that use AI agents knowledge graphs telemetry analytics and automation frameworks to proactively detect diagnose and remediate production issues.
  • Architect AI-powered production intelligence solutions that correlate logs metrics traces configuration data deployment events CMDB cloud services and dependency mappings to generate actionable operational insights.
  • Establish architectural patterns for AI-assisted incident triage impact analysis service dependency intelligence and automated remediation workflows.
  • Lead the adoption of Agentic AI frameworks MCP (Model Context Protocol) integrations and enterprise AI platforms to improve developer productivity and operational effectiveness.
  • Define governance frameworks for AI usage in production operations including explainability auditability security risk management and human-in-the-loop controls.
  • Partner with Data Engineering Platform Engineering and Application teams to develop AI-enabled reliability use cases and operational copilots.
  • Establish mechanisms to continuously evaluate and improve AI model effectiveness hallucination mitigation operational accuracy and reliability of AI-driven decision making.

Required Qualifications:
  • Proven experience in an architectural role designing solutions for reliability scalability and performance
  • Deep understanding and practical application of SRE principles (SLIs/SLOs error budgets toil reduction automation incident management postmortems)
  • Experience designing and implementing AIOps AI-powered observability or intelligent automation solutions to improve incident detection root cause analysis operational efficiency and service reliability.
  • Working knowledge of Generative AI AI agents MCP (Model Context Protocol) RAG architectures and enterprise AI platforms with experience building or operationalizing AI-enabled engineering tools in production environments.
  • Expertise in cloud computing platforms (e.g. AWS) including infrastructure networking and security services
  • Strong experience with containerization and orchestration technologies (Kubernetes Docker serverless computing)
  • Solid experience designing and implementing observability solutions (e.g. Dynatrace Prometheus Grafana ELK/EFK Stack Jaeger OpenTelemetry)
  • Strong programming/scripting skills (e.g. Python Go Bash) for automation and tool development
  • Excellent analytical problem-solving and strategic thinking skills.
  • Strong communication collaboration and leadership skills with the ability to influence technical direction across teams

Preferred Qualifications:
  • Experience designing and implementing chaos engineering practices and platforms



Staffworxs is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive workplace for all employees regardless of race color religion gender sexual orientation national origin age disability or veteran status.