Site Reliability Engineer (SRE)
Orlando, FL - USA
Job Summary
Site Reliability Engineer (SRE)
Location: Orlando FL (Hybrid)
Job Summary
STAFFXPERT LLC is seeking a Site Reliability Engineer (SRE) on behalf of our client in Orlando FL. This role is responsible for ensuring the reliability availability scalability and performance of critical applications and platforms. The ideal candidate will have strong experience in production operations automation cloud technologies observability and incident management with a focus on improving system stability and operational excellence in a high-availability environment.
Key Responsibilities
-
Lead incident response efforts including troubleshooting service restoration and issue resolution for production environments.
-
Conduct root cause analysis (RCA) and implement preventive measures to improve system reliability.
-
Define monitor and enhance service health through observability practices SLIs and SLOs.
-
Design and maintain monitoring alerting logging and tracing solutions.
-
Automate operational processes and workflows to improve efficiency and reduce manual effort.
-
Optimize system performance scalability and resiliency across cloud and on-premises environments.
-
Support capacity planning disaster recovery failover strategies and resilience testing.
-
Collaborate with engineering infrastructure and operations teams to implement reliability best practices.
-
Participate in on-call rotations and major incident management activities.
Required Qualifications
-
Bachelors degree in Computer Science Information Technology Engineering or a related field or equivalent work experience.
-
Proven experience in Site Reliability Engineering DevOps Platform Engineering or Production Support.
-
Strong knowledge of cloud platforms such as AWS and/or Azure.
-
Hands-on experience with containerization and orchestration technologies including Kubernetes and Docker.
-
Experience with monitoring and observability tools such as Prometheus Grafana Splunk Datadog ELK or similar platforms.
-
Proficiency in scripting and automation using Python Bash PowerShell or comparable languages.
-
Experience with Infrastructure as Code (IaC) tools such as Terraform or Ansible.
-
Strong understanding of CI/CD pipelines and deployment automation.
-
Knowledge of distributed systems system performance tuning and reliability engineering concepts.
-
Excellent analytical troubleshooting and problem-solving skills.
Preferred Qualifications
-
Experience supporting large-scale mission-critical production environments.
-
Familiarity with SLI SLO and error budget frameworks.
-
Experience with security compliance and vulnerability remediation practices.
-
Industry certifications related to AWS Azure Kubernetes DevOps or Site Reliability Engineering.