Senior SRE Senior Site Reliability Engineer (SRE)
Orlando, FL - USA
Job Summary
High-Priority!
| 243352 | Site Reliability Engineer - Observability & Resilience Local to HUBs specific locations (Glendale Orlando Seattle) |
| RECRUITER ADDITIONAL REQUIREMENT NOTES : | Orlando FL - Recruiter Focus: Target senior SRE candidates with strong experience in reliability engineering incident management SLO/SLI implementation using Nobl9 Kubernetes observability (OpenTelemetry Grafana Cloud AppDynamics) and AWS Well-Architected Framework reviews. Prioritize candidates who have led automation chaos engineering RCA-driven reliability improvements and large-scale production resilience initiatives. |
| JOB TITLE : | Senior SRE |
| SKILL CATEGORY : | Cloud: AWS |
| REQUIRED SKILLS : | Site Reliability Engineering (SRE) & Kubernetes Operations |
| WORK LOCATION : | Orlando FL |
| ONSITE / REMOTE : | Hybrid |
| SALARY : | $100000 - $150000 Yearly |
| Contract / Direct Hire : | |
| DURATION : | Full Time |
| MUST BE INCLUDED WITH SUBMITTAL : |
|
| This opportunity is competitive and the required turnaround time for quality talent is rather slim. With that please confirm whether or not youll have talent available for our review over the next 24-72 hours. Please feel free to reach out if you need me to clarify the qualification criteria or the scope of responsibilities. | |
| JOB DESCRIPTION : | Job Title: Senior Site Reliability Engineer (SRE) Overview / Summary We are seeking a Site Reliability Engineer (SRE) with 8-10 years of experience to drive reliability observability and resilience improvements across critical systems. This is a high-impact front-line operations role focused on real-time incident response proactive prevention continuous automation and reliability engineering for Tier-1 business-critical applications. Key Responsibilities Drive automation initiatives to improve system performance and operational efficiency. Reliability Assessment & Engineering Conduct application reliability assessments using established reliability frameworks. Incident Management & RCA Analyze incident trends using CSI or equivalent incident management platforms. Service Level Management Define and implement SLIs. Cloud & Platform Reliability Review cloud architectures against AWS Well-Architected Framework principles. Kubernetes & Infrastructure Reliability Review Kubernetes cluster health and workload configurations. Observability & Monitoring Design and improve enterprise observability strategies. Application Performance Engineering Conduct dependency mapping and architecture reviews. Chaos Engineering & Resilience Testing Design and execute chaos engineering experiments using Gremlin or Harness Chaos Engineering. Automation & Self-Healing Identify repetitive operational tasks suitable for automation. Required Qualifications 8-10 years of experience in Site Reliability Engineering. #LI-ST1 #LI-Hybrid #Hiring |
Swathi Goutham
swathi@
Skanda Solutions LLC
105 Raider Boulevard Suite 205 Hillsborough NJ 08844
| | |
This email is not subject to a legally binding commitment. The information transmitted is intended only for the person or entity to which it is addressed and may contain confidential and / or privileged material. Any review retransmission dissemination or other use of or taking of any action in reliance upon this information by persons or entities other than the intended recipient is prohibited. If you received this in error please contact the sender and delete the material from any computer.
Required Skills:
KUBERNETES SPLUNK AWS