Senior AWS Site Reliability Engineers 2850
Guadalajara - Mexico
Job Summary
Seeking Experienced Senior AWS Site Reliability Engineers for Exciting Projects Remote in Mexico
We are looking for skilled Senior Site Reliability Engineers with a minimum of 5 years of experience to join a dynamic team within a leading organization. This role involves supporting and improving cloud operations for microservice-based platforms with a focus on production reliability incident response cloud infrastructure automation observability Kubernetes operations and CI/CD workflows across AWS and Azure environments.
- Own and improve the reliability of cloud-based services and supporting infrastructure.
- Participate in on-call rotations and support production systems outside normal business hours.
- Lead incident response activities including triage escalation mitigation and service restoration.
- Drive blameless postmortems and ensure corrective actions are tracked to closure.
- Design implement and maintain Infrastructure as Code using Terraform and tools such as Atlantis.
- Manage and enhance GitOps and deployment workflows using ArgoCD and related CI/CD tools.
- Support and improve cloud and container platforms across AWS and Azure.
- Manage Kubernetes-based workloads containers virtual servers and distributed systems.
- Build automation to reduce manual effort and improve operational efficiency.
- Configure and improve monitoring alerting logging diagnostics and observability.
With over 5 years of experience as a Senior Site Reliability Engineer you must be proficient in the following technical skills:
- Strong hands-on experience with AWS and Azure cloud platforms.
- Strong experience with Terraform for Infrastructure as Code (IaC).
- Experience with Atlantis ArgoCD or similar infrastructure and deployment automation tools.
- Strong hands-on experience with Docker and Kubernetes.
- Experience designing maintaining and troubleshooting complex CI/CD pipelines.
- Strong production support experience including incident management Root Cause Analysis (RCA) postmortems and runbook creation.
- Strong observability experience including monitoring alerting logging diagnostics and performance analysis.
- Good understanding of cloud networking security access controls and InfoSec practices.
- Experience with version control branching merging pull requests and conflict resolution.
- Understanding of cloud cost optimization and resource utilization.
- Experience with microservice-based platforms.
- Experience with Datadog CloudWatch Grafana Prometheus Splunk AppDynamics or similar tools.
- Scripting or programming experience using Python Bash Go or Java.
- Experience with SLI/SLO/SLA error budgets capacity planning and resilience engineering.
- Experience with disaster recovery testing and production readiness reviews.
- Prior experience mentoring junior engineers or leading technical troubleshooting.
- Bachelors degree or higher.
- Fluent in English (Advanced).
- Excellent communication empathy commitment leadership teamwork and a proactive attitude.
- Remote work from Mexico.
- Preferred hybrid model in Guadalajara Jalisco with expected onsite attendance 2 days per week.
- Work hours Monday to Friday 09:00 18:00.
- Advanced English skills are mandatory and only residents of Mexico.
- Attractive Salary Premium Benefits
- Performance bonuses grocery coupons and savings are found.
- Aguinaldo premium vacations and vacations paid
- SGMM Medical insurance family and Life insurance.
Candidates must include their compensation expectations in their applications and resumes in English.
Interested Apply now through this link:
Required Experience:
Senior IC