Senior Kubernetes-focused SRE
Job Location:
Dallas, TX - USA
Monthly Salary:
Not provided by the employer
Posted:
14 August 2026 (21 days ago)
Application Deadline:
11 November 2026
Vacancies:
1 Vacancy
Job Summary
About the Role
We are seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale.
Key Responsibilities
- Build automation and operational tools using Java Python and to improve efficiency scalability and platform operations.
- Leverage AI and Generative AI technologies (Gemini Llama Mistral Qwen etc.) to automate alert analysis incident response operational workflows and runbook execution.
- Implement API and microservices reliability solutions using Apigee/Apigee X REST APIs GraphQL gateways traffic routing canary deployments and failover strategies.
- Manage Kubernetes platforms across GKE and Rancher RKE2 including cluster administration performance tuning and troubleshooting.
- Ensure platform reliability and high availability by supporting active-active deployments disaster recovery readiness and multi-datacenter Kubernetes environments.
- Develop observability and monitoring capabilities using tools such as Splunk Grafana Datadog and AppDynamics to meet reliability and performance objectives.
- Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability security incident management and continuous improvement.
Core Technical Skills
- Site Reliability Engineering (SRE) Reliability availability incident management SLO/SLI monitoring and operational excellence.
- Kubernetes Platform Engineering 5 years of Strong hands-on experience with GKE and Rancher RKE2 multi-cluster management troubleshooting and performance optimization.
- Cloud & Infrastructure Automation Strong experience in GCP Terraform Helm GitHub CI/CD and production-grade automation.
- Software Development 5 years of Advanced programming skills in Python and Java ( preferred for integrations and automation workflows).
- Observability & Monitoring Splunk Grafana Datadog AppDynamics alerting and platform health monitoring.
- API & Microservices Engineering Apigee/Apigee X REST APIs GraphQL traffic routing canary deployments and failover strategies.
- AI-Driven Operations (AIOps) Applying LLMs such as Gemini Llama Mistral and Qwen for alert analysis incident triage automation and operational workflows.