Enter a job title or keyword

Senior Kubernetes-focused SRE


Job Location:

Dallas, TX - USA

Monthly Salary: Not provided by the employer
Posted: 14 August 2026 (21 days ago)
Application Deadline: 11 November 2026
Vacancies: 1 Vacancy

Job Summary

About the Role

We are seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale.

Key Responsibilities

  • Build automation and operational tools using Java Python and to improve efficiency scalability and platform operations.
  • Leverage AI and Generative AI technologies (Gemini Llama Mistral Qwen etc.) to automate alert analysis incident response operational workflows and runbook execution.
  • Implement API and microservices reliability solutions using Apigee/Apigee X REST APIs GraphQL gateways traffic routing canary deployments and failover strategies.
  • Manage Kubernetes platforms across GKE and Rancher RKE2 including cluster administration performance tuning and troubleshooting.
  • Ensure platform reliability and high availability by supporting active-active deployments disaster recovery readiness and multi-datacenter Kubernetes environments.
  • Develop observability and monitoring capabilities using tools such as Splunk Grafana Datadog and AppDynamics to meet reliability and performance objectives.
  • Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability security incident management and continuous improvement.

Core Technical Skills

  • Site Reliability Engineering (SRE) Reliability availability incident management SLO/SLI monitoring and operational excellence.
  • Kubernetes Platform Engineering 5 years of Strong hands-on experience with GKE and Rancher RKE2 multi-cluster management troubleshooting and performance optimization.
  • Cloud & Infrastructure Automation Strong experience in GCP Terraform Helm GitHub CI/CD and production-grade automation.
  • Software Development 5 years of Advanced programming skills in Python and Java ( preferred for integrations and automation workflows).
  • Observability & Monitoring Splunk Grafana Datadog AppDynamics alerting and platform health monitoring.
  • API & Microservices Engineering Apigee/Apigee X REST APIs GraphQL traffic routing canary deployments and failover strategies.
  • AI-Driven Operations (AIOps) Applying LLMs such as Gemini Llama Mistral and Qwen for alert analysis incident triage automation and operational workflows.