[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes
Department:
Job Summary
About the Role
This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack.
This is not general infrastructure and it is not a front-end developer role. The strongest candidates will have production SRE experience across Kubernetes operations observability runtime troubleshooting JVM / Java service troubleshooting Splunk and incident ownership.
What Youll Do
- Support the deployment operation and reliability of production services running on Kubernetes.
- Monitor service health and investigate production incidents across distributed applications.
- Participate in on-call support incident response root cause analysis postmortems and reliability improvements.
- Troubleshoot application runtime networking and service-to-service issues in collaboration with engineering teams.
- Support CI/CD GitOps-based deployments observability and production monitoring.
- Work within a client-directed backlog and established priorities.
Qualifications :
Required Qualifications
- 5 years of experience in Site Reliability Engineering DevOps Platform Engineering Production Engineering or a closely related role including strong recent hands-on experience supporting Kubernetes-based production services.
- 3 years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations including deployment scaling rollout / rollback resource tuning and service-to-service troubleshooting
- Strong production incident response experience including on-call runbooks postmortems and paging hygiene
- Splunk experience for log aggregation search and production troubleshooting
- Prometheus and Grafana experience specifically building alert rules and dashboards not only using existing dashboards
- CI/CD and infrastructure-as-code for containerized deployments including Helm and GitOps tools such as ArgoCD or Flux
- Strong Linux and networking fundamentals including DNS load balancing TCP / HTTP HTTP/2 and Kubernetes networking
- Production troubleshooting experience across and JVM/Java services with strong depth in at least one runtime environment. Experience may include heap snapshots CPU profiling event-loop and memory analysis as well as JVM GC log analysis thread dumps JVM tuning and Java service latency investigation.
- Service-to-service authentication experience including mTLS certificate rotation certificate format conversion and JWT-based service authentication
Additional Information :
Nice to Have
- Web Components / Lit experience to perform first-level debugging of UI-related issues
- Server-side rendering or isomorphic runtime experience
- Canary rollout / multi-version production operations
- Distributed tracing and request-context correlation
- KEDA or event-driven autoscaling
- Experience with enterprise platform integration layers
What We Offer
- Competitive salary and laptop
- Professional development and training opportunities
- Work with cutting-edge cloud and container technologies
- Flexible work arrangements and collaborative team environment
- Impact on organization-wide digital transformation initiatives
Remote Work :
Yes
Employment Type :
Full-time
About Company
Software Mind develops solutions that make an impact for companies around the globe. Tech giants & unicorns, transformative projects, emerging technologies and limitless opportunities these are a few words that describe an average day for us. Building cross-functional engineering te ... View more