[MLA] Senior Site Reliability Engineer (SRE) – Kubernetes
Job Summary
Project the aim youll have
We are the AI Experience Framework team that builds the platform powering ServiceNows AI-first user interfaces - an SSR runtime (karuna) built on Lit and server-rendered web components running behind a multi-tier proxy/HTTP2 routing chain with sharded V8 isolate pools paired with a ServiceNow Glide/Java platform layer (karuna-glide) that supplies metadata ACLs and service artifacts. This role owns production reliability for that stack end to end: Kubernetes deployment and operations observability and hands-on troubleshooting of both the and JVM sides of the system - not generalist infrastructure work.
Position how youll contribute
- Support the deployment operation and reliability of production services running on Kubernetes.
- Monitor service health and investigate production incidents across distributed applications.
- Participate in on-call support incident response root cause analysis postmortems and reliability improvements.
- Troubleshoot application runtime networking and service-to-service issues in collaboration with engineering teams.
- Support CI/CD GitOps-based deployments observability and production monitoring.
- Work within a client-directed backlog and established priorities.
Qualifications :
Expectations the experience you need
- 5 years of experience in Site Reliability Engineering DevOps Platform Engineering Production Engineering or a closely related role including strong recent hands-on experience supporting Kubernetes-based production services.
- 3 years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations including deployment scaling rollout / rollback resource tuning and service-to-service troubleshooting
- Strong production incident response experience including on-call runbooks postmortems and paging hygiene
- Splunk experience for log aggregation search and production troubleshooting
- Prometheus and Grafana experience specifically building alert rules and dashboards not only using existing dashboards
- CI/CD and infrastructure-as-code for containerized deployments including Helm and GitOps tools such as ArgoCD or Flux
- Strong Linux and networking fundamentals including DNS load balancing TCP / HTTP HTTP/2 and Kubernetes networking
- Production troubleshooting experience across and JVM/Java services with strong depth in at least one runtime environment. Experience may include heap snapshots CPU profiling event-loop and memory analysis as well as JVM GC log analysis thread dumps JVM tuning and Java service latency investigation.
- Service-to-service authentication experience including mTLS certificate rotation certificate format conversion and JWT-based service authentication
- Very good spoken and written English.
Additional skills the edge you have
- Web Components / Lit experience to perform first-level debugging of UI-related issues
- Server-side rendering or isomorphic runtime experience
- Canary rollout / multi-version production operations
- Distributed tracing and request-context correlation
- KEDA or event-driven autoscaling
- Experience with enterprise platform integration layers
Additional Information :
Our offer professional development personal growth:
- Flexible employment and remote work
- International projects with leading global clients
- International business trips
- Non-corporate atmosphere
- Language classes
- Internal & external training
- Private healthcare and insurance
- Multisport card
- Well-being initiatives
Position at: Software Mind
Remote Work :
Yes
Employment Type :
Full-time
About Company
Software Mind develops solutions that make an impact for companies around the globe. Tech giants & unicorns, transformative projects, emerging technologies and limitless opportunities these are a few words that describe an average day for us. Building cross-functional engineering te ... View more