Enter a job title or keyword

[MLA] Senior Site Reliability Engineer AI Experience Framework

Software Mind


Job Location:

Kraków - Poland

Monthly Salary: Not provided by the employer
Posted: 14 August 2026 (30+ days ago)
Application Deadline: 11 November 2026
Vacancies: 1 Vacancy

Job Summary

Project the aim youll have 

We are the AI Experience Framework team that builds the platform powering ServiceNows AI-first user interfaces - an SSR runtime (karuna) built on Lit and server-rendered web components running behind a multi-tier proxy/HTTP2 routing chain with sharded V8 isolate pools paired with a ServiceNow Glide/Java platform layer (karuna-glide) that supplies metadata ACLs and service artifacts. This role owns production reliability for that stack end to end: Kubernetes deployment and operations observability and hands-on troubleshooting of both the and JVM sides of the system - not generalist infrastructure work. 

Position how youll contribute

  • Own Kubernetes deployment and operational health for Framework services including scaling rollout/rollback strategy and resource tuning 
  • Build and maintain production observability - Grafana dashboards and Prometheus alerting rules - across the SSR runtime and the Glide platform layer 
  • Diagnose and resolve production incidents: event-loop stalls heap growth V8 isolate exhaustion and isolate-pool scheduling issues under concurrent versioned traffic (vN/vN-1) 
  • Diagnose and resolve JVM production incidents on the Glide/Java side: GC pressure thread dumps and platform-service latency 
  • Own incident response for the team: runbooks on-call rotation postmortems and paging hygiene 
  • Drive CI/CD and infrastructure-as-code for Kubernetes manifests/Helm and deployment pipelines 
  • Partner with the framework engineering team to identify reliability gaps before they become incidents - capacity planning load testing chaos/failure-injection where useful 
  • Represent production reliability concerns in architecture reviews for new framework capabilities 

Qualifications :

Expectations the experience you need

  • Production operations/SRE experience including hands-on Kubernetes deployment scaling and incident response 
  • Direct operational experience troubleshooting in production: reading heap snapshots and CPU profiles diagnosing event-loop blocking and understanding process/worker isolation models (V8 isolates or equivalent sandboxing) 
  • Direct operational experience troubleshooting JVM-based services in production: GC log analysis thread dump analysis and JVM tuning 
  • Hands-on experience building and maintaining Prometheus alerting rules and Grafana dashboards from scratch not just consuming existing ones 
  • Strong Linux/networking fundamentals: DNS load balancing TCP/HTTP semantics (including HTTP/2) and debugging service-to-service networking inside Kubernetes 
  • Experience with CI/CD and infrastructure-as-code for containerized deployments (Helm GitOps tooling such as ArgoCD/Flux or equivalent) 
  • Track record owning on-call rotations writing runbooks and driving postmortems that lead to real reliability improvements 
  • Experience with Splunk for log aggregation search and production troubleshooting. 
  • Hands-on experience with in-memory caching systems (Valkey/Redis) in production key design TTL/eviction tuning and tenant-scoped cache invalidation  
  • Experience managing service-to-service mTLS - certificate issuance rotation and format conversion (e.g. PKCS/BCFKSPEM) - plus JWT-based service authentication 
  • Very good spoken and written English. 

Additional skills the edge you have

  • Familiarity with server-side rendering architectures and the specific failure modes of isomorphic runtimes (markup mismatches browser-API leakage into server code) 
  • Experience operating multi-version/canary rollout strategies (two live app versions served concurrently) 
  • Working knowledge of the ServiceNow Glide platform or a comparable enterprise platform integration layer 
  • Experience with distributed tracing and request-context correlation across service boundaries 
  • Familiarity with event-driven autoscaling (e.g. KEDA ScaledObjects driven by PromQL triggers) as a complement to standard HPA-based scaling 

Additional Information :

Our offer professional development personal growth:

  • Flexible employment and remote work  
  • International projects with leading global clients 
  • International business trips  
  • Non-corporate atmosphere 
  • Language classes 
  • Internal & external training 
  • Private healthcare and insurance  
  • Multisport card 
  • Well-being initiatives 

Position at: Software Mind


Remote Work :

Yes


Employment Type :

Full-time


About Company

Company Logo

Software Mind develops solutions that make an impact for companies around the globe. Tech giants & unicorns, transformative projects, emerging technologies and limitless opportunities – these are a few words that describe an average day for us. Building cross-functional engineering te ... View more

View Profile View Profile