Enter a job title or keyword

[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

Software Mind


Job Location:

Montreal - Canada

Monthly Salary: Not provided by the employer
Posted: 14 August 2026 (30+ days ago)
Application Deadline: 11 November 2026
Vacancies: 1 Vacancy

Department:

Software Development

Job Summary

About the Role

This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack.

This is not general infrastructure and it is not a front-end developer role. The strongest candidates will have production SRE experience across Kubernetes operations observability runtime troubleshooting JVM / Java service troubleshooting Splunk and incident ownership.

What Youll Do

  • Support the deployment operation and reliability of production services running on Kubernetes.
  • Monitor service health and investigate production incidents across distributed applications.
  • Participate in on-call support incident response root cause analysis postmortems and reliability improvements.
  • Troubleshoot application runtime networking and service-to-service issues in collaboration with engineering teams.
  • Support CI/CD GitOps-based deployments observability and production monitoring.
  • Work within a client-directed backlog and established priorities.

Qualifications :

Required Qualifications

  • 5 years of experience in Site Reliability Engineering DevOps Platform Engineering Production Engineering or a closely related role including strong recent hands-on experience supporting Kubernetes-based production services.
  • 3 years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations including deployment scaling rollout / rollback resource tuning and service-to-service troubleshooting
  • Strong production incident response experience including on-call runbooks postmortems and paging hygiene
  • Splunk experience for log aggregation search and production troubleshooting
  • Prometheus and Grafana experience specifically building alert rules and dashboards not only using existing dashboards
  • CI/CD and infrastructure-as-code for containerized deployments including Helm and GitOps tools such as ArgoCD or Flux
  • Strong Linux and networking fundamentals including DNS load balancing TCP / HTTP HTTP/2 and Kubernetes networking
  • Production troubleshooting experience across and JVM/Java services with strong depth in at least one runtime environment. Experience may include heap snapshots CPU profiling event-loop and memory analysis as well as JVM GC log analysis thread dumps JVM tuning and Java service latency investigation.
  • Service-to-service authentication experience including mTLS certificate rotation certificate format conversion and JWT-based service authentication

Additional Information :

Nice to Have

  • Web Components / Lit experience to perform first-level debugging of UI-related issues
  • Server-side rendering or isomorphic runtime experience
  • Canary rollout / multi-version production operations
  • Distributed tracing and request-context correlation
  • KEDA or event-driven autoscaling
  • Experience with enterprise platform integration layers

What We Offer

  • Competitive salary and laptop
  • Professional development and training opportunities
  • Work with cutting-edge cloud and container technologies
  • Flexible work arrangements and collaborative team environment
  • Impact on organization-wide digital transformation initiatives

Remote Work :

Yes


Employment Type :

Full-time


About Company

Company Logo

Software Mind develops solutions that make an impact for companies around the globe. Tech giants & unicorns, transformative projects, emerging technologies and limitless opportunities – these are a few words that describe an average day for us. Building cross-functional engineering te ... View more

View Profile View Profile