[8SN] Senior Site Reliability Production Support Engineer

Software Mind


Job Location:

Montreal - Canada

Monthly Salary: Not Disclosed
Posted on: 10 hours ago
Vacancies: 1 Vacancy

Department:

Software Development

Job Summary

About the Role

We are looking for a Senior Site Reliability / Production Support Engineer to support the deployment operations and ongoing reliability of a production UI service running on Kubernetes.

This role is focused on maintaining highly available cloud-native applications troubleshooting production issues and improving operational excellence. You will work closely with engineering teams to monitor service health investigate incidents and ensure reliable service delivery.

While this role supports a UI-based service it is not a frontend development position. Basic knowledge of Web Components is sufficient to perform first-level debugging when necessary.
 

What Youll Do

  • Support deployment operations and ongoing maintenance of a production service running on Kubernetes.
  • Monitor application health availability and performance.
  • Investigate and resolve production incidents using logs monitoring and debugging tools.
  • Perform log analysis using Splunk to identify root causes and troubleshoot service issues.
  • Collaborate with software engineers to improve service reliability and operational efficiency.
  • Participate in incident response and production support activities.
  • Assist with first-level debugging of UI-related issues involving Web Components.
  • Contribute to continuous improvements in automation monitoring and operational processes.
  • Support CI/CD pipelines and cloud-native deployment practices.

Qualifications :

Required Qualifications

  • 5 years of experience in Site Reliability Engineering DevOps Platform Engineering or Production Operations.
  • Strong hands-on experience with Kubernetes in production environments.
  • Experience supporting cloud-native applications.
  • Experience monitoring production systems and troubleshooting complex incidents.
  • Strong knowledge of Splunk for log analysis and debugging.
  • Experience working in Linux environments.
  • Understanding of networking fundamentals and distributed systems.
  • Experience collaborating with software engineering teams to resolve production issues.
  • Strong troubleshooting and root cause analysis skills.
  • Excellent written and spoken English (B2).

Additional Information :

Preferred Qualifications

  • Experience with CI/CD pipelines.
  • Experience with cloud platforms such as AWS Azure or GCP.
  • Familiarity with container technologies such as Docker.
  • Exposure to observability tools (Prometheus Grafana OpenTelemetry etc.).
  • Basic understanding of Web Components and frontend architecture.
  • Experience supporting high-availability enterprise SaaS platforms.
  • Knowledge of infrastructure automation or Infrastructure as Code (Terraform Helm Ansible etc.) is a plus.

What We Offer

  • Competitive salary and laptop
  • Professional development and training opportunities
  • Work with cutting-edge cloud and container technologies
  • Flexible work arrangements and collaborative team environment
  • Impact on organization-wide digital transformation initiatives

Remote Work :

Yes


Employment Type :

Full-time

About the RoleWe are looking for a Senior Site Reliability / Production Support Engineer to support the deployment operations and ongoing reliability of a production UI service running on Kubernetes.This role is focused on maintaining highly available cloud-native applications troubleshooting produc...

About Company

Company Logo

Software Mind develops solutions that make an impact for companies around the globe. Tech giants & unicorns, transformative projects, emerging technologies and limitless opportunities – these are a few words that describe an average day for us. Building cross-functional engineering te ... View more

View Profile View Profile