Enter a job title or keyword

SaaS Monitoring Engineer

HelloKindred


Job Location:

London - UK

Monthly Salary: Not provided by the employer
Posted: 14 July 2026 (30+ days ago)
Application Deadline: 11 October 2026
Vacancies: 1 Vacancy

Job Summary

Anticipated Contract End Date/Length: December 18 2026
Work Set Up: Hybrid (60% Office - 40% Home)
Clearance Required: BPSS Eligibility Required

Our client in the Information Technology and Services industry is looking for a SaaS Monitoring Engineer to join a growing cloud operations team. This role is responsible for designing implementing and maintaining monitoring solutions that ensure the health performance availability and reliability of Software-as-a-Service (SaaS) platforms. The successful candidate will play a key role in proactive issue detection operational excellence incident management and observability initiatives while developing centralized monitoring dashboards that provide real-time visibility into service health and performance across distributed cloud environments.

Responsibilities:

  • Design implement and maintain monitoring frameworks for SaaS applications and cloud infrastructure.
  • Monitor system health availability latency error rates resource utilization and overall platform performance across distributed environments.
  • Enhance observability through the implementation and optimization of logs metrics and traces using modern monitoring platforms.
  • Develop and maintain a centralized console dashboard that provides a real-time view of SaaS service health and operational metrics.
  • Configure dashboard views to deliver actionable insights on service uptime API performance incident alerts dependency status and business-critical metrics.
  • Integrate data from multiple monitoring and operational sources into unified visualization platforms.
  • Optimize dashboard usability and reporting capabilities for engineering operations and leadership stakeholders.
  • Establish intelligent alerting mechanisms to detect anomalies performance degradation and service disruptions.
  • Investigate incidents identify root causes and implement corrective and preventive measures.
  • Collaborate with DevOps and engineering teams during incident response activities and post-incident reviews.
  • Automate monitoring processes alert escalation workflows and operational response procedures.
  • Refine alert thresholds and monitoring configurations to reduce noise and improve signal accuracy.
  • Implement predictive monitoring techniques to proactively identify potential service issues and outages.
  • Partner with software engineering DevOps and product teams to incorporate monitoring requirements throughout the development lifecycle.
  • Analyze SaaS performance trends and provide reporting on operational risks and service reliability.
  • Document monitoring strategies configurations standards and best practices.

Qualifications :

  • Bachelors degree in Computer Science Engineering or a related field or equivalent practical experience.
  • Proven experience monitoring cloud-based applications SaaS platforms or distributed systems.
  • Strong understanding of distributed systems architecture microservices and cloud platforms including AWS Azure or GCP.
  • Hands-on experience with monitoring observability and visualization tools such as Grafana Prometheus ELK Stack Splunk Datadog Azure Monitor or similar technologies.
  • Experience designing and developing interactive dashboards and console-based monitoring solutions.
  • Proficiency in scripting or programming languages such as Python Bash or Go.
  • Familiarity with containerization and orchestration technologies including Docker and Kubernetes.
  • Strong analytical troubleshooting and problem-solving capabilities.
  • Demonstrated ability to manage incidents identify root causes and implement sustainable solutions.
  • Excellent verbal and written communication skills with the ability to translate technical findings into actionable business insights.
  • Experience working in collaborative fast-paced technology environments.
  • Knowledge of Site Reliability Engineering (SRE) principles and practices preferred.
  • Familiarity with CI/CD pipelines and DevOps methodologies preferred.
  • Exposure to AIOps solutions or machine learning-based monitoring technologies preferred.
  • Cloud platform or monitoring technology certifications preferred.
  • Ability to maintain a proactive approach to operational excellence and continuous improvement.
  • Eligibility to obtain and maintain BPSS clearance.

Additional Information :

Please submit a CV/resume (mandatory) along with your application.

All your information will be kept confidential according to EEO guidelines.

Candidates must be legally authorized to live and work in the country where the position is based without requiring employer sponsorship.

HelloKindred is committed to fair transparent and inclusive hiring practices. We assess candidates based on skills experience and role-related requirements.

We appreciate your interest in this opportunity. While we review every application carefully only candidates selected for an interview will be contacted.

HelloKindred is an equal opportunity employer. We welcome applicants of all backgrounds and do not discriminate on the basis of race colour religion sex gender identity or expression sexual orientation age national origin disability veteran status or any other protected characteristic under applicable law.


Remote Work :

No


Employment Type :

Contract


About Company

Who is HelloKindred?HelloKindred are specialists in staffing marketing, creative and technology roles, offering a range of talent solutions that can be delivered on-site, remotely or hybrid.Our vision is to make work accessible and people’s lives better. We do this by disrupting tradi ... View more

View Profile View Profile