Enter a job title or keyword

Cloud Systems Engineer Site Reliability

TherapyNotes.com


Job Location:

Philadelphia, PA - USA

Yearly Salary: USD 110000 - 150000
Posted: 1 October 2026 (19 hours ago)
Application Deadline: 29 December 2026
Vacancies: 1 Vacancy

Job Summary

Description

About Us

TherapyNotes is the go-to superhero for behavioral health Practice Management and EHR software! Our top-notch SaaS solution handles scheduling billing documenting telehealth and more so clinicians can focus on awesome patient care.

Were a dynamic team of pros who love to innovate and push the envelope keeping our software cutting-edge. Join us and lets revolutionize behavioral health software together while making a real difference!

About The Position

We are seeking a Site Reliability Engineer to improve the reliability and operability of the production services and shared platforms supporting our growing 247 SaaS environment. In this role you will apply software and systems engineering practices to improve availability performance scalability resilience observability incident response and operational automation. You will partner with software development infrastructure database security and other technology teams to establish measurable reliability goals reduce operational toil and ensure services are supportable throughout their lifecycle. If you are passionate about building reliable systems solving complex production problems and driving continuous improvement we want to hear from you.



Requirements
  • BS degree in Information Systems Engineering or equivalent experience.
  • 5 years of engineering experience in Systems Engineering Cloud or Platform Engineering DevOps Software Engineering and/or SRE.
  • Experience designing and operating production systems using cloud-based compute storage networking and containerization technologies; Azure and Kubernetes preferred.
  • Strong Linux systems and networking fundamentals with experience troubleshooting complex distributed systems in production.
  • Expertise with an observability platform; Datadog experience strongly preferred. Experience with Prometheus Grafana New Relic or equivalent platforms is also valuable.
  • Experience with scripting and operational automation using tools such as Bash PowerShell or Python along with infrastructure-as-code and configuration-management practices.
  • Experience participating in production on-call rotations incident response root cause analysis and post-incident improvement.
  • Experience working in Agile/DevOps environments and operating production services using ITSM practices where applicable.
  • Prior software development experienceor experience investigating application behavior through code logs and distributed tracesis a plus
Responsibilities
  • Own and continuously improve how we use Datadog to make reliability visible and actionable across metrics logs traces dashboards monitors alerts and service-level views.
  • Design implement and maintain high-availability high-throughput data- and compute-intensive critical systems supporting a growing 247 SaaS platform.
  • Partner with service owners to define and improve reliability through meaningful SLIs SLOs error budgets actionable alerting and operational-readiness practices.
  • Participate in and help drive incident management for production events serving as an incident commander or technical responder as needed. Coordinate triage service restoration escalation communication incident documentation root cause analysis and completion of corrective actions.
  • Partner with development teams to investigate issues across the infrastructure and application layers using metrics logs distributed traces and code-level context.
  • Improve deployment safety and service resilience through automated validation recovery and rollback capabilities reliability testing and analysis of system failure modes.
  • Partner with other technical leaders to ensure all newly introduced systems are supportable and maintainable by both development and operations.
  • Provide escalated technical guidance and support to other technology teams throughout the organization.
  • Provide on-call coverage for production support and other duties as required.
  • Ensure supported systems and operational activities comply with organizational security HIPAA and operating policies.
  • Identify and eliminate repetitive operational toil using Bash PowerShell Python or Ansible. Manage infrastructure as code using Terraform/OpenTofu and configuration automation using Ansible.


Benefits
  • Competitive salary - $110000-$150000
  • Employer sponsored health dental vision life and disability insurance
  • Retirement plan with company contribution
  • Annual company profit sharing
  • Personal development/training budget
  • Open collaborative work environment
  • Extensive 2-week onboarding plan
  • Comprehensive mentorship program

Equal Opportunity Employer Statement & Applicant Rights
TherapyNotes LLC is an Equal Opportunity Employer and does not discriminate based on race color religion sex national origin age disability genetic information or any other protected status under federal state or local law. We are committed to providing a workplace free of discrimination and more information about your rights under federal employment laws please review the following:

If you require a reasonable accommodation during the application process please contact .

9/17/2026


Required Experience:

IC


About Company

Company Logo

Who We Are We're 180+ strong and still growing! Our team of developers, QA engineers, IT experts, customer success specialists, and business professionals are dedicated to providing our customers with the best EMR software experience. With thousands of active users, TherapyNotes is th ... View more

View Profile View Profile