Staff Observability Engineer
Posted:
31 May 2026 (30+ days ago)
Application Deadline:
28 August 2026
Vacancies:
1 Vacancy
Job Summary
Job Description Summary
As a Staff Software Engineer (Observability) you will be responsible for defining and implementing the observability strategy across PCS Digital Solutions Cloud Applications.Job Description
Roles and Responsibilities
In this role you will:
- Define and evolve theobservability vision and roadmapfor PCS DS applications
- Design and implement/integrate standardized observability frameworks(metrics logs traces events profiling).
- Collaborate with platform SRE and product teams toinstrument servicesusing OpenTelemetry and other modern observability tooling.
- Build and maintaindashboards alerts and SLOsthat reflect both technical and business health indicators.
- Evaluate integrate and optimize observability agents (e.g. Prometheus Fluent bit OTEL and other agents).
- Design self-remediation solutions leveraging observability tooling.
- Implement Best Practices for using GenAI tools of Observability platforms.
- Lead / contribute toincident analysis and postmortem reviews driving improvements in system resilience and observability coverage.
- Conduct Operational Readiness Reviews (ORRs) to validate monitoring alerting and rollback strategies before go-live.
- Ensure observability practices align withhealthcare compliance standards(e.g. HIPAA GDPR HITRUST).
- Mentor engineers and promote aculture of observability-first development.
Required Qualifications
- Bachelors or masters degree in computer science Engineering or a related technical field.
- 10 years of experience in software engineering SRE or platform engineering roles.
- 4 years of experience in contributing in observability solutions incloud-native environments(Kubernetes microservices serverless).
- Deep expertise inobservability pillars(metrics logs traces) and tools like OpenTelemetry Prometheus Grafana Datadog Dynatrace etc.
- Strong programming/scripting skills (e.g. Go Python Bash Terraform).
- Experience withdistributed tracingSLO/SLI frameworks andincident response workflows.
- Deep expertise in distributed systems microservices and cloud platforms (AWS Azure GCP).
- Experience with AI-powered anomaly detection automated incident response and cost optimization for observability at scale.
- Familiarity with SRE practices chaos engineering
- Excellent communication and collaboration skills.
Desired Characteristics
- Experience inhealthcare or regulated industries.
- Knowledge ofdata privacy and compliance(HIPAA HITRUST).
- Experience withcost optimizationandtelemetry data governance.
- Contributions to open-source observability projects.
Additional Information
Relocation Assistance Provided: No
Required Experience:
Staff IC
About Company
GE HealthCare provides digital infrastructure, data analytics & decision support tools helps in diagnosis, treatment and monitoring of patients