Senior Observability Engineer
Job Location:
Woodland Hills, CA - USA
Monthly Salary:
Not provided by the employer
Posted:
26 June 2026 (30+ days ago)
Application Deadline:
23 September 2026
Vacancies:
1 Vacancy
Job Summary
Descriptions:
Customer is seeking a seasoned Observability expert who doesnt just manage dashboards but actively lives and breathes telemetry this role Personnel will elevate customer observability maturity across infrastructure applications and business transactions.
Personnel will own design and optimize the following core domains:
1. Operations & Noise Reduction
Alert-to-Incident Signal Optimization: Analyze and optimize our Alert-to-Incident noise ratio (targeting a baseline better than 10:1). Drive the evolution from chaotic alerting to high-fidelity actionable incident creation.
Dynamic Baselining & Anomaly Detection: Shift the paradigm away from rigid static thresholds. Implement dynamic baseline that intelligently accounts for time-of-day day-of-week and seasonal traffic patterns.
2. Guardrails Standards & Observability-as-Code
Observability-as-Code (OaC): Drive the maturity of our telemetry infrastructure by ensuring all dashboards alerts SLOs and monitor configurations are defined versioned and deployed as code.
CI/CD Instrumentation Gates: Establish and enforce automated instrumentation compliance gates within our deployment pipelines to ensure code is observable before it hits production.
Fleet Health Management: Centrally manage version and monitor the health of our Open Telemetry (OTel) collectors and agent fleets.
3. Advanced Diagnostics & Next-Gen Tech
Automated Root Cause Analysis (RCA): Implement platform capabilities that automatically surface probable root cause the moment an incident fire.
Change & Deployment Correlation: Ensure all deployments configuration changes feature flag toggles and database migrations are automatically annotated on dashboards and correlated to active incident timelines.
GenAI/LLM-Assisted Triage: Evaluate and adopt GenAI/LLM capabilities for advanced log pattern explanation and accelerated incident troubleshooting.
4. Telemetry Architecture & Data Strategy
Cloud-Native & Third-Party Monitoring: Ensure deep telemetry integration across cloud-managed services (AWS/Azure/GCP EKS/AKS Lambda RDS) and critical third-party SaaS dependencies (e.g. Guidewire Salesforce Earnix Uniphore payment gateways).
Lakehouse & Data Pipeline Integration: Architect pipelines to export raw telemetry data to our data Lakehouse (S3/ADLS) to power advanced ML pipelines and predictive analytics.
Predictive Capacity Analytics: Leverage the observability platform for capacity forecastingpredicting utilization trends for CPU memory queue depth and storage before saturation occurs.
Log Standardization: Drive org-wide standards for log structure and serialization to ensure seamless cross-platform parsing and querying.
5. Culture SLOs & Business Impact
End-to-End Business Transaction Tracing: Map and trace complex multi-service customer journeys (e.g. policy quote bind pay) to provide full-context business transaction visibility.
SLO/SLA Governance: Define implement and track Service Level Objectives (SLOs) across all production services.
Developer Empowerment & Self-Service: Democratize observability by fostering a proactive culture where developers instrument their own services during active development backed by standardized self-service health dashboards.
Monitoring logging tracing design (metrics logs traces)
Dashboarding alerting and telemetry pipelines
Observability platform design & optimization
Root Cause Analysis (RCA) incident analysis
SLO / SLI / SLA definition and error budgets
Strong understanding of AWS / Azure / GCP environments PennyMac - SRE Word
Expertise in:
Microservices architecture
Distributed systems & event-driven systems
High availability & scalability patterns
CI/CD pipelines (GitLab Jenkins) West - Excel
Infrastructure as Code (Terraform CloudFormation) PennyMac - SRE Word
Containerization (Docker Kubernetes troubleshooting) West - Excel
Release observability & rollback readiness
Advanced / Differentiator Skills
AIOps / AI-driven observability RE: Outlook
Predictive alerting / anomaly detection
Observability cost optimization
Chaos engineering basics
API & integration observability
Skills: AI Agents
Experience Required: 10 & Above
Customer is seeking a seasoned Observability expert who doesnt just manage dashboards but actively lives and breathes telemetry this role Personnel will elevate customer observability maturity across infrastructure applications and business transactions.
Personnel will own design and optimize the following core domains:
1. Operations & Noise Reduction
Alert-to-Incident Signal Optimization: Analyze and optimize our Alert-to-Incident noise ratio (targeting a baseline better than 10:1). Drive the evolution from chaotic alerting to high-fidelity actionable incident creation.
Dynamic Baselining & Anomaly Detection: Shift the paradigm away from rigid static thresholds. Implement dynamic baseline that intelligently accounts for time-of-day day-of-week and seasonal traffic patterns.
2. Guardrails Standards & Observability-as-Code
Observability-as-Code (OaC): Drive the maturity of our telemetry infrastructure by ensuring all dashboards alerts SLOs and monitor configurations are defined versioned and deployed as code.
CI/CD Instrumentation Gates: Establish and enforce automated instrumentation compliance gates within our deployment pipelines to ensure code is observable before it hits production.
Fleet Health Management: Centrally manage version and monitor the health of our Open Telemetry (OTel) collectors and agent fleets.
3. Advanced Diagnostics & Next-Gen Tech
Automated Root Cause Analysis (RCA): Implement platform capabilities that automatically surface probable root cause the moment an incident fire.
Change & Deployment Correlation: Ensure all deployments configuration changes feature flag toggles and database migrations are automatically annotated on dashboards and correlated to active incident timelines.
GenAI/LLM-Assisted Triage: Evaluate and adopt GenAI/LLM capabilities for advanced log pattern explanation and accelerated incident troubleshooting.
4. Telemetry Architecture & Data Strategy
Cloud-Native & Third-Party Monitoring: Ensure deep telemetry integration across cloud-managed services (AWS/Azure/GCP EKS/AKS Lambda RDS) and critical third-party SaaS dependencies (e.g. Guidewire Salesforce Earnix Uniphore payment gateways).
Lakehouse & Data Pipeline Integration: Architect pipelines to export raw telemetry data to our data Lakehouse (S3/ADLS) to power advanced ML pipelines and predictive analytics.
Predictive Capacity Analytics: Leverage the observability platform for capacity forecastingpredicting utilization trends for CPU memory queue depth and storage before saturation occurs.
Log Standardization: Drive org-wide standards for log structure and serialization to ensure seamless cross-platform parsing and querying.
5. Culture SLOs & Business Impact
End-to-End Business Transaction Tracing: Map and trace complex multi-service customer journeys (e.g. policy quote bind pay) to provide full-context business transaction visibility.
SLO/SLA Governance: Define implement and track Service Level Objectives (SLOs) across all production services.
Developer Empowerment & Self-Service: Democratize observability by fostering a proactive culture where developers instrument their own services during active development backed by standardized self-service health dashboards.
Monitoring logging tracing design (metrics logs traces)
Dashboarding alerting and telemetry pipelines
Observability platform design & optimization
Root Cause Analysis (RCA) incident analysis
SLO / SLI / SLA definition and error budgets
Strong understanding of AWS / Azure / GCP environments PennyMac - SRE Word
Expertise in:
Microservices architecture
Distributed systems & event-driven systems
High availability & scalability patterns
CI/CD pipelines (GitLab Jenkins) West - Excel
Infrastructure as Code (Terraform CloudFormation) PennyMac - SRE Word
Containerization (Docker Kubernetes troubleshooting) West - Excel
Release observability & rollback readiness
Advanced / Differentiator Skills
AIOps / AI-driven observability RE: Outlook
Predictive alerting / anomaly detection
Observability cost optimization
Chaos engineering basics
API & integration observability
Skills: AI Agents
Experience Required: 10 & Above