SRE Observability SLO Engineer
Mexico City - Mexico
Job Summary
Roles and Responsibilities
Telemetry Standards & Architecture
Implement organization-wide telemetry standards covering metrics logs and distributed traces across all GridOS SaaS services.
Implement metrics collection for Kubernetes-hosted services (EKS/Rancher) including pod-level namespace-level and cluster-level metrics.
Working with the SRE Lead and SRE Platform Engineers help define and implement data retention policies cardinality budgets and telemetry cost controls to keep observability economically sustainable.
Publish and maintain an Observability Runbook library covering onboarding alert tuning and dashboard standards for Platform SRE and Production DevOps teams.
SLO Definition Tooling & Governance
Partner with product engineering Platform SRE and customer stakeholders to define meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs) per product and customer tier.
Build and maintain SLO tooling error budget burn-rate alerts burn-rate dashboards and automated SLO compliance reports.
Govern the SLO review cycle: facilitate monthly SLO reviews identify reliability risks early and drive prioritization of reliability work with the SRE Lead.
Translate SLOs into SLAs for customer-facing commitments in coordination with the SRE Team Lead.
Dashboards & Alerting
Design and build operational dashboards covering availability latency error rates and saturation (the Golden Signals) for every GridOS SaaS product.
Implement alert policies with noise-reduction practices: symptom-based alerting multi-window burn-rate rules and alert deduplication.
Create executive-level dashboards for SRE leadership and customer-facing uptime/availability reports aligned to contractual SLAs.
Establish and maintain alert routing escalation policies and on-call schedules in coordination with the incident response workflow.
Synthetic Monitoring
Design and implement a synthetic monitoring plan covering critical user journeys for each GridOS SaaS product and customer environment.
Build synthetic checks for API health UI flows and integration endpoints using AWS CloudWatch Synthetics or equivalent tooling.
Define alerting thresholds for synthetic monitors and integrate them into the broader incident detection pipeline.
Continuous Improvement Cadence
After v1.0 delivery transition into a roadmap-aligned improvement cycle: expand coverage for new features tune alert signal-to-noise and retire stale monitors.
Conduct periodic observability health reviews to identify gaps in coverage reduce MTTD (Mean Time to Detect) and improve MTTR (Mean Time to Resolve).
Collaborate with the Production DevOps engineer on FinOps validation correlate infrastructure cost metrics with performance and reliability data.
Required Experience
23 years in SRE observability engineering or infrastructure reliability roles.
Fluent in English.
Experience with at least one major observability platform Datadog Grafana Prometheus AWS CloudWatch Dynatrace or New Relic.
Decent understanding of distributed systems telemetry: metrics (Prometheus/CloudWatch) structured logging (CloudWatch
Logs Insights ELK) and distributed tracing (OpenTelemetry AWS X-Ray).
Experience with Kubernetes observability kube-state-metrics node exporters Helm deployed monitoring stacks and namespace-level resource metrics.
Proficiency in at least one query/visualization language: PromQL Splunk SPL Datadog Query Language or CloudWatch Logs Insights query syntax.
Experience enabling monitoring alerts to provide visibility to system health.
Scripting skills in Python and/or Bash for automation of monitoring configuration and report generation.
Key Skills and Technologies
Cloud Technologies - AWS Cloud Infrastructure - EKS RDS MSK S3 EC2 EBS SQS etc. Kubernetes - EKS Rancher
Deployment and Configuration Tools - Ansible Chef or Puppet
Observability tools and technology - Datadog Splunk NewRelic etc.
Alerting and notification - AWS and Azure alerting notification
Scripting - Go Python Groovy Bash
Linux Administration Skills
Nice to Have
Familiarity with OpenTelemetry (OTel) for vendor-agnostic instrumentation.
Experience with synthetic monitoring tools AWS CloudWatch Synthetics Datadog Synthetics or Catchpoint.
Experience in regulated industries energy utilities healthcare where compliance grade audit trails are required.
AWS certifications: CloudWatch / Observability specialty Solutions Architect Associate or Professional.
Education Qualification
Bachelors Degree in Computer Science
Personal Attributes:
Critical thinker; able to quickly adapt to changing environments
A hacker or tinkerer at heart
Risk taker not afraid to think outside the box or challenge the status quo
Emotional Intelligence ability to influence up and out and the ability to work independently
Must be a team player with a strong desire to win
Passionate about continuously learning
Highly organized and efficient; able to balance competing priorities and execute accordingly
Strong oral and written communication skills.
Relocation Assistance Provided: Yes
Required Experience:
IC
About Company
GE Vernova's Asset Performance Management software can help you increase asset reliability, minimize costs and reduce operational risks. View a demo today.