Automation SRE Engineer AIOps & Observability
Posted:
21 June 2026 (30+ days ago)
Application Deadline:
18 September 2026
Vacancies:
1 Vacancy
Job Summary
OP is partnering with a globally renowned leader in media entertainment and consumer experiences to secure a talented Automation SRE Engineer - AIOps & Observability. The team s mission is to eliminate reactive operations through high-fidelity telemetry and AI/ML-driven insight. As a hands-on contributor the Automation SRE designs observability frameworks drives SRE practices and develops the automation that connects signals to action - accelerating incident prevention detection and resolution across the global technology landscape.
Responsibilities of Role:
- Observability Platform Engineering
- Design build and maintain enterprise observability platforms spanning the full range of telemetry signals including metrics logs distributed traces events and profiles - across applications infrastructure and network domains.
- Implement and operate observability tooling including LogicMonitor Datadog Grafana Prometheus OpenTelemetry Splunk AppDynamics (AppD) and similar platforms across infrastructure application and network domains.
- Instrument services applications and infrastructure with telemetry collectors exporters and agents using standards such as OpenTelemetry (OTel).
- Build and maintain dashboards for various personas SLO/SLI reports and alerting configurations that provide accurate actionable signals with minimal noise.
- Define and manage Service Level Objectives (SLOs) Service Level Indicators (SLIs) and error budgets for critical enterprise services spanning applications infrastructure and platform layers.
- Drive alert quality initiatives - tuning thresholds eliminating alert fatigue and building runbook-linked notification workflows.
- Data Engineering for Operations
- Build and operate telemetry data pipelines that ingest normalize enrich and route operational data at enterprise scale.
- Design data schemas and models for operational metrics events logs and traces to support analytics and AIOps use cases.
- Integrate observability data sources with data lakes streaming platforms and analytics tools to enable advanced operational reporting.
- Ensure data quality retention policies access controls and governance for operational datasets.
- Develop self-service data products and APIs that surface operational insights to engineering and operations teams.
- AIOps & Intelligent Automation
- Design and deploy AIOps capabilities including anomaly detection predictive alerting event correlation and noise suppression.
- Build auto-remediation workflows that trigger from observability signals reducing mean time to repair (MTTR) and manual operational effort.
- Apply ML/AI models to operational data for root cause correlation incident trend prediction and capacity forecasting.
- Integrate AIOps platforms (e.g. Dynatrace BigPanda ServiceNow ITOM or equivalent) with the observability stack to enable unified event management and automated triage across IT domains.
- Explore and apply Generative AI and LLM-based tooling to operational use cases including incident summarization intelligent alert triage and on-call assistant capabilities.
- Continuously evaluate and improve the accuracy and coverage of AIOps detections through feedback loops and model retraining.
- Site Reliability Engineering (SRE) Practices
- Champion SRE principles across the organization: toil reduction error budget policy capacity planning reliability reviews and a culture of blameless continuous improvement.
- Lead post-incident reviews (PIRs) and blameless postmortems; identify systemic improvement opportunities and track remediation actions to closure.
- Partner with platform and application engineering teams to embed reliability practices into CI/CD pipelines and service design.
- Develop and maintain runbooks operational playbooks and on-call response documentation for observability platform services.
- Participate in on-call rotation to support the observability platform and drive continuous improvement of operational procedures.
- Automation & Infrastructure as Code
- Develop automation scripts tools and integrations using Python Go or similar languages to reduce manual operational effort.
- Implement observability-as-code practices: manage dashboards alerts SLOs and monitors through version-controlled configuration (e.g. Terraform).
- Maintain CI/CD pipelines for continuous delivery of observability stack changes configuration updates and tooling enhancements.
- Contribute to shared automation libraries reusable modules and internal developer tools used across the broader Services and Platforms organization.
- Collaboration & Business Partnership
- Work closely with Network Engineering Platform Engineering Cloud Security and Application teams to align observability coverage with business priorities.
- Present observability metrics SLO performance and reliability trends to leadership and stakeholders in clear non-technical terms where appropriate.
- Actively engage in code reviews architectural discussions and knowledge-sharing sessions with globally distributed team members.
- Mentor junior team members and contribute to a culture of engineering excellence curiosity and continuous learning.
Must Haves (Years of Experience languages programs tools etc.):
- 2 5 years of experience in Site Reliability Engineering DevOps Platform Engineering or a related technical discipline.
- Hands-on experience with one or more enterprise observability platforms (e.g. Datadog Grafana Prometheus AppD Splunk New Relic Logic Monitor or equivalent).
- Proficiency in Python and/or Go for scripting automation and tooling development; comfort working in a CLI/Linux environment.
- Solid understanding of telemetry fundamentals: metrics collection and aggregation structured logging and distributed tracing.
- Experience with cloud platforms (AWS Azure or GCP) and containerized workloads (Docker Kubernetes).
- Experience building or maintaining CI/CD pipelines and using version control (Git) for infrastructure and configuration management.
- Familiarity with Infrastructure-as-Code (IaC) tools such as Terraform Ansible or Pulumi for managing observability configurations.
- Demonstrated experience defining and managing SLOs SLIs and error budgets in a production environment.
- Strong analytical troubleshooting and problem-solving skills with a data-driven approach to reliability.
- Effective written and verbal communication skills; ability to collaborate across globally distributed cross-functional teams.
Education:
- Required: Bachelor s degree in Computer Science Engineering Information Systems or a related technical field.
- Preferred: Advanced degree or equivalent hands-on professional experience in SRE observability or data engineering.
Nice To Haves:
- Experience with OpenTelemetry (OTel) for end-to-end instrumentation collector configuration and telemetry pipeline design.
- Knowledge of AIOps platforms or experience applying ML/AI to operational data (anomaly detection event correlation forecasting).
- Experience with data streaming and pipeline technologies (e.g. Kafka Flink Spark Streaming or cloud-native equivalents).
- Familiarity with network observability: streaming telemetry (gNMI/gRPC) SNMP flow data (NetFlow/IPFIX) and network performance monitoring.
- Experience with chaos engineering tools and practices (e.g. Chaos Monkey Gremlin) to proactively validate system resilience.
- Experience operating in a 24x7 on-call environment and familiarity with incident management practices (ITSM/ServiceNow PagerDuty).
- Relevant certifications: AWS/Azure/GCP cloud certifications Kubernetes (CKA/CKAD) Datadog Fundamentals Google SRE certification or equivalent.
- Prior experience in large-scale enterprise media or entertainment technology environments.
- Experience with GenAI / LLM-based operational tooling for use cases such as automated incident summarization intelligent runbook generation or AI-assisted root cause analysis.
- Familiarity with CMDB / service catalog integration (e.g. ServiceNow CMDB) to enrich observability data with topology ownership and dependency context.
- Knowledge of eBPF-based observability tooling (e.g. Cilium Pixie Falco) for deep kernel-level and container telemetry without code instrumentation.
Why Join Us
At our organization technology powers the magic. As part of our team you ll help build and maintain the systems that support our iconic brands and global operations. We offer a collaborative work environment opportunities for growth and a culture that celebrates diversity creativity and innovation.
At our organization technology powers the magic. As part of our team you ll help build and maintain the systems that support our iconic brands and global operations. We offer a collaborative work environment opportunities for growth and a culture that celebrates diversity creativity and innovation.