Digital Technology Senior Specialist – Observability & AI Ops
Job Summary
DT Senior Specialist Observability & AI Ops
Would you like to help shape and implement our Digital Technology teams strategic direction
Are you passionate about helping improve observability and digital operations
Join our Digital Technology team!
We operate at the heart of Baker Hughes digital transformation journey. Our team delivers enterprise observability and AIOps capabilities that help technology teams detect issues earlier troubleshoot faster and improve service performance across cloud infrastructure and application environments.
Partner with the best
As an AIOps & Observability Engineer you will support the implementation enhancement onboarding and day-to-day operations of observability platforms with a focus on Elastic Stack capabilities and practical SRE-driven operational outcomes.
As a Senior AI Ops Engineer you will be responsible for:
- Implement and manage the life cycle of enterprise observability platform based on elastic tech stacks not limited to Kibana Logstash Beats Elastic Agent Fleet Elastic APM components etc.
- Onboard full suite of 20000 plus devices under observability umbrella
- Onboarding infrastructure cloud services applications and platforms into observability and monitoring solutions.
- Building and maintaining dashboards visualizations alerts and operational reports for technology teams.
- Configuring log metric trace uptime and APM data collection across supported environments.
- Assisting with data ingestion parsing enrichment and retention activities.
- Supporting incident investigation troubleshooting and root cause analysis using observability data.
- Collaborating with cloud infrastructure application and SRE teams to improve system reliability and service visibility.
- Contributing to automation initiatives using scripting Infrastructure as Code and repeatable deployment practices.
- Participating in observability platform upgrades patching performance tuning and operational support activities.
- Creating and maintaining runbooks knowledge articles dashboards standards and operational runbooks related documentation.
- Contributing to continuous improvement of observability practices monitoring coverage and operational readiness.
Fuel your passion
- Have 7 years SRE/DevOps experience in enterprise-scale or mission-critical environments
- Have 5 years Cloud / Application / Platform operations and administration (AWS Azure hybrid or multi-cloud)
- Have 5 years Automation CI/CD and scripting proficiency (Python Bash PowerShell Ruby or equivalent)
- Have 5 years - Exposure to containers and cloud-native platforms such as Docker Kubernetes Prometheus or Grafana.
- Have 3 years Proven experience administering Elastic Observability platforms across the full lifecycle including deployment maintenance upgrades patching and capacity scaling.
Preferred qualifications
- AWS or Azure Associate-level certification or equivalent practical cloud operations experience.
- Elastic Certified Engineer or equivalent observability platform certification
- Familiarity with infrastructure as code (GitHub Actions CloudFormation Terraform Ansible) for repeatable automation
- Exposure to cloud-native observability frameworks (Open Telemetry service meshes)
- Experience documenting runbooks playbooks and consumption guides for SMEs
- Process knowledge such as Agile and/or ITIL
Must have Technical Skills
- Strong background in observability platforms ( stack preferred: Elasticsearch Kibana Logstash Beats Elastic APM and Fleet/Elastic Agent)
- Telemetry Fundamentals:Strong practical understanding of fundamental observability concepts including the collection and analysis oflogs metrics traces and synthetic monitoring.
- Experience with administration of leading observability platforms (Grafana Graylog Splunk Sumo Logic Tanzu or open-source equivalents) including lifecycle management (Kubernetes Docker Prometheus and Grafana installation patching upgrades scaling).
- Strong knowledge of distributed infrastructure domains (network servers VMs AWS Azure databases) from an observability perspective.
- Proven ability to design and tune scalable ingestion pipelines for diverse globally distributed data sources.
- Flexibility to adapt and evolve observability standards per domain needs while ensuring consistency across the enterprise.
- Operational ownership mindset accountable for uptime reliability and lifecycle management of the hosted observability platform.
- Incident management and SRE practices: monitoring alerting troubleshooting root cause analysis and postmortems.
- Proficiency in automation and scripting (Python Bash PowerShell Ruby etc.) for ingestion upgrades and operational tasks.
- Familiarity with REST APIs and tools like Postman plus DevOps constructs (GitHub Jenkins CI/CD pipelines serverless technologies).
- Configuration Proficiency:Demonstrated proficiency in managing system configurations usingYAML-based configurations.
- Knowledge of infrastructure as code (Terraform Ansible) for repeatable automation.
- Strong understanding of AWS/Azure services relevant to ingestion and enrichment (e.g. Kinesis Event Hub Lambda Functions).
- Data Presentation:Extensive experience inperformance optimization advanceddata visualization and sophisticateddashboardingusing Kibana (or similar platforms).
- Ability to document and enable SMEs with clear runbooks onboarding guides and consumption standards.
- Collaboration skills to guide Infra domain SMEs on onboarding data sources and consuming observability outputs.
Good-to-Have
- Exposure to cloud-native observability frameworks (Open Telemetry service meshes).
- Basic understanding of shared infrastructure services (DNS DHCP Active Directory SSL load balancing).
- Experience with hybrid deployments (on-prem cloud) and handling data sovereignty/regulatory considerations in global industries.
- Security monitoring and compliance awareness relevant to IT domains.
- Contribution to observability best practices (dashboards alerting strategies operational standards).
- Mentoring and cross-training skills to expand in-house observability expertise.
- Process knowledge like Agile and/or ITIL.
- AWS Professional-level certification or the ability to demonstrate equivalent functional knowledge and expertise.
Desired Characteristics
- Strong analytical and troubleshooting skills.
- Clear written and verbal communication skills.
- Ability to collaborate with application infrastructure cloud and operations teams.
- Self-motivated with a passion for learning new technologies.
- Customer-focused mindset and commitment towards operational excellence.
- Ability to work independently and as part of a globally distributed team.
- Proactive approach to maintaining documentation fostering knowledge transfer and supporting continuous service improvement.
Work in a way that works for you
We recognize that everyone is different and that the way in which people want to work and deliver at their best is different for everyone this role we can offer the following flexible working patterns:
- Working flexible hours - flexing the times when you work in the day to help you fit everything in and work when you are the most productive
Working with us
Our people are at the heart of what we do at Baker Hughes. We know we are better when all of our people are developed engaged and able to bring their whole authentic selves to work. We invest in the health and well-being of our workforce train and reward talent and develop leaders at all levels to bring out the best in each other.
Working for you
Our inventions have revolutionized energy for over a century. But to keep going forward tomorrow we know we have to push the boundaries today. We prioritize rewarding those who embrace change with a package that reflects how much we value their input. Join us and you can expect:
- Contemporary work-life balance policies and wellbeing activities
- Comprehensive private medical care options
- Safety net of life insurance and disability programs
- Tailored financial programs
- Additional elected or voluntary benefits
Required Experience:
Senior IC
About Company
Baker Hughes (NYSE: BKR) is an energy technology company that provides solutions for energy and industrial customers worldwide. Built on a century of experience and with operations in over 120 countries, our innovative technologies and services are taking energy forward – making it sa ... View more