Senior Engineer – ML

Albertsons


Job Location:

Pleasanton, CA - USA

Monthly Salary: Not Disclosed
Posted on: 8 hours ago
Vacancies: 1 Vacancy

Job Summary

Description

Why choose us

Are you ready to take the next step in your career Join us for an exciting opportunity at Albertsons Companies where innovation and customer service go hand-in-hand!

At Albertsons Companies we are looking for someone whos not just seeking a job but someone who wants to make an this role youll have the opportunity to lead innovate and contribute to the growth of a company that values great service and lasting customer relationships. This position offers the chance to work in a fast-paced dynamic environment thats constantly evolving.

This role is an individual contributor position responsible for designing developing fine-tuning and operationalizing AI/ML capabilities for the AIOps platform within the Observability product. The candidate will work closely with the Lead Engineer SRE teams Platform team Data Ingestion team Platform DevOps team Visualization team and other portfolio teams.

As part of the AIOps Platform team you will design and build intelligent systems that improve observability incident response and operational efficiency. This includes developing machine learning models for forecasting anomaly prediction alert classification event intelligence and causal analysis as well as building AI agents and multi-agent workflows for RCA summarization investigative assistance and SRE productivity use cases.

This position will be based out of Phoenix Arizona or Pleasanton CA.

Main responsibilities:

  • Design develop and productionize AI/ML capabilities for the AIOps platform to support intelligent observability and operational decision-making.
  • Build machine learning systems for time-series forecasting anomaly prediction incident prediction alert classification noise reduction and event correlation.
  • Develop causal ML and statistical inference solutions to identify likely root causes dependency impacts and relationships across systems and services.
  • Create and fine-tune models for incident intelligence use cases such as forecasting service degradation capacity risk prediction alert prioritization and anomaly explanation.
  • Design feature pipelines and model training workflows using telemetry log metric trace topology and incident data.
  • Build intelligent RCA summarization capabilities using LLMs and agentic frameworks such as LangChain and LangGraph.
  • Develop AI agents and multi-agent systems for use cases such as Multi-Agent RCA SRE Assistant remediation guidance incident triage and operational knowledge retrieval.
  • Design prompt orchestration reasoning workflows retrieval pipelines tool usage patterns and memory/context handling for AI agents.
  • Integrate AI/ML services with observability platforms event systems knowledge bases CMDB incident management tools and automation platforms.
  • Collaborate with platform and engineering teams to build scalable model-serving and agent-serving architectures.
  • Define and implement evaluation frameworks for model quality agent effectiveness hallucination reduction relevance and operational usefulness.
  • Ensure AI/ML systems are scalable reliable explainable and aligned with enterprise security governance and responsible AI practices.
  • Build and maintain APIs and microservices for model inference online scoring batch predictions and agent orchestration.
  • Partner with SREs observability engineers and product stakeholders to translate operational pain points into ML and AI-driven solutions.
  • Continuously improve model performance feature quality inference latency agent reliability and business impact through experimentation and monitoring.
  • Establish engineering best practices for ML development prompt engineering evaluation model deployment testing versioning and documentation.
  • Support production incident analysis for AI/ML services and drive root cause identification and remediation for model or agent failures.
  • Create technical documentation covering model design feature logic training pipelines evaluation metrics deployment architecture and agent workflows.
  • Drive innovation in AI-enabled observability causal intelligence and agentic SRE workflows to enhance the value of the AIOps platform.

We are searching for someone with the following skills:

  • Strong experience designing and building AI/ML systems for real-world production use cases.
  • Solid hands-on experience with Python and common ML frameworks and libraries such as scikit-learn XGBoost PyTorch TensorFlow Pandas and NumPy.
  • Experience building machine learning solutions for forecasting anomaly detection prediction classification clustering ranking or recommendation problems.
  • Strong understanding of time-series modeling techniques for forecasting and operational prediction use with alert classification incident prediction event deduplication prioritization or
  • signal correlation in observability or IT operations contexts.
  • Knowledge of causal ML causal inference graph-based reasoning and dependency-aware analysis techniques for RCA and impact analysis.
  • Hands-on experience building AI applications using LLM frameworks such as LangChain and LangGraph.
  • Experience designing AI agents or multi-agent systems for reasoning summarization task orchestration troubleshooting or assistant workflows.
  • Strong understanding of prompt engineering RAG architecture embeddings vector stores tool calling memory handling and agent evaluation techniques.
  • Experience integrating LLM systems with enterprise tools APIs knowledge repositories and operational systems.
  • Experience building backend services and APIs for AI/ML model inference and agent orchestration.
  • Good understanding of observability data such as logs metrics traces topology incidents and alerts.
  • Experience with data engineering concepts including feature engineering data preprocessing model pipelines and batch or streaming inference.
  • Familiarity with graph databases such as Neo4j and their use in dependency mapping causal analysis and knowledge-driven AI systems.
  • Experience with REST APIs microservices architecture Docker Kubernetes and cloud-native deployment patterns.
  • Familiarity with CI/CD MLOps model lifecycle management experiment tracking and model versioning practices.
  • Knowledge of OpenTelemetry monitoring systems and observability platforms is highly desirable.
  • Strong understanding of software engineering fundamentals system design and scalable architecture patterns.
  • Strong analytical and problem-solving skills with the ability to convert ambiguous operational problems into measurable AI/ML solutions.
  • Excellent communication and collaboration skills to work with SREs platform engineers product owners and business stakeholders.
  • Self-driven mindset with strong curiosity innovation and the ability to learn and apply emerging AI techniques effectively.

We believe the successful candidate has these qualifications and experience:

  • Bachelors degree in computer science Information Systems Engineering Data Science Artificial Intelligence or a related field or equivalent practical experience.
  • 6 to 10 plus years of overall experience in software engineering machine learning or AI system development.
  • 3 plus years of hands-on experience building and deploying machine learning systems in production.
  • Strong experience in Python-based AI/ML development is required.
  • Experience working on observability monitoring or AIOps-related platforms is strongly preferred.
  • Experience building LLM-powered applications AI agents or multi-agent workflows for enterprise use cases is highly preferred.
  • Experience in AIOps Observability SRE IT operations or incident management domains.
  • Experience applying AI/ML to RCA anomaly explanation incident summarization service health prediction or remediation recommendations.
  • Familiarity with knowledge graphs and graph-based ML techniques for dependency-aware intelligence.
  • Experience using vector databases and retrieval frameworks for enterprise search and agentic applications.
  • Experience integrating AI services with tools such as ServiceNow Grafana Prometheus Splunk AppDynamics or similar platforms.
  • Familiarity with MCP-based client or agent integrations is a plus.

We also provide a variety of benefits including:

  • Competitive wages paid weekly
  • Access to up to 50% of your earned wages before payday via our partnership with Stream
  • Associate discounts
  • Health and financial well-being benefits for eligible associates (Medical Dental 401k and more!)
  • Time off (vacation holidays sick pay). For eligibility requirements please visit myACI Benefits
  • Leaders invested in your training career growth and development
  • An inclusive work environment with talented colleagues who reflect the communities we serve

Our Values Click below to view video: ACI Values

A copy of the full job description can be made available to you.

#LI-MF1




Required Experience:

Senior IC

DescriptionWhy choose usAre you ready to take the next step in your career Join us for an exciting opportunity at Albertsons Companies where innovation and customer service go hand-in-hand!At Albertsons Companies we are looking for someone whos not just seeking a job but someone who wants to make an...

About Company

Company Logo

Albertsons Companies is at the forefront of the revolution in retail. Committed to innovation and fostering a culture of belonging, our team is united with a unique purpose: to bring people together around the joys of food and to inspire well-being. We want talented individuals to b ... View more

View Profile View Profile