Enter a job title or keyword

Lead AI Operations Engineer

Johnson & Johnson


Job Location:

milan - Italy

Yearly Salary: EUR 48600 - 77395
Posted: 19 July 2026 (30+ days ago)
Application Deadline: 2 January 2027
Vacancies: 1 Vacancy

Job Summary

At Johnson & Johnsonwe believe health is everything. Our strength in healthcare innovation empowers us to build aworld where complex diseases are prevented treated and curedwhere treatments are smarter and less invasive andsolutions are our expertise in Innovative Medicine and MedTech we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow and profoundly impact health for more at .

As guided by Our Credo Johnson & Johnson is responsible to our employees who work with us throughout the world. We provide an inclusive work environment where each person is considered as an individual. At Johnson & Johnson we respect the diversity and dignity of our employees and recognize their merit.

Job Function:

Technology Product & Platform Management

Job Sub Function:

Technical Product Management

Job Category:

Scientific/Technology

All Job Posting Locations:

Lisbon Portugal Madrid Spain Milano Italy

Job Description:

We are recruiting for a Lead AI Operations Engineer based in Milan - Italy ; Madrid ; Spain or Lisbon ; Portugal


The AI Operations Engineer is responsible for shaping designing implementing and continuously improving the enterprise capabilities required to operate AI applications and AI agents safely reliably transparently and cost-effectively at scale. The role combines hands-on AI platform engineering with Site Reliability Engineering DevSecOps LLMOps AgentOps FinOps security and compliance practices.

The AI Operations Engineer builds reusable operational capabilities across observability runtime controls cost management auditability incident response and production support. The role works closely with the Agent Factory AI Engineering Data Platforms Cloud Infrastructure Cybersecurity Privacy Risk Quality and Responsible AI stakeholders.

The role does not own the end-to-end lifecycle management of agents or AI products. Agent design development functional evaluation release content product evolution and retirement decisions remain with the Agent Factory and the relevant AI product teams. The AI Operations Engineer provides the shared operational platform telemetry controls and guardrails that enable those teams to run AI solutions in production.


Key Responsibilities

AI Observability & Production Reliability

  • Design and implement end-to-end observability for AI applications and agents including prompts responses model calls tool calls retrieval steps decision paths latency failures token consumption and session context.
  • Establish common telemetry and distributed tracing across agent workflows APIs data services vector stores model endpoints and external tools.
  • Build operational dashboards and alerts covering availability latency errors reliability quality signals policy violations consumption and service health.
  • Define service-level indicators service-level objectives error budgets alert thresholds and operational readiness criteria for production AI services.
  • Enable trace-based debugging incident reconstruction and controlled session replay while protecting confidential or sensitive information in logs.
  • Monitor retrieval quality data freshness model and prompt regressions anomalous agent loops degraded tool performance and unexpected runtime behavior.
  • Lead technical root-cause analysis for AI platform and runtime incidents and convert findings into preventive controls automation and engineering improvements.

LLMOps & AgentOps Platform Enablement

  • Engineer reusable pipelines templates and controls for configuration prompt model and agent-component versioning across environments.
  • Implement automated technical gates for deployment readiness including integration tests regression checks operational validation security checks and observability coverage.
  • Enable controlled rollout patterns such as canary releases feature flags model or provider routing fallback strategies and technical rollback mechanisms.
  • Provide common operational tooling that supports multiple models frameworks clouds and agent patterns without creating a separate operating process for each solution.
  • Integrate functional evaluation signals supplied by the Agent Factory or AI product teams into deployment gates and runtime monitoring while functional quality ownership remains with those teams.
  • Maintain reusable runbooks reference implementations engineering standards and paved-road patterns for production operation and support.

AI FinOps & Consumption Efficiency

  • Create transparent metering allocation and showback capabilities by product agent workflow model environment and business unit where the required identifiers are available.
  • Monitor token consumption model utilization repeated or runaway loops retrieval overhead infrastructure usage and cost per successful transaction or workflow.
  • Implement budgets thresholds anomaly alerts and runtime guardrails to detect and contain unexpected consumption.
  • Partner with Agent Factory and product teams to optimize model selection routing context size caching batching retries and tool usage while preserving agreed quality and compliance requirements.
  • Define operational unit economics and provide evidence for capacity planning optimization priorities and platform investment decisions.

Security Governance & Compliance Engineering

  • Embed security privacy Responsible AI and compliance controls into the shared AI runtime and operational toolchain in line with enterprise policies and approved risk frameworks.
  • Implement identity role-based access control least privilege managed identities secrets management and segregation of duties for agents tools services and operators.
  • Engineer runtime controls for prompt injection jailbreak attempts unauthorized tool use excessive permissions data leakage unsafe execution paths and anomalous access patterns.
  • Design privacy-aware logging retention redaction and access patterns for prompts responses memory traces and audit evidence.
  • Provide auditable records linking versions configurations identities actions approvals policy decisions and operational outcomes.
  • Automate policy checks and evidence collection where feasible partnering with Cybersecurity Privacy Quality Legal Risk and Responsible AI stakeholders for control definition and approval.
  • Support threat modelling security testing incident response remediation and continuous control improvement for the AI platform.

Operational Service Management & Enablement

  • Define operating processes for monitoring support incident management problem management change management escalation and service recovery.
  • Create production-readiness checklists service acceptance criteria on-call runbooks escalation paths recovery procedures and business-continuity requirements.
  • Clarify operational handoffs and accountability across AI Operations the Agent Factory AI product teams and underlying platform owners.
  • Drive automation that reduces manual reconstruction of agent behavior and shortens detection diagnosis containment and recovery times.
  • Produce clear technical documentation and enable engineering teams to adopt approved operational patterns through practical coaching and reusable examples.

Stakeholder Management

  • Act as the technical bridge between the Agent Factory AI product teams platform engineering cloud infrastructure data platforms architecture cybersecurity privacy compliance and service management functions.
  • Translate operational risk and control requirements into implementable technical capabilities and platform standards.
  • Influence engineering teams to adopt common observability reliability cost security and compliance patterns.

Governance Risk & Compliance

  • Ensure shared AI Operations capabilities support applicable enterprise requirements for security privacy Responsible AI auditability and regulatory compliance.
  • Maintain traceability of operational controls exceptions evidence ownership and remediation actions.

Metrics & Continuous Improvement

  • Define and monitor KPIs for service reliability observability coverage incident performance cost efficiency control coverage audit evidence and platform adoption.
  • Use operational evidence to prioritize automation reliability improvements cost optimization and risk reduction.
  • Measure the proportion of production AI services covered by approved telemetry alerts runbooks cost controls and security guardrails.

Qualifications

Required

  • Bachelors or Masters degree in Computer Science Engineering Artificial Intelligence or a related discipline or equivalent relevant experience.
  • 5 years of hands-on experience in Platform Engineering Site Reliability Engineering DevOps DevSecOps MLOps or production software engineering including responsibility for live services.
  • Practical experience operating Machine Learning Generative AI or agentic systems in production beyond proofs of concept or prompt engineering.
  • Hands-on capability in Python APIs infrastructure as code and CI/CD practices using technologies such as Terraform GitHub Actions or Azure DevOps.
  • Experience with cloud-native platforms containers orchestration and enterprise cloud services with Azure experience strongly valued.
  • Experience implementing observability using logs metrics distributed traces dashboards and alerts ideally using OpenTelemetry or equivalent standards.
  • Strong understanding of LLM application patterns including RAG vector stores model gateways tool calling agent memory and multi-agent orchestration.
  • Working knowledge of identity and access management secrets management secure logging threat modelling privacy by design and audit controls.
  • Ability to turn ambiguous operational and control requirements into reusable technical capabilities standards automation and runbooks.
  • Strong stakeholder management technical communication and documentation skills.
Preferred
  • Experience working in regulated industries with formal security privacy quality validation risk or audit requirements.
  • Hands-on experience with Azure OpenAI Azure AI Foundry Azure Monitor Application Insights Microsoft Fabric or comparable cloud services.
  • Experience with AI observability or evaluation platforms such as Langfuse LangSmith Arize MLflow or equivalent technologies.
  • Experience operating heterogeneous model providers multi-cloud services or multiple agent frameworks.
  • Knowledge of FinOps practices cloud cost allocation budgeting and unit-economics measurement for AI workloads.
  • Experience defining service-level objectives on-call models incident management and operational-readiness gates for enterprise platforms.

Required Skills:

Preferred Skills:

Agile Product Development Analytical Reasoning Coaching Collaboration Competitive Landscape Analysis Critical Thinking Customer Alignment Demand Forecasting Human-Computer Interaction (HCI) Organizing Product Development Product Improvements Product Strategies Requirements Analysis Research and Development Software Development Life Cycle (SDLC) Software Development Management Stakeholder Management Technical Credibility Technical Writing Technologically Savvy

The anticipated base pay range for this position is:

48.60000 - 77.39500

Benefits:

In addition to base pay we offer the following benefits*: an annual bonus with set target (% of pay) depending on pay grade / location where the actual amount is based on the employees and companies performance of the previous calendar year or sales commissions. Moreover we offer vacation days parental leave for a minimum of 12 weeks bereavement leave caregiver leave volunteer leave well-being reimbursement programs for financial physical and mental health. We also offer service anniversary and recognition awards and subject to the terms of their respective plans employees - and in some locations eligible dependents - can participate in several insurance plans. For more information visit Employee benefits Supporting well-being & career growth Johnson & Johnson Careers.

*This is for informative purposes only. Amounts and actual benefits may vary by location and are subject to change.


Required Experience:

IC


About Company

Company Logo

About Johnson & Johnson A t Johnson & Johnson, we believe good health is the foundation of vibrant lives, thriving communities and forward progress. That’s why for more than 130 years, we have aimed to keep people well at every age and every stage of life. Today, as the world’s larges ... View more

View Profile View Profile