AI Infrastructure & Platform Operations Engineer

Mirantis


Job Location:

Poznań - Poland

Monthly Salary: Not Disclosed
Posted on: 2 days ago
Vacancies: 1 Vacancy

Job Summary

We are building a European AI Infrastructure & Platform Operations team responsible for operating large-scale AI infrastructure environments powered by NVIDIA GPUs high-performance networking Kubernetes and next-generation platform technologies.

The team is responsible for ensuring the availability performance and operational stability of critical AI infrastructure platforms deployed across multiple datacenters. Working at the intersection of infrastructure networking and platform operations you will help support the environments that power modern AI workloads.

This is an opportunity to work with some of the latest technologies in AI infrastructure while contributing to the evolution of AI-powered operational services through platforms such as k0rdent AI.

Responsibilities:

  • Monitor operate and support production AI infrastructure platforms.
  • Investigate and resolve infrastructure networking hardware and platform-related incidents.
  • Support NVIDIA GPU infrastructure and associated platform services.
  • Monitor and troubleshoot Kubernetes-based environments.
  • Investigate performance availability and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams hardware vendors datacenter personnel and service delivery teams to resolve technical issues.
  • Participate in incident response root cause analysis and operational improvement activities.
  • Contribute to improvements in monitoring observability automation and operational processes.
  • Maintain operational documentation runbooks and knowledge articles.

Qualifications :

Required Skills & Experience:

  • 3 years of experience in infrastructure operations platform operations network operations site reliability engineering cloud operations datacenter operations or related technical roles.
  • Strong Linux administration and troubleshooting skills.
  • Good understanding of networking concepts and experience diagnosing infrastructure-related issues.
  • Working knowledge of Kubernetes in production environments.
  • Experience supporting production infrastructure and services.
  • Strong analytical and problem-solving skills.
  • Experience working within structured operational and incident management processes.
  • Excellent communication and collaboration skills.
  • Ability to work within a shift-based operational environment.

Preferred Qualifications:

Experience in one or more of the following areas is highly desirable:

  • NVIDIA GPU infrastructure and accelerated computing platforms.
  • InfiniBand networking and NVIDIA UFM.
  • Kubernetes platform operations.
  • AI infrastructure or HPC environments.
  • Site Reliability Engineering (SRE) or Platform Engineering.
  • Observability platforms such as Grafana Prometheus ELK or OpenTelemetry.
  • Infrastructure automation technologies and Infrastructure-as-Code practices.
  • Large-scale distributed systems and production platforms.

Additional Information :

What does Mirantis offer you

  • Work with some of the most advanced AI infrastructure environments in production today.
  • Gain exposure to NVIDIA GPU technologies Kubernetes platforms and high-performance networking environments.
  • Help define how next-generation AI infrastructure is operated and supported.
  • Be part of a team shaping the future of AI-powered operations through k0rdent AI.
  • Join a growing organisation investing heavily in AI infrastructure and platform services.

It is understood that Mirantis Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to 

By submitting your resume you consent to the processing and storage of your personal data in accordance with applicable data protection laws for the purposes of considering your application for current and future job opportunities.

#LI-REMOTE

We are a Leader for Container Management in G2 (#2 after AWS)!


Remote Work :

Yes


Employment Type :

Full-time

We are building a European AI Infrastructure & Platform Operations team responsible for operating large-scale AI infrastructure environments powered by NVIDIA GPUs high-performance networking Kubernetes and next-generation platform technologies.The team is responsible for ensuring the availability p...

About Company

Mirantis is an open cloud company that helps organizations achieve digital self determination by giving them complete control over their strategic infrastructure. The company combines intelligent automation and cloud-native expertise for managing and operating virtual machines, contai ... View more

View Profile View Profile