MLOps Engineer
Job Summary
This role requires the candidate to work on-site in Tokyo Japan.
Client Overview
Our client is a global technology solutions provider specializing in Artificial Intelligence (AI) Cloud Computing Data Engineering and Digital Transformation (DX). The organization partners with enterprise clients across multiple industries to design deploy and operate scalable AI platforms that support machine learning advanced analytics and business automation.
As enterprise adoption of AI continues to accelerate the company is expanding its MLOps Engineering team to build production-grade machine learning infrastructure automate AI deployment pipelines and optimize cloud-native AI platforms using modern DevOps and MLOps technologies.
Job Role
The MLOps Engineer will design implement and manage enterprise machine learning infrastructure that enables efficient model development deployment monitoring and lifecycle management. Working closely with Data Scientists AI Engineers and Cloud Engineers this role will build scalable MLOps pipelines automate machine learning workflows and ensure AI models are securely deployed and maintained in production environments.
This position is ideal for engineers who are passionate about cloud infrastructure DevOps Kubernetes and production-scale AI systems.
Key Responsibilities
- Design build and maintain cloud-based machine learning infrastructure using AWS Microsoft Azure or Google Cloud Platform (GCP).
- Develop and optimize containerized AI applications using Docker and Kubernetes (EKS AKS or GKE).
- Build and manage Infrastructure as Code (IaC) using Terraform or CloudFormation.
- Design and maintain CI/CD pipelines for machine learning applications using GitHub Actions or similar tools.
- Develop and automate machine learning workflows using Airflow Kubeflow MLflow Vertex AI Pipelines SageMaker Pipelines or equivalent platforms.
- Deploy machine learning models into production using scalable serving frameworks such as KServe or SageMaker Endpoints.
- Monitor infrastructure and model performance using Prometheus Grafana Datadog and other observability tools.
- Optimize cloud infrastructure performance resource utilization operational costs and automate model retraining processes.
Candidate Requirements
- Bachelors degree in Computer Science Software Engineering Information Technology Artificial Intelligence Data Engineering or a related discipline.
- Minimum 3 years of experience in Software Engineering DevOps Cloud Engineering MLOps or related technical roles.
- Strong programming experience using Python.
- Hands-on experience with at least one major cloud platform: AWS Microsoft Azure or Google Cloud Platform (GCP).
- Practical experience with Docker Kubernetes and container orchestration in production environments.
- Experience implementing Infrastructure as Code (IaC) using Terraform or CloudFormation.
- Experience designing and maintaining CI/CD pipelines using GitHub Actions or similar DevOps tools.
- Working knowledge of machine learning concepts model deployment and AI production environments.
- Familiarity with MLflow Kubeflow Airflow Vertex AI SageMaker Prometheus Grafana or Datadog is highly preferred.
- Experience with Linux system administration and cloud cost optimization is an advantage.
- Strong analytical thinking automation mindset and problem-solving skills.
- Business-level Japanese proficiency (JLPT N2 or above) is mandatory.
- Professional English communication skills are preferred.
- Candidates eligible for relocation and visa sponsorship are encouraged to apply.