Senior MLOps ML Platform Engineer
Job Summary
- Build and maintain ML training orchestration pipelines across hourly daily and weekly schedules
- Implement retries backfills and idempotent execution mechanisms
- Design and support model registry workflows including versioning lineage evaluation gates and promotion processes
- Develop isolated per-advertiser model environments with namespace and configuration separation
- Build scalable refresh pipelines and publishing workflows for serving infrastructure
- Implement shadow mode and champion/challenger deployment strategies
- Develop monitoring and alerting for ML-specific metrics including feature drift prediction drift train/serve skew and calibration decay
- Ensure reproducibility of ML workflows using containerized environments pinned dependencies and data snapshots
- Monitor training and scoring costs across tenants
- Collaborate with DevOps and SRE engineers on CI/CD and infrastructure automation
- Prepare operational documentation and platform handover materials
Qualifications :
- 5 years of experience in MLOps ML platform engineering or infrastructure engineering supporting production ML systems
- Strong Python skills and experience building platform-level tooling and automation
- Hands-on experience with Kubernetes and Docker
- Experience building CI/CD pipelines for ML workloads
- Hands-on production experience with MLflow Kubeflow Airflow Argo Workflows Vertex Pipelines or similar orchestration and ML lifecycle platforms
- Experience with ML platforms and model lifecycle tools such as Vertex AI MLflow or Kubeflow
- Strong understanding of ML observability including drift detection train/serve skew monitoring and incident response
- Experience designing or supporting multi-tenant ML systems and isolated model environments
- Experience working with cloud platforms preferably GCP
- Experience with infrastructure-as-code tools such as Terraform
- Experience with Linux environments
- Understanding of the ML lifecycle and productionization processes
- Upper-Intermediate English level or higher
WILL BE A PLUS
- Experience with feature stores and feature consistency management
- Experience with large-scale batch scoring systems operating under freshness SLAs
- Familiarity with experiment tracking platforms and evaluation gates
- Experience with on-premises Kubernetes or bare-metal Linux infrastructure
- Knowledge of DVC lakeFS or other data versioning tools
- Experience with Bigtable Redis Aerospike or similar low-latency serving databases
- GPU scheduling and training cost optimization experience
- Familiarity with SOC 2 ISO 27001 or GDPR-related compliance requirements
Additional Information :
PERSONAL PROFILE
- Strong ownership mindset and focus on operational reliability
- Ability to work independently in complex distributed systems environments
- Strong collaboration and communication skills
- Analytical thinking with attention to scalability and maintainability
- Comfortable working in fast-paced product-oriented environments
Remote Work :
Yes
Employment Type :
Full-time
About Company
At Sigma Software, we are involved with the clients team to contribute to the design and development of a technical solution for their tokenized domain reservation platform. We started by assigning a software architect to design the smart contracts and integrate blockchain into the s ... View more