Platform Engineer – MLOps
Santa Clara County, CA - USA
Job Summary
About the Role
We are looking for a Platform Engineer with an MLOps focus to build and operate the platform capabilities that enable data science teams to develop deploy and manage machine learning workloads reliably and at scale.
Youll work across ML lifecycle tooling CI/CD MLflow Databricks and Unity Catalog while also supporting the AWS infrastructure underpinning the Lakehouse. Some client engagements extend into hybrid or on-premises Kubernetes-based ML platforms (feature platforms such as Chalk or notebook environments such as Domino) - prior exposure to these is a plus not a requirement. The role combines platform engineering automation governance observability and collaboration with data science and data engineering teams.
Infrastructure as Code (Terraform/Terragrunt) and hands-on AWS experience are valuable but not core requirements. The fundamentals are Databricks MLOps and Python. Youll report into the Data & Platform Engineering team.
Key Responsibilities
MLOps & ML Lifecycle Enablement
- Build and maintain CI/CD pipelines for ML workloads testing packaging and promotion of models and pipelines (Databricks Asset Bundles Git-based workflows).
- Own ML lifecycle tooling: experiment tracking model registry versioning and promotion gates (MLflow or equivalent).
- Define development-to-production promotion procedures for models and pipelines including validation gates rollback and an audit trail of changes.
- Implement model deployment and serving patterns including batch inference and real-time endpoints using Databricks Model Serving or equivalent technologies with rollback options.
- Build and maintain observability for ML workloads covering pipeline health model/data drift performance latency and cost.
- Partner with data scientists to productionize notebooks and experiments into governed reliable and repeatable pipelines.
Databricks & Cloud Platform Support
- Administer and evolve Databricks workspaces Unity Catalog metastores and catalogs (e.g. bronze/silver/gold) with proper access controls.
- Support the AWS infrastructure underpinning the Lakehouse particularly IAM networking S3 and related platform services - working alongside platform engineers with deeper AWS expertise.
Security Reliability & Collaboration
- Apply least-privilege access and Unity Catalog governance across data and ML assets.
- Monitor and tune Databricks jobs and compute for cost and performance.
- Maintain lightweight architecture documentation and runbooks for deployment troubleshooting and production support.
- Partner with data engineering and data science teams to define and maintain platform standards automation and documentation.
- Translate ML workload requirements into reliable governed and reusable platform capabilities.
Requirements
- 5 years of relevant experience in platform DevOps MLOps ML engineering or related engineering roles.
- Hands-on experience with the ML lifecycle: experiment tracking model registry versioning and deployment (MLflow or equivalent).
- Experience building CI/CD pipelines (Git-based workflows) for ML or data workloads.
- Working experience with Databricks platform capabilities including workspaces Unity Catalog compute jobs/workflows permissions and environment configuration.
- Solid Python scripting for automation and pipeline tooling.
- Ability to work closely with data scientists and data engineers to translate ML requirements into reliable platform capabilities.
- Fluency in English (written and spoken).
Nice to Have
- Hands-on experience with AWS particularly IAM networking S3 and related platform services.
- Infrastructure as Code experience (Terraform Terragrunt).
- Familiarity with Databricks Asset Bundles Delta Live Tables or Databricks Workflows.
- Exposure to feature stores or feature platforms (e.g. Chalk) or real-time model serving.
- Experience with notebook-based data science environments such as Domino Data Lab.
- Experience with on-premises or hybrid Kubernetes environments and Git-based deployment workflows using GitLab.
- Databricks (ML Associate) or AWS certifications.
- Observability tooling (CloudWatch Prometheus/Grafana) and cost-optimization practices.
About Opplane
Opplane specializes in providing advanced data-focused solutions for financial services telecommunication and reg-tech to accelerate their digital transformation journey. Opplane leadership team is comprised of Silicon Valley serial entrepreneurs and experienced executives. Its expertise comes from years of specific industry experience at some of the worlds top companies such as PayPal Xerox Parc Amazon Wells Fargo SoFi in the areas of product management data technology data governance data privacy security machine learning and risk management.
Why Opplane
Global & Multicultural Diverse perspectives global collaboration (US Portugal India and Singapore offices)
Startup Energy Fast-moving impact-driven environment
Ownership Mindset Engineers own what they build
Collaborative & Friendly Open curious and supportive culture