Senior MLOps Engineer I
San Francisco, CA - USA
Job Summary
About Us:
Zeitview is the leading intelligent aerial imaging company for high-value infrastructure providing businesses with actionable real-time insights to recover revenue reduce risk and improve build quality. We serve customers in the solar wind insurance construction real estate and critical infrastructure industries. Trusted by the largest enterprises in the world Zeitview is active in over 70 countries. Our mission is to accelerate the global transition to renewable energy and sustainable infrastructure through advanced inspection solutions. Take a look at our latest achievements here!
About the Role:
As the Senior MLOps Engineer I you will help turn the models built by our ML Scientists Data Scientists and Perception Engineers into reliable production-grade services. Youll work on the infrastructure pipelines and tooling that take a model or an LLM/agent-backed workflow from a research notebook to a fully monitored deployment running across multiple industry verticals including our model registry deployment pipelines and the cloud infrastructure our AI/ML platform depends on.
This role sits at the intersection of R&D Software Engineering and DevOps. You will work daily with our R&D team to understand what a model needs to run in production (compute data inputs versioning post-processing) and youll partner closely with the Platform and DevOps teams to provision the infrastructure permissions and deployment pathways that make it possible. Youll also contribute to broader automation initiatives helping provide the deployment visibility and pipeline reliability that let R&D Software Product and Ops teams move in lockstep.
The day-to-day will include maintaining and extending our model registry building and debugging deployment pipelines and cloud infrastructure and setting up model and pipeline monitoring and testing. You will also troubleshoot issues such as failed deployments permissions errors or inconsistent environments. Youll also help shape and document standards for how models move from staging to production. Perhaps most importantly you will serve as a key communicator ensuring R&D goals and challenges are well understood by Software Engineering and DevOps teams.
Responsibilities:
- Partner with Scientists: Work directly and iteratively with ML Scientists Data Scientists and Perception Engineers to translate experimental research-oriented code into dependable scalable production services without slowing down their research velocity.
- Cross-Functional Collaboration: Coordinate with DevOps and Software Engineering teams on infrastructure requests and shared data pipeline needs and support broader automation initiatives and team goals.
- Model Registry Deployment & Release Management: Maintain and improve model registry and deployment pipelines and help implement safer release practices (e.g. shadow deployments rollback procedures) to reduce risk.
- Cloud Infrastructure & CI/CD: Build maintain and troubleshoot cloud infrastructure and CI/CD pipelines that ML workloads run on working closely with Engineering and DevOps teams on shared tooling infrastructure-as-code and cost optimization for compute-heavy workloads.
- Monitoring Drift & Reproducibility: Implement monitoring and observability for models and pipelines in production help R&D track model performance and drift over time and support experiment tracking and dataset/model versioning.
- Ongoing Maintenance & Platform Support: Keep deployed ML systems healthy over time with dependency and infrastructure upgrades capacity and cost management data pipeline upkeep and retraining or redeployment support and extend support as needs evolve.
- Standards & Documentation: Help define and document conventions for model versioning deployment promotion and model documentation/lineage and build tools to allow scientists and engineers to self-serve.
Qualifications:
The following describes the qualifications for this position. Successful candidates are expected to meet most but not all of these requirements.
- Bachelors degree in Computer Science Software Engineering Data Engineering or a related field; typically 4 years of professional experience in MLOps ML platform engineering or infrastructure engineering supporting machine learning teams.
- Solid applied knowledge of MLOps practices with the ability to work independently across varied production scenarios and escalate only genuinely complex or ambiguous problems.
- Demonstrated experience working directly with researchers or ML scientists. You understand research workflows and can translate them into reliable services and productionized models without becoming a bottleneck. You serve as a key link communicating R&D goals and challenges to Software Engineering and DevOps teams.
- Strong Python skills and solid software engineering fundamentals (testing code review version control)
- Hands-on experience with a major cloud platform (e.g. AWS) infrastructure-as-code (Terraform) CI/CD tooling (Github Actions) and containerization/orchestration (e.g. Docker Kubernetes)
- Experience building and operating production ML pipelines and model registries including model versioning and safer release practices (e.g. canary deployments rollbacks) across environments as well as coordinating moderately complex cross-functional infrastructure or deployment projects.
- Experience building feedback loops from production back into training data capturing human corrections as labels and turning retraining into a repeatable pipeline. Familiarity with experiment tracking dataset/model versioning and model documentation practices that support reproducible auditable ML workflows is a plus.
- Familiarity with computer vision or geospatial ML pipelines
- Nice to have Experience operating LLM/Agentic systems in production evaluation harness prompt/tool/retrieval versioning tracing token cost optimization
- Nice to have Experience building data pipelines against relational databases (e.g. PostgreSQL) and API/GraphQL data layers (e.g. Hasura) and integrating external/third-party APIs into production workflows.
Whats Included:
- Feel great about your work as you join a leading mission-driven intelligent aerial imaging company - our goal is to accelerate the global transition to renewable energy and sustainable infrastructure and you personally will play a large part in making this happen!
- Base salary range of $170000 - $180000 USD
- Target annual bonus
- Eligibility for stock options
- Your choice of multiple medical insurance plans including options with an HSA and 100% coverage of the premium for yourself and your dependents
- 100% paid dental and vision insurance
- Unlimited PTO: We mean it when we say we prioritize work-life balance and mental health
- Autonomy and upward mobility
- Diverse equitable and inclusive culture: a place where your voice matters
This role has a base salary range of $170000 - $180000 USD plus a target annual bonus and eligibility for stock options. Actual compensation may vary based on experience skills and location within the USA.
Zeitview is proud to be an equal opportunity employer. At Zeitview we believe in cultivating an environment where our team members can bring their authentic whole selves to work. Encouraging identity and belonging is one of the many aspects of our culture that makes us stronger as an organization and drives innovation. We are committed to building and delivering a diverse inclusive and equitable workforce that includes age color sex disability national origin race religion or veteran status that is representative of the world around us where all individuals are treated with respect and dignity - and to act swiftly if this value is ever threatened. We are constantly striving to be better and we continue to take strategic steps to advance representation.
We also provide reasonable accommodation for qualified individuals with disabilities and for seriously held religious beliefs in accordance with applicable law.
Required Experience:
Senior IC
About Company
For global customers in energy and infrastructure, Zeitview builds advanced inspection software that delivers fast, accurate insights, lowers costs, and improves asset performance and longevity. We are second to none at partnering with our customers to achieve flexible, long-term solu ... View more