Senior DevOps Engineer
Job Summary
Senior DevOps / Cloud Platform Engineer ML & AI Infrastructure
Job Summary
We are looking for a Senior DevOps / Cloud Platform Engineer with strong experience in AWS Kubernetes CI/CD infrastructure automation and ML/AI infrastructure to design deploy manage and optimize cloud infrastructure supporting our machine learning and AI services.
The ideal candidate will have hands-on experience with AWS EKS SageMaker Bedrock Docker Kubernetes Terraform Helm GitHub Actions Databricks Elasticsearch and self-hosted LLM deployments. This role will work closely with Data Engineering Machine Learning and Software Engineering teams to build reliable scalable secure and cost-efficient platforms for ML services across development and production environments.
Key Responsibilities
AWS & Kubernetes Infrastructure
- Design deploy and administer AWS infrastructure supporting ML and AI workloads.
- Manage Amazon EKS clusters including cluster provisioning upgrades scaling networking and troubleshooting.
- Work with AWS SageMaker AWS Bedrock EKS ECS and related AWS services.
- Configure and manage Kubernetes Ingress controllers such as NGINX and AWS ALB.
- Manage Cloudflare Tunnels DNS Cloudflare configuration networking and security.
- Troubleshoot application networking compute and infrastructure issues across AWS and Kubernetes environments.
- Implement best practices for security reliability availability and scalability.
CI/CD & Azure-to-AWS Migration
- Build and maintain CI/CD pipelines for ML and AI services.
- Develop and manage GitHub Actions and Azure DevOps pipelines using YAML.
- Migrate repositories and CI/CD workflows from Azure DevOps to GitHub/AWS.
- Automate build test containerization deployment and release processes.
- Establish deployment strategies across development staging and production environments.
ML Service Deployment
- Deploy and manage ML services across AWS EKS/ECS and SageMaker.
- Build and maintain Docker containers and Kubernetes deployments.
- Manage environment segregation and configuration across Dev QA and Production.
- Develop and maintain Kubernetes manifests and Helm charts.
- Troubleshoot ML service deployment networking scaling and runtime issues.
App Runner to EKS Migration
- Lead migration of existing services from AWS App Runner to Amazon EKS.
- Containerize applications and develop Kubernetes manifests/Helm charts.
- Design appropriate Kubernetes architecture networking ingress scaling and deployment strategies.
- Ensure minimal service disruption during migration and establish operational best practices on EKS.
Self-Hosted LLM & AI Infrastructure
- Deploy and manage self-hosted Large Language Models and inference services.
- Work with model serving frameworks such as vLLM.
- Design containerized infrastructure for GPU-based model serving and inference.
- Manage model versions deployments configurations and rollback strategies.
- Support migration of ML services from managed APIs/services to self-hosted models.
- Work with engineering teams on API integration and inference infrastructure.
Databricks Administration
- Administer Databricks workspaces clusters permissions and access controls.
- Manage cluster configuration policies and resource utilization.
- Support LMI Insights and related ML/AI workloads.
- Troubleshoot Databricks infrastructure and connectivity issues.
- Implement appropriate security and access-control practices.
Elasticsearch Infrastructure
- Design deploy and manage Elasticsearch clusters.
- Perform cluster sizing scaling configuration and performance optimization.
- Manage indices mappings retention and data lifecycle requirements.
- Support Kibana configuration dashboards and troubleshooting.
- Monitor Elasticsearch health capacity and performance.
Monitoring Reliability & Auto-Scaling
- Implement monitoring and observability for Kubernetes AWS and ML services.
- Use Prometheus Grafana and AWS CloudWatch for monitoring and alerting.
- Configure Kubernetes HPA/VPA and other auto-scaling mechanisms.
- Establish proactive alerting for infrastructure and application health.
- Perform capacity planning and resource optimization.
- Identify opportunities for AWS infrastructure and compute cost optimization.
Infrastructure as Code & Automation
- Build and maintain infrastructure using Terraform.
- Develop reusable Terraform modules for AWS and Kubernetes infrastructure.
- Manage Kubernetes deployments using Helm charts.
- Automate infrastructure provisioning configuration deployments and operational tasks.
- Maintain infrastructure documentation and deployment standards.
Cross-Team Collaboration
- Partner closely with Data Engineering ML Engineering Data Science and Software Engineering teams.
- Understand data pipelines SQL APIs and ML service architecture sufficiently to troubleshoot end-to-end workflows.
- Coordinate infrastructure requirements for new ML models and services.
- Participate in production incident resolution root-cause analysis and continuous improvement.
- Establish engineering standards around deployment monitoring security and operational ownership.
Required Skills & Experience
- 5 years of experience in DevOps Cloud Infrastructure SRE or Platform Engineering.
- Strong hands-on experience with AWS.
- Strong experience administering Amazon EKS and Kubernetes in production.
- Hands-on experience with:
- AWS EKS
- AWS SageMaker
- AWS Bedrock
- AWS ECS
- AWS App Runner
- Kubernetes
- Docker
- NGINX / AWS ALB Ingress
- Cloudflare / Cloudflare Tunnels
- Strong experience with Terraform and Helm.
- Strong experience developing CI/CD pipelines using GitHub Actions and/or Azure DevOps.
- Strong YAML scripting and Git experience.
- Experience migrating CI/CD pipelines and repositories from Azure to AWS/GitHub.
- Experience deploying and operating ML/AI services.
- Experience with self-hosted LLM/model serving preferably vLLM.
- Experience with GPU-based workloads is highly desirable.
- Experience with Databricks administration.
- Experience managing Elasticsearch and Kibana.
- Experience with Prometheus Grafana and CloudWatch.
- Strong understanding of Kubernetes HPA/VPA networking ingress DNS and service discovery.
- Strong understanding of cloud networking fundamentals.
- Experience with production troubleshooting monitoring capacity planning and cost optimization.
- Strong understanding of security IAM secrets management and access control.
Preferred / Nice-to-Have Skills
- Experience supporting Generative AI / LLM platforms.
- Experience with GPU infrastructure and NVIDIA/CUDA environments.
- Experience with model lifecycle and model version management.
- Experience migrating workloads between managed cloud services and Kubernetes.
- Experience with AWS networking such as VPC load balancers security groups and Route 53.
- Experience with API gateways and microservice architectures.
- Experience with Python or shell scripting for infrastructure automation.
- Experience working with Data Engineering and ML teams in a production environment.
What Youll Own
- AWS ML/AI infrastructure
- EKS cluster administration and upgrades
- ML service deployment and production operations
- CI/CD automation
- App Runner EKS migration
- Self-hosted LLM infrastructure and vLLM
- Databricks platform administration
- Elasticsearch infrastructure
- Monitoring and auto-scaling
- Terraform and Helm-based infrastructure automation
- Cloud cost reliability and performance optimization
Ideal Candidate
The ideal candidate is a hands-on infrastructure engineer who can independently take an ML/AI service from containerization CI/CD AWS infrastructure EKS deployment monitoring scaling production support.
They should be comfortable working across both traditional DevOps infrastructure and modern AI/ML infrastructure and should be able to collaborate closely with Data Engineering and ML teams while taking ownership of the underlying platform.
#LI-Onsite
Our Culture:
Spark Greatness. Shatter Boundaries. Share Success. Are you ready Because here right now is where the future of work is happening. Where curious disruptors and change innovators like you are helping communities and customers enable everyone anywhere to learn grow and advance. To be better tomorrow than they are today.
Who We Are:
At Cornerstone we believe in AI that works in the service of people amplifying their judgment to drive high-performing future-ready organizations forward. Cornerstone Workforce AI the intelligence platform for workforce readiness brings together workforce and labor market data into a proprietary Cornerstone People Graph translating signals into intelligence targeting learning where it matters developing critical skills and surfacing hidden talent. Delivered as an open enterprise platform across whatever application your people work in every day Cornerstone Workforce AI is built for scale security and trust with certified AI guardrails. As an industry leader Cornerstone is helping approximately 7000 organizations 140M users across 186 countries build continuous workforce readiness.
Check us out on LinkedIn Comparably Glassdoor and Facebook!
Required Experience:
Senior IC
About Company
Cornerstone and our intelligence platform for workforce readiness deliver the insights leaders want, the skills and learning people need, and AI agents that make action easy.