Software Engineer SRE (Site Reliability Engineer)

NetApp


Job Location:

Bengaluru - India

Monthly Salary: Not Disclosed
Posted on: 16 hours ago
Vacancies: 1 Vacancy

Job Summary

Job Summary

Were looking for a Site Reliability Engineer focused on building and operating the data and AI/ML infrastructure platform that powers NetApps cloud-native data services. Youll work at the intersection of software engineering and infrastructure operations designing systems for reliability driving automation and ensuring our platforms meet the highest availability standards for customers worldwide.

This is an infrastructure-focused SRE role youll own the reliability of large-scale Kubernetes clusters (including GPU workloads) streaming data pipelines (Kafka) and analytical compute infrastructure (Spark Dremio) across hybrid-cloud and multi-cloud environments.

Job Requirements

  • 5 years in SRE DevOps Platform Engineering or Infrastructure Engineering roles
  • Extensive experience with Linux (RHEL/CentOS) including shells filesystems kernel tuning networking and performance optimization
  • Deep expertise with Kubernetes at scale including cluster administration troubleshooting networking storage RBAC and lifecycle management (on-premises and Rancher Kubernetes)
  • Hands-on experience operating GPU workloads on Kubernetes including NVIDIA GPU Operator device plugins scheduling and resource management
  • Strong experience managing Confluent Kafka in production including operations monitoring performance tuning and disaster recovery
  • Experience operating Apache Spark and/or Dremio including cluster management job scheduling scaling and performance optimization
  • Proficiency in Infrastructure as Code using Terraform Helm and GitOps workflows with ArgoCD/FluxCD
  • Proficiency in scripting and automation using Shell Ansible and Python with a strong automation-first mindset
  • Experience with scheduling and orchestration tools such as cron jobs and Apache Airflow
  • Deep familiarity with monitoring and observability tools including Dynatrace Grafana and Prometheus
  • Solid understanding of SQL and NoSQL databases including operations backup and monitoring
  • Experience designing and maintaining CI/CD pipelines and release processes
  • Expertise in AWS cloud platforms and hybrid-cloud integration
  • Strong systems thinking with an understanding of how infrastructure design choices impact failure modes scalability and recovery
  • Strong incident management skills and post-mortem facilitation experience
  • Excellent written communication skills for design documents runbooks post-mortems and operational documentation

Nice to Have

  • Knowledge of Generative AI tools and frameworks including the application of AI-based predictive analytics and automation in infrastructure operations
  • Familiarity with ML platforms such as Kubeflow MLflow and Ray as well as AI/ML training infrastructure
  • Experience with Kafka Streams ksqlDB or Apache Flink

Education

  • 5-8 years of relevant experience.
  • Bachelor of Science Degree in Computer Science Electrical Engineering or a related field; a Masters Degree is preferred.

Required Experience:

IC

Job Summary Were looking for a Site Reliability Engineer focused on building and operating the data and AI/ML infrastructure platform that powers NetApps cloud-native data services. Youll work at the intersection of software engineering and infrastructure operations designing systems for reliabilit...

About Company

Company Logo

At NetApp, our top priority is the health and safety of our event attendees and employees, including every community around the world being impacted by COVID-19. As a result, we have decided to reimagine our annual NetApp INSIGHT Paris and Berlin events to be fully digital. We’re als ... View more

View Profile View Profile