Enter a job title or keyword

Lead Site Reliability Engineer Expert (Palo Alto & Versa SDWAN Experience)

Sita


Job Location:

Delhi - India

Monthly Salary: Not provided by the employer
Posted: 27 August 2026 (24 days ago)
Application Deadline: 24 November 2026
Vacancies: 1 Vacancy

Job Summary

Description External

WELCOME TO SITA

At SITA we keep airports moving airlines flying smoothly and borders open. Our technology and communication innovations power the success of the global air travel industry.

Youll find us in 95% of international airports working closely with over 2500 transportation and government clients. Each partnership brings unique challenges and we thrive on delivering fresh solutions and cutting-edge tech to keep operations running like clockwork. We dont just move the world forward-were proud to be recognized as a Great Place to Work by 79% of our employees and certified in most of our growing locations. Here we feel empowered supported and inspired to grow.

Are you ready to love your job

The adventure begins right here with you at SITA.

ABOUT THE ROLE & TEAM

Responsible for ensuring highly reliable scalable and resilient production systems across cloud and onprem environments. Ensures high availability disaster recovery readiness and continuous improvement of service performance. Leads automation initiatives for provisioning deployment monitoring and selfhealing to reduce manual effort and improve stability. Owns the event catalog operational readiness and reliability engineering practices to prevent recurrence of incidents and strengthen system resilience. Drives collaboration across Product Engineering T&E ICE and Service Support Architects to ensure providergrade reliability and seamless operational integration of new releases.

WHAT YOULL DO

  • Reliability Engineering
    • Design & maintain resilient systems ensuring high availability scalability and fault tolerance.
    • Ensure effective Disaster Recovery (DR) failover strategies and resilience engineering across environments.
    • Improve platform reliability observability and performance across cloud and onpremises systems.
    • Establish and maintain SLIs SLOs and error budgets to measure and govern service reliability.
    • Take ownership of production availability capacity planning performance tuning and longterm reliability initiatives.

Automation DevOps & NetOps

    • Drive automation for infrastructure provisioning deployment monitoring and operational workflows.
    • Develop and implement autoremediation and selfhealing solutions to reduce manual intervention.
    • Manage CI/CD pipelines and Infrastructure as Code (IaC) frameworks for secure repeatable deployments.
    • Implement and manage zerodowntime deployment strategies (bluegreen canary rolling).
    • Support containerized and cloudnative platforms including Kubernetes Docker and distributed systems.
    • Support NetOps tooling and network observability ensuring visibility into network performance events and operational health.

Incident Problem & Event Management

    • Perform incident management production troubleshooting and lead RCA/PMIR (Postmortem) for critical outages.
    • Proactively identify reliability gaps performance bottlenecks and operational risks.
    • Optimize incident event and problem management processes to reduce MTTR and improve operational efficiency.
    • Define and maintain the event catalog thresholds and remediation workflows.
    • Develop event response protocols and ensure teams are trained for rapid incident handling.

Observability & Monitoring

    • Build and maintain observability solutions using monitoring logging tracing and alerting platforms.
    • Implement APM distributed tracing and proactive alerting to detect issues early.
    • Integrate network telemetry and NetOps monitoring tools into the overall observability stack.
    • Collaborate with stakeholders to improve event coverage and postevent learning.
    • Experience with AIassisted observability anomaly detection and predictive alerting.

Deployment & Operational Readiness

    • Own the quality of new release deployments for the PSO.
    • Conduct operational readiness assessments and manage deployment risk.
    • Ensure supportability for new applications platform releases and infrastructure changes.
    • Coordinate with internal/external stakeholders to drive continuous service improvement.

CrossFunctional Collaboration

    • Work closely with Development Platform Engineering Product T&E ICE and Service Support Architects to embed reliability best practices.
    • Collaborate with vendors and engineering teams to enhance system reliability and operational excellence.
    • Support new product productization as SGS technical expert and ensure operational readiness.
  • business value

Who you are :

Education and Professional Qualifications:

Bachelors degree in Computer Science Information Technology Engineering or a related field. Masters degree preferred for senior roles.
Relevant certifications such as ITIL CCNP/CCIE Palo Alto Security SASE SDWAN Juniper Mist/Aruba CompTIA Security or Certified Kubernetes Administrator (CKA).
Certifications in cloud platforms (AWS Azure Google Cloud) or DevOps methodologies.
Certifications in automation and IaC tools (Ansible Terraform).
Certifications in observability and monitoring platforms (Dynatrace Prometheus Grafana ELK).
Certifications in ServiceNow Jira or other operational tooling.

Experience:

8 years in IT operations service management or infrastructure reliability including roles such as Site Reliability Engineer Problem Manager or DevOps Engineer.
Strong experience with high availability systems resilience engineering and DR readiness.
Deep expertise in RCA incident management PMIR and implementing permanent fixes for recurring issues.
Hands on experience with CI/CD automation IaC and self healing/auto remediation workflows.
Proficiency in observability platforms (APM logging tracing alerting) and integrating network telemetry / NetOps monitoring.
Experience defining and governing SLIs SLOs and error budgets to improve service reliability.
Experience with Kubernetes containerized workloads and distributed systems.
Experience managing deployments operational readiness risk assessments and improving event/problem management processes.
Strong cross functional collaboration with Development Operations Engineering Product T&E ICE and SSA.
Familiarity with cloud platforms scalable architectures and zero downtime deployment strategies.

Technical Skills:

  • Cloud Infrastructure AWS/Azure Linux virtualization HA/DR architecture.
  • Automation & IaC Ansible Terraform CI/CD pipelines selfhealing workflows.
  • Observability & Monitoring APM logging tracing alerting Dynatrace Prometheus Grafana ELK.
  • NetOps Monitoring network telemetry event monitoring and operational visibility tools.
  • Containerization & Orchestration Docker Kubernetes distributed systems.
  • Deployment & Release Engineering zerodowntime strategies (bluegreen canary) operational readiness.
  • Programming & Scripting Python Bash PowerShell for automation and tooling.
  • Reliability Engineering SLIs/SLOs error budgets capacity planning performance tuning.

Qualifications External

WHAT WE OFFER

Were all about diversity. We operate in 200 countries and speak 60 different languages and cultures. Were really proud of our inclusive environment. Our offices are comfortable and fun places to work and we make sure you get to work from home too. Find out what its like to join our team and take a step closer to your best life ever.

Flex Week: Work from home up to 2 days/week (depending on your teams needs)

Flex Day: Make your workday suit your life and plans.

Flex-Location: Take up to 30 days a year to work from any location in the world.

Employee Wellbeing: We have got you covered with our Employee Assistance Program (EAP) for you and your dependents 24/7 365 days/year. We also offer Champion Health - a personalized platform that supports a range of wellbeing needs.

Professional Development: At SITA we believe growth fuels innovation. Our learning ecosystem offers access to world-class platforms and programs designed to help you thrive. From LinkedIn Learning Microsofts Enterprise Skills Initiative and Airport Council International -available to all employees-to specialized solutions like Pluralsight for technology upskilling Harvard Business Publishing for people leadership Stanford for strategic development and many others we align learning opportunities with your Development Plan and our business priorities. Your development journey is supported every step of the way.

Competitive Benefits: Competitive benefits that make sense with both your local market and employment status.

SITA is an Equal Opportunity Employer. We value a diverse support of our Employment Equity Program we encourage women aboriginal people members of visible minorities and/or persons with disabilities to apply and self-identify in the application process.


Required Experience:

IC


About Company

Company Logo

At SITA we lead one of the most exciting and advanced industries in the world. With us, there are no limits for people looking to explore the edges of possibility and beyond. We are the world’s leading specialist in air transport communications and information technology. Around the ... View more

View Profile View Profile