Site Reliability Engineer Expert Specialist (Must have strong experience in Windows Server, Azure and AKS)
Job Summary
At SITA we keep airports moving airlines flying smoothly and borders open. Our technology and communication innovations power the success of the global air travel industry.
Youll find us in 95% of international airports working closely with over 2500 transportation and government clients. Each partnership brings unique challenges and we thrive on delivering fresh solutions and cutting-edge tech to keep operations running like clockwork. We dont just move the world forward-were proud to be recognized as a Great Place to Work by 79% of our employees and certified in most of our growing locations. Here we feel empowered supported and inspired to grow.
Are you ready to love your job
The adventure begins right here with you at SITA.
PURPOSE
Ensure high product performance reliability and stability by proactively supporting products resolving root causes of incidents and implementing improvements that prevent recurrence. This role focuses on event management automation service deployment and operational integration to improve efficiency and collaboration across service operations.
What will you do
- Build and maintain reliable support systems to ensure high availability and strong product performance.
- Manage complex operational cases incidents and root cause analysis to deliver permanent fixes.
- Define and maintain event catalogs alerts thresholds and remediation actions.
- Implement automation for provisioning monitoring deployment self-healing and recovery.
- Collaborate with Product Engineering Service Architecture and Operations teams to improve service readiness availability and performance.
- Support customer success initiatives through reporting documentation communication materials and process improvements.
- Contribute to knowledge management resources such as FAQs training materials and operational guidance.
- Apply data governance standards monitor data quality and act as a subject matter expert for data-related queries.
Experience
Bachelors degree in computer science Information Technology Engineering or a related field.
5 years of experience in IT operations service management or infrastructure management including roles such as Site Reliability Engineer Problem Manager or DevOps Manager.
Proven experience managing high-availability systems and ensuring operational reliability with Azure/Windows Environments.
Extensive experience in root cause analysis (RCA) incident management and developing permanent solutions for recurring service disruptions.
Hands-on experience with CI/CD pipelines automation system performance monitoring and infrastructure as code (IaC).
Strong background in collaborating with cross-functional teams (Development Operations Engineering etc.) to improve operational processes and service delivery.
Experience managing deployments conducting risk assessments and optimizing event and problem management processes.
Familiarity with cloud technologies containerization and scalable architectures including zero-downtime deployment strategies.
Development Experience
Operating Systems:
Strong hands on expertise with Windows Servers (AD GPO DNS DHCP).
Solid Problem Management and troubleshooting skills.
Working knowledge of Unix/Linux (Red Hat).
PowerShell Scripting.
Cloud:
Strong Experience with Azure and AWS.
Strong Knowledge and skills in AKS and on prem Kubernetes.
Automation experience including CI/CD pipelines and exposure to Terraform.
Virtualisation:
Hands on experience with VMware (VCF VCD).
Storage:
Experience managing SAN and NAS.
Backup:
Background with Veeam and Veritas backup solutions.
Networking:
CCNA level networking knowledge.
Databases:
Skills in SQL and MongoDB for restore operations and performance tuning.
Observability & Monitoring
Strong experience with enterprise monitoring and observability platforms.
Hands-on experience with Prometheus Grafana Dynatrace Azure Monitor and AWS CloudWatch.
Expertise in infrastructure application and service monitoring across on-premises and cloud environments.
Experience designing proactive monitoring alerting and capacity management solutions to improve reliability and reduce incidents.
Knowledge of log management metrics dashboards distributed tracing and root cause analysis.
Experience defining and monitoring SLIs SLOs and SLAs to support Site Reliability Engineering (SRE) practices.
Strong focus on operational excellence platform stability performance optimisation and incident reduction through observability-driven insights.
Core Competencies: (Problem Management)
Conduct thorough problem investigations and root cause analyses to diagnose recurring incidents and service disruptions.
Coordinate with Incident Management teams and collaborate with PSOs and Engineering/Product teams to implement permanent solutions.
Monitor the effectiveness of problem resolution activities and provide regular reporting to ensure continuous improvement.
Qualifications
Bachelors degree in computer science Information Technology or related field.
Minimum 3-5 years of experience in platform and network administration with L2/L3 support on Windows Servers/Azure .
Strong problem-solving and analytical skills.
Excellent communication and collaboration abilities.
Relevant certifications: VMware AWS/Azure MCSE RHSA certifications are highly recommended.
CompTIA Security or Certified Kubernetes Administrator (AKS OR CKA). Certifications in cloud platforms (AWS Azure Google Cloud) or DevOps methodologies (e.g. Certified DevOps Professional).
WHAT WE OFFER
Were all about diversity. We operate in 200 countries and speak 60 different languages and cultures. Were really proud of our inclusive environment. Our offices are comfortable and fun places to work and we make sure you get to work from home too. Find out what its like to join our team and take a step closer to your best life ever.
Flex Week: Work from home up to 2 days/week (depending on your teams needs)
Flex Day: Make your workday suit your life and plans.
Flex-Location: Take up to 30 days a year to work from any location in the world.
Employee Wellbeing: We have got you covered with our Employee Assistance Program (EAP) for you and your dependents 24/7 365 days/year. We also offer Champion Health - a personalized platform that supports a range of wellbeing needs.
Professional Development: At SITA we believe growth fuels innovation. Our learning ecosystem offers access to world-class platforms and programs designed to help you thrive. From LinkedIn Learning Microsofts Enterprise Skills Initiative and Airport Council International -available to all employees-to specialized solutions like Pluralsight for technology upskilling Harvard Business Publishing for people leadership Stanford for strategic development and many others we align learning opportunities with your Development Plan and our business priorities. Your development journey is supported every step of the way.
Competitive Benefits: Competitive benefits that make sense with both your local market and employment status.
SITA is an Equal Opportunity Employer. We value a diverse support of our Employment Equity Program we encourage women aboriginal people members of visible minorities and/or persons with disabilities to apply and self-identify in the application process.
Required Experience:
IC
About Company
At SITA we lead one of the most exciting and advanced industries in the world. With us, there are no limits for people looking to explore the edges of possibility and beyond. We are the world’s leading specialist in air transport communications and information technology. Around the ... View more