Site Reliability Engineer SAP Business AI Platform (BAIP)
Job Summary
We help the world run better
At SAP we keep it simple: you bring your best to us and well bring out the best in you. Were builders touching over 20 industries and 80% of global commerce and we need your unique talents to help shape whats next. The work is challenging but it matters. Youll find a place where you can be yourself prioritize your wellbeing and truly belong. Whats in it for you Constant learning skill growth great benefits and a team that wants you to grow and succeed.
Meet the team
The CF Backing Services SRE team is part of the SAP Business AI Platform organization responsible for the operational excellence and reliability of business-critical platform services used by thousands of SAP customers globally. Our scope covers services running on both Cloud Foundry and Kubernetes across production landscapes worldwide.
We apply SRE principles in practice - not as a philosophy but as daily engineering work. We own SLOs we run Chaos Days we automate toil and we treat reliability as a feature. We also embrace AI-first tooling as a core part of how we work - from AI-assisted incident response to building and contributing to our own internal SRE tooling.
We are looking for an engineer ready to grow into a well-rounded SRE - someone who is curious technically solid and motivated to take real ownership of the services they support.
What youll build
Live Site Operations
- Youll participate in hotline and on-call rotation responding to incidents and SLO violations across supported services
- Youll investigate production issues with deep technical analysis - log analysis distributed tracing Kubernetes debugging service dependency mapping
- Youll contribute to RCA creation and post-incident follow-up including tracking and closing action items
- Youll participate in Chaos Days and fire drills to proactively test service resilience
Reliability Engineering
- Youll monitor service behavior through SLOs SLIs and the 4 Golden Signals - and act on what you find
- Youll identify and drive improvements to alerting quality - reduce noise increase signal eliminate false positives
- Youll contribute to the teams automation and tooling - scripts Recommended Actions runbooks and internal tools that reduce toil
- Youll support the onboarding of new services into SRE scope - documentation monitoring setup KT sessions access verification
Collaboration and Growth
- Youll work closely with development teams on reliability topics improvements and incident learnings
- Youll contribute to knowledge sharing within the team - KT sessions Show&Tells documentation
- Youll use and contribute to AI tooling actively - AI-assisted workflows and team-built tools
- Youll participate in compliance activities and follow internal processes and procedures
What We Work With
This is not an exhaustive list - it reflects our actual daily environment:
Platform and Infrastructure
- Kubernetes (K8s) Helm Istio - tools for deployment traffic management and debugging
- ArgoCD - GitOps-based continuous delivery for Kubernetes workloads
- Cloud Foundry - active stack hosting business-critical platform services
- AWS GCP Azure - multi-cloud landscape coverage
- Terraform Concourse Jenkins - infrastructure as code and CI/CD pipelines
- Vault Gardener Kyma
- Linux - primary operating environment across all infrastructure
Observability
- Dynatrace - primary observability platform for metrics traces and alerting
- ELK Stack (Elasticsearch Logstash Kibana) - log aggregation search and analysis
- Prometheus Grafana - supplementary monitoring
- Custom alerting layer for SLO violation tracking
Development and Automation
- Python Bash - primary scripting languages for automation and tooling
- GitHub Jira - version control and task management
AI Tooling
- Joule - SAP internal AI assistant integrated into daily workflows
- Claude Code GitHub Copilot - standardized AI-first development workflow
- Perplexity - used for research and technical investigation
- Team-built AI skills workflows and tools distributed through our internal AI marketplace
What you bring
Must have:
- Solid Linux/Unix foundation - comfortable in a terminal understanding of processes file systems and system administration
- Networking fundamentals - TCP/IP DNS HTTP/S load balancing basic network troubleshooting
- Genuine curiosity about how distributed systems fail and how to make them more resilient
- Ability to communicate clearly and precisely - in incidents in writing in cross-team discussions
- Fluency in English - our team and stakeholders are international
- A team-first attitude - we share on-call responsibility and we help each other
Strong advantage:
- Experience or solid knowledge of Kubernetes - deploying debugging understanding what goes wrong
- Familiarity with GitOps and tools like ArgoCD
- Hands-on experience with Python or Bash scripting
- Experience with log analysis using ELK or similar stacks
- Familiarity with observability concepts - SLOs SLIs alerting distributed tracing
- Experience with cloud platforms (AWS GCP or Azure)
- Understanding of CI/CD pipelines and infrastructure as code
- Experience with databases - particularly PostgreSQL
- Experience with incident management - on-call rotations incident response workflows and RCA processes
- Strong troubleshooting and problem-solving skills - ability to analyze symptoms diagnose root causes and implement effective solutions in distributed systems
Mindset we value more than any specific skill:
- You ask why is this breaking not just how do I close this ticket
- You look for patterns across incidents not just fixing individual issues
- You automate things that repeat - you do not accept permanent toil
- You use AI tools actively - you see them as a multiplier not a threat
- You flag problems early and communicate proactively - no surprises
Where you belong
- Real ownership of business-critical platform services from day one - you will not be a ticket processor
- A team that invests in knowledge sharing - KT sessions mentoring and structured onboarding
- On-call rotation with special compensation
- Active AI tooling adoption - you will be working with some of the most advanced AI-assisted SRE workflows in the organization
- An international collaborative environment with direct access to the development teams building the services you support
- Clear career progression path with regular feedback cycles and transparent promotion criteria
Education
BSc in Computer Science Engineering or a related field - or equivalent practical experience. We value demonstrated capability over academic credentials.
We know that strong candidates often do not apply because they feel they do not meet every requirement. If you are curious technically grounded and genuinely motivated to learn - apply. We have seen people grow into excellent SREs from very different starting points.
Bring out your best
SAP innovations help more than four hundred thousand customers worldwide work together more efficiently and use business insight more effectively. Originally known for leadership in enterprise resource planning (ERP) software SAP has evolved to become a market leader in end-to-end business application software and related services for database analytics intelligent technologies and experience management. As a cloud company with two hundred million users and more than one hundred thousand employees worldwide we are purpose-driven and future-focused with a highly collaborative team ethic and commitment to personal development. Whether connecting global industries people or platforms we help ensure every challenge gets the solution it deserves. At SAP you can bring out your best.
We win with inclusion
SAPs culture of inclusion focus on health and well-being and flexible working models help ensure that everyone regardless of background feels included and can run at their best. At SAP we believe we are made stronger by the unique capabilities and qualities that each person brings to our company and we invest in our employees to inspire confidence and help everyone realize their full potential. We ultimately believe in unleashing all talent and creating a better world.
SAP is committed to the values of Equal Employment Opportunity and provides accessibility accommodations to applicants with physical and/or mental disabilities. If you are interested in applying for employment with SAP and are in need of accommodation or special assistance to navigate our website or to complete your application please send an e-mail with your request to Recruiting Operations Team:
For SAP employees: Only permanent roles are eligible for the SAP Employee Referral Program according to the eligibility rules set in the SAP Referral Policy. Specific conditions may apply for roles in Vocational Training.
Qualified applicants will receive consideration for employment without regard to their age race religion national origin ethnicity gender (including pregnancy childbirth et al) sexual orientation gender identity or expression protected veteran status or disability in compliance with applicable federal state and local legal requirements.
Successful candidates might be required to undergo a background verification with an external vendor.
AI Usage in the Recruitment Process
For information on the responsible use of AI in our recruitment process please refer to our Guidelines for Ethical Usage of AI in the Recruiting Process.
Please note that any violation of these guidelines may result in disqualification from the hiring process.
Requisition ID: 458694 Work Area: Software-Development Operations Expected Travel: 0 - 10% Career Status: Professional Employment Type: Regular Full Time Additional Locations: #LI-Hybrid #LI-NM6
Required Experience:
IC
About Company
SAP started in 1972 as a team of five colleagues with a desire to do something new. Together, they changed enterprise software and reinvented how business was done. Today, as a market leader in enterprise application software, we remain true to our roots. That’s why we engineer soluti ... View more