IB931 Site Reliability Engineer
Job Summary
Position Summary:
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale distributed fault-tolerant systems enabling online ordering for thousands of restaurants across multiple brands. SRE ensures that Inspire Digital Platform (IDP) services have reliability uptime appropriate to users needs and a fast rate of improvement. Additionally SREs will keep an ever-watchful eye on our systems capacity and performance. SRE is also responsible to perform regular capacity planning exercise. Much of our software development focuses on optimizing existing systems building infrastructure and eliminating toil through automation.
Responsibility:
Technical:
- Review current workload patterns understand the business case and prioritize areas of weakness within the platform through log and metric investigation as well as application profiling.
- Work with senior engineering and testing team members to build tools and recommend testing strategies for problem prevention detection.
- Employ deep troubleshooting skills to improve the availability performance and security to ensure services are designed with 24/7 availability and operational readiness and rigor.
- Perform in depth postmortem on production incidents to assess effective business impact and for Engineering to learn from these.
- Create Dashboards and alerts for Monitoring the IDP platform define key metrics and service level indicators and ensure relevant metric data is collected to create actionable alerts for SRE and Network Operation Center.
- Participate in the 24/7 on call rotation.
- Automate toil by building software and automation for seamless application deployment and third-party tool integration.
- Ensure the platform holds a high degree of reliability at least three 9s.
- Define non-functional requirements as part of the product lifecycle to influence the new designs standards and methods for scalable highly available distributed systems
- own technically intricate issues that cross between DevOps Databases Networking Code Infrastructure and people; drive them to satisfactory completion.
- Provide recommendations and feedback in design reviews and review sessions.
Education:
- 4-yeardegree in computer science Information Technology or related field
Experience:
- Minimum 5 years of experience as a Software Engineer Platform SRE or Devops engineer supporting large scale SAAS Production B2C or B2B Cloud Platforms.
- Hands-on problem-solving and troubleshooting
Knowledge and skills:
- Minimum 5 years of experience as a Software Engineer Platform SRE or Devops engineer supporting large scale SAAS Production B2C or B2B Cloud Platforms.
- Hands on Java application development/support (code) experience
- Hands on Azure Cloud experience particularly with AKS API management Azure Cache for Redis Azure Blob Storage Cosmo DB Service Bus Azure Functions.
- Proficiency in monitoring APM and profiling tools New Relic Splunk Prometheus Grafana.
- Working experience with containers Kubernetes and Helm.
- Functional knowledge of Cloud Network Firewalls Ingress and Egress controllers Service Mesh and
- experience with Auth0 Secret management and Cloudflare CDN Load Balancer Cache Firewall worker features.
- Experience with ArgoCD GitLab CICD Terraform Infrastructure as Code.
- Strong communication skills and ability to explain technical concepts clearly
- A willingness to dive into understanding debugging and improving any layer of the stack
Technical Skills:
- Level of competency 3 on a scale of 5 for skills mentioned below.
Cloud Provider: Azure:
- Core Services: Elasticpool SQL Application Gateway API Management (APIM) Key Vaults AKS (Azure Kubernetes Service) VMSS (Virtual Machine Scale Sets) VM
- Networking: NSG (Network Security Groups) Private Endpoints Private Linked Service VNet Subnets WAF (Web Application Firewall) GeoReplication
- Storage: Storage Accounts
- Messaging and Events: EventHub EventGrid Azure Service Bus (Namespaces Queues Topics)
- Identity and Security: Managed Identities/Workload Identities Private DNS Auth0
Containerization and Orchestration:
- Kubernetes (K8s): For container orchestration
- Helm: For Kubernetes package management
- Docker: For containerization
Monitoring and Observability:
- New Relic / Splunk
Automation and Scripting:
- PowerShell
- Python
Required Skills:
Application Support Java Microservice along with SRE core responsibilities i.e Observability Kubernetes and any monitoring tolls (Splunk New Relic & Dynatrace)