Senior SRE Engineer
Job Summary
We are seeking a Senior SRE Engineer with strong technical authority to influence design and operational decisions across architecture engineering security and operations teams. The ideal candidate is a pragmatic problem-solver who remains calm and methodical under pressure balancing reliability security cost and delivery speed while clearly communicating complex technical concepts to diverse audiences. This role requires onsite presence at the client office three times a week.
- Design and operate observability platforms including monitoring logging and alerting systems
- Manage metrics logs APM and alerting through Datadog
- Apply SRE principles such as SLOs error budgets incident management and reliability engineering
- Collaborate with security teams to uphold cloud security principles
- Apply cloud security best practices to ensure system reliability
- Manage security incidents and drive mitigation and remediation measures
- Monitor optimize and control costs associated with ML models and API-based AI services
- Operate and maintain Kubernetes and containerized platforms
- Manage AWS infrastructure including EKS ECS EC2 networking IAM and managed services
- 6-9 years of hands-on technical experience in SRE Platform Engineering Infrastructure or related roles
- Strong experience with AWS services such as EKS ECS EC2 networking IAM and managed services
- Deep hands-on experience with Kubernetes and containerized platforms
- Proven experience designing and operating observability platforms including monitoring logging and alerting
- Hands-on experience with Datadog for metrics logs APM and alerting
- Strong understanding of SRE principles including SLOs error budgets incident management and reliability engineering
- Solid understanding of cloud security principles and experience collaborating with security teams
- Extensive hands-on experience in cloud security engineering applying best practices to ensure system reliability managing security incidents and driving mitigation and remediation measures
- Experience or working knowledge of FinOps practices including monitoring optimizing and controlling costs associated with ML models and API-based AI services
- Experience supporting multi-cloud or hybrid environments
- Exposure to Infrastructure as Code such as Terraform and CloudFormation
- Experience in large-scale complex or regulated environments
- Knowledge of vector databases and RAG architecture for building internal SRE knowledge assistants
- Knowledge of Generative AI and LLM platforms such as Claude and Amazon Bedrock
Required Experience:
Senior IC