Gen Ai Infra
Job Summary
Hi All
Please share the profiles on priority as per the below JD for GEN AI Developer & Tech Lead-Infra:
Grade: 4C/5A
Location: All BSL except Mumbai
Enterprise AI Platform Team Job Description:
The Enterprise AI Platform Team builds and maintains the secure scalable AWS foundation for agentic AI workloads. This team owns infrastructure provisioning gateways security and observability to enable self-service agent development.
Key Responsibilities:
Provision and manage core AWS infrastructure including Bedrock Agent Core runtime and CSP AWS services
Implement identity management with IAM Cognito Okta integration for enterprise SSO
Build platform components: memory stores policy enforcement engines browser tooling
Deploy and operate gateways: MCP Gateway (AgentCore/Kong) LLM Gateway (multimodel routing) Bedrock integration
Manage containers and compute: Docker builds ECR registries EKS clusters EC2 instances
Author IaC with Terraform modules and maintain Jenkins/Terraform Enterprise CI/CD pipelines
Set up RAG data infrastructure using PostgreSQL and Redshift
Develop MCP server tools for APIs and browser integrations
Enable multi-LLM support (AWS Nova OpenAI Anthropic Claude/Opus Gemini)
Implement comprehensive observability: CloudWatch X-Ray Splunk New Relic dashboards and alerting
Design platform APIs supporting REST SOAP GraphQL OAuth2 streaming protocols and A2A communication
Required Skills & Experience:
5 years AWS platform engineering with Bedrock AgentCore EKS/ECR experience
Expert Terraform IaC (modules state management AWS providers) and Jenkins pipeline development
Deep gateway experience (Kong MCP servers LLM routing rate limiting guardrails)
Strong observability implementation across CloudWatch/X-Ray/Splunk/New Relic stacks
Enterprise identity patterns (IAM roles/SCPs Cognito federation Okta SSO)
Distributed systems design for multi-tenant AI platforms with fault tolerance and compliance
Nice to Have:
Previous enterprise AI platform builds (Bedrock Guardrails AgentCore gateway deployments)
Multi-model LLM gateway operations (model routing cost optimization SLAs)