AI Reliability Engineer (SRE) – GenAI & LLM Infrastructure
Posted:
26 June 2026 (30+ days ago)
Application Deadline:
23 September 2026
Vacancies:
1 Vacancy
Job Summary
Hi
Hope all is well
Please find the job description given below and let me know your interest.
Position: AI Reliability Engineer (SRE) GenAI & LLM Infrastructure
Location: Tampa FL (Onsite)
Duration : 6 Months
Job Summary
We are seeking an experienced AI Reliability Engineer (SRE) to design deploy and maintain highly available infrastructure supporting GenAI and Large Language Model (LLM) workloads. The ideal candidate will have strong expertise in Kubernetes AI infrastructure observability automation and modern GenAI frameworks.
Key Responsibilities:- Design scale and maintain reliable infrastructure for LLM training fine-tuning and inference workloads.
- Architect and implement agentic AI systems for alert triage root cause analysis and self-healing operations.
- Manage GPU orchestration cluster health and compute utilization across Kubernetes environments.
- Define and monitor AI-specific SLOs and SLIs such as Time-to-First-Token (TTFT) latency and cost-per-query metrics.
- Support vector databases and RAG pipelines to ensure reliability and performance.
- Implement security controls and guardrails to protect LLM applications from prompt injection hallucinations and data leakage.
- Automate operational processes and improve infrastructure reliability.
- Strong experience with Kubernetes (EKS/GKE).
- Expertise in Infrastructure as Code using Terraform.
- Experience with CI/CD pipelines and DevOps practices.
- Proficiency in Python or Go.
- Hands-on experience with LLM frameworks such as LangChain LlamaIndex or AutoGen.
- Experience with vector databases including Pinecone Milvus Qdrant or pgvector.
- Knowledge of monitoring tools such as Datadog Prometheus and OpenTelemetry.
- Strong understanding of AI infrastructure automation and observability.
- Experience implementing semantic caching solutions.
- Knowledge of AI cost optimization strategies.
- Experience converting operational runbooks into automated workflows.
- Background supporting large-scale AI and machine learning platforms.
Nitin papnai
US Technical Recruiter
Contact : 1 Ext : 111