Enter a job title or keyword

AI Reliability Engineer (SRE) – GenAI & LLM Infrastructure

DMS Vision Inc


Job Location:

Tampa, FL - USA

Monthly Salary: Not provided by the employer
Posted: 26 June 2026 (30+ days ago)
Application Deadline: 23 September 2026
Vacancies: 1 Vacancy

Job Summary

Hi
Hope all is well

Please find the job description given below and let me know your interest.
Position: AI Reliability Engineer (SRE) GenAI & LLM Infrastructure

Location: Tampa FL (Onsite)
Duration : 6 Months

Job Summary

We are seeking an experienced AI Reliability Engineer (SRE) to design deploy and maintain highly available infrastructure supporting GenAI and Large Language Model (LLM) workloads. The ideal candidate will have strong expertise in Kubernetes AI infrastructure observability automation and modern GenAI frameworks.

Key Responsibilities:
  • Design scale and maintain reliable infrastructure for LLM training fine-tuning and inference workloads.
  • Architect and implement agentic AI systems for alert triage root cause analysis and self-healing operations.
  • Manage GPU orchestration cluster health and compute utilization across Kubernetes environments.
  • Define and monitor AI-specific SLOs and SLIs such as Time-to-First-Token (TTFT) latency and cost-per-query metrics.
  • Support vector databases and RAG pipelines to ensure reliability and performance.
  • Implement security controls and guardrails to protect LLM applications from prompt injection hallucinations and data leakage.
  • Automate operational processes and improve infrastructure reliability.
Required Skills:
  • Strong experience with Kubernetes (EKS/GKE).
  • Expertise in Infrastructure as Code using Terraform.
  • Experience with CI/CD pipelines and DevOps practices.
  • Proficiency in Python or Go.
  • Hands-on experience with LLM frameworks such as LangChain LlamaIndex or AutoGen.
  • Experience with vector databases including Pinecone Milvus Qdrant or pgvector.
  • Knowledge of monitoring tools such as Datadog Prometheus and OpenTelemetry.
  • Strong understanding of AI infrastructure automation and observability.
Preferred Qualifications:
  • Experience implementing semantic caching solutions.
  • Knowledge of AI cost optimization strategies.
  • Experience converting operational runbooks into automated workflows.
  • Background supporting large-scale AI and machine learning platforms.
--
Nitin papnai
US Technical Recruiter
Email :
Contact : 1 Ext : 111