AI Architect Santa Clara, CA Fulltime
Job Location:
Santa Clara County, CA - USA
Monthly Salary:
Not provided by the employer
Posted:
15 September 2026 (3 hours ago)
Application Deadline:
13 December 2026
Vacancies:
1 Vacancy
Job Summary
Job Title: AI Architect
Location: Santa Clara CA
Duration: Fulltime
Key Requirements:Looking for a strong customer-facing resource with excellent communication and stakeholder management skills. Bay Area local candidates are highly preferred.
- Will work on theintelligence layer formultipleprograms owns all model quality RAG accuracy prompt engineering and AI safety acrossapplications
- Socratic tutor persona adaptive learning recommendation engine multi-modal AI (text and voice) RAG evaluation framework and feedback loop into retrieval
- 6-LLM call chain orchestration (NeMoGuardrails intent classification query rewriting RAG synthesis) and compatibility check logic
- Production-grade AI quality from launch this is not a research or prototyping role; accuracy thresholds latency requirements and safety guardrails must pass InfoSec adversarial testing before Release 1
Required Skills
- Total IT 15 Years
- 4-7 years of software engineering with at least 2 years focused on LLM application development in production not research not demos not internal tools with 10 users
- Has shipped an LLM-powered feature or product to production where real users depend on the accuracy and the engineer owns the quality metrics
- Has owned an AI safety or guardrails implementation for a customer-facing product not just added an off-the-shelf filter; designed and tested the safety layer
- Has built RAG evaluation pipelines and used them to make go/no-go release decisions accuracy gating is part of the workflow.
- Has profiled and optimized a multi-step LLM call chain for latency
LLM Application Development
- LLM prompt engineering system prompts few-shot examples chain-of-thought instruction following Expert Must-have
- Multi-step LLM chain orchestration LangChain LlamaIndex or custom orchestration Expert Must-have
- Multi-turn conversation design context window management conversation summarization session memory Advanced Must-have
- Streaming LLM response handling token-by-token streaming partial response rendering Advanced Must-have
- Model selection and benchmarking matching model size to task; balancing latency cost and accuracy Advanced Must-have
RAG Pipeline Design & Quality
- RAG pipeline design chunking strategy embedding model selection retrieval configuration Expert Must-have
- Vector similarity search tuning index parameters similarity thresholds retrieval depth Advanced Must-have
- Reranking cross-encoder rerankers relevance scoring Advanced Must-have
- RAG evaluation frameworks RAGAS TruLens or equivalent; automated eval pipelines Advanced Must-have
- Hybrid search combining dense vector retrieval with BM25 or keyword search Proficient Nice to have
AI Safety & Guardrails
- Prompt injection detection and mitigation Advanced Must-have
- Jailbreak testing and red-teaming LLM systems Advanced Must-have
- Content safety classifier integration Advanced Must-have
- Hallucination detection and mitigation strategies Advanced Must-have
- Topical control enforcing scope boundaries on LLM responses Advanced Must-have
Evaluation & Production Quality
- Automated evaluation pipeline design test set curation metric selection regression detection Advanced Must-have
- A/B evaluation methodology for prompt and model changes Proficient Must-have
- Latency profiling for LLM call chains identifying bottlenecks across multi-step pipelines Proficient Must-have
- Feedback loop design user signal collection signal-to-retrieval-weight integration Proficient Must-have
- Production model monitoring accuracy drift detection quality degradation alerting Proficient Must-have
Development
- Python ML/AI application development async programming Expert Must-have
- API design for AI services streaming endpoints error handling timeout management Advanced Must-have
- Embedding model operations model selection batch embedding index updates Advanced Must-have
Nice to Have
- Adaptive learning systems or personalization engine experience
- Knowledge graph integration with RAG
- Multi-agent orchestration patterns
- ServiceNow API integration
- Prior experience building AI products on NVIDIA infrastructure
Regards
Rajesh
Arrowminds Inc