Enter a job title or keyword

AI Architect Santa Clara, CA Fulltime

Arrowminds Inc


Job Location:

Santa Clara County, CA - USA

Monthly Salary: Not provided by the employer
Posted: 15 September 2026 (3 hours ago)
Application Deadline: 13 December 2026
Vacancies: 1 Vacancy

Job Summary

Job Title: AI Architect

Location: Santa Clara CA

Duration: Fulltime

Key Requirements:Looking for a strong customer-facing resource with excellent communication and stakeholder management skills. Bay Area local candidates are highly preferred.

  • Will work on theintelligence layer formultipleprograms owns all model quality RAG accuracy prompt engineering and AI safety acrossapplications
  • Socratic tutor persona adaptive learning recommendation engine multi-modal AI (text and voice) RAG evaluation framework and feedback loop into retrieval
  • 6-LLM call chain orchestration (NeMoGuardrails intent classification query rewriting RAG synthesis) and compatibility check logic
  • Production-grade AI quality from launch this is not a research or prototyping role; accuracy thresholds latency requirements and safety guardrails must pass InfoSec adversarial testing before Release 1

Required Skills

  • Total IT 15 Years
  • 4-7 years of software engineering with at least 2 years focused on LLM application development in production not research not demos not internal tools with 10 users
  • Has shipped an LLM-powered feature or product to production where real users depend on the accuracy and the engineer owns the quality metrics
  • Has owned an AI safety or guardrails implementation for a customer-facing product not just added an off-the-shelf filter; designed and tested the safety layer
  • Has built RAG evaluation pipelines and used them to make go/no-go release decisions accuracy gating is part of the workflow.
  • Has profiled and optimized a multi-step LLM call chain for latency

LLM Application Development

  • LLM prompt engineering system prompts few-shot examples chain-of-thought instruction following Expert Must-have
  • Multi-step LLM chain orchestration LangChain LlamaIndex or custom orchestration Expert Must-have
  • Multi-turn conversation design context window management conversation summarization session memory Advanced Must-have
  • Streaming LLM response handling token-by-token streaming partial response rendering Advanced Must-have
  • Model selection and benchmarking matching model size to task; balancing latency cost and accuracy Advanced Must-have

RAG Pipeline Design & Quality

  • RAG pipeline design chunking strategy embedding model selection retrieval configuration Expert Must-have
  • Vector similarity search tuning index parameters similarity thresholds retrieval depth Advanced Must-have
  • Reranking cross-encoder rerankers relevance scoring Advanced Must-have
  • RAG evaluation frameworks RAGAS TruLens or equivalent; automated eval pipelines Advanced Must-have
  • Hybrid search combining dense vector retrieval with BM25 or keyword search Proficient Nice to have

AI Safety & Guardrails

  • Prompt injection detection and mitigation Advanced Must-have
  • Jailbreak testing and red-teaming LLM systems Advanced Must-have
  • Content safety classifier integration Advanced Must-have
  • Hallucination detection and mitigation strategies Advanced Must-have
  • Topical control enforcing scope boundaries on LLM responses Advanced Must-have

Evaluation & Production Quality

  • Automated evaluation pipeline design test set curation metric selection regression detection Advanced Must-have
  • A/B evaluation methodology for prompt and model changes Proficient Must-have
  • Latency profiling for LLM call chains identifying bottlenecks across multi-step pipelines Proficient Must-have
  • Feedback loop design user signal collection signal-to-retrieval-weight integration Proficient Must-have
  • Production model monitoring accuracy drift detection quality degradation alerting Proficient Must-have

Development

  • Python ML/AI application development async programming Expert Must-have
  • API design for AI services streaming endpoints error handling timeout management Advanced Must-have
  • Embedding model operations model selection batch embedding index updates Advanced Must-have

Nice to Have

  • Adaptive learning systems or personalization engine experience
  • Knowledge graph integration with RAG
  • Multi-agent orchestration patterns
  • ServiceNow API integration
  • Prior experience building AI products on NVIDIA infrastructure

Regards

Rajesh

Arrowminds Inc