AI Engineer – Algorithm Evaluation & Agentic Systems
Sunnyvale, CA - USA
Job Summary
Within the DAQ team our core mission is to evaluate and elevate advanced visual technologies. As a key member of this group you will lead the benchmarking and integration of state-of-the-art models for image and video understanding. Rather than focusing on core model training you will apply your deep CV and ML expertise to rigorously test models in applied settings uncover edge-case failure modes and architect advanced agentic systems. If you are passionate about AI safety robust evaluation and building autonomous multi-modal workflows that bridge experimentation with production wed love to hear from you.
Algorithm Evaluation u0026 Benchmarking: Design build and scale comprehensive evaluation pipelines. You will be responsible for both holistic end-to-end system evaluation and granular component-level testing to rigorously measure model capabilities on complex image and video understanding Failure Analysis: Leverage your CV and ML background to dive deep into model outputs identifying root causes of visual hallucinations temporal inconsistencies in video and edge-case Architecture: Build deploy and evaluate agentic workflows that utilize these vision models to autonomously solve multi-step user problems (e.g. video summarization visual search). You will heavily utilize component-level evaluation to isolate and triage exactly which parts of the agentic workflow (e.g. tool selection memory retrieval visual reasoning) are succeeding or Data Curation: Lead the strategy for curating high-quality schematized datasets and ground-truth benchmarks specifically tailored for evaluating multi-modal -Functional Collaboration: Partner closely with the core model training teams. You will provide them with actionable data-driven insights and metrics to guide the next iteration of model training and fine-tuning.
MS and a minimum of 3 years relevant industry experience n3 years of applied experience in Machine Learning Computer Vision or AI System EvaluationnSolid ML Foundation: Deep understanding of core Machine Learning principles including probability statistics data distributions and model bias/variance. You can apply statistical rigor to ensure evaluation metrics are meaningful and Vision Expertise: Deep theoretical and practical understanding of Computer Vision (CV) and Vision-Language Models (VLMs). You must understand how Vision Transformers (ViTs) spatial-temporal modeling and image/video processing work under the hood to effectively evaluate Evaluation Skills: Proven track record of defining robust metrics/KPIs and designing rigorous evaluation frameworks for generative AI or foundation models. Deep experience with custom benchmark creation automated regression testing LLM/VLM-as-a-judge methodologies and human-in-the-loop Systems: Experience building and evaluating LLM/VLM-powered agents including tool use multi-step reasoning planning and memory management Analysis: Strong intuition for probing ML models to discover edge cases hallucinations and performance bottlenecks in constrained environments. Be able to translate findings into actionable improvement Excellence: Strong proficiency in Python and experience with deep learning frameworks (PyTorch) for running inference extracting embeddings and building scalable evaluation pipelines.
Demonstrated ability to lead technical evaluation strategies end-to-end drive architectural decisions for testing infrastructure and mentor engineers. nStrong foundation in statistics including hypothesis testing confidence intervals and experimental designnKnowledge of reinforcement learning planning or decision-making systemsnExperience evaluating multi-modal or multi-agent systemsnPrior work on AI reliability safety or benchmarking
Required Experience:
Unclear Seniority
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more