AI Data Platform Engineer
Cupertino, CA - USA
Job Summary
Design build and maintain scalable AI data platforms services and APIs that support and enable AI model development and data ingestion transformation and publishing pipelines for structured unstructured and multimodal AI-ready datasets through ground truth creation data curation annotation workflows dataset versioning and metadata data quality frameworks validation pipelines observability and evaluation metrics to ensure trusted AI and implement Retrieval-Augmented Generation (RAG) pipelines embedding workflows vector database integrations and metadata services for enterprise AI scalable platform capabilities for managing the end-to-end AI data lifecycle including ground truth dataset creation dataset versioning metadata and lineage management automated data quality validation governance and secure publishing of AI-ready with AI/ML engineers software engineers product teams and domain experts to define AI data requirements and deliver production-ready data platform scalability reliability performance security and cost across cloud-native engineering best practices for AI data architecture platform design automation testing monitoring and operational emerging AI technologies and continuously improve platform capabilities that enable GenAI agentic AI and embodied AI solutions.
Bachelors or Masters degree in Computer Science Software Engineering Data Engineering or a related field.n5 Experience designing and building scalable data platforms and distributed programming skills in Python and SQL with proficiency in Java or Scala with Airflow Kubeflow or MLflow to build and orchestrate scalable AI data building scalable batch and streaming data pipelines using Spark (PySpark) Kafka Airflow and Ray with proficiency in Pandas and modern data lake/lakehouse architectures (e.g. Iceberg Delta Lake).nHands-on experience with AI data engineering including ground truth dataset creation data curation annotation pipelines dataset versioning and metadata implementing data validation quality frameworks observability and AI dataset of RAG architectures embedding generation vector databases and AI data preparation for LLMs and agentic with cloud platforms (AWS Azure or GCP) Kubernetes Docker CI/CD and Infrastructure as understanding of distributed systems APIs microservices and enterprise integration communication collaboration and technical leadership skills.
Experience building platforms supporting GenAI Agentic AI or Embodied AI with multimodal datasets knowledge graphs AI evaluation frameworks or vector search with enterprise data governance lineage metadata management and AI working with manufacturing operational IoT or industrial data ability to lead technical initiatives and mentor engineers.
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more