AI Data & Knowledge Engineer
Cupertino, CA - USA
Job Summary
As an AI Data u0026 Knowledge Engineer you will develop infrastructure systems services and tools for automating sales processes. Were looking for an exceptional engineer that lives at the intersection of development operations data and systems engineering to build solutions for large-scale continuous data transformation and delivery. This role will specifically focus on building and maintaining data pipelines for both structured and unstructured data enabling the development and deployment of AIML models.
Responsible for the development and design of data pipelines and data knowledge layers for agentic AI and implement data models for a semantic layer that integrates analytics data from multiple sources in an efficient and effective manner. nDesign and buildscalable data and knowledge layersthat power chatbots and other agentic -ready data pipelines and knowledge layersencompassing document ingestion parsing metadata tagging embeddings indexing with vector scalable architecture forsemantic and hybrid search knowledge graphs to enable contextually accurate text-to-SQL mechanisms forincremental synchronization of data to knowledge updatesso agent responses are current and and operating distributed data systems from SQL/NoSQL databases Vector search and withAnalytics and Data Science teamsto translate business requirements intoreliable actionable knowledge layersthat support AI agent development and deliver targeted business with internal business partners internal technology resources (database system networking) external vendors and partners. nPlay an active role in the development and maintenance of user documentation including data models mapping rules and data dictionaries. nEnsure data quality and accuracy by developing data validation and reconciliation processes. nBuild and maintain data pipelines for ingesting processing and transforming unstructured data sources such as customer feedback social media data or sales call recordings. nDevelop data quality monitoring and validation processes specifically for AIML datasets including identifying and addressing data bias. nWork with data scientists to understand data requirements for AIML model training and deployment ensuring data is available in the appropriate format and quality. nImplement data governance policies and procedures to ensure the responsible and ethical use of data in AIML applications.
Experience designing and building knowledge layers for AI systems including knowledge graphs RAG pipelines and vector databases to ground LLM-driven applications in accurate structured unstructured and retrievable enterprise modeling enterprise knowledge and metadata within semantic layers to represent business entities attributes and their relationships.n5 years of experience in designing building and maintaining scalable data solutions for large-scale in SQL and development experience with cloud database environments like Snowflake Redshift in programming languages like Python Java R and open-source frameworks for distributed processing like Hadoop and building data pipelines to ingest transform and continuously synchronizestructured and unstructured enterprise datafrom multiple -on experience using development tools in a modern cloud data stack for code management versioning using Git CI/CD tools automation and orchestration using Apache Airflow or others and monitoring u0026 alerting. nExperience with Cloud platforms AWS Azure or Google Cloud.
Experience architecting and developing data pipelines through ETL tools API integration with on-premise and cloud-based building ontology-based semantic layer including a business ontology of sales concepts a technical ontology of data sources and schemas and execution traces that provide feedback for continuous understanding of LLM evaluation and AI quality tooling retrieval metrics and observability to improve application working with unstructured and Semi-structured data sets (e.g. JSON Parquet PDF text images audio video)nExperience with data governance and observability tools; for example DataHub Collibra nExperience articulating and translating business questions into data solutions and proven ability to lead development projects from start to knowledge of web standards relating to REST HTTP JSON etc. nExperience with data labeling and annotation tools and processes. nFamiliarity with AI/ML model development lifecycle and data needs for training and to balance competing priorities long-term projects and ad hoc requirements. nAbility to work in a fast-paced dynamic constantly evolving business environment.
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more