AI Engineer (Data Guardrails & LLM Ingestion Pipelines)
Job Summary
This is a remote position.
We are looking for a highly skilled AI Engineer to design and build robust data ingestion cleaning validation and LLM enhancement pipelines that power our AI applications. You will transform raw unstructured data into high-quality AI-ready datasets while implementing guardrails that ensure accuracy consistency and reliability.
Design and develop scalable data ingestion pipelines for structured and unstructured data.
Build automated data cleaning normalization and preprocessing workflows.
Develop AI-powered enrichment pipelines using LLMs (OpenAI Claude Gemini etc.).
Implement data quality validation and AI guardrails.
Develop prompt engineering workflows for data transformation.
Build document processing pipelines for PDFs Word documents CSVs websites and APIs.
Develop Retrieval-Augmented Generation (RAG) pipelines.
Create evaluation frameworks for LLM quality and accuracy.
Build ETL/ELT workflows for AI-ready datasets.
Integrate vector databases for semantic search.
Monitor pipeline performance cost latency and data quality.
Collaborate with cross-functional teams to deliver production AI systems.
Python (Expert)
SQL
Git
OpenAI API
Anthropic Claude API
Google Gemini API
Prompt Engineering
Function Calling
Structured Outputs
LangChain
LlamaIndex
DSPy (Preferred)
PydanticAI (Nice to Have)
Pandas
Polars
ETL/ELT Pipelines
Apache Airflow (Preferred)
Data Validation Frameworks
Pinecone
Weaviate
Qdrant
ChromaDB
FAISS
Docker
Kubernetes (Preferred)
AWS / Azure / GCP
Linux
PostgreSQL
MongoDB
Redis
Experience building production-grade AI systems.
Strong understanding of RAG architectures.
Experience implementing AI guardrails and hallucination mitigation.
Experience with OCR and document parsing.
Experience with embedding models and semantic search.
Knowledge of data governance and security best practices.
Build scalable ingestion pipelines.
Deliver automated data cleaning and LLM enhancement workflows.
Implement AI guardrails to improve output quality.
Develop evaluation pipelines for LLM performance.
Contribute to a production-ready AI platform.
Required Skills:
We are looking for a highly skilled Senior Social Media Scraping & Automation Engineer to build a scalable platform for collecting and processing publicly available data from social media platforms and other web sources. You will design reliable scraping infrastructure browser automation API integrations and data pipelines capable of processing large volumes of data efficiently.