Technical Lead Data Engineer (Data&AI)
Job Summary
Lead Data Engineer
We are looking for a Lead Data Engineer who combines hands-on multi-platform expertise with strong leadership in data architecture pipelines and CI/CD. This role requires a versatile engineer with deep technical skills across modern data platforms (such as Databricks Snowflake AWS and Azure) an understanding of MLOps/DevOps practices and the ability to guide a high-performing team in building scalable production-ready data solutions. You will not be limited to a single platform but will leverage a diverse toolkit to solve complex data challenges.
Pipeline & Architecture: Lead hands-on development of scalable ETL/ELT pipelines data models and integration frameworks to process high-volume (billions of records) structured and unstructured retail data.
Multi-Platform Engineering: Design develop and optimize data processing applications across multiple platforms including Databricks (Spark/Delta Lake) Snowflake AWS or Azure.
Data Integration & Orchestration: Build and manage robust data pipelines using Apache Airflow for orchestration and Airbyte for seamless data integration and movement.
Data Processing: Architect and implement robust solutions for Change Data Capture (CDC) large-scale batch processing and low-latency real-time/streaming data processing.
API Management: Work extensively with external APIs for data ingestion as well as design create and manage internal REST APIs to serve data to downstream applications and users.
AI-Augmented Deliverables: Actively leverage AI assistants to conceptualize design and accelerate the development of data pipelines and everyday engineering tasks.
DevOps & CI/CD: Own and evolve CI/CD pipelines (Git workflows automated testing release cycles secrets management documentation). Guide DevOps-oriented deployments utilizing Dockerized applications Kubernetes orchestration and monitoring/logging tools (Splunk Datadog Dynatrace).
MLOps Alignment: Collaborate with Data Scientists on data readiness for ML projects and ensure alignment with ML lifecycle stages (data prep feature engineering model deployment).
Governance & Leadership: Establish and enforce best practices in data governance data quality metadata and security. Mentor team members through peer reviews knowledge sharing and technical leadership.
Innovation: Stay ahead of industry trends in MLOps observability and GenAI introducing relevant tools and practices.
Experience: 5 years of experience in Data Engineering.
Data Lakes & Warehouses: Mandatory expertise in designing building and managing large-scale Data Warehouses and Data Lakes from the ground up.
Data Processing Paradigms: Extensive hands-on experience working with Change Data Capture (CDC) mechanisms complex batch processing and real-time/streaming data processing.
Platform Expertise: Proven expertise in more than one major cloud data platform/ecosystem (e.g. Databricks Snowflake AWS Analytics Azure Data Engineering).
SQL Mastery: Advanced proficiency in writing optimizing and debugging complex SQL queries for large-scale data processing and analytics.
Programming: Strong programming skills in Python (async threading decorators advanced I/O).
APIs: Strong proficiency in interacting with third-party APIs and hands-on experience creating and managing REST APIs (using frameworks like FastAPI Flask or similar).
Tooling: Deep hands-on experience with workflow orchestration (Apache Airflow) and data integration platforms (Airbyte).
AI-Assisted Engineering: Mandatory capability to use AI coding assistants and tools to design pipelines write code and enhance day-to-day productivity.
Data Architecture: Experience with data modeling (e.g. Delta Lake or Snowflake architecture) and scalable ETL/ELT design.
DevOps/CI/CD: Hands-on experience with Git-based CI/CD (GitLab preferred) and a working knowledge of Docker & Kubernetes for deployment and scaling.
MLOps: Understanding of MLOps concepts including data preparation model lifecycle registries and monitoring.
Soft Skills: Strong problem-solving skills with the ability to design for scale and performance coupled with excellent collaboration communication and leadership skills.
Customer Data Platform (CDP): Experience working with building or implementing CDPs to unify customer data across systems.
Experience in the retail domain or other large-scale data-heavy environments.
Familiarity with streaming frameworks (Kafka Spark Streaming etc.).
Agentic Pipeline Development: Experience or strong interest in building agentic pipelines using LLMs for dynamic data orchestration and automation.
Knowledge of model observability tools and ML deployment pipelines.
Exposure to GenAI concepts (vector embeddings vector databases RAG).
Required Experience:
Senior IC