Data Engineer AI JavaPythonSpark
Job Summary
Role Overview:
We are looking for a Senior/Lead Data & GenAI Engineer to design build and maintain scalable production-grade data and AI platforms.
The role combines software engineering distributed data processing cloud-native technologies and Generative AI. You will work across teams drive technical initiatives and build reusable libraries and frameworks that enable reliable scalable and testable systems.
Key Responsibilities:
Develop test and maintain high-quality production-ready software.
Design and implement large-scale data pipelines and distributed processing systems.
Build scalable cloud-native services and platforms using modern engineering practices.
Provide technical leadership for cross-team initiatives and complex engineering projects.
Design and develop reusable libraries frameworks and platform components.
Optimize distributed data processing workloads for performance scalability and reliability.
Work with data platforms including Databricks Apache Spark and Snowflake.
Develop and deploy applications using Python and/or Java.
Build and operate containerized workloads using Kubernetes and cloud-native technologies.
Design and implement GenAI/LLM-based applications and services.
Work with frameworks such as LangChain and LangGraph for LLM orchestration and agentic workflows.
Collaborate with data scientists software engineers architects and product teams.
Establish engineering best practices around testing observability reliability and deployment.
Required Experience:
5 years of professional software/data engineering experience.
Strong hands-on experience with Python and/or Java.
Strong experience with Apache Spark and distributed data processing.
Experience with Databricks and/or modern lakehouse platforms.
Experience with Snowflake or comparable cloud data warehouses.
Practical experience with Kubernetes and cloud-native technologies.
Experience designing and maintaining large-scale data pipelines.
Strong understanding of distributed systems scalability and production engineering.
Experience developing ML/AI or GenAI applications.
Experience with LLM-based applications RAG AI agents or LLM orchestration.
Familiarity with LangChain LangGraph or similar GenAI frameworks.
Strong software engineering fundamentals including testing code quality and system design.
Nice to Have:
Experience with AWS Azure or GCP.
Experience with streaming technologies such as Kafka.
Experience with Delta Lake / Lakehouse architecture.
Experience building RAG pipelines and vector-search solutions.
Experience with LLM evaluation observability and productionization.
Experience with AI agents tool calling and multi-step workflows.
Experience building internal developer platforms frameworks or reusable engineering libraries.
Experience leading cross-functional or cross-team technical initiatives.
Ideal Candidate Profile:
The strongest candidate is not purely a Data Engineer and not purely an ML Engineer.
We are looking for someone who combines:
Software Engineering Data Engineering Cloud/Platform Engineering GenAI
Typical backgrounds may include:
Senior Data Engineer
Lead Data Engineer
Senior Software Engineer Data
Data Platform Engineer
Senior Cloud Data Engineer
AI/ML Platform Engineer
Senior ML Engineer with strong data engineering experience
GenAI Engineer with strong distributed-data/platform experience
Data & AI Architect / Technical Lead
Core Technology Stack:
Languages: Python Java
Data: Apache Spark Databricks Snowflake Delta Lake
Cloud/Platform: Kubernetes Docker AWS/Azure/GCP cloud-native technologies
GenAI/ML: LLMs RAG LangChain LangGraph AI agents vector search
Engineering: Distributed systems APIs CI/CD automated testing observability scalability
About Company
At Infotree, meeting your career needs is a top priority. Client satisfaction is largely dependent on the resources we can provide, and we take pride in our delivery. We have a supportive team in place to give quality people a chance to grow and challenge themselves in their roles whi ... View more